Replicating and Extending Published Research using Meta Quest 3 V4 Headsets

Figure 1: Delayed-onset design
Figure 5: Feature trajectories, Experiment 2

Figures 1 and 5 from Stoner and Blanc (2010). Figure 5 shows both conventional depiction of delayed-onset design and the feature-trajectory depiction that allows easier appreciation of both spatiotemporal dot continuity and feature history. We initally wanted to replicate Stoner and Blanc's findings for the standard (i.e. no feature swaps) conditions and the motion/color swap conditions. Figure 5 A and C show the no motion/color swap conditions and Figure F and H show the motion/color swap conditions.

Stoner & Blanc (2010) presented observers with two overlapping fields of randomly distributed dots rotating in opposite directions (clockwise and counterclockwise) within a circular aperture 4.0° in diameter. Dot density was 5 dots/deg², yielding ~63 dots per field; each dot subtended 0.03° (1 pixel). Rotation speed was 81°/sec. One field appeared 750 ms before the other — the delayed-onset field served as the attentional cue. After 300 ms of dual rotation, one field briefly translated (2.26°/sec, 40 ms, 3 frames at 75 Hz) in one of 8 equally-spaced directions. Roughly 40–55% of the translating dots moved coherently; the remainder were distributed across the other 7 directions. Observers judged translation direction. On swap trials the non-translating field reversed rotation direction (motion swap), the two fields exchanged colors (color swap), or both. Stimuli were presented on a Trinitron CRT at 75 Hz; observers sat in a chin/forehead rest at 57 cm.

VRDots (Meta Quest 3 replication) (`Exp_StonerBlanc_Replication`, `StonerBlanc_Replication_v1`) preserves the core design as closely as the hardware permits. The aperture is 2.0° radius (4.0° diameter), matching S&B exactly. Dot density is 5 dots/deg², yielding 63 dots per field — the same count as the original. Each dot subtends 0.08°, larger than S&B's single-pixel dots owing to Quest 3 pixel pitch. Rotation (81°/sec) and translation (2.26°/sec) speeds are identical. Translation duration is 44 ms (4 frames at 90 Hz), closely matching S&B's 40 ms (3 frames at 75 Hz). Coherence is 50% per field. Stimuli are rendered binocularly at 90 Hz at 2 m virtual viewing distance, with both fields on the same depth plane (zero disparity). The point fixation of S&B is replaced by a ring-and-crosshair target scaled to subtend the same angular region.

The principal stimulus differences are display medium (flat CRT → stereoscopic VR headset), dot size (0.03° → 0.08°), coherence structure (40–55% random → 50% fixed), fixation type (point → ring-and-crosshair, exclusion radius 0.396°), and refresh rate (75 → 90 Hz). Aperture size, dot density, rotation and translation speeds, delayed onset, and trial timing are matched to the original.

Live stimulus demos — motion at ½ speed, translation duration 4× actual (176 ms vs 44 ms). Red field is always-on; green field has delayed onset (750 ms) and serves as the attention cue. On cued trials the green (delayed-onset) field translates; on uncued trials the red (always-on) field translates.

No Swap — Cued

Red field onset precedes green by 750 ms

No Swap — Uncued

Both fields onset simultaneously

Color + Motion Swap — Cued

Features exchange during translation

Color + Motion Swap — Uncued

Swap without temporal onset cue

Figure 7: Results of Experiment 2

Figure 7 from Stoner and Blanc (2010) shows data for the four different swap conditions: No motion/color swaps (A and C), Motion swap (D and B), Color swap (E and G), and Motion/Color swap (H and F). We replicated A,C,H, and F conditions and collected data with Quest 3 headset.

VRDots replication of S&B Experiment 2

Quest headset data from one subject (GS). We qualitatively replicated Stoner and Blanc's findings for both the standard (no swaps) delayed onset design and for the motion/color swap design: Performance (motion direction discrimination) was better for cued vs uncued conditions for both swap conditions. These results imply that its the delayed dots that enjoy the processing advantage regardless of global feature configuration.

Çatak et al. (2022)

Çatak et al. (2022) used the same basic paradigm as Stoner & Blanc (2010) with one key difference in translation duration: 133 ms (0.30° displacement) rather than S&B's 40 ms. Aperture was 3.3° diameter (1.65° radius), dot density 5 dots/deg² (~43 dots/field), dot size 0.05°. Rotation (81°/s) and translation (2.26°/s) speeds, delayed onset (750 ms), and pre-translation period (300 ms) were identical to S&B. Stimuli were presented on a 60 Hz CRT at 57 cm. Importantly, motion and color swaps were tested as separate conditions (not combined). Fifteen naïve observers participated. Key results: no-swap cueing advantage = +20.2 percentage points (pp); motion swap reduced cueing by ~49% (+10.4 pp, p = .018); color swap by ~34% (+13.4 pp, p = .049).

VRDots Çatak replication (`Exp_SubfieldSwap_CatekExact_NDb`, `SubfieldSwap_CatekExact_NDb_v1`) preserves the aperture (1.65° radius), dot density (5 dots/deg², 43 dots/field), rotation and translation speeds, and delayed onset. Translation duration is 80 ms — shorter than Çatak's 133 ms and closer to S&B's 40 ms. Motion and color swaps are applied simultaneously as a single combined condition rather than separately. Dot size is 0.08° (larger, due to Quest 3 pixel pitch). Stimuli are rendered binocularly at 90 Hz at 2 m virtual viewing distance.

The principal differences are display medium (CRT → VR), translation duration (133 → 80 ms), swap structure (separate motion/color → combined), dot size (0.05° → 0.08°), and refresh rate (60 → 90 Hz). The VRDots replication yields a no-swap cueing advantage of +14.6 pp () and a motion/color-swap cueing advantage of +11.5 pp () — replicating the direction and approximate magnitude of Çatak's motion-swap result (+10.4 pp) and the ~20–50% reduction in cueing.

Live stimulus demos — cued trials only, motion at ½ speed, translation duration 4× actual (320 ms vs 80 ms). Aperture 1.65° radius, 43 dots/field. The green field has delayed onset and translates; the red field is always-on and competes. The four conditions differ in what happens at translation onset: N = nothing changes; M = competing field reverses rotation direction (and stays reversed); C = both fields exchange colors; MC = both simultaneously.

No Swap

N — cued

Motion Swap

M — cued

Color Swap

C — cued

Motion+Color

MC — cued

Çatak et al. (2022) — Figure 3

Çatak et al. (2022) Figure 3: Behavioral results

Çatak et al. (2022) Figure 3: Behavioral results for the no-swap (N), motion-swap (M), and color-swap (C) conditions. Bars show percent correct for cued and uncued trials; error bars are 95% CIs across observers (n=15). The cueing advantage (cued > uncued) is present in all three conditions, indicating that neither the motion-swap nor the color-swap eliminates the delayed-onset advantage. The VRDots replication (right) tests a combined motion+color swap (MC) at Çatak-matched aperture and density parameters.

VRDots Replication

VRDots Çatak replication (SubfieldSwap_CatekExact_NMoCol_v1, aperture 1.65° radius, 43 dots/field, 80 ms translation). Five sessions, observer G.S. Bars show mean % correct ± 95% CI (normal approximation, n = 320 per arm). Blue = cued (delayed-onset field translates); orange = uncued. Brackets show cueing effect (Δpp) with two-proportion z-test significance.

0255075chance% Correct*No SwapΔ = +10.0 pp (p = .010)***Motion SwapΔ = +15.3 pp (p < .001)***Color SwapΔ = +17.5 pp (p < .001)***M+C SwapΔ = +14.7 pp (p < .001)Cued (delayed onset)Uncued

2³ Factor Analysis — OLS linear probability model

Each trial is assigned to three ±1 factors: F1 = onset dot cue (delayed vs. non-delayed field translates); F2 = color on translating field (delayed-onset color vs. non-delayed); F3 = competing rotation direction (non-delayed/older vs. delayed/newer). Together these form a full 2³ factorial — all eight cells are present and the factors are orthogonal. N = 2,560 trials. Effects reported as full ±1 contrast (2β × 100 pp) from an OLS linear probability model.

TermEffect (pp)95% CIp
F1: onset dot cue+14.4[+10.6, +18.2]< .001***
F2: color identity-1.7[-5.5, +2.1].372n.s.
F3: competing rotation-0.6[-4.4, +3.2].746n.s.
F1 × F2-1.6[-5.3, +2.2].417n.s.
F1 × F3+1.4[-2.4, +5.2].465n.s.
F2 × F3+1.3[-2.5, +5.0].516n.s.
F1 × F2 × F3-2.0[-5.8, +1.8].292n.s.

Only the onset dot cue (F1) is significant. Neither the color identity of the translating field (F2) nor the direction of the competing rotation (F3) affects accuracy, and neither interacts with F1. The cueing advantage is fully accounted for by which field initiated motion with a temporal delay — consistent with surface-based attentional selection and replicating the Çatak et al. (2022) and Stoner & Blanc (2010) findings in VR.

Dot Density

How does the cueing effect depend on dot density? Using four conditions at fixed aperture (3.5° radius, 7° diameter) and varying only dot density — ~1.6, ~4.5, 13, and 26 dots/deg²/field — the cueing advantage is strikingly flat across an 8× range (~1.6–13 dots/deg²/field), then drops ~10 pp at the highest density tested (26 dots/deg²/field). Critically, it is the cued arm that falls at the highest density, not the uncued arm that rises. This is consistent with a V1 receptive-field mixing account: at very high density, multiple dots from both fields fall within single RFs, degrading the cued field's neural segregation. The effect remains highly significant even at 26 dots/deg²/field.

Demos below show the cued condition (green field = delayed onset, translates) at each density. Aperture, rotation speed, translation speed, and timing are identical across conditions.

VRDots

~1.6 dots/deg²/field — cued

HighDens

~4.5 dots/deg²/field — cued

Peak

13 dots/deg²/field — cued

UltraHigh

26 dots/deg²/field — cued

0255075chance% Correct***VRDots~1.6 dots/deg²Δ = +34.8 pp (p < .001)***HighDens~4.5 dots/deg²Δ = +33.6 pp (p < .001)***Peak13 dots/deg²Δ = +34.8 pp (p < .001)***UltraHigh26 dots/deg²Δ = +25.0 pp (p < .001)CuedUncued

All four conditions: aperture 3.5° radius, 80 ms translation, 750 ms delayed onset, N (no swap), single observer (G.S.). Bars show mean % correct ± 95% CI (normal approximation, n = 256 per arm). Blue = cued (delayed-onset field translates); orange = uncued. Brackets show cueing effect (Δpp) with two-proportion z-test significance. Chance = 12.5% (1/8 directions). Densities quoted per condition are estimates — nominal dot count divided by aperture area (N/π·3.5²) — and do not yet account for any central fixation exclusion in the Unity assets; the exclusion would raise them ~1.3%.

Motion+Color Swap at High Density

Does the object-based cueing advantage survive motion+color swaps at high dot density? Two experiments tested N (no swap) and MC (motion+color swap) conditions at the two highest density levels from the parametric series, using identical aperture (3.5° radius) and timing.

At Peak density (13 dots/deg²/field) the cueing advantage survives the MC swap robustly — consistent with the Çatak et al. replication at lower density. But at UltraHigh density (26 dots/deg²/field) the MC swap completely abolishes cueing: the cued and uncued arms become statistically indistinguishable. This is a striking density × swap interaction. One session per condition; n = 256 per arm.

0255075chance% CorrectPeak13 dots/deg²UltraHigh26 dots/deg²***NΔ = +38.7 ppp < .001***MCΔ = +16.8 ppp < .001***NΔ = +27.3 ppp < .001n.s.MCΔ = +0.8 ppp = .854CuedUncued

Aperture 3.5° radius, 80 ms translation, 750 ms delayed onset, single observer (G.S.), one session per condition (n = 256 per arm). N = no swap; MC = simultaneous motion+color swap at translation onset. Blue = cued; orange = uncued. Chance = 12.5%. The N condition cueing advantage drops from +38.7 pp to +27.3 pp across densities; the MC advantage drops from +16.8 pp to +0.8 pp (n.s.). N-condition values differ slightly from the parametric series above because these are separate sessions.

Modeling

Re-implementing the published accounts, then testing them against the VR data

Valdés-Sosa et al. (2000) introduced a transparent-motion design that has spawned numerous subsequent psychophysical and neurophysiological studies of object-based attention*. Taken together, these studies appear to provide some of the best evidence for object-based selection (Reynolds, Alborzian, & Stoner, 2003; Çatak, Özkan, Kafalıgönül, & Stoner, 2022). Although there are some design differences in the various studies inspired by Valdés-Sosa and colleagues' original study, all of these studies used stimuli composed of two superimposed counter-rotating dot fields. One of the two dot fields is “cued”, either endogenously (e.g., fixation point color indicating the color of the field to be attended if the dots are differently colored) or exogenously (e.g., by a delayed-onset of one dot field, see below). The rotations of the dot fields are interrupted by brief translations (one or two translations, depending upon the design), and subjects are asked to report the direction of those translations. These translations consist of a subset of the dots moving coherently in typically one of eight directions. To discourage tracking of individual dots, the remaining dots of the translating dot field translate in randomly chosen directions. Numerous studies using this basic design have repeatedly found that subjects judge translations of the cued dot field more accurately than translations of the uncued dot field (Khoe et al., 2005; López et al., 2004; Mitchell et al., 2003; Mitchell, Stoner, & Reynolds, 2004; Pinilla et al., 2001; Reynolds et al., 2003; Rodríguez et al., 2002; Stoner & Blanc, 2010; Valdés-Sosa, Cobo, & Pinilla, 2000). Several of these studies also reported electrophysiological (ERP) correlates of the cueing advantage (Valdés-Sosa, Bobes, Rodríguez, & Pinilla, 1998; Pinilla et al., 2001; López, Rodríguez, & Valdés-Sosa, 2004; Rodríguez & Valdés-Sosa, 2006; Çatak et al., 2022).

As observed by Stoner & Blanc (2010), for the object-based account to hold for these experiments, the attentional cue — whether endogenous, exogenous, or both — must first privilege the subset of overlapping dots rotating in one direction and then privilege those same dots when they translate. The visual system must therefore "bind" the cued rotation to the subsequent translation in an object-specific manner. The mechanism of such binding has never been identified, nor has any directly applicable model been advanced. In this section, we first question the need for an object-based account. We demonstrate that established feature-based biased-competition and normalization models appear to account for previous results. We then show that these models cannot, however, account for the results of "feature-swap" experiments devised by Stoner and Blanc. We then consider computational models of incremental complexity that might account for those results as well as other findings attributed to object-based attention in transparent-motion displays.

* The term "surface-based attention" is sometimes used interchangeably in this literature, emphasizing perceptual surfaces rather than discrete objects as the units of selection; the two framings make similar predictions in transparent-motion paradigms.

Biased-competition circuit (Reynolds et al. 1999) mapped onto MT/MST and V1

Figure 4. Inputs feeding the biased-competition circuit. (A) The directional time-varying input into the model for a cued, no-swap trial (see Figure 3). (B) The three direction-selective model neurons of "Stage 1" receive this excitatory directional input. We can think of these three neurons as belonging to a single MT hypercolumn (see Figure 2). These three neurons each have a directional preference corresponding to one of the three predominant directions within the collective receptive field of that MT hypercolumn: Down, Right, and Up. The responses of these three Stage 1 neurons adapt over time, mediated by recurrent inhibition from the I unit and modeled by the equations shown. These adapted responses then feed into the Stage 2 translation detector R₍TD₎, with R₍T₎ providing excitatory (W⁺) input and R₁ and R₂ providing inhibitory (W⁻) shunting input. Taken together, these interactions implement the biased-competition computation of Reynolds et al. (1999). In this cued example, Down appears first and is therefore more adapted by translation onset, while Up arrives later. Hence the translation competes with a weaker inhibitory input than occurs with the uncued condition. In our simple model, we take the output of this translation detector as our measure of how well the model detects the translation (see Figure 5).

Feature trajectories, no-swap and motion-swap — Stoner & Blanc (2010) Figs. 1B & 4

The stimulus in feature coordinates — replication of Stoner & Blanc (2010), Figs. 1B & 4. Each dot field is a trajectory through three motion tracks (clockwise rotation, translation, counter-clockwise rotation). Following S&B's convention, line style = field identity (dotted = the test field that translates; solid = the competitor), line colour = dot colour, and vertical position = type of motion. The same (dotted) test field translates in every condition — only the onset (which field is delayed) and the swap differ. No swap (left column): the test field briefly notches into TRANS and returns to its rotation. Motion swap (right column): at translation the two fields exchange rotation directions — the test field keeps its identity (still dotted) but ends on the opposite track. The consequence is decisive: after a swap, a cued-swap trial's motion content matches an uncued-no-swap trial, forcing any direction-only model to predict that the cueing advantage reverses.

Directional inputs S(t) across all four trial types

Directional inputs S(t) — what actually drives the model, across all four trial types. The feature trajectories re-expressed as the binary drive to each motion channel (rotation drive = 50, translation = 25; Mode 1). The test field (dotted) drives its current rotation track plus the brief TRANS pulse, with a gap where its rotation is interrupted by the translation; the competitor (solid) drives the other rotation track. Under a motion swap the test field keeps its identity (dotted) but its drive shifts to the opposite rotation track. These per-channel drives are the model's complete input — and the stage at which the two surfaces' identities collapse into direction-only signals.

Model responses across all four trial types — cued advantage and its reversal under motion swap

Model responses across all four trial types. Top row: the two rotation channels — dotted marks the channel interrupted by translation (it shows the brief gap), solid is the other. Middle row: the detector's two inputs — the summed normalization pool I = R₁ + R₂ (black) and the excitatory translation drive E (blue, dotted). Bottom row: the translation-detector response R₍TD₎ with its peak. Under no swap the cued field gives the larger response (61 vs 46 — a +33% advantage). Under motion swap the cued-swap and uncued-swap responses are identical to the opposite no-swap conditions, so the advantage reverses (46 vs 61, −25%). The behavioral data show no such reversal — the model's clean falsification, and the motivation for surface-identity selection.

Colour-swap feature trajectories — identity (line style) is invariant; the model is colour-blind

Colour swap — and why the model is colour-blind. Colour (the dots' red/green) is carried by line colour; field identity by line style. A colour swap flips the line colour at translation while the line style — the identity — is untouched, and it can be combined with a motion swap or applied on its own, fully independently. Our current model has no colour input: it is deliberately colour-blind. That is acceptable because colour differences between the surfaces have been shown not to be necessary for the surface-selection effect — the cueing advantage survives when the two fields share the same colour. Colour is shown here only to make the identity-versus-feature dissociation explicit (line style stays fixed while colour and motion vary); it changes none of the model's predictions.