The EyeRobot 2.0 authors present an active-gaze framework for fine-grained bimanual manipulation using a single stereo camera, aiming to give robots a task-focused visual signal without wrist-mounted cameras. For robotics developers, the research suggests that movable stereo viewpoints may offset some of the performance loss associated with removing wrist cameras, based on the paper’s reported physical and simulated trials.
Active visual fixation
The system swivels two eye viewpoints to center their gaze on a 3D fixation point. It then processes the images foveally, assigning more visual tokens to image centers where the selected task-relevant features should appear. The authors call this approach Active Visual Fixation.
Gaze and gripper policies
EyeRobot 2.0 separates gaze control into a low-level servoing policy conditioned on a goal object and a target selector that chooses fixation goals as the task advances. Both are trained with reinforcement learning on real-world data; the target selector co-trains with a behavior-cloning gripper policy. The paper also expresses gripper information in a fixation-relative SE(3) frame, which the authors say makes the action distribution more compact for learning.
Trial design
The evaluation drew on teleoperation data from seven real-world tasks and six simulated tasks. The authors report more than 1,000 physical trials and 1,800 simulated trials, comparing the system with passive stereo and ego-plus-wrist-camera policies trained on the same data.
Reported results
According to the paper, standard passive stereo saw real-world success fall from 52% to 27% when wrist cameras were removed. EyeRobot 2.0 outperformed passive stereo by 40% in real-world tests and 20% in simulation. With clear wrist views, it reported 69% success versus 64% for the ego-plus-wrist baseline; when grasped objects occluded wrist cameras, the reported figures were 48% versus 22%.
Research status
These results are reported in an arXiv submission rather than an independently verified product benchmark. The paper was submitted on 2 October 2026 and is classified by arXiv under Robotics and Artificial Intelligence. Source: arXiv
Definition. Active Visual Fixation is EyeRobot 2.0's method of directing stereo viewpoints toward a 3D fixation point and concentrating visual processing near the image center.
| Condition | Reported result |
|---|---|
| Passive stereo, wrist cameras removed | Real-world success fell from 52% to 27% |
| EyeRobot 2.0 versus passive stereo | 40% higher in real-world tests; 20% higher in simulation |
| Clear wrist views | 69% EyeRobot 2.0 versus 64% ego-plus-wrist baseline |
| Occluded wrist views | 48% EyeRobot 2.0 versus 22% ego-plus-wrist baseline |
Key takeaways
- The system uses movable stereo viewpoints to focus on task-relevant 3D points.
- Its gaze controller combines a low-level servoing policy with a target selector trained using reinforcement learning.
- The evaluation used teleoperation data from seven real-world tasks and six simulated tasks.
- The paper reports over 1,000 physical trials and 1,800 simulated trials.
- When wrist cameras were occluded, reported success was 48% for EyeRobot 2.0 versus 22% for the ego-plus-wrist baseline.
- The findings are from an arXiv submission, not an independently verified product benchmark.
FAQ
What is EyeRobot 2.0?
EyeRobot 2.0 is an active-gaze framework for fine-grained bimanual manipulation that uses a single stereo camera rather than wrist-mounted cameras.
How does Active Visual Fixation work?
It swivels two eye viewpoints toward a 3D fixation point and allocates more visual tokens near image centers where task-relevant features are expected.
How did EyeRobot 2.0 compare with passive stereo?
The paper reports that EyeRobot 2.0 outperformed passive stereo by 40% in real-world tests and 20% in simulation.
What happened when wrist cameras were occluded?
The paper reports 48% success for EyeRobot 2.0 and 22% for the ego-plus-wrist baseline when grasped objects blocked wrist-camera views.