Fisheye-VLA: Decoupling Coverage and Acuity for Manipulation with a Single Fisheye Camera

Ziang Ren1,2, Zike Yan1,2, Raymond Zhang3,4, Xuguo He3, Zhongyu Li1,2,*
1Hong Kong Embodied AI Lab, 2The Chinese University of Hong Kong, 3DeepCybo, 4University of Washington

*Corresponding author

Fisheye-VLA overview

A fixed fisheye camera provides a global view and two interaction-centered policy views over a wide workspace.

Fisheye-VLA uses one fixed fisheye camera to preserve the full workspace while rendering two end-effector-centered perspective crops for interaction detail—without physical wrist cameras.

Overview Video

System Demonstrations

Robot execution, policy observations, and human data collection.

Robot Manipulation

Shelf-to-Box Retrieval

Retrieving an object from an elevated shelf and transferring it to the box.

Toy-to-Shelf Placement

Picking up the toy and placing it on the elevated target shelf.

Wide-Workspace Fruit Transfer

Selecting and transferring fruit across the expanded tabletop workspace.

Bin-to-Conveyor Transfer

Placing fruit onto the moving conveyor under the same global/local policy.

Observation & Data Collection

Policy Input

Global + Local Policy Views

The fisheye overview and two tracked end-effector crops used by the policy.

Human Data Collection

Handheld Demonstration Collection

Tracked handheld gripper poses reuse the same global/local rendering interface.

Abstract

Manipulation requires both broad scene awareness and detailed local feedback, yet conventional camera rigs provide them through separate front and wrist cameras. We present Fisheye-VLA, a visual interface that brings these capabilities together using a single passive fisheye. A global view preserves the workspace, while local perspective crops direct detail toward the interaction.

A controlled re-rendering study finds that end-effector-centered views capture most of the estimated benefit of a much larger candidate pool. Calibrated end-effector projection and motion lead track the crops, while a shared ray encoding preserves their spatial meaning as they move. Integrated with a pretrained VLA, the interface achieves 84% and 82% success in the two expanded tabletop regions and supports shelf and conveyor manipulation without physical wrist cameras.

Method

Fisheye-VLA pipeline from end-effector poses and a fisheye image through local crop rendering, ray encoding, and action prediction.
Current end-effector poses predict viewing centers. Calibrated inverse perspective sampling renders two local views, while the global image preserves scene context. Source-ray features place all visual tokens in a shared geometric reference frame before action prediction.

Results

84%Expanded region R2
82%Expanded region R3
92.5%Full-candidate gain captured by two EE crops
1.3 msMedian crop rendering time
Nested tabletop regions and endpoint-error reduction for alternative crop placements.
End-effector-centered views account for most of the estimated reduction in endpoint error.
Execution sequences for tabletop, shelf, and conveyor manipulation tasks.
The same global/local interface supports four task families across different spatial and timing demands.

Acknowledgments

This work was in part supported by the InnoHK initiative of the Innovation and Technology Commission of the Hong Kong Special Administrative Region Government via the Hong Kong Centre for Logistics Robotics. We also thank DeepCybo for its support, and Hualong Liu, Peize Li, and Zhekai Wang for their assistance.

BibTeX

@article{ren2026fisheyevla,
  title   = {Fisheye-VLA: Decoupling Coverage and Acuity for Manipulation with a Single Fisheye Camera},
  author  = {Ren, Ziang and Yan, Zike and Zhang, Raymond and He, Xuguo and Li, Zhongyu},
  journal = {arXiv preprint},
  year    = {2026}
}