Shelf-to-Box Retrieval
Retrieving an object from an elevated shelf and transferring it to the box.
Robot execution, policy observations, and human data collection.
Retrieving an object from an elevated shelf and transferring it to the box.
Picking up the toy and placing it on the elevated target shelf.
Selecting and transferring fruit across the expanded tabletop workspace.
Placing fruit onto the moving conveyor under the same global/local policy.
The fisheye overview and two tracked end-effector crops used by the policy.
Tracked handheld gripper poses reuse the same global/local rendering interface.
Manipulation requires both broad scene awareness and detailed local feedback, yet conventional camera rigs provide them through separate front and wrist cameras. We present Fisheye-VLA, a visual interface that brings these capabilities together using a single passive fisheye. A global view preserves the workspace, while local perspective crops direct detail toward the interaction.
A controlled re-rendering study finds that end-effector-centered views capture most of the estimated benefit of a much larger candidate pool. Calibrated end-effector projection and motion lead track the crops, while a shared ray encoding preserves their spatial meaning as they move. Integrated with a pretrained VLA, the interface achieves 84% and 82% success in the two expanded tabletop regions and supports shelf and conveyor manipulation without physical wrist cameras.
This work was in part supported by the InnoHK initiative of the Innovation and Technology Commission of the Hong Kong Special Administrative Region Government via the Hong Kong Centre for Logistics Robotics. We also thank DeepCybo for its support, and Hualong Liu, Peize Li, and Zhekai Wang for their assistance.
@article{ren2026fisheyevla,
title = {Fisheye-VLA: Decoupling Coverage and Acuity for Manipulation with a Single Fisheye Camera},
author = {Ren, Ziang and Yan, Zike and Zhang, Raymond and He, Xuguo and Li, Zhongyu},
journal = {arXiv preprint},
year = {2026}
}