VersaCamVLACamera-Configurable VLA Policies for Robotic Manipulation

Boyao Han1,2* Chen Shi1* Jingjing Qian1 Zhuotao Tian2,3 Li Jiang1,2†

1 The Chinese University of Hong Kong, ShenzhenCUHK-Shenzhen
2 Shenzhen Loop Area InstituteSLAI3 Harbin Institute of Technology, ShenzhenHIT-Shenzhen

* Equal contribution    † Corresponding author

Abstract

Vision-language-action (VLA) policies are powerful foundations for robotic manipulation, but they often depend on a fixed camera configuration. Changing a camera's position or adding a new view can disrupt the visual interface learned during training. We introduce VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning. A unified scene-token interface maps a variable set of calibrated RGB views into a fixed-size representation. Multi-signal target-view prediction preserves visual and spatial structure, while Wrist-Augmented Pose Sampling uses natural wrist-camera motion to provide viewpoint diversity. A lightweight spatial encoder then supplies compact scene tokens to a pretrained VLA as an additional visual condition. At deployment, the policy requires neither explicit 3D reconstruction nor novel-view rendering. Experiments on RoboTwin 2.0, LIBERO, and a real dual-arm robot demonstrate robust manipulation across unseen camera poses and different numbers of input views.

VersaCamVLA handles different camera configurations, including varying camera counts and poses.

Method

VersaCamVLA first learns a scene-token interface from multi-view RGB and camera geometry, then conditions a pretrained VLA with compact scene tokens.
Scene representation learning and policy learning are separated by a fixed-size scene-token interface.

Learning a camera-configurable representation. A scene encoder combines RGB observations and camera rays into shared scene tokens. A training-only decoder predicts RGB, semantic, and edge targets across available views. Wrist-Augmented Pose Sampling varies the source camera subset, using wrist motion to broaden pose coverage within existing demonstrations.

Connecting scene tokens to actions. The pretrained scene encoder is frozen and its decoder removed. A lightweight spatial encoder compresses the representation into 192 supplementary tokens for π0.5, alongside the policy's native visual, language, and proprioceptive inputs. The scene-token pathway supports varying camera counts while preserving a consistent input size for the action policy.

Results

Benchmark performance and camera-configuration results from the paper. All values are task success rates (%).

Benchmark performance

Table 1. Simulation results on RoboTwin 2.0. Clean and domain-randomized (DR) settings across 16 tasks. Bold indicates the best performance in each column.

Scroll horizontally to see all columns.

TaskACTDPDP3OpenVLA-OFTπ₀.₅VersaCamVLA
CleanDRCleanDRCleanDRCleanDRCleanDRCleanDR
Blocks Ranking RGB0000200038116712
Blocks Ranking Size30102170124356
Hanging Mug6019035112044104
Move Stapler Pad000070005785
Open Laptop72153077182090709662
Open Microwave7207809125235643265146
Pick Diverse Bottles802805519036214712
Place A2B Left10303015042245417
Place A2B Right206043011024194527
Place Burger Fries67080075230048318149
Place Dual Shoes0040902025245125
Place Empty Cup3802707512007718127
Place Object Basket102104935053267635
Scan Object308029181643515
Stack Blocks Three500030001850433
Stack Bowls Three37051057429053246123
Average (%)19.690.0623.690.0039.942.5615.193.5635.8821.6352.5623.00

Robustness to camera-pose variations

Table 2. Robustness to camera-pose variations. On RoboTwin 2.0, both policies are trained under the nominal camera setup and evaluated with camera poses perturbed in azimuth, distance, and pitch over a workspace-centered spherical region, probing generalization to unseen viewpoints.

MethodRoboTwin 2.0 Clean
Seen PoseUnseen Pose
π₀.₅48.8832.88
VersaCamVLA53.4452.63

Flexibility across input camera counts

Table 3. Effect of input camera count. On RoboTwin 2.0 Clean, we train two π0.5 baselines separately with 3 and 4 views and a single VersaCamVLA policy. We evaluate VersaCamVLA with 3, 4, 5, and 6 input views and each π0.5 baseline at its matched view count, keeping camera viewpoints aligned during testing to isolate the effect of camera count.

MethodRoboTwin 2.0 Clean
3V4V5V6V
π₀.₅35.8848.88××
VersaCamVLA52.5653.4453.1355.63

V denotes input views; × denotes settings not reported.

Camera Pose

RoboTwin policy executions at 4× speed. Head, left wrist, right wrist, and Agent views are synchronized across the top row, above a close-up third-person view. The moving camera and its trajectory are shown in red.

Camera Count

One policy with 3, 4, 5, or 6 input cameras. The input views appear above the cropped third-person execution. Colored trajectories identify the additional external cameras. All videos are shown at 4× speed.

Real-world

Front-view recordings of three bimanual tasks, shown at 4× speed.

Pick CubeFront
Stack CubesFront
Insert TubeFront

BibTeX

If you find this work useful, please cite:

@inproceedings{han2026versacamvla,
  title     = {VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation},
  author    = {Han, Boyao and Shi, Chen and Qian, Jingjing and Tian, Zhuotao and Jiang, Li},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026},
  eprint    = {2610.12451},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  url       = {https://arxiv.org/abs/2610.12451}
}