VersaCamVLACamera-Configurable VLA Policies for Robotic Manipulation
* Equal contribution † Corresponding author
Abstract
Vision-language-action (VLA) policies are powerful foundations for robotic manipulation, but they often depend on a fixed camera configuration. Changing a camera's position or adding a new view can disrupt the visual interface learned during training. We introduce VersaCamVLA, a camera-configurable framework that decouples camera-set representation from action learning. A unified scene-token interface maps a variable set of calibrated RGB views into a fixed-size representation. Multi-signal target-view prediction preserves visual and spatial structure, while Wrist-Augmented Pose Sampling uses natural wrist-camera motion to provide viewpoint diversity. A lightweight spatial encoder then supplies compact scene tokens to a pretrained VLA as an additional visual condition. At deployment, the policy requires neither explicit 3D reconstruction nor novel-view rendering. Experiments on RoboTwin 2.0, LIBERO, and a real dual-arm robot demonstrate robust manipulation across unseen camera poses and different numbers of input views.
Method

Learning a camera-configurable representation. A scene encoder combines RGB observations and camera rays into shared scene tokens. A training-only decoder predicts RGB, semantic, and edge targets across available views. Wrist-Augmented Pose Sampling varies the source camera subset, using wrist motion to broaden pose coverage within existing demonstrations.
Connecting scene tokens to actions. The pretrained scene encoder is frozen and its decoder removed. A lightweight spatial encoder compresses the representation into 192 supplementary tokens for π0.5, alongside the policy's native visual, language, and proprioceptive inputs. The scene-token pathway supports varying camera counts while preserving a consistent input size for the action policy.
Results
Benchmark performance and camera-configuration results from the paper. All values are task success rates (%).
Benchmark performance
Table 1. Simulation results on RoboTwin 2.0. Clean and domain-randomized (DR) settings across 16 tasks. Bold indicates the best performance in each column.
Scroll horizontally to see all columns.
| Task | ACT | DP | DP3 | OpenVLA-OFT | π₀.₅ | VersaCam | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Clean | DR | Clean | DR | Clean | DR | Clean | DR | Clean | DR | Clean | DR | |
| Blocks Ranking RGB | 0 | 0 | 0 | 0 | 2 | 0 | 0 | 0 | 38 | 11 | 67 | 12 |
| Blocks Ranking Size | 3 | 0 | 1 | 0 | 2 | 1 | 7 | 0 | 12 | 4 | 35 | 6 |
| Hanging Mug | 6 | 0 | 19 | 0 | 35 | 1 | 12 | 0 | 4 | 4 | 10 | 4 |
| Move Stapler Pad | 0 | 0 | 0 | 0 | 7 | 0 | 0 | 0 | 5 | 7 | 8 | 5 |
| Open Laptop | 72 | 1 | 53 | 0 | 77 | 1 | 82 | 0 | 90 | 70 | 96 | 62 |
| Open Microwave | 72 | 0 | 78 | 0 | 91 | 25 | 23 | 56 | 43 | 26 | 51 | 46 |
| Pick Diverse Bottles | 8 | 0 | 28 | 0 | 55 | 1 | 9 | 0 | 36 | 21 | 47 | 12 |
| Place A2B Left | 1 | 0 | 3 | 0 | 30 | 1 | 5 | 0 | 42 | 24 | 54 | 17 |
| Place A2B Right | 2 | 0 | 6 | 0 | 43 | 0 | 11 | 0 | 24 | 19 | 45 | 27 |
| Place Burger Fries | 67 | 0 | 80 | 0 | 75 | 2 | 30 | 0 | 48 | 31 | 81 | 49 |
| Place Dual Shoes | 0 | 0 | 4 | 0 | 9 | 0 | 2 | 0 | 25 | 24 | 51 | 25 |
| Place Empty Cup | 38 | 0 | 27 | 0 | 75 | 1 | 20 | 0 | 77 | 1 | 81 | 27 |
| Place Object Basket | 1 | 0 | 21 | 0 | 49 | 3 | 5 | 0 | 53 | 26 | 76 | 35 |
| Scan Object | 3 | 0 | 8 | 0 | 29 | 1 | 8 | 1 | 6 | 4 | 35 | 15 |
| Stack Blocks Three | 5 | 0 | 0 | 0 | 3 | 0 | 0 | 0 | 18 | 50 | 43 | 3 |
| Stack Bowls Three | 37 | 0 | 51 | 0 | 57 | 4 | 29 | 0 | 53 | 24 | 61 | 23 |
| Average (%) | 19.69 | 0.06 | 23.69 | 0.00 | 39.94 | 2.56 | 15.19 | 3.56 | 35.88 | 21.63 | 52.56 | 23.00 |
Robustness to camera-pose variations
Table 2. Robustness to camera-pose variations. On RoboTwin 2.0, both policies are trained under the nominal camera setup and evaluated with camera poses perturbed in azimuth, distance, and pitch over a workspace-centered spherical region, probing generalization to unseen viewpoints.
| Method | RoboTwin 2.0 Clean | |
|---|---|---|
| Seen Pose | Unseen Pose | |
| π₀.₅ | 48.88 | 32.88 |
| VersaCamVLA | 53.44 | 52.63 |
Flexibility across input camera counts
Table 3. Effect of input camera count. On RoboTwin 2.0 Clean, we train two π0.5 baselines separately with 3 and 4 views and a single VersaCamVLA policy. We evaluate VersaCamVLA with 3, 4, 5, and 6 input views and each π0.5 baseline at its matched view count, keeping camera viewpoints aligned during testing to isolate the effect of camera count.
| Method | RoboTwin 2.0 Clean | |||
|---|---|---|---|---|
| 3V | 4V | 5V | 6V | |
| π₀.₅ | 35.88 | 48.88 | × | × |
| VersaCamVLA | 52.56 | 53.44 | 53.13 | 55.63 |
V denotes input views; × denotes settings not reported.
Camera Pose
RoboTwin policy executions at 4× speed. Head, left wrist, right wrist, and Agent views are synchronized across the top row, above a close-up third-person view. The moving camera and its trajectory are shown in red.
Camera Count
One policy with 3, 4, 5, or 6 input cameras. The input views appear above the cropped third-person execution. Colored trajectories identify the additional external cameras. All videos are shown at 4× speed.
Real-world
Front-view recordings of three bimanual tasks, shown at 4× speed.
BibTeX
If you find this work useful, please cite:
@inproceedings{han2026versacamvla,
title = {VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation},
author = {Han, Boyao and Shi, Chen and Qian, Jingjing and Tian, Zhuotao and Jiang, Li},
booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
year = {2026},
eprint = {2610.12451},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2610.12451}
}