VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation
Conference on Neural Information Processing Systems (NeurIPS), 2026, Poster
VersaCamVLA enables pretrained VLA policies to handle varying camera counts and unseen camera poses through a unified scene-token interface. Learned from posed RGB views, its compact scene tokens improve manipulation robustness on RoboTwin, LIBERO, and real robots without explicit 3D reconstruction or novel-view rendering at deployment.
