GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
Published in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Highlight, 2026
GeoPredict is a geometry-aware VLA framework that injects predictive kinematic trajectories and 3D Gaussian geometry priors as training-time supervision, enabling precise 3D reasoning for robotic manipulation without extra inference cost. It consistently outperforms strong VLA baselines on RoboCasa Human-50, LIBERO, and real-world geometry-intensive tasks.
Recommended citation: Jingjing Qian, Boyao Han, Chen Shi, Lei Xiao, Long Yang, Shaoshuai Shi, Li Jiang. (2026). "GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation." CVPR. (Highlight).
Download Paper
