Boyao Han

I am an M.Phil. student at the School of Data Science, The Chinese University of Hong Kong, Shenzhen, advised by Prof. Li Jiang. Before that, I studied at the College of Computer Science and Electronic Engineering, Hunan University.

My research interests include vision-language-action (VLA) models, AI agents, and 3D computer vision.

Feel free to contact me for discussion or collaboration!

Boyao Han

News

Publications

VersaCamVLA framework: scene-token interface learning and scene-token-conditioned VLA policy learning

VersaCamVLA: Camera-Configurable VLA Policies for Robotic Manipulation

Boyao Han*, Chen Shi*, Jingjing Qian, Zhuotao Tian, Li Jiang * Equal contribution.

Conference on Neural Information Processing Systems (NeurIPS), 2026, Poster

VersaCamVLA enables pretrained VLA policies to handle varying camera counts and unseen camera poses through a unified scene-token interface. Learned from posed RGB views, its compact scene tokens improve manipulation robustness on RoboTwin, LIBERO, and real robots without explicit 3D reconstruction or novel-view rendering at deployment.

Overview of the GeoPredict architecture

GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation

Jingjing Qian, Boyao Han, Chen Shi, Lei Xiao, Long Yang, Shaoshuai Shi, Li Jiang

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Highlight

GeoPredict improves 3D reasoning in VLA manipulation by predicting robot trajectories and future workspace geometry. Its training-only geometric supervision boosts simulated and real-world performance without adding 3D decoding at inference.

Memory Forcing: Spatio-temporal Memory for Consistent Scene Generation on Minecraft

Junchao Huang, Xinting Hu, Boyao Han, Shaoshuai Shi, Zhuotao Tian, Tianyu He, Li Jiang

Conference on Neural Information Processing Systems (NeurIPS), 2026, Poster

Memory Forcing equips autoregressive video diffusion with spatio-temporal memory, balancing exploration of new Minecraft scenes with consistent reconstruction of revisited regions for more coherent long-horizon generation.

Overview of the PointSLAM++ framework

PointSLAM++: Robust Dense Neural Gaussian Point Cloud-based SLAM

Xu Wang*, Boyao Han*, Xiaojun Chen, Ying Liu, Ruihui Li * Equal contribution.

AAAI Conference on Artificial Intelligence (AAAI), Poster

PointSLAM++ combines hierarchical neural Gaussians, progressive pose optimization, and adaptive Gaussian density to improve camera tracking and structural consistency under noisy RGB-D input, producing accurate real-time reconstruction and photorealistic rendering.

Research Experience

Education