Since August 2026, I have been an R&D Engineer at Morphi Robot, working closely with Tai Wang. I received my PhD from Zhejiang University in July 2026, where I was a member of the APRIL Lab under the guidance of Yong Liu.
From March 2025 to March 2026, I was a visiting researcher at the Department of Computer Science, University of California, Los Angeles, advised by Bolei Zhou. From 2023 to 2025, I was a research intern in ADLab at Shanghai AI Laboratory. Before that, I received my bachelor’s degree from Jianxing Campus of Zhejiang University of Technology in 2021, advised by Li Yu.
My research focuses on general-purpose embodied intelligence, with the goal of bringing robots into homes everywhere and freeing people from repetitive and labor-intensive work. You can find my publications on Google Scholar.

News
-
🎉 Paper From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation is accepted by CoRL 2026 !
-
🎉 Paper AURA: Multimodal Shared Autonomy for Real-World Urban Navigation is accepted by CVPR 2026 !
-
🎉 Paper From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning is accepted by ICLR 2026 !
-
🎉 Paper Learning Sidewalk Autopilot from Multi-Scale Imitation with Corrective Behavior Expansion is accepted by ICRA 2026 !
-
🎉 Paper LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking is accepted by TNNLS 2025 !
-
🎉 Paper DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving is accepted by ICCV 2025 !
-
🎉 Paper CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking is accepted by ACMMM 2025 !
-
🎉 Paper L2COcc: Lightweight Camera-Centric Semantic Scene Completion via Distillation of LiDAR Model is accepted by IROS 2025 !
-
🎉 Paper Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving is accepted bys AAAI Oral.
-
🎉 Paper LiCROcc: Teach Radar for Accurate Semantic Occupancy Prediction using LiDAR and Camerar is accepted by 2025 IEEE Robotics and Automation Letters (RAL) !
-
🎉 Paper Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving is accepted by NeurIPS 2024 !
-
🎉 Paper Riders: radar-infrared depth estimation for robust sensing is aceepted by 2024 IEEE Transactions on Intelligent Transportation (TITS) !
-
🎉 Paper FMCW Radar on LiDAR map localization in structural urban environments is aceepted by 2024 Journal of Field Robotics (JFR) !
-
🎉 Paper Geo-localization with transformer-based 2D-3D match network is aceepted by 2023 IEEE Robotics and Automation Letters (RAL) !
Selected Publications
All publications →* denotes equal contribution.
-
Monocular 3D Occupancy Perception for Robots on Sidewalks via Hybrid 2D-3D Learning
BibTeX
@misc{ma2026monocular3doccupancyperception, title = {Monocular 3D Occupancy Perception for Robots on Sidewalks via Hybrid 2D-3D Learning}, author = {Ma, Yukai and Lin, Joe and Liu, Liu and He, Honglin and Ricketts, Lulu and Squicciarini, Brad and Liu, Yong and Zhou, Bolei}, year = {2026}, eprint = {2606.19122}, archiveprefix = {arXiv}, primaryclass = {cs.RO}, demo = {https://vail-ucla.github.io/walkocc/}, } -
From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation
BibTeX
@inproceedings{he2026imitationalignmenthumanpreferenceflow, title = {From Imitation to Alignment: Human-Preference Flow Policies for Long-Horizon Sidewalk Navigation}, author = {He, Honglin and Liu, Zhizheng and Ma, Yukai and Zhou, Bolei}, booktitle = {Conference on Robot Learning}, year = {2026}, eprint = {2606.12603}, archiveprefix = {arXiv}, primaryclass = {cs.RO}, demo = {https://vail-ucla.github.io/FlowPilot/}, } -
SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation
BibTeX
@misc{wang2026sparseworldenhancingendtoendautonomous, title = {SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation}, author = {Wang, Ruoyu and Wang, Jingke and Ma, Yukai and Huang, Yuehao and Lei, Shuangming and Xu, Guanglin and Ye, Aixue and Liu, Yong}, year = {2026}, eprint = {2605.24354}, archiveprefix = {arXiv}, primaryclass = {cs.CV}, demo = {https://wryzju.github.io/SparseWorld/}, } -
VEOcc: Voxel-Centric Online Semantic Occupancy Prediction For Embodied Scene Understanding
BibTeX
@article{wang2026veocc, title = {VEOcc: Voxel-Centric Online Semantic Occupancy Prediction For Embodied Scene Understanding}, author = {Wang, Ruoyu and Liu, Yong and Tao, Sheng and Lin, Yuhang and Ma, Yukai}, journal = {arXiv preprint arXiv:2605.25059}, year = {2026}, demo = {https://wryzju.github.io/VEOcc/}, } -
AURA: Multimodal Shared Autonomy for Real-World Urban Navigation
BibTeX
@misc{ma2026auramultimodalsharedautonomy, title = {AURA: Multimodal Shared Autonomy for Real-World Urban Navigation}, author = {Ma, Yukai and He, Honglin and Song, Selina and Wu, Wayne and Zhou, Bolei}, year = {2026}, eprint = {2604.01659}, archiveprefix = {arXiv}, primaryclass = {cs.RO}, demo = {https://vail-ucla.github.io/aura/}, } -
Drive-Cascade: Autoregressive Occupancy to LiDAR and Video Synthesis
BibTeX
@inproceedings{lei2026drive, title = {Drive-Cascade: Autoregressive Occupancy to LiDAR and Video Synthesis}, author = {Lei, Shuangming and Huang, Yuehao and Yi, Yao and Xie, Yijia and Wang, Jingke and Wang, Ruoyu and Lv, Jiajun and Xu, Guanglin and Ye, AiXue and Liu, Bingbing and others}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition}, pages = {4552--4561}, year = {2026}, demo = {https://summersray.github.io/Drive-Cascade/}, } -
Learning Sidewalk Autopilot from Multi-Scale Imitation with Corrective Behavior Expansion
BibTeX
@misc{he2026learningsidewalkautopilotmultiscale, title = {Learning Sidewalk Autopilot from Multi-Scale Imitation with Corrective Behavior Expansion}, author = {He, Honglin and Ma, Yukai and Squicciarini, Brad and Wu, Wayne and Zhou, Bolei}, year = {2026}, eprint = {2603.22527}, archiveprefix = {arXiv}, primaryclass = {cs.RO}, demo = {https://vail-ucla.github.io/MIMIC}, } -
From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning
BibTeX
@article{he2025seeingexperiencingscalingnavigation, title = {From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement Learning}, author = {He*, Honglin and Ma*, Yukai and Wu, Wayne and Zhou, Bolei}, journal = {Arxiv}, year = {2026}, publisher = {Arxiv}, demo = {https://metadriverse.github.io/s2e/}, } -
X-scene: Large-scale driving scene generation with high fidelity and flexible controllability
BibTeX
@article{yang2026x, title = {X-scene: Large-scale driving scene generation with high fidelity and flexible controllability}, author = {Yang, Yu and Liang, Alan and Mei, Jianbiao and Ma, Yukai and Liu, Yong and Lee, Gim Hee}, journal = {Advances in Neural Information Processing Systems}, volume = {38}, pages = {104415--104451}, year = {2026}, demo = {https://x-scene.github.io/}, } -
CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking
BibTeX
@inproceedings{huang2025cogddn, title = {CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking}, author = {Huang, Yuehao and Liu, Liang and Lei, Shuangming and Ma, Yukai and Su, Hao and Mei, Jianbiao and Zhao, Pengxiang and Gu, Yaqing and Liu, Yong and Lv, Jiajun}, booktitle = {Proceedings of the 33rd ACM International Conference on Multimedia}, pages = {5237--5246}, year = {2025}, demo = {https://yuehaohuang.github.io/CogDDN/}, } -
L2COcc: Lightweight Camera-Centric Semantic Scene Completion via Distillation of LiDAR Model
BibTeX
@article{wang2025l2cocc, title = {L2COcc: Lightweight Camera-Centric Semantic Scene Completion via Distillation of LiDAR Model}, author = {Wang, Ruoyu and Ma, Yukai and Yao, Yi and Tao, Sheng and Li, Haoang and Zhu, Zongzhi and Liu, Yong and Zuo, Xingxing}, booktitle = {Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)}, year = {2025}, demo = {https://studyingfufu.github.io/L2COcc/}, } -
Driving in the occupancy world: Vision-centric 4d occupancy forecasting and planning via world models for autonomous driving
BibTeX
@inproceedings{yang2025driving, title = {Driving in the occupancy world: Vision-centric 4d occupancy forecasting and planning via world models for autonomous driving}, author = {Yang, Yu and Mei, Jianbiao and Ma, Yukai and Du, Siliang and Chen, Wenqing and Qian, Yijie and Feng, Yuxiang and Liu, Yong}, booktitle = {Proceedings of the AAAI Conference on Artificial Intelligence}, volume = {39}, number = {9}, pages = {9327--9335}, year = {2025}, demo = {https://drive-occworld.github.io/}, } -
LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking
BibTeX
@article{ma2025leapvad, title = {LeapVAD: A Leap in Autonomous Driving via Cognitive Perception and Dual-Process Thinking}, author = {Ma, Yukai and Wei, Tiantian and Zhong, Naiting and Mei, Jianbiao and Hu, Tao and Wen, Licheng and Yang, Xuemeng and Shi, Botian and Liu, Yong}, journal = {arXiv preprint arXiv:2501.08168}, year = {2025}, demo = {https://pjlab-adg.github.io/LeapVAD/}, } -
LiCROcc: Teach radar for accurate semantic occupancy prediction using lidar and camera
BibTeX
@article{ma2024licrocc, title = {LiCROcc: Teach radar for accurate semantic occupancy prediction using lidar and camera}, author = {Ma, Yukai and Mei, Jianbiao and Yang, Xuemeng and Wen, Licheng and Xu, Weihua and Zhang, Jiangning and Zuo, Xingxing and Shi, Botian and Liu, Yong}, journal = {IEEE Robotics and Automation Letters}, year = {2024}, publisher = {IEEE}, demo = {https://hr-zju.github.io/LiCROcc/}, } -
DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving
Abstract
This paper presetns DriveArena, the first high-fidelity closed-loop simulation system designed for driving agents navigating in real scenarios. DriveArena features a flexible, modular architecture, allowing for the seamless interchange of its core components: Traffic Manager, a traffic simulator capable of generating realistic traf- fic flow on any worldwide street map, and World Dreamer, a high-fidelity conditional generative model with infinite autoregression. This powerful synergy empowers any driving agent capable of processing real-world images to navigate in DriveArena simulated environment. The agent perceives its surroundings through images generated by World Dreamer and output trajectories; then these trajectories are fed into Traffic Manager, achieving realistic interactions with other vehicles and producing a new scene lay- out. Finally, the latest scene layout is relayed back into World Dreamer, perpetuating the simulation cycle. This iterative process fosters closed-loop exploration within a highly realistic environment, providing a valuable platform for developing and evaluating driving agents across diverse and challenging scenarios. DriveArena signifies a substantial leap forward in leveraging generative image data for the driving simulatior, opening insights for closed-loop autonomous driving.
BibTeX
@article{yang2024drivearena, title = {DriveArena: A Closed-loop Generative Simulation Platform for Autonomous Driving}, author = {Yang*, Xuemeng and Wen*, Licheng and Ma*, Yukai and Mei*, Jianbiao and Li*, Xin and Wei*, Tiantian and Lei, Wenjie and Fu, Daocheng and Cai, Pinlong and Dou, Min and Shi, Botian and He, Liang and Liu, Yong and Qiao, Yu}, journal = {arXiv preprint arXiv:2408.00415}, year = {2024}, demo = {https://pjlab-adg.github.io/DriveArena/}, } -
Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving
Abstract
Autonomous driving has advanced significantly due to sensors, machine learning, and artificial intelligence improvements. However, prevailing methods struggle with intricate scenarios and causal relationships, hindering adaptability and interpretability in varied environments. To address the above problems, we introduce LeapAD, a novel paradigm for autonomous driving inspired by the human cognitive process. Specifically, LeapAD emulates human attention by selecting critical objects relevant to driving decisions, simplifying environmental interpretation, and mitigating decision-making complexities. Additionally, LeapAD incorporates an innovative dual-process decision-making module, which consists of an Analytic Process (System-II) for thorough analysis and reasoning, along with a Heuristic Process (System-I) for swift and empirical processing. The Analytic Process leverages its logical reasoning to accumulate linguistic driving experience, which is then transferred to the Heuristic Process by supervised fine-tuning. Through reflection mechanisms and a growing memory bank, LeapAD continuously improves itself from past mistakes in a closed-loop environment. Closed-loop testing in CARLA shows that LeapAD outperforms all methods relying solely on camera input, requiring 1-2 orders of magnitude less labeled data. Experiments also demonstrate that as the memory bank expands, the Heuristic Process with only 1.8B parameters can inherit the knowledge from a GPT-4 powered Analytic Process and achieve continuous performance improvement.
BibTeX
@article{mei2024continuously, title = {Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving}, author = {Mei*, Jianbiao and Ma*, Yukai and Yang, Xuemeng and Wen, Licheng and Cai, Xinyu and Li, Xin and Fu, Daocheng and Zhang, Bo and Cai, Pinlong and Dou, Min and others}, journal = {Advances in Neural Information Processing Systems (NeurIPS)}, year = {2024}, demo = {https://leapad-2024.github.io/LeapAD/}, } -
RadarCam-Depth: Radar-Camera Fusion for Depth Estimation with Learned Metric Scale
BibTeX
@inproceedings{10610929, author = {Li*, Han and Ma*, Yukai and Gu, Yaqing and Hu, Kewei and Liu, Yong and Zuo, Xingxing}, booktitle = {2024 IEEE International Conference on Robotics and Automation (ICRA)}, title = {RadarCam-Depth: Radar-Camera Fusion for Depth Estimation with Learned Metric Scale}, year = {2024}, volume = {}, number = {}, pages = {10665-10672}, keywords = {Point cloud compression;Image coding;Accuracy;Three-dimensional displays;Robot vision systems;Estimation;Radar}, doi = {10.1109/ICRA57147.2024.10610929}, } -
RIDERS: Radar-Infrared Depth Estimation for Robust Sensing
BibTeX
@article{10623522, author = {Li*, Han and Ma*, Yukai and Huang, Yuehao and Gu, Yaqing and Xu, Weihua and Liu, Yong and Zuo, Xingxing}, journal = {IEEE Transactions on Intelligent Transportation Systems}, title = {RIDERS: Radar-Infrared Depth Estimation for Robust Sensing}, year = {2024}, volume = {25}, number = {11}, pages = {18764-18778}, keywords = {Radar;Radar imaging;Estimation;Cameras;Measurement;Laser radar;Accuracy;Autonomous driving;Infrared imaging;Multisensor systems;Depth estimation;radar perception;infrared camera;multi-sensor fusion}, doi = {10.1109/TITS.2024.3432996}, } - JFR 2024
FMCW Radar on LiDAR map localization in structural urban environments
BibTeX
@article{ma2024fmcw, title = {FMCW Radar on LiDAR map localization in structural urban environments}, author = {Ma*, Yukai and Li*, Han and Zhao, Xiangrui and Gu, Yaqing and Lang, Xiaolei and Li, Laijian and Liu, Yong}, journal = {Journal of Field Robotics}, volume = {41}, number = {3}, pages = {699--717}, year = {2024}, publisher = {Wiley Online Library}, } -
Geo-Localization With Transformer-Based 2D-3D Match Network
BibTeX
@article{10168166, author = {Li*, Laijian and Ma*, Yukai and Tang, Kai and Zhao, Xiangrui and Chen, Chao and Huang, Jianxin and Mei, Jianbiao and Liu, Yong}, journal = {IEEE Robotics and Automation Letters}, title = {Geo-Localization With Transformer-Based 2D-3D Match Network}, year = {2023}, volume = {8}, number = {8}, pages = {4855-4862}, keywords = {Laser radar;Point cloud compression;Feature extraction;Three-dimensional displays;Satellites;Location awareness;Global Positioning System;Geo-localization;2D-3D match;SLAM}, doi = {10.1109/LRA.2023.3290526}, } -
RoLM: Radar on LiDAR Map Localization
BibTeX
@inproceedings{10161203, author = {Ma, Yukai and Zhao, Xiangrui and Li, Han and Gu, Yaqing and Lang, Xiaolei and Liu, Yong}, booktitle = {2023 IEEE International Conference on Robotics and Automation (ICRA)}, title = {RoLM: Radar on LiDAR Map Localization}, year = {2023}, volume = {}, number = {}, pages = {3976-3982}, keywords = {Location awareness;Laser radar;Automation;Autonomous systems;Robot sensing systems;Cameras;Robustness}, doi = {10.1109/ICRA48891.2023.10161203}, }
Robots I Have Worked With
Robotic platforms I have worked with, from delivery and humanoid robots to custom-built navigation systems.
AgileX UMR
Mobile robot chassis
Coco Robotics
Sidewalk delivery robot
Morphi Robot
Wheeled dual-arm humanoid robot
Multi-Sensor Scooter
Self-built outdoor navigation data-collection platform for multi-sensor fusion
Sound-Beacon Smart Car
Self-built competition robot · Sound Beacon Category
National First Prize · 15th National University Student Smart Car Competition
Unitree Go2
Quadruped robot
Education
- 2021.09 – 2026.07Zhejiang University
PhDAPRIL Lab · Advisor Yong Liu - 2017.09 – 2021.07Zhejiang University of Technology
Bachelor's DegreeJianxing Campus · Advisor Li Yu
Experience
- 2026.08 – presentMorphi Robot
R&D EngineerWorking closely with Tai Wang - 2025.03 – 2026.03University of California, Los Angeles
Visiting ResearcherDepartment of Computer Science · Advisor Bolei Zhou - 2023.10 – 2025.03Shanghai AI Laboratory
Research InternADLab