Research areas include Vision-Language Action (VLA), Vision-and-Language Navigation (VLN), World Action Model (WAM), and AI Agent.
Background
2026-Present: Master of Science, Nanyang Technological University, specializing in Machine Learning, under supervision of Prof. Lihua Xie.
2021-2025: Bachelor of Engineering (Honours), Chongqing University of Technology, majoring in Computer Science and Technology, under supervision of Prof. Xiaofei Zhu.
Internship Experience
Lenovo (Shanghai) Information Technology Co., Ltd. (LR Robotics Algorithm) — Embodied AI Research Intern (Robotics) 2026.05-2026.08
Using the Lenovo Morningstar MX robot as a platform, an intelligent agent system, AgenticVLN, was constructed, which is a collaborative system of NavAgent (the task planning mastermind) and DualVLN (a slow-fast dual-system navigation). NavAgent, based on a LangGraph state machine, is responsible for global task parsing, memory storage, telemetry self-checking, and anomaly diagnosis. DualVLN is responsible for specific motion generation, using System2 (2Hz, fine-tuned QwenVL-2.5) to predict 2D pixel sub-targets, and System1 (30Hz, NextDiT) to output high-frequency physical trajectories. SR increased by 10.1%, SPL increased by 6.7%, NE decreased by 17.6%, and StR decreased by 32.5%.
Project Experience
DirectVLA: Development of a High-Frequency, Low-Memory Real-Time VLA Strategy for a Simplified Direct Mapping Architecture 2026.06-Present
Addressing the bottlenecks of traditional VLA, which relies on LLM with billions of parameters, resulting in large memory requirements (>10 GB) and high latency (~11 Hz), we spearheaded the development of DirectVLA, a real-time strategy with a minimalist direct mapping architecture. The project abandons the V→L→A path, constructing a V+L→A model based on DINOv3 and BERT, combining bidirectional text-to-image interaction with non-autoregressive continuous action block decoding. On a single RTX 4090, the model achieves an ultra-low latency of ~30ms with only 0.22B parameters and 0.9GB of memory. High robustness closed-loop verification was completed on a single LIBERO arm (98.7%), RoboTwin dual arms (60.2%), and the xMate SR3 real-world device (up to 92.5%).
OmniVLN: Open-Set Instance-Level 3D Semantic Mapping with Rotating LiDAR and Panoramic Vision Co-first Author, IEEE RA-L (SCI Q2 TOP, Under Review)
To address the challenges of spatial reasoning deficiencies and context bloat in embodied agents, an autonomous navigation agent architecture based on a LangGraph state graph and a 5-layer dynamic scene graph environment world model is constructed. The system integrates continuous coherent topology-based room allocation and VLM dual-graph pruning to refine spatial representations. At the decision-making end, an innovative ego-centric 3D 8-neighborhood observation model and a multi-resolution spatial attention cue mechanism are implemented, combined with Actor-Critic self-reflection and spatial toolchain enhancements to achieve a long-range zero-shot proactive perception closed loop. Experiments show that the agent’s decision success rate reaches 93.18% (with a 22.73% improvement in viewpoint-dependent accuracy) and inference token consumption is reduced by 69.98%.
A Review of Data Fusion and Deep Learning Models for Multimodal Sentiment Analysis. First Author, Published to In Proceedings of E-commerce and Artificial Intelligence (ECAI 2024)
Honors & Awards
- Outstanding graduates of Chongqing, 2025
- Outstanding Bachelor’s Thesis Award, 2025
- National Scholarship, 2024
- Second Prize in the National Undergraduate Data Statistics and Analysis Competition, 2023
- Third Prize in the National Finals of the Team Programming Ladder Competition, 2023