We released AerialVLN-Verified and AerialVLN-S-Verified, two high-quality datasets curated with ViRC-7B and 150 hours of human verification.
Multimodal AI · Research & Engineering
Lihong Wang
Multimodal Algorithm Engineer at Ant Phecda Lab, Ant Group
My work focuses on multimodal reasoning, reinforcement learning, and agentic vision-language models, with an emphasis on reliable training data and interpretable reasoning processes.

Recent updates
News
Research releases, publications, and milestones.
Multimodal Algorithm Engineer
Ant Phecda Lab (蚂蚁天玑实验室), Ant Digital Technologies, Ant Group
Building multimodal reasoning systems and reinforcement learning pipelines for vision-language models.
M.E. in Artificial Intelligence
Jilin University
Multimodal reasoning research in the Intelligent Content Learning Group.
National Second Prize
The Interpretable Auxiliary Diagnosis of Scoliosis Based on Learnable B-Splines
5th China Graduate AI Innovation Competition
First Prize, South China Division
Privacy Protection System Based on Image Encryption and Invisible Watermarking
2022 C4-Network Technology Challenge, A-ST Track
National Third Prize
Multimedia Data Security Protection System for Medical Diagnosis
15th National College Student Information Security Contest
Connect
Let’s build more capable multimodal systems.
For research discussions, open-source collaboration, or professional opportunities, feel free to get in touch.