I'm a first-year PhD student at the School of Computer Science, University of Sydney, supervised by Prof. Chang Xu. Before starting my PhD in October 2025, I spent over three years as an Algorithm Researcher at Huawei Noah's Ark Lab (Feb 2022 – Sep 2025), where I worked on efficient vision models, diffusion-based image generation, and large vision-language models. I received my B.Eng. from Taiyuan University of Technology in 2021.
My past work has covered efficient object detection, structure-aware diffusion models for low-level vision, and positional encoding for large vision-language models. Currently, my research focuses on World Models, Vision-Language Models, and Agents — I'm hoping to explore, through the two modalities of vision and language, how models can learn to think and understand the world.
I'm mainly interested in World Models, Vision-Language Models, and Agents — hoping to explore, through the two directions of vision and language, how models can learn to think and understand the world.
A fully autonomous AI research system with 20+ specialized agents that orchestrate end-to-end ML research — from literature survey and hypothesis generation, through GPU experiment execution, to conference-ready paper writing — with zero human intervention. Built on a dual-loop architecture: inner loops refine individual projects, while outer loops let the system learn and evolve from past research cycles.
1,073 citations · h-index 8 · i10-index 8 — full list on Google Scholar. * denotes equal contribution.