I am Zhangyang Qi (Nickname: Alex Chi, Chinese name: 戚张扬). I received my Ph.D. in computer science from The University of Hong Kong (HKU) in June 2026, advised by Prof. Hengshuang Zhao and Prof. Yizhou Yu. Before that, I received my B.Eng. from Harbin Institute of Technology (HIT) in July 2022.
I am currently a Research Intern at Kimi (Moonshot AI), working on multimodal large language models, where I contributed to Kimi K3 and PerceptionBench. Previously, I was a Research Intern at Shanghai AI Laboratory, supervised by Jiaqi Wang and Tong Wu.
My research interest includes multimodal language models for 3D scene understanding and interactions.
- Multimodal large language models
- Visual perception and evaluation
- 3D scene understanding
- Video language models
- 3D point language models
I welcome any inquiries to reach out to me via WeChat: qi-zhangyang or LinkedIn. Attached are my English Resume and Chinese Resume for your reference.
🔥 News
- 2026.07: 🎉🎉 PerceptionBench is released, a benchmark for atomic visual perception in MLLMs.
- 2026.07: 🎉🎉 Kimi K3 is released, with open weights available for download.
- 2026.07: 🎉🎉 The VLMEvalKit technical report is updated on arXiv, now covering 450+ models and 330+ benchmarks.
- 2026.06: 🎉🎉 Received my Ph.D. in Computer Science from HKU.
- 2026.02: 🎉🎉 Joined Kimi (Moonshot AI) as a Research Intern.
- 2026.01: 🎉🎉 GPT4Scene has been accepted by ICLR 2026.
- 2025.12: 🎉🎉 GGBench has been accepted by AAAI 2026.
- 2025: 🎉🎉 GPT4Point++ has been accepted by IEEE TPAMI.
- 2025.06: 🎉🎉 VLN-R1 is released on arXiv.
- 2024.03: 🎉🎉 GPT4Point has been accept by CVPR 2024.
- 2023.10: 🎉🎉 OCBEV has been accept by 3DV 2024.
- 2022.09: 🎉🎉 Join HKU as a Ph.D. student.
- 2022.07: 🎉🎉 Got bachelor’s degree from HIT with Top Ten Outstanding Students and Outstanding Graduate.
🌐 Experiences

Kimi (Moonshot AI), Beijing, China · 2026.02 – Present
- Research Intern
- Research on multimodal large language models, contributing to Kimi K3 and PerceptionBench.

Shanghai AI Laboratory, Shanghai, China · 2022.07 – 2026.01
- Research Intern, Supervisors: Jiaqi Wang, Tong Wu
- Research on 3D and video language models, developing the GPT4Point, GPT4Point++, and GPT4Scene.
- Curated training data for InternLM-XComposer series and V3Det dataset.

Tencent PCG, Shenzhen, China · 2021.12 – 2022.05
- Research Intern
- Built CLIP-based cross-modal alignment via contrastive learning for image-text matching.
- Designed joint training paradigms enhancing embedding alignment in multimodal retrieval.
📖 Educations
- 2022.09 - 2026.06, Ph.D. in Computer Science, The University of Hong Kong (HKU).
- 2018.08 - 2022.07, Bachelor in Information Engineering, Harbin Institute of Technology (HIT).
📄 Tech Reports

Kimi K3: Open Frontier Intelligence
Kimi Team (Zhangyang Qi, Contributor)
- An open-weight, natively multimodal agentic model: 2.8T total parameters (104B activated) in a Mixture-of-Experts architecture, a 1M-token context window, and native vision via MoonViT-V2.
- Released in July 2026 with full open weights, the largest open-weight model released to date.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Kimi Team (Zhangyang Qi, Contributor)
- A benchmark that separates perception from reasoning: 3,000 verified questions covering ten atomic perceptual capabilities, distilled from real model failures across 42 existing benchmarks.
- The ten capabilities are visual relation, counting, attribute, depth & 3D, localization, comparison, fine-grained recognition, context integration, OCR, and hallucination.
- Across frontier MLLMs, no model reaches 60% accuracy, and perception-related hallucination is the weakest capability on average.

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models
Haodong Duan, Xinyu Fang, Junming Yang, Xiangyu Zhao, Zerun Ma, Yuxuan Qiao, Mo Li, Tianhao Liang, Lin Zhu, Amit Agarwal, Xiaozhe Li, Shengyuan Ding, Jiazi Bu, Ziyu Liu, Zhangyang Qi, et al., Pan Zhang, Jiaqi Wang, Dahua Lin, Kai Chen
- The open-source evaluation toolkit behind OpenCompass, supporting 450+ multi-modality models and 330+ benchmarks behind a single unified interface.
- I contributed the spatial-grounding evaluations, including DA-2K, ERQA, and RefSpatialBench.
📝 Publications

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
Zhangyang Qi, Zhixiong Zhang, Ye Fang, Jiaqi Wang, Hengshuang Zhao
- The first to utilize a video-based large language model for indoor scene understanding.

Game Ground Bench: Probing the Limits of LVLMs in Complex Semantic Grounding Across Game Universes
Zhangyang Qi, Jinsong Li, Hongjian Wu, Jiaqi Wang, Hengshuang Zhao
- GGBench, a large-scale cross-genre benchmark for visual grounding in interactive game environments: 10 genres from card games to first-person shooters and RPGs, 1,335 test images requiring reasoning over game mechanics and UI rather than direct noun-to-object matching.
- We further propose Game-R1, a Grounded Reinforcement Policy Optimization (GRPO) method that reaches strong few-shot generalization and outperforms both open- and closed-source LVLMs on GGBench.

GPT4Point: A Unified Framework for Point-Language Understanding and Generation
Zhangyang Qi, Ye Fang, Zeyi Sun, Xiaoyang Wu, Tong Wu, Jiaqi Wang, Dahua Lin, Hengshuang Zhao
- The first object-level 3D point cloud multimodal large language model, unifying both point cloud understanding and generation tasks.
- Extended to GPT4Point++, published in IEEE TPAMI 2025.

Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images
Zhangyang Qi, Yunhan Yang, Mengchen Zhang, Long Xing, Xiaoyang Wu, Tong Wu, Dahua Lin, Xihui Liu, Jiaqi Wang, Hengshuang Zhao
- Our work introduces a novel framework for 3D object generation and editing, leveraging dual-view image manipulation.

OCBEV: Object-Centric BEV Transformer for Multi-View 3D Object Detection
Zhangyang Qi, Jiaqi Wang, Xiaoyang Wu, Hengshuang Zhao
- An object-centric BEV (Bird’s Eye View) autonomous driving 3D object detection framework, achieving performance improvements on the nuScenes dataset with half the training data.
Other Publications
-
GPT4Point++: Advancing Unified Point-Language Understanding and Generation, IEEE TPAMI 2025.
Zhangyang Qi, Ye Fang, Zeyi Sun, Xiaoyang Wu, Tong Wu, Jiaqi Wang, Dahua Lin, Hengshuang Zhao.
The journal extension of GPT4Point, replacing the two-stage pipeline with a unified end-to-end training scheme. -
VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning, Arxiv 2025. [Paper] [Code]
Zhangyang Qi, Zhixiong Zhang, Yizhou Yu, Jiaqi Wang, Hengshuang Zhao.
An end-to-end framework turning egocentric video streams into continuous navigation actions via GRPO-based reinforcement fine-tuning with a time-decayed reward.
🎖 Awards
- Hong Kong PhD Fellowship Scheme (HKPFS), 2022.
- HKU Presidential Scholarship (HKUPS), 2022.
- Top Ten Students of Harbin Institute of Technology, 2021.
- National Scholarship, 2020.
💻 Professional Services
- Conference reviewer: CVPR’24,25, ICCV’25.
- Teaching assistant: DASC7606: Deep Learning (Graduate course @ HKU), 2023 Spring, 2024 Spring, 2024 Fall.