Zhihang Fu · 付志航
Post-training Qwen for real-world impact.
I am an advanced algorithm engineer at the Token Foundry, Alibaba Group. I received my Ph.D. from Zhejiang University in 2019, advised by Prof. Yaowu Chen, and my B.Eng. from Chu Kochen Honors College, Zhejiang University.
工作于阿里ATH,Token Foundry,致力于通过 Qwen 后训练,在实际业务场景中落地应用、创造价值。2019 年博士毕业于浙江大学,本科毕业于浙江大学竺可桢学院。
The Path to Recursive Self-Improving Agents
Foundation, Framework, and Future Directions
Can agent systems move beyond manual refinement and autonomously improve themselves? Our new survey studies self-improving agent systems — systems that transform experience and evaluation feedback into persistent updates to their own components — and charts a roadmap toward recursive self-improvement (RSI).

🪜 An L1–L5 Capability Grading Standard for Self-Improvement
📐 Formal Foundation
Formal definitions of agent-system self-improvement and RSI, graded from manual improvement (L1) to general recursive self-improvement (L5).
🧩 Unified Framework
Analyzes core components, dependencies, and improvement procedures of self-improving agent systems in one research framework.
🗂️ Living Taxonomy
A community-maintained literature map spanning harness, data system, trainer self-improvement, and cross-component co-improvement.
🧭 Future Roadmap
Open problems on long-horizon evaluation, modifiable infrastructure, generalizability, safety, and human–agent co-improvement.
🎓️ Research Interest
My research focuses on LLM, building key components of the post-training pipeline to deeply optimize Planning & Tool-Use capabilities at both model and system levels, reducing hallucination and accurately completing user tasks. I lead a team that has delivered domain-specific LLM solutions across verticals including international sports events and education.
当前从事大语言模型算法研究,致力于通过构建后训练过程中的关键环节:Data/Trajectory Curation、Training Strategy、Reasoning Scaffolding、Fine-grained Benchmarking,深度优化模型和链路的Planning & ToolUse能力,降低幻觉并精准完成用户任务。
我带领团队在奥运会国际赛事、国内教育等场景均有行业化落地。米兰冬奥基于阿里千问打造奥运官方大模型,并在奥运会期间上线olympics.com官网,服务全球用户。

🎖️ We are hiring! Looking for talented interns and full-time researchers in LLM research and applications. Learn more →
🔥 News
- Aug 2026Our survey The Path to Recursive Self-Improving Agents released — L1–L5 capability grading, a unified framework, and a living literature map for self-improving agent systems. [Paper]
- Apr 2026Interact-RAG accepted by ICLR'26 — Breaking black-box RAG paradigm, enabling LLM agents to actively manipulate retrieval.
- Dec 2025Thinking Speed Control accepted by NeurIPS'25 Spotlight — First dynamic fast/slow thinking switch for reasoning models.
- May 2025StruXGPT2 accepted by ACL'25 — Achieving 100% knowledge injection performance with only 5% training corpus.
- May 2025ROPO accepted by ICML'25 — Noise-robust preference alignment without external models.
- Feb 2025ROUTE accepted by ICLR'25 — Multi-task collaborative Text-to-SQL for open-source LLMs.
📜 Professional Services
Programme Committee: NeurIPS (2023–2025), ICLR (2023–2024), ICML (2023–2024), CVPR (2024), AAAI (2024), KDD (2023)
Journal Reviewer: IEEE TIP, TCSVT
