I am Xiao Feng, a Ph.D. student at TMLR group of Hong Kong Baptist University, fortunate to be advised by Prof. Bo Han, and working with Prof. Jiangchao Yao. My research focuses on trustworthy agentic intelligence, with the goal of building agents we can trust to solve complex problems. My work spans three directions:

  • Systems: How can we build trustworthy agentic systems across different models, harnesses, and algorithms? [AlphaApollo]
  • Methods: How can we enable agents to learn and solve problems reliably? [RewardFlow] [MAD-M²]
  • Understanding and Benchmarking: How do agents reason, and where are their capabilities and limitations? [LoT] [AR-Bench]

E-mail: xiaofeng [at] comp.hkbu.edu.hk

📰 News

  • 2026.09, RewardFlow was accepted to NeurIPS 2026. See you in Sydney!
  • 2025.09, I started my Ph.D. at the TMLR Group, Hong Kong Baptist University.

📖 Education and Experience

  • 2025.09 - present, Ph.D. Student, TMLR Group, Hong Kong Baptist University, advised by Prof. Bo Han.
  • 2021.09 - 2022.11, MSc in Artificial Intelligence, University of Southampton.
  • 2017.09 - 2021.06, B.E. in Computer Science, Lanzhou University.

📝 Selected Publications

* Co-first author, ✉️ Corresponding author.

RewardFlow framework: state graph construction and reward propagation

RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with LLMs.
Xiao Feng, Bo Han✉️, Zhanke Zhou, Jiaqi Fan, Jiangchao Yao, Ka Ho Li, Dahai Yu, Michael Ng
NeurIPS 2026. [paper] [code]

sym

AlphaApollo: A System for Deep Agentic Reasoning.
Zhanke Zhou, Chentao Cao, Xiao Feng, Xuan Li, Zongze Li, Xiangyu Lu, Jiangchao Yao,
Weikai Huang, Tian Cheng, Jianghangfan Zhang, Tangyu Jiang, Linrui Xu, Yiming Zheng,
Brando Miranda, Tongliang Liu, Sanmi Koyejo, Masashi Sugiyama, Bo Han✉️
Technical Report. [paper] [code]

sym

From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
Zhanke Zhou*, Xiao Feng*, Zhaocheng Zhu, Jiangchao Yao, Sanmi Koyejo, Bo Han✉️
ICML 2025. [paper] [code] [slides] [poster] [EN-video] [CN-video] [CN-blog]

sym

Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models.
Zhanke Zhou*, Zhaocheng Zhu*, Xuan Li*, Mikhail Galkin, Xiao Feng, Sanmi Koyejo, Jian Tang, Bo Han✉️
ICLR 2026. [paper] [code] [tutorial] [slides] [poster] [twitter]

Co-rewarding framework: data-side cross-reference and model-side self-distillation

Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models.
Zizhuo Zhang, Jianing Zhu, Xinmu Ge, Zihua Zhao, Zhanke Zhou, Xuan Li, Xiao Feng, Jiangchao Yao, Bo Han✉️
ICLR 2026. [paper] [OpenReview] [code]

sym

Multi-Agent Debate with Memory Masking (MAD-M²).
Hongduan Tian, Xiao Feng, Ziyuan Zhao, Xiangyu Zhu, Rolan Yan, Bo Han✉️
ICLR 2026. [paper] [code]

🎖 Awards

  • ICML Gold Reviewer Award, 2026.
  • HKBU PhD Transdisciplinary Research Scholarship Scheme, 2025.
  • Excellent Research Bronze Award of TMLR Group, 2024-2025.
  • Industry Collaboration Award of TMLR Group, 2024-2025.

💻 Services

  • Conference Reviewer: NeurIPS, ICLR, ICML, AAAI, ACL, AISTATS.
  • Journal Reviewer: TPAMI, JAIR, TNNLS, NEUNET.

© 2026 Xiao Feng | Last Update: 2026.10