Baolin Peng

Principal Researcher and Research Manager
Microsoft Research, Redmond

I am a Principal Researcher and Research Manager at Microsoft Research, working on AI agents, reinforcement learning, and recursive self-improvement. I am interested in how agents can use experience, search, and feedback to improve their reasoning and future learning.

My research explores how agents can learn from interaction and use world-model simulation to anticipate the consequences of their actions. I study how planning and exploratory learning can turn test-time experience into improved policies.

My systems work supports these questions through reusable agent environments, harness-native training, and long-horizon credit assignment. I also work on visual web agents, world-model learning, and multimodal verification.

Open-source

Public implementations from MSR-Orchard, alongside Orchard, Dyna-Mind, and ExACT.

  • Orchard

    SHARED FOUNDATION

    Reusable sandbox environments, a Python SDK, and agentic modeling recipes across software engineering, browser navigation, and assistant tasks.

  • OpenForge-RL

    HARNESS-NATIVE TRAINING

    A rollout engine, environment implementations, and training integration for agents operating inside their deployment harnesses. Initial release; see the repository roadmap.

  • trace

    CREDIT ASSIGNMENT

    The official TRACE implementation, with turn-level reward computation, training launchers, a reference-scoring helper, and released training data.

Browse MSR-Orchard

Selected publications

Selected publications, newest years first. Paper and code links are listed together. Full list.

  1. OpenForge RL: Train Harness-native Agents in Any Environment

    Xiao Yu, Baolin Peng, Ruize Xu, Hao Zou, et al.

    July 2026 · arXiv:2607.21557

    Training agents within their deployment harnesses, using reusable environments and a shared reinforcement-learning workflow.

  2. TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents

    Leitian Tao, Baolin Peng, Wenlin Yao, Tao Ge, et al.

    July 15, 2026 · arXiv:2607.13988

    Dense credit assignment for multi-turn agents, deriving turn-level rewards without training an additional critic.

  3. OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents

    Rui Yang, Qianhui Wu, Yuxi Chen, Hao Bai, Wenlin Yao, Hao Cheng, Baolin Peng, et al.

    June 1, 2026 · arXiv:2606.02031

    An end-to-end online RL framework for visual agents interacting with live websites.

  4. Orchard: An Open-Source Agentic Modeling Framework

    Baolin Peng, Wenlin Yao, Qianhui Wu, Hao Cheng, Xiao Yu, Rui Yang, et al.

    May 14, 2026 · arXiv:2605.15040

    A shared environment service and open training recipes for software engineering, browser, and personal-assistant agents.

  5. GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL

    Rui Yang, Qianhui Wu, Zhaoyang Wang, Hanyang Chen, Ke Yang, Hao Cheng, Huaxiu Yao, Baolin Peng, et al.

    February 25, 2026 · arXiv:2602.22190

    Post-training for GUI agents that balances reasoning, visual grounding, and partially verifiable action rewards.

  6. Reinforcement World Model Learning for LLM-based Agents

    Xiao Yu, Baolin Peng, Ruize Xu, Yelong Shen, et al.

    February 5, 2026 · arXiv:2602.05842

    Learning action-conditioned world models by aligning simulated outcomes with actual environment transitions.

  7. Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

    Reuben Tan, Baolin Peng, Zhengyuan Yang, Hao Cheng, et al.

    December 3, 2025 · arXiv:2512.03438 · Updated July 2026

    Adaptive verification of responses, grounding, and reasoning to provide richer training signals for multimodal agents.

  8. ThetaEvolve: Test-time Learning on Open Problems

    Yiping Wang, Shao-Rong Su, Zhiyuan Zeng, Eva Xu, Liliang Ren, Xinyu Yang, Zeyi Huang, Xuehai He, Luyao Ma, Baolin Peng, et al.

    November 28, 2025 · arXiv:2511.23473

    Combining program evolution and test-time reinforcement learning for open optimization problems.

  9. SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?

    Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu, Jianfeng Gao

    November 2025 · EMNLP 2025 (conference publication)

    A benchmark testing whether simulated users can reliably substitute for people in multi-turn assistant evaluation.

  10. Dyna-Mind: Learning to Simulate from Experience for Better AI Agents

    Xiao Yu, Baolin Peng, Michel Galley, Hao Cheng, Qianhui Wu, Janardhan Kulkarni, Suman Nath, Zhou Yu, Jianfeng Gao

    October 10, 2025 preprint; ICLR 2026

    ReSim grounds reasoning in simulated futures drawn from real experience; Dyna-GRPO improves policies using outcome rewards and intermediate interaction feedback.

  11. Dyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI Agents

    Xiao Yu, Baolin Peng, Ruize Xu, Michel Galley, Hao Cheng, Suman Nath, Jianfeng Gao, Zhou Yu

    May 31, 2025 preprint; updated October 2025

    Integrating world-model simulation with reasoning and action through imitation learning and alternating world-model and policy training.

  12. Latent Action Pretraining from Videos

    S. Ye, J. Jang, B. Jeon, S. J. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, et al.

    ICLR 2025

  13. ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning

    Xiao Yu, Baolin Peng, Vineeth Vajipey, Hao Cheng, Michel Galley, Jianfeng Gao, Zhou Yu

    October 2, 2024 preprint; updated February 2025

    Reflective tree search gathers exploration experience, which is transferred into the agent through exploratory learning.

  14. Instruction Tuning with GPT-4

    Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, Jianfeng Gao

    2023 · arXiv:2304.03277

  15. Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models

    Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, Jianfeng Gao

    NeurIPS 2023

  16. Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback

    Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, et al.

    2023 · arXiv:2302.12813

    Repository contains project materials; an implementation is not included.

  17. Spoken Language Understanding Using Long Short-Term Memory Neural Networks

    K. Yao, B. Peng, Y. Zhang, D. Yu, G. Zweig, Y. Shi

    IEEE SLT 2014

Contact

For collaboration and relevant AI research opportunities, contact baolinpeng [at] microsoft [dot] com.