I am a Principal Researcher and Research Manager at Microsoft Research, working on AI agents, reinforcement learning, and recursive self-improvement. I am interested in how agents can use experience, search, and feedback to improve their reasoning and future learning.
My research explores how agents can learn from interaction and use world-model simulation to anticipate the consequences of their actions. I study how planning and exploratory learning can turn test-time experience into improved policies.
My systems work supports these questions through reusable agent environments, harness-native training, and long-horizon credit assignment. I also work on visual web agents, world-model learning, and multimodal verification.
Open-source
Public implementations from MSR-Orchard, alongside Orchard, Dyna-Mind, and ExACT.
Orchard
SHARED FOUNDATIONReusable sandbox environments, a Python SDK, and agentic modeling recipes across software engineering, browser navigation, and assistant tasks.
OpenForge-RL
HARNESS-NATIVE TRAININGA rollout engine, environment implementations, and training integration for agents operating inside their deployment harnesses. Initial release; see the repository roadmap.
trace
CREDIT ASSIGNMENTThe official TRACE implementation, with turn-level reward computation, training launchers, a reference-scoring helper, and released training data.
Selected publications
Selected publications, newest years first. Paper and code links are listed together. Full list.
OpenForge RL: Train Harness-native Agents in Any Environment
Training agents within their deployment harnesses, using reusable environments and a shared reinforcement-learning workflow.
TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents
Dense credit assignment for multi-turn agents, deriving turn-level rewards without training an additional critic.
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
An end-to-end online RL framework for visual agents interacting with live websites.
Orchard: An Open-Source Agentic Modeling Framework
A shared environment service and open training recipes for software engineering, browser, and personal-assistant agents.
GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL
Post-training for GUI agents that balances reasoning, visual grounding, and partially verifiable action rewards.
Reinforcement World Model Learning for LLM-based Agents
Learning action-conditioned world models by aligning simulated outcomes with actual environment transitions.
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
Adaptive verification of responses, grounding, and reasoning to provide richer training signals for multimodal agents.
ThetaEvolve: Test-time Learning on Open Problems
Combining program evolution and test-time reinforcement learning for open optimization problems.
SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants?
A benchmark testing whether simulated users can reliably substitute for people in multi-turn assistant evaluation.
Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
ReSim grounds reasoning in simulated futures drawn from real experience; Dyna-GRPO improves policies using outcome rewards and intermediate interaction feedback.
Dyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI Agents
Integrating world-model simulation with reasoning and action through imitation learning and alternating world-model and policy training.
Latent Action Pretraining from Videos
ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory Learning
Reflective tree search gathers exploration experience, which is transferred into the agent through exploratory learning.
Instruction Tuning with GPT-4
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models
Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback
Repository contains project materials; an implementation is not included.
Spoken Language Understanding Using Long Short-Term Memory Neural Networks
Contact
For collaboration and relevant AI research opportunities, contact baolinpeng [at] microsoft [dot] com.