About Experiences Publications CV

Biography

My research focuses on self-improving agents and autonomous research, reinforcement learning for agentic LLM post-training, and decision-making. Download my CV.

Education

  • Ph.D. candidate, Computer Science, University of Illinois Urbana-Champaign (2024.08 - present).

  • MPhil, Artificial Intelligence, The Hong Kong University of Science and Technology (2021.09 - 2024.08).

  • B.S., Statistics, University of Science and Technology of China (2017.09 - 2021.06).

Research Experience

  • Research Intern, Meta MSL / FAIR (2026.05 - present).

    • Managers: Sainbayar Sukhbaatar and Jason Weston.

    • Learning agent skills and file-system memory from experience for continual learning and synthetic data generation.

    • Distilling agent reflections into reusable task-authoring and rubric-design skills.

    • Investigating compute-efficient training for autonomous research agents on long-horizon, multi-turn tasks.

  • Applied Scientist Intern, Amazon (2025.05 - 2026.05).

    • Hosts: Dr. Zhou Yu and Dr. Ziji Zhang.

    • Developed Adaptive Layerwise Perturbation (ALP) for robust off-policy RL under system noise and training-inference mismatch, with applications to mathematical and tool-integrated reasoning.

    • Proposed the Process Consistency Filter (PROF) and developed PROF-GRPO to improve answer accuracy and intermediate reasoning quality.

  • Ph.D. Research, University of Illinois Urbana-Champaign (2024.09 - present).

  • MPhil Research, The Hong Kong University of Science and Technology (2021.09 - 2024.08).

    • Advisor: Prof. Tong Zhang.

    • Formulated RLHF as a reverse-KL-regularized contextual bandit problem and developed statistically efficient algorithms with finite-sample guarantees.

    • Developed uncertainty-weighted algorithms for corruption-robust reinforcement learning across online and offline settings.

  • Visiting Research Scholar, University of California, Los Angeles (2023.08 - 2023.12).

    • Host: Prof. Quanquan Gu.

    • Studied reinforcement learning algorithms robust to adversarial corruption in online and offline decision-making.

Professional Service

  • Conference reviewer: ICML, NeurIPS, ICLR, AISTATS.

  • Journal reviewer: JMLR, Machine Learning, Artificial Intelligence.

Honors and Awards

  • Gold Prize for Outstanding Student Scholarship (1/40), 2020.09.

  • Bronze Prize for Outstanding Student Scholarship, 2019.09 and 2018.09.

Skills

  • Programming and ML: Python, PyTorch, C++.

  • Developer tools: Git, Docker.