About Experiences Publications CV

Chenlu Ye

I am on the job market.I am seeking full-time positions. My interests include self-improving agents, autonomous research, and reinforcement learning for LLMs.View CVEmail me

Hi! My name is Chenlu Ye (叶晨璐). I am a Ph.D. candidate in Computer Science at UIUC, advised by Prof. Tong Zhang.

My research focuses on self-improving agents and autonomous research, reinforcement learning for agentic LLM post-training, and decision-making. I study how agents learn from experience and improve their skills and memory, building on my work in LLM reasoning, RLHF, and robust reinforcement learning.

I am currently a research intern at Meta MSL / FAIR, working with Sainbayar Sukhbaatar and Jason Weston. Previously, I was an applied scientist intern at Amazon and received my MPhil from HKUST and B.S. from USTC.
Chenlu Ye
chenluy3[AT]illinois.edu

Selected Publications (Full)

* denotes equal contribution.

PreprintAdaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RLChenlu Ye*, Xuanchang Zhang*, Yifan Hao*, Zhou Yu, Ziji Zhang, Abhinav Gullapalli, Hao Chen, Jing Huang, Tong ZhangPreprintLayerwise policy perturbations for robust off-policy RL under system noise and training-inference mismatch.blogPDFCode

PreprintReinforce-Ada: An Adaptive Sampling Framework under Non-linear RL ObjectivesWei Xiong*, Chenlu Ye*, Baohao Liao*, Hanze Dong*, Xinxing Xu, Christof Monz, Jiang Bian, Nan Jiang, Tong ZhangPreprintAdaptive generation-budget allocation across prompts to preserve informative and diverse RL training signals.PDFCode

PreprintBeyond Correctness: Harmonizing Process and Outcome Rewards through RL TrainingChenlu Ye, Zhou Yu, Ziji Zhang, Hao Chen, Narayanan Sadagopan, Jing Huang, Tong Zhang, Anurag BeniwalPreprintProcess-outcome consistency for RL data selection, improving final-answer accuracy and intermediate reasoning quality.PDF

PreprintSelf-Rewarding Correction for Mathematical ReasoningWei Xiong*, Hanning Zhang*, Chenlu Ye*, Lichang Chen, Nan Jiang, Tong ZhangPreprintPDF

ICML 2025
Spotlight
Catoni Contextual Bandits are Robust to Heavy-tailed RewardsChenlu Ye, Yujia Jin, Alekh Agarwal, Tong ZhangICML 2025 SpotlightPaper

NeurIPS 2024Online Iterative Reinforcement Learning from Human Feedback with General Preference ModelChenlu Ye*, Wei Xiong*, Yuheng Zhang*, Hanze Dong*, Nan Jiang, Tong ZhangNeurIPS 2024PDFCode

ICML 2024Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-ConstraintWei Xiong*, Hanze Dong*, Chenlu Ye*, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang, Tong ZhangICML 2024Paper

Experiences

Education

See my biography for research experience, professional service, and honors, or download my CV.