强化学习算法实习生(LLM Agent 方向)
上海
校招
实习
技术 - 算法
职位 ID:A87241
职位描述
1、面向 agentic 任务(多轮工具调用、代码修复 / SWE、网页搜索与信息检索、长程任务规划等)开展 LLM 的RL 后训练研究; 2、研究并改进面向多轮、长 horizon、稀疏奖励场景的 RL 算法(GRPO / PPO 及其变体),包括优势估计与 credit assignment、off-policy 修正、探索与熵控制、训练稳定性等问题; 3、参与 agent 任务的训练数据与评测设计:构造训练任务,设计可验证的奖励信号,分析 rollout 轨迹质量; 4、独立设计并执行实验:数据构造与筛选、消融实验、训练曲线与 rollout 轨迹分析,定位 reward hacking、训练崩溃等问题并给出改进方案,撰写技术报告;有机会产出论文或开源成果; 5、跟踪 agent RL 前沿工作并快速复现,在内部训练框架中落地验证。
职位要求
1、硕士及以上学历在读,计算机、人工智能、数学等相关专业;能持续实习 6 个月优先; 2、扎实的机器学习基础,熟悉 Transformer 架构,熟悉强化学习基本理论(策略梯度、PPO、GRPO 等),理解 LLM 后训练(SFT / RLHF / RLVR / OPD)的流程; 3、良好的 coding 能力与良好的实验素养:熟练使用 Python 和 PyTorch,能从训练指标和 rollout 样本中发现问题并得出可信结论; 4、自驱力强,能独立推进研究课题; 5、优秀的沟通与协作能力:能清晰地表达实验思路、结论和遇到的问题,与团队成员高效协作。 加分项 1、有 LLM RL 训练(RLHF / RLVR / agent RL)的实际经验; 2、有强化学习理论研究背景(策略优化、探索、credit assignment、off-policy 学习等),或相关实习; 3、在 NeurIPS / ICML / ICLR / ACL 等会议发表过相关论文,或有高质量开源项目; 4、ACM-ICPC / Kaggle 等竞赛获奖经历。
投递

ModelBest established in August 2022, is an Artificial Intelligence technology company headquartered in Beijing, China. The company is deeply rooted in the field of general AI, with a focus on the innovation and application transformation of large-scale models. ModelBest boasts a prestigious founding R&D team from the Tsinghua in the AI field. Leveraging numerous cutting-edge technologies in natural language processing, the company is currently constructing a large-scale pre-trained model library and corresponding tools, aiming to standardize the technology and applications of large models.
Based on the capabilities of the CPM series of large models, ModelBest has successfully facilitated the intelligent upgrade and efficiency enhancement of various industries. Moreover, the company has released a multi-modal dialog assistant with hundreds of billions of parameters, “Luca", to the public. Adhering to the core philosophy of "intelligence encompassing all", we look forward to exploring and expanding unknown possibilities with global partners, aiming to ensure AI technology serves humanity for a better life in a safe and inclusive manner, laying a solid foundation for the advent of the AGI world.
OpenBMB (Open Lab for Big Model Base,https://github.com/OpenBMB), founded by ModelBest Inc & TsinghuaNLP, aims to build foundation models and systems towards AGI.