AI 评测工程师(大模型 / 智能体方向)
北京、上海、深圳
社招
全职
技术
职位 ID:A14257
职位描述
大模型项目最难的问题之一,是如何判断它在真实业务里到底有没有变好。你将负责建立从数据集、任务环境、评测方法到上线监控的完整评测体系,让团队能用事实回答“模型哪里好、哪里不够好、下一步该怎么改”。你的评测对象不只是单轮问答,还包括长文档理解、知识库问答、工具调用、多轮对话、长程任务、多智能体协同、视觉理解和端侧推理。你会与算法、产品、研发和客户团队一起定义真实任务,把一次项目中的难例沉淀成下一轮模型和产品迭代的燃料。 岗位职责 1. 负责大语言模型、多模态模型、智能体和行业 AI 应用的评测方案设计,明确任务、数据、指标、基线和验收标准。 建设覆盖能力、可靠性、安全性、可解释性和工程性能的评测集,支持问答、长文档、RAG、工具调用、报告生成、视觉理解等任务。 2. 设计人工评测、自动评测、模型评审、规则校验和抽样复核流程,提升评测结果的稳定性、一致性和可复现性。 建立智能体任务环境和轨迹分析方法,评估任务完成率、步骤有效性、工具调用正确率、证据引用、错误恢复和长程执行能力。 3. 负责准确率、召回率、事实一致性、引用可追溯性、幻觉率、响应时延、吞吐、显存及成本等指标分析,建设回归评测、版本对比、难例挖掘和线上质量监控机制。 4. 与算法、研发和产品团队协作定位问题根因,提出模型、数据、提示词、流程或系统优化建议,并面向客户项目输出评测报告和验收材料。
职位要求
投递

ModelBest established in August 2022, is an Artificial Intelligence technology company headquartered in Beijing, China. The company is deeply rooted in the field of general AI, with a focus on the innovation and application transformation of large-scale models. ModelBest boasts a prestigious founding R&D team from the Tsinghua in the AI field. Leveraging numerous cutting-edge technologies in natural language processing, the company is currently constructing a large-scale pre-trained model library and corresponding tools, aiming to standardize the technology and applications of large models.
Based on the capabilities of the CPM series of large models, ModelBest has successfully facilitated the intelligent upgrade and efficiency enhancement of various industries. Moreover, the company has released a multi-modal dialog assistant with hundreds of billions of parameters, “Luca", to the public. Adhering to the core philosophy of "intelligence encompassing all", we look forward to exploring and expanding unknown possibilities with global partners, aiming to ensure AI technology serves humanity for a better life in a safe and inclusive manner, laying a solid foundation for the advent of the AGI world.
OpenBMB (Open Lab for Big Model Base,https://github.com/OpenBMB), founded by ModelBest Inc & TsinghuaNLP, aims to build foundation models and systems towards AGI.