大模型预训练数据算法工程师
北京、上海
社招
全职
互联网 / 电子 / 网游
职位描述
围绕大语言模型预训练,探索并构建前沿的大规模预训练数据处理管线;研究不同数据源的数据处理方法,将有效策略落地为可规模化运行的生产方案,最终通过预训练实验完成模型效果验证;岗位工作覆盖数据算法与数据工程两个方向,既关注数据处理方法本身,也关注方法在真实数据规模上的可靠落地。 工作职责: 1. 面向网页、代码、PDF/文档、多语言等不同类型的数据,研究数据构造、提取、解析、清洗、去重、质量评估、过滤和数据选择等方法。 2. 分析不同数据源的数据特征和质量问题,通过统计分析、数据抽样、错误案例及必要的实验评估处理效果。 3. 结合预训练模型的训练与评测结果,理解不同数据处理方式对模型效果的影响,并持续迭代数据方案。 4. 设计、开发和优化大规模预训练数据处理管线,关注数据正确性、处理效率、资源成本、稳定性和可复现性。 5. 根据实际问题持续完善相关的数据算法、处理工具和工程能力,并将有效方案沉淀为可复用的数据处理流程。 6. 与预训练算法、数据基础设施及其他相关团队协作,推动数据方案在实际训练任务中的应用。
职位要求
投递

MiniMax is a leading global technology company and one of the pioneers of large language models (LLMs) in Asia. Our mission is to build a world where intelligence thrives with everyone.
MiniMax develops proprietary LLMs across various modalities, including a trillion-parameter MoE model, a speech model with low latency and native support for major Asian languages, and a state-of-the-art text-to-speech and text-to-video models. Experience it now at https://hailuoai.com/
Leveraging these multi-modality general-purpose models, the MiniMax API Platform offers enterprises and developers secure, flexible, and reliable API services, enabling the rapid deployment of AI applications.