数据治理专家(大模型数据 / 行业知识数据方向)
北京、上海、深圳
社招
全职
技术
职位 ID:A237044
职位描述
政企 AI 的效果,往往从数据开始,也在数据里持续变好。你将负责把分散在文档、系统、业务流程和现场反馈中的数据,建设成可理解、可管理、可追溯、可用于训练和应用迭代的数据资产。你的工作会覆盖数据标准、数据目录、质量规则、文档解析、知识切分、标注体系、数据安全、权限管理、难例回流和持续学习闭环。你需要理解行业语义,也需要把治理规则落实到工具和流程中,让数据真正服务于模型训练、知识库问答、智能体运行和客户交付。 岗位职责 1. 负责政企行业数据治理体系建设,制定数据分类分级、元数据、数据目录、质量标准、生命周期和使用规范。 2. 面向政策法规、公文材料、科研文献、设备资料、SOP、工艺数据、业务记录和多模态数据,设计采集、清洗、解析、切分、标注和入库流程。 3. 建立适用于大模型训练、RAG、知识库和智能体的高质量数据集,管理数据来源、版本、授权、加工过程和结果追溯。 4. 设计数据质量规则和检查机制,关注完整性、准确性、一致性、时效性、重复率、可引用性和敏感信息处理。 5. 协同产品、算法、评测和交付团队完成数据需求分析、标注规范制定、样本验收、难例挖掘和数据闭环,推动一次性交付形成持续学习能力。 6. 参与数据脱敏、权限控制、审计留痕和私有化部署的数据方案设计,沉淀治理工具、操作手册、行业数据模板和质量报告。
职位要求
投递

ModelBest established in August 2022, is an Artificial Intelligence technology company headquartered in Beijing, China. The company is deeply rooted in the field of general AI, with a focus on the innovation and application transformation of large-scale models. ModelBest boasts a prestigious founding R&D team from the Tsinghua in the AI field. Leveraging numerous cutting-edge technologies in natural language processing, the company is currently constructing a large-scale pre-trained model library and corresponding tools, aiming to standardize the technology and applications of large models.
Based on the capabilities of the CPM series of large models, ModelBest has successfully facilitated the intelligent upgrade and efficiency enhancement of various industries. Moreover, the company has released a multi-modal dialog assistant with hundreds of billions of parameters, “Luca", to the public. Adhering to the core philosophy of "intelligence encompassing all", we look forward to exploring and expanding unknown possibilities with global partners, aiming to ensure AI technology serves humanity for a better life in a safe and inclusive manner, laying a solid foundation for the advent of the AGI world.
OpenBMB (Open Lab for Big Model Base,https://github.com/OpenBMB), founded by ModelBest Inc & TsinghuaNLP, aims to build foundation models and systems towards AGI.