ShengShu

大模型infra工程师 (VLM训练)

ShengShu  •  Onsite  •  19 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

大模型infra工程师 (VLM训练)
北京
全职
互联网 / 电子 / 网游
职位描述
负责支撑百亿/千亿级 VLM 从零开始的预训练。直面万卡级集群的通信瓶颈与显存墙,深入分布式训练内核,从第一性原理出发设计并落地业界领先的并行策略与通信调度方案。
1. 分布式并行架构设计与演进:主导设计并持续优化 4D 并行(DP/PP/TP/CP)策略。深入理解并落地 DeepEP、Dual Pipe、ZeRO 等核心机制,在多机多卡环境下实现通信与计算的高度重叠,逼近集群线性加速比。
2. 集群通信效率极致优化:从拓扑亲和性与通信原语调度层面,对分布式训练中的集体通信模式进行深度调优。通过计算通信重叠、通信分组优化等手段,将集群通信效率逼近硬件带宽上限。
3. 显存与算力利用率的极致压榨:围绕 MFU(Model FLOPS Utilization) 和显存利用率两大核心指标,持续进行系统级调优。结合 Context Parallelism 等并行策略,在多卡环境下实现显存、带宽、算力三者之间的最优平衡。
4. 训练稳定性与全链路 Profiling:建立系统级性能看护体系,通过端到端 Profiling 精准定位慢节点、通信拥塞及负载不均等系统性瓶颈。从集群调度、存储 I/O 到计算图执行进行全链路分析,保障数月级长稳训练及故障快速恢复。
职位要求
1. 原理级认知:对分布式训练有极深的理解。不仅熟悉主流并行策略的工程实现,更能从数学与系统层面推导通信量、显存峰值与算力之间的制约关系,具备从第一性原理出发设计或改造并行方案的能力。
2. 工程实战:精读过 Megatron-LM 或 DeepSpeed 核心源码。有从零构建或深度定制 3D/4D 并行框架的实际经验,对 PyTorch 分布式模块的底层执行逻辑有系统性认知。
3. 系统级性能调优能力:具备大规模集群下的端到端调优经验,能够从计算图、内存分配、网络通信、存储 I/O 等多维度进行联合优化。对 MFU 的拆解与提升有系统方法论。
4. 异构计算认知:熟悉现代 GPU 架构特性(SM 调度、Tensor Core、显存层级),对异构计算场景下的负载均衡有实战经验。
5. 加分项:
- 有千卡/万卡级以上 VLM 预训练成功落地经验
- 有 Megakernel 开发或深度优化经验
- 有多框架(PyTorch/JAX)或多芯片(华为昇腾/AMD MI 系列)适配经验
投递
ShengShu

About ShengShu

We are the first General World Model ("GWM") company globally that built a GWM unifying the digital world and the physical world. We are dedicated to building a unified intelligence framework capable of modeling, reasoning, predicting and acting upon the underlying rules that govern both digital and physical worlds. Guided by first principles thinking, we use visual and auditory information, which naturally encodes the physical world, to train our foundation world model and replicate the human process of perceiving, simulating and interacting with the world, and ultimately, to enable AGI that connects the digital world with the physical world.

Industry
Unknown
Company Size
11-50 employees
Headquarters
Beijing, CN
Year Founded
2023
Social Media