MiniMax

大模型算力负责人

MiniMax  •  Onsite  •  2 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

大模型算力负责人高优

上海、北京、杭州

社招

全职

研发 - 基础架构

职位描述

  1. 算力中心规划与建设(核心职责) — Lead 大模型训练/推理场景下的算力中心端到端建设:包括高性能网络拓扑设计(IB/RoCE 组网方案)、集群组网与超节点架构设计、服务器选型规划,从源头规避系统性故障风险,交付高可用、高性能的 AI 基础设施 2. 平台与硬件持续演进 — 参与 GPU 算力池化与调度、基础设施管理平台(资产/容量/故障自愈/可观测性)建设,跟踪 GPU/网络/存储技术演进并推动硬件选型与架构优化,持续降低 TCO

职位要求

  1. 5 年以上云计算/IDC/服务器硬件/网络基础设施相关经验,深入理解计算机体系结构 2. 在以下方向中至少一个有深入的设计或研发经验:GPU 服务器架构设计、高速网络(IB/RoCE/NVLink/NVSwitch)、交换机研发、超节点架构设计 3. 熟悉主流 AI 训练/推理基础设施生态(NVIDIA DGX/HGX/GB200 NVL、集合通信、NCCL 等),了解大模型训练对基础设施的核心需求 4. 具备跨团队项目推动经验和良好的沟通领导力,能协调硬件、网络、软件等多方团队 加分项 1. 有万卡级 AI 算力集群的规划、建设或运营经验,主导过算力中心从 0 到 1 的落地 2. 有交换机研发、超节点架构设计、或服务器整机架构设计经验 3. 有新一代 AI 芯片/加速卡的适配设计经验(如华为昇腾/海思、国产 GPU 等) 4. 有头部云厂商或 AI 公司基础设施团队背景

投递

MiniMax

About MiniMax

MiniMax is a leading global technology company and one of the pioneers of large language models (LLMs) in Asia. Our mission is to build a world where intelligence thrives with everyone.

MiniMax develops proprietary LLMs across various modalities, including a trillion-parameter MoE model, a speech model with low latency and native support for major Asian languages, and a state-of-the-art text-to-speech and text-to-video models. Experience it now at https://hailuoai.com/

Leveraging these multi-modality general-purpose models, the MiniMax API Platform offers enterprises and developers secure, flexible, and reliable API services, enabling the rapid deployment of AI applications.

Industry
IT & Software
Company Size
51-200 employees
Headquarters
Singapore, SG
Year Founded
2022
Social Media