大模型推理加速算法工程师
上海
全职
芯片板块
职位描述
1、核心技术研发:负责大语言模型(LLM)在推理阶段的架构级优化与加速,解决长文本(Long Context)推理及高并发场景下的性能瓶颈。 2、前沿技术落地:探索并实现大模型推理领域的最新前沿加速技术,包括但不限于 KV Cache 压缩/量化、Context 压缩、Special Token 深度定制、Layer Skipping(层跳过)、推测解码(Speculative Decoding)等。 3、推理框架协同:深入底层计算架构,配合或优化主流大模型推理框架(如 vLLM, TensorRT-LLM, DeepSpeed-FastGen, LMDeploy 等),提升系统吞吐量(Throughput)并降低首字延迟(TTFT)。 4、软硬件协同设计:针对不同硬件平台(如 NVIDIA GPU、算力芯片等)进行极致的算力算子调优,优化内存带宽与显存占用。
职位要求
1、教育背景:计算机、电子信息、自动化或数学等相关专业硕士及以上学历。 2、专业经验:2年以上 大模型(LLM/VLM)相关研发经验,其中至少 1 年深入聚焦于大模型推理加速或底层吞吐调优。 3、深度理解LLM架构:熟练掌握 Transformer 架构原理,对 Attention 机制、注意力矩阵计算以及内存访问瓶颈有深刻理解。 4、加速技术常识:深入研究过并实际落地过以下至少两项技术: (1)显存与上下文优化:KV Cache Eviction/Quantization、PageAttention、长文本 Context 压缩/表征提取。 (2)架构与算子加速:Layer Skipping/Early Exiting 动态路由机制、Special Token 优化策略、FlashAttention/FlashDecoding 算子调优。 5、工程能力:精通 Python 和 C++/CUDA 编程,熟悉 PyTorch 框架,具备优秀的 Code Review 习惯。 6、框架经验:熟练使用或阅读过 vLLM、TensorRT-LLM 等至少一种开源推理框架的核心源码,有二次开发经验者优先。 加分项 在开源社区(如 vLLM, Hugging Face)有高质量 PR 贡献或维护知名加速开源项目者优先。 在 OSDI, SOSP, ASPLOS, ISCA, MLSys, NeurIPS, ICML 等系统或机器学习顶级会议发表过论文者优先。 拥有高性能计算(HPC)、并行计算(数据并行、张量并行、流水线并行)实际落地经验者优先。
投递

XPeng is a leading Chinese Smart EV company that designs, develops, manufactures, and markets Smart EVs that appeal to the large and growing base of technology-savvy middle-class consumers. Its mission is to drive Smart EV transformation with technology and data, shaping the mobility experience of the future. In order to optimize its customers’ mobility experience, XPeng develops in-house its full-stack advanced driver-assistance system technology and in-car intelligent operating system, as well as core vehicle systems including powertrain and the electrical/electronic architecture. XPeng is headquartered in Guangzhou, China. In 2021, the Company established its European headquarters in Amsterdam, along with other dedicated offices in Copenhagen, Munich, Oslo, and Stockholm.The Company’s Smart EVs are mainly manufactured at its plant in Zhaoqing and Guangzhou,Guangdong province.
For more information, please visit https://heyxpeng.com.