ShengShu

流式视频agent实习生

ShengShu  •  Onsite  •  17 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

流式视频agent实习生
北京
实习
研发 - 算法
职位描述
1、参与国内领先的实时交互式视频生成模型 Vidu S 的研发,负责视频生成 Agent 的设计与落地,使其能够准确理解用户意图,并实时转化为驱动视频生成模型的控制指令。
2、负责面向交互式视频生成场景的 VLM 训练、评测与迭代优化,持续提升模型的多模态理解、指令遵循和交互决策能力;
3、构建 Agent 的多轮对话、意图识别、上下文管理及工具调用能力,打通“用户输入 → 意图解析 → 指令生成 → 视频反馈”的端到端链路;
4、设计并优化 Agent 的编排与决策机制,提升其在流式、低延迟场景下的响应速度、引导准确性与运行稳定性;
5、与算法、Infra及产品团队紧密协作,推进 Agent、VLM 与视频生成模型的联合优化、系统集成及产品落地。
职位要求
1、计算机、人工智能等相关专业,扎实的工程与系统基础,熟练掌握 Python;
2、熟悉 LLM / Agent 的基本原理与开发,有实际构建对话系统或 Agent 应用的经验;
3、对 VLM(视觉语言模型)、视频理解有一定了解,具备参与相关模型训练的意愿与基础;
4、具备较强的问题定位与工程落地能力;
5、加分项:有多模态、视频/图像生成方向经验;有 VLM、视频理解相关的训练或研究经验;熟悉主流 Agent 框架与 LLM 应用编排;有实时/流式交互系统经验,或顶会论文、显著开源贡献;
6、每周到岗5 天,最短实习时长3 个月,是否可转正:是,是否远程:否。
投递
ShengShu

About ShengShu

We are the first General World Model ("GWM") company globally that built a GWM unifying the digital world and the physical world. We are dedicated to building a unified intelligence framework capable of modeling, reasoning, predicting and acting upon the underlying rules that govern both digital and physical worlds. Guided by first principles thinking, we use visual and auditory information, which naturally encodes the physical world, to train our foundation world model and replicate the human process of perceiving, simulating and interacting with the world, and ultimately, to enable AGI that connects the digital world with the physical world.

Industry
Unknown
Company Size
11-50 employees
Headquarters
Beijing, CN
Year Founded
2023
Social Media