ShengShu

多模态数据算法工程师 (视频)

ShengShu  •  Onsite  •  18 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

多模态数据算法工程师 (视频)
北京
实习
研发 - 算法
职位描述
1、视频预处理与解码:负责原始视频的元数据探测、转码、HDR→SDR、黑边检测、音视频分离对齐。优化解码路径(FFmpeg / Decord / GPU decoder),保证大规模任务的吞吐与稳定性;
2、镜头与场景切分:设计并落地 shot / scene 切分:镜头边界检测、视觉特征场景合并、按边界切片、jump-cut 检测。建立边界 F1 / mIoU / 长度合规等评估,持续压误切、漏切;
3、视觉理解与结构化标注:接入 VLM 做视频 caption、语义打标、实体提取;接入开放词表检测、字幕擦除等视觉算子。保证标注 schema、token 用量、失败重试可统计、可回放。
职位要求
1、算法竞赛获奖经历,或为显著开源贡献者优先,计算机视觉 / 机器学习等相关专业,本科及以上,具有视频理解、视频处理或大规模数据生产经验;
2、精通 Python;熟悉 PyTorch 推理落地(加载、batch、混合精度、OOM 降级)。
3、熟悉视频容器与编解码:FFmpeg / ffprobe、H.264/H.265、PTS/DTS、转码与抽帧;有过长视频切分或直播流(如 MPEG-TS)踩坑经验更佳;
4、做过至少一类视觉模型生产化:CLIP / DINOv2 / 检测 / OCR / 质量评估 / VLM 打标中的若干项;
5、加分项:镜头边界 / 场景分割的论文或生产经验,有人工标注对齐、failure mode 分析;VLM 视频打标:prompt 设计、结构化输出、限流与 token 统计;字幕检测与擦除、人脸/实体检测、参考视频数据构建。
6、每周到岗5 天,最短实习时长3 个月,是否可转正:是,是否远程:否。
投递
ShengShu

About ShengShu

We are the first General World Model ("GWM") company globally that built a GWM unifying the digital world and the physical world. We are dedicated to building a unified intelligence framework capable of modeling, reasoning, predicting and acting upon the underlying rules that govern both digital and physical worlds. Guided by first principles thinking, we use visual and auditory information, which naturally encodes the physical world, to train our foundation world model and replicate the human process of perceiving, simulating and interacting with the world, and ultimately, to enable AGI that connects the digital world with the physical world.

Industry
Unknown
Company Size
11-50 employees
Headquarters
Beijing, CN
Year Founded
2023
Social Media