Hiring Research Scientists to help define and advance the core technical direction of an early-stage AI company focused on multimodal representation learning
This role is for a researcher who has trained models natively across multiple modalities, audio, images, video, text, and other structured inputs, and understands how richer representations can improve downstream reasoning and action. You’ll work directly with the founder, help set research direction from the earliest stage, and have significant autonomy over what gets explored and built.
What You’ll Own
- Lead research in multimodal representation learning
- Train and improve models across audio, image, video, text, and related modalities
- Explore architectures that learn unified representations across multiple input types
- Help define research priorities and long-term technical direction
- Design experiments, evaluate model behavior, and identify promising research paths
- Work across ambiguous, greenfield research problems with substantial autonomy
- Translate research insights into systems that can ultimately support real-world AI products
What We’re Looking For
- 3–6 years of relevant research experience, with flexibility for exceptional senior candidates
- Deep expertise in multimodal representation learning
- Track record of training models natively across multiple modalities
- Strong understanding of modern deep learning architectures, training techniques, and evaluation
- Experience conducting frontier-level research at a leading AI lab, research organization, or similarly high-bar environment
- Research judgment strong enough to independently propose and drive new technical directions
- Comfortable working in a very early-stage, high-intensity environment with significant ambiguity
- High ownership and the ability to operate without a predefined research roadmap
Strong Green Flags
- Direct experience training omni-models or native multimodal models
- Research spanning combinations of video, audio, vision, language, and structured signals
- Experience at leading multimodal or video-generation organizations such as Luma, Runway, Pika, or comparable frontier AI teams
- Published research in multimodal learning, representation learning, generative modeling, video models, or adjacent fields
- Prior research or technical leadership experience
- Experience taking research ideas from hypothesis through large-scale training and rigorous evaluation
- Exceptional technical or competitive achievement, including IOI, IMO, quantitative research, or other world-class competitive backgrounds
The strongest candidate has worked directly on models that learn across several modalities rather than simply combining independently trained models at inference time. You should understand the challenges of cross-modal representation learning, model training, conditioning, data quality, evaluation, and scaling—and be interested in pushing those systems further.
Strong multimodal researchers from adjacent areas, particularly video generation and multimodal foundation models, are also highly relevant.