Models are what they eat. But a large portion of training compute is wasted training on data that are already learned, irrelevant, or even harmful, leading to worse models that cost more to train and deploy.
At DatologyAI, we’ve built a state of the art data curation suite to automatically curate and optimize petabytes of data to create the best possible training data for your models. Training on curated data can dramatically reduce training time and cost ( 7-40x faster training depending on the use case), dramatically increase model performance as if you had trained on >10x more raw data without increasing the cost of training, and allow smaller models with fewer than half the parameters to outperform larger models despite using far less compute at inference time, substantially reducing the cost of deployment. For more details, check out our recent research on synthetic data scaling ( BeyondWeb) and pretraining with domain-specific data ( The Finetuner’s Fallacy).
We raised a total of $57.5M in two rounds, a Seed and Series A. Our investors include Felicis Ventures, Radical Ventures, Amplify Partners, Microsoft, Amazon, and AI visionaries like Geoff Hinton, Yann LeCun, Jeff Dean, and many others who deeply understand the importance and difficulty of identifying and optimizing the best possible training data for models. Our team has pioneered this frontier research area and has the deep expertise on both data research and data engineering necessary to solve this incredibly challenging problem and make data curation easy for anyone who wants to train their own model on their own data.
This role is based in Redwood City, CA. We are in office 4 days a week.
You'll shape the way users experience our platform from the ground up.
Our users are ML engineers and researchers, and the workflows they run are long, expensive, and technically dense. Your job is to make those systems legible: designing the configuration surfaces, telemetry, and visualizations that let an engineer parameterize a curation pipeline, inspect its behavior at runtime, and evaluate the quality of the dataset it produces.
This role asks for high visual craft, and the ability to take a genuinely complex feature and define a design for it that feels obvious. You'll own the visual quality of what lands in production and partner closely with engineers to get it there. You'll work with leadership, engineering, research, and customers, with a rare opportunity to influence product direction at an early stage.
Create and design key user-facing features and components of our data curation platform end to end, from concept to production
Define the design principles for an enterprise platform, and set the bar for clarity and confidence in workflows where the stakes are high
Establish and evolve an intuitive, scalable design system, style guide, and illustration language so the product scales consistently
Deliver high quality assets—icons, illustrations, and motion—that hold up across the product
Create wireframes, prototypes, and high-fidelity designs that balance user needs with real technical constraints
Conduct user research, usability testing, and analysis to understand pain points and validate design solutions
Collaborate with engineering and research teams to define both long-term strategy and short-term deliverables
Help establish our design culture as an early design team member
Technical workflows. Front-end UI for setting up, launching, monitoring, and intervening in long-running curation and training jobs
Data visualization. Making dataset composition, quality, and cost legible at a glance
Onboarding and integration. Flows that get customers' data into the platform
5+ years of experience designing digital products, with a portfolio showcasing strong UX, UI, and interaction design skills
High visual craft, with the ability to produce finished assets yourself—iconography, illustration, and interface detail that ship
A track record of motion design that is beautiful and intuitive, clarifying what's happening rather than decorating it
You've created intuitive, scalable design systems and style guides from scratch, and evolved them as a product grew
Experience simplifying complex technical concepts into intuitive user experiences
Experience designing data-dense interfaces and data visualization
Experience designing with cross-functional partners for multiple use cases and distinct user types—enterprise technical users alongside ML researchers and scientists
Excellent communication skills with the ability to articulate design decisions to technical and non-technical stakeholders
Proficiency with Figma and other prototyping software
Comfortable working in an ambiguous, fast-paced startup environment
Ability to balance short-term deliverables with long-term vision
Self-motivated problem solver who can work independently while collaborating effectively with cross-functional teams
You implement your own work—comfortable in HTML and CSS, and shipping front-end code alongside engineers
Experience working with AI tools to prototype and implement designs
Experience designing platform workflows for ML engineers and researchers
Experience designing for technical or enterprise users
Experience working as an early designer at a startup
At DatologyAI, we are dedicated to rewarding talent with competitive salary and meaningful equity. The salary for this position ranges from $180,000 to $250,000.
Starting pay is based on job-related skills, experience, qualifications, and interview performance.
Benefits:
100% covered health benefits (medical, vision, and dental).
401(k) plan with a generous 4% company match.
Unlimited PTO policy
Paid Parental Leave of 12 weeks, plus 6 months of WFH flexibility.
Annual $2,000 wellness stipend.
Annual $1,000 learning and development stipend.
Daily lunches and snacks are provided in our office!
Relocation assistance for employees moving to the Bay Area.

DatologyAI builds tools to automatically select the best data on which to train deep learning models. Our tools leverage cutting-edge research—much of which we perform ourselves—to identify redundant, noisy, or otherwise harmful data points. The algorithms that power our tools are modality-agnostic—they’re not limited to text or images—and don’t require labels, making them ideal for realizing the next generation of large deep learning models. Our products allow customers in nearly any vertical to train better models for cheaper.