San Francisco · in person · $250k-$400k + equity
Applying here considers you for research roles across Vizcom. Roles are defined by the person, not the posting.
Vizcom is where design teams at Nike, GM, New Balance, and Hasbro do their daily work: sketch, render, color and material, 3D, export. The render was never the point. The point is the physical thing. We call it pencil to product. We're a Series B company with over $52M raised.
More than 700,000 designers have worked in Vizcom, and every session leaves a trail: candidates picked, outputs promoted into the design, regions masked and renamed, batches kept or thrown away. That trail is the most valuable thing we make that isn't the product itself. It's also, today, more archaeology than asset. Only a fraction of what happens in a session reaches training grade, and the strongest post-training work in the world is increasingly won or lost on exactly this kind of data. Your job is to turn the trail into a machine.
One problem to take home: a design session is a branching tree, not a sequence. Designers fork, backtrack, and abandon whole directions on the way to the thing they keep. The judgment lives in the shape of that tree, and today we only capture single steps.
As an ML data engineer here, you'll build the flywheel itself: the system that turns professional design work into training-grade preference data, and training results back into a better product. You'll work beside the researchers consuming what you build, inside the product code where the signals are born, and against the warehouse where they land. Your customer sits one desk away, and you'll feel it within days when a dataset you shipped lets them ask a question nobody could ask before.
This is not a support role, and it is not offline ETL. The pipelines you design run through a live canvas that professional teams use every day, under enterprise agreements with some of the most protective brands in the world. Capturing more without breaking trust or performance is the craft. If you want to train models and not build the systems that feed them, this isn't the seat. If you believe the next advances get won in the data, it is.
We think a dataset is a product: it has users, versions, and a quality bar. A training result should reproduce from a dataset fingerprint months later, and "where did this example come from" should always have an answer. That standard is rare. Here, it's the job.
The capture surface: what the product records, designed with product engineers and shipped in product code.
The pipeline from canvas to warehouse to training set: clean, versioned, reproducible from a fingerprint.
Dataset contracts and lineage: every example traceable to its origin, every result replayable months later.
The boundary system: what's capturable under which enterprise agreements, consent and isolation designed in from the start, sanitized slices for research and work trials.
Collection instruments: when historical signal runs out, the mechanisms that gather explicit feedback designers actually want to use.
The honest encoding of ambiguity: unpicked is not disliked, abandoned is not rejected, and the data should say so.
This list is a charter, not a week one to-do. Nobody runs all of it at once, and the sequencing is yours to argue for.
Days 1 to 30: map. Walk the event surface, the warehouse, and the research log. Know what exists, what's missing, and which questions the data currently can't answer.
Days 30 to 60: ship. Take one new signal end to end: instrumented in product, landed in the warehouse, versioned, and in a researcher's hands.
Days 60 to 90: contract. Write the data standard everything after runs on: versioning, fingerprints, boundaries.
Have built training-data or large-scale data pipelines that other people depended on.
Have an experimentalist mindset
You've trained models yourself, so you know what training actually consumes.
You've run labeling, human-feedback, or evaluation ops.
You've worked with data under real privacy or contractual constraints and enjoyed the puzzle.
And the disposition we keep coming back to: you look at product exhaust and see evidence.
A dataset you'll know better than anyone alive: the recorded decisions of more than 700,000 designers, growing every day.
Researchers as your users, one desk away. Your work changes what they can try within the same week.
A field position that barely exists anywhere else: every frontier lab is starved for professional preference data. You'd be the person who makes it exist.
A direct line to the founders, including a CEO who trained as a transportation designer at Honda. The judgment in your pipelines is in the building.
Research-log culture: what we learned, not what we worked on. Negative results are celebrated.
Nothing scales on vibes. A result repeats before it earns compute.
We publish what we learn: writeups and showcases, negative results included.
We hire through paid work trials on real problems with real data, not LeetCode.
Visa sponsorship: yes.

Building tools that shorten the distance between having ideas and bringing them to life.
https://linktr.ee/vizcom_