Vi

Staff Data Scientist

Vi  •  Boston, MA (Onsite)  •  17 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

Role Summary

Vi Engage puts predictive models into the daily operations of the largest health systems and health plans in the country — driving care navigation, specialty capture, and the workflows that follow from them. This role builds the modeling engine those deployments run on.

You will own the pipelines that turn longitudinal claims, EHR, lab, and online behavioral intent data into predictions about which patients face upcoming healthcare utilization and which are high-propensity to enroll in preventative care programs. Not one model for one customer — the config-driven machinery that trains, selects, scores, and delivers models for any customer, with automatic feature engineering and model selection doing the work that bespoke engineering does today.

The measure of this work is generalization. A model that lifts enrollment for one health plan is a good result; a pipeline that reproduces that lift for the next twenty without per-customer engineering is the product. You set the modeling standard that Forward Deployed Data Scientists build their customer deployments against, and you are accountable for the improvements that hold across all of them.

What You'll Own

  • ML pipelines; Built for scale. Config-driven pipelines that train, score, and deliver predictions from large longitudinal datasets.
  • Mass customization. Automatic feature engineering and model selection that produces a customer-specific model without customer-specific work.
  • Cross-customer generalization. Find the modeling improvements that hold up across every deployment.
  • Building blocks for the field. Implementing and maintaining the DS components that Forward Deployed Data Scientists compose into running customer deployments, so the next deployment is faster than the last.
  • Productization. Turn pilots and one-off proofs into product capabilities that survive their fourth customer.

What We're Looking For

  • Modeling longitudinal data in production. You have personally shipped uplift, survival, or propensity models whose output changed how an organization spent money or who it reached. This is the capability we screen hardest on.
  • POC to product. You have taken a pilot or proof of concept and turned it into something repeatable that survived a second, third, and fourth customer. You know which parts of a one-off are the product and which parts are the customer.
  • Real data science depth. Segmentation, campaign optimization, and the identification strategy behind an uplift estimate. You can defend a modeling choice to a skeptical internal analytics team, evaluate a model honestly, and say when a simpler approach is the right answer.
  • Python and ML engineering. Fluent in Python and the working stack — pandas, sklearn, PySpark, airflow. You build the pipeline, not just the model inside it.
  • MLOps. Model tracking and deployment tooling — mlflow, SageMaker, or comparable
  • AWS and Cloud. You don’t need to hand-off to a dedicated engineer. You can get your models running at scale in the cloud using our AWS stack: S3, Glue, EMR, MWAA, SageMaker.

Nice To Have

  • Healthcare or life sciences domain knowledge — claims, EHR, HL7/FHIR, lab data, or population health analytics
  • Familiarity with HIPAA and healthcare compliance and data governance frameworks
  • Experience designing pilots and efficacy studies that tie model performance to a business outcome
  • Experience building internal platforms or frameworks that other engineers build on top of

What This Role Is Not

This is an applied, in-production role. It is not a research position — the work is measured by pipelines that run and models that hold up across customers, not by novelty. It is primarily not a customer-facing role (Applied Data Scientists own the customer accounts) but you may interface with design partner clients on occasion. This is not a management role — you will be hands-on architecting and building this product with the team.

Vi

About Vi

The enterprise AI platform powering health

Vi has already helped 190M+ people live healthier lives, generated $2B+ in enterprise ROI, and accelerated development of 50+ world-changing drugs.


Our platform transforms raw data into business ROI and better health outcomes—driving measurable impact across the entire healthcare value chain.

At the heart is the Vi Data Web—a privacy-safe network of 190M+ de-identified patient and member profiles enriched with behavioral and consumer data covering 96% of U.S. households. This data moat powers our suite of AI applications:

Vi Activate — Precision targeting to activate new patients

Vi Engage — Predictive & personalized interactions with patients and providers

Vi Operate — An agentic suite built to driver operational excellence.

All insights flow through Vi Pulse, our real-time ROI dashboard that connects prediction to measurable impact.

With Vi, health enterprises get AI on Day 1—no disruption to IT/cloud stacks.

Our 4X ROI model means we win only when our partners do — delivering financial and health outcomes across healthcare, Biopharma, and wellness.

We are building the infrastructure for the future of health—where data, AI, and enterprises work seamlessly together as one powerful ecosystem.

Industry
IT & Software
Company Size
51-200 employees
Headquarters
New York, NY
Year Founded
2016
Website
vi.co
Social Media