
We are looking for a highly skilled Data Engineer with Real World Data (RWD) experience to build and manage end-to-end data pipelines for multimodal healthcare datasets. The ideal candidate will work at the intersection of data engineering, analytics, and healthcare research, transforming complex healthcare data into analysis-ready assets that support advanced analytics, AI/ML initiatives, and evidence-generation studies. This role requires expertise in large-scale healthcare data processing, data harmonization, cloud platforms, and modern data engineering practices.
Responsibilities
Data Engineering & Pipeline Development
• Design, develop, and maintain scalable ETL/ELT pipelines for large healthcare and real world datasets.
• Build and manage data ingestion, transformation, harmonization, and analytics layers.
• Implement data quality frameworks, governance controls, lineage tracking, and monitoring.
• Manage data lifecycle processes across raw, curated, and analytics-ready environments.
• Work with structured and unstructured healthcare datasets from multiple sources.
Healthcare Data Harmonization
• Harmonize heterogeneous healthcare data sources and coding systems into standardized formats.
• Map and transform clinical terminologies including
o SNOMED CT
o ICD-10
o LOINC
o RxNorm
o CPT/HCPCS
• Support implementation of common data models such as OMOP and FHIR.
Analytics & Study Support
• Support data feasibility assessments and data quality evaluations.
• Collaborate with epidemiologists, biostatisticians, data scientists, and business stakeholders.
• Develop reusable data assets, cohorts, and model-ready datasets.
• Enable advanced analytics and AI/ML use cases through reliable data engineering practices.
Application & Platform Development
• Contribute to analyst-facing applications, dashboards, and self-service data products.
• Support development of data products using modern workflow automation and AI assisted engineering approaches.
• Provide guidance on efficient querying and optimization of large longitudinal datasets.
Requirements
Data Engineering
• Strong experience with Python, SQL, Spark / PySpark
• Experience building production-grade ETL/ELT pipelines.
• Strong understanding of data modelling concepts - Star schema, Snowflake schema, Normalization and denormalization
• Experience with metadata management, lineage, monitoring, and data governance.
Platforms & Technologies
Experience in one or more of the following - Palantir Foundry, Databricks, Snowflake, AWS or equivalent cloud platforms, HPC environments
• Containerized workloads Software Engineering Practices
• Git
• CI/CD pipelines
• Unit testing and automation
• Performance monitoring and optimization
Domain Expertise
Candidates should have working knowledge of Healthcare Real World Data (RWD), Claims data, Electronic Health Records (EHR), Registries, Patient-reported outcomes, Wearables and digital health datasets
Understanding of study feasibility, observational research, and healthcare analytics workflows is highly desirable.
AI & Automation Experience
Preferred experience with Large Language Models (LLMs), AI-assisted data engineering, Agentic workflows, Data profiling and automated data quality assessments, integration of ML outputs into production data pipelines
Qualification / Requirement
• 4-8 years of experience in data engineering, healthcare analytics, or real-world data platforms.
• Experience working with large-scale healthcare datasets in regulated environments.
• Strong problem-solving and analytical skills.
• Excellent stakeholder communication capabilities.
• Formal educational qualifications are flexible; relevant experience and expertise are valued.
Nice to Have
• Experience with multimodal healthcare datasets (clinical, omics, imaging, genomics, proteomics, microbiome, etc.).
• Hands-on experience implementing OMOP/FHIR at scale.
• Experience building self-service applications and data products for business users.
• Familiarity with federated data networks and data quality frameworks.
What We're Looking For
• Systems thinker who can work with complex and evolving datasets.
• Strong collaboration skills across technical and business teams.
• Agile mindset with a focus on delivery.
• Commitment to data privacy and ethics.

We help science-driven organisations innovate for a better future through our full range of specialist scientific informatics services.
If your company has Laboratory/R&D, Manufacturing or Trial operations and strives to push boundaries, our informatics services are designed for you.
We deliver a full range of services including Digital Transformation Coaching, Business Analysis, Program & Project Management, Managed Services, Validation, and Custom Development & Integrations.
These services are supported by specialist expertise in Chem & Bioinformatics, Data Science, Cloud, HPC, Instrument Data, MultiOmics, Data Modeling, Semantics, and AI&ML.
We speak the language of science and technology and support our customers and partners to discover, develop, manufacture and test products and solutions that support global health and wellbeing. The only way we can do this is with our people.
We believe that the scientific community’s most important duty to the world is to stay curious. And this is what we do at Zifo.
We strive to stay curious, day in and day out. Asking the right questions, and listening to provide the right solutions.
Contact us at info@zifornd.com.