The Geneva Foundation

Lead AI Data Engineer (Hybrid)

The Geneva Foundation  •  $155k - $193k/yr  •  Bethesda, MD (Hybrid)  •  2 hours ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

About The Position

The Lead AI Data Engineer serves as a scientific, technical, and managerial lead for data-heavy research projects. The Lead AI Data Engineer will overlap with a multidisciplinary team of government and contract researchers, academic experts and consumers. The Lead AI Data Engineer will provide technical and management support, oversee project execution, and provide key guidance on data architecture and infrastructure. The team that the Lead AI Data Engineer oversees is responsible for the curation and maintenance of large data pipelines, developing ETL pipelines, defining schemas, identifying bottlenecks, and deriving variables /features directly used in machine learning / AI model development. The Lead AI Data Engineer serves as a versatile position who designs, builds, and maintains the entire data ecosystem. As project aims evolve, collaborators join, and team processes change, an ideal candidate is flexible and can adapt quickly. The ideal candidate will also be mindful of data privacy / security and will adhere to data governance and data management best practices.

This is a full time hybrid remote position working at the Walter Reed National Military Medical Center in Bethesda, MD that will require working in office/on site at least 2 days per week. Background checks will be administered.

About The Program

Sleep Physiology Modeling Project: Sleep & Wearables Operational Readiness for Research & Defense (SWORD) Lab

Salary Range

$155,000 - $193,000. Salaries are determined based on several factors including external market data, internal equity, and the candidate’s related knowledge, skills, and abilities for the position.

Qualifications

  • PhD in a relevant field (e.g., Computer Science, Engineering, Data Science, Biomedical Engineering) required
  • 8+ years experience with multimodal data analysis and data pipeline engineering required
  • Proven experience with multivariate signal processing (e.g., time-series biosensor data)
  • Hands-on experience with relational (SQL) and non-relational (NoSQL) databases
  • Hands-on experience with version control systems (eg, Git) and demonstrated ability to work in (and lead) a collaborative coding environment
  • Solid problem-solving and analytical skills to address complex technical challenges
  • Hands-on experience using Google Cloud Platform (GCP) cloud infrastructure (or equivalent), including setting up and managing cloud-native data warehouses (eg, BigQuery), storage, and compute resources
  • Ability to translate high-level scientific hypotheses into scalable engineering solutions and data products
  • Ability to work in a fast-paced, multidisciplinary, multi-site (sometimes asynchronous) team environment
  • Preferred qualifications: Strong knowledge of sleep science and hands-on experience with handling data from consumer wearable devices (eg, actigraphy, PPG, EEG); familiarity with machine learning workflows, including model development, model tuning, and deploying models at scale; Leadership and/or project management experience with the ability to oversee a team of people ingesting data

Management Responsibilities

  • Foster a collaborative coding and research environment, driving skill development for junior and mid-level data engineers and analysts across the data pipeline
  • Serve as the primary technical liaison to senior management, translating high-level research aims into actionable objectives
  • Communicate team progress, bottlenecks, and milestones

Responsibilities

  • Produce clean, well-documented, efficient code across the entire stack
  • Design, develop, and deploy robust, scalable applications (both front-end interfaces and back-end data pipelines) to support large-scale research
  • Lead the engineering workflows to acquire, ingest, and clean multimodal datasets, ensuring efficient storage, retrieval, and processing of massive datasets (+1million records)
  • Architect and maintain scalable infrastructure to support advanced machine learning models using physiological features and sleep microarchitectures
  • Optimize application performance and scalability through performance tuning, code refactoring, and database optimization techniques
  • Stay updated with emerging industry trends, academic literature, and technologies in data engineering, cloud architecture, and machine learning to continuously improve the lab’s technical capabilities
  • Ensure data integrity throughout engineering workflows. Assist in the preparation of Standard Operating Procedures (SOPs), analytical frameworks, and technical documentation
  • Maintain open communication with leadership, advise on technical processes, and curate progress reports
  • Provide technical support and oversight to team members with less experience
  • Assist in regulatory support, Data Sharing Agreements, and other project documentation
The Geneva Foundation

About The Geneva Foundation

The Geneva Foundation is a 501(c)3 nonprofit established in 1993 with the purpose to ensure optimal health for service members and the communities they serve. This purpose is accomplished through our mission to advance military medicine through research, development, and education.

Geneva has 30 years of experience in delivering full spectrum scientific, technical, and program management expertise in the areas of federal grants, federal contracts, industry sponsored clinical trials, and educational services.

Industry
Biotech & Life Sciences
Company Size
201-500 employees
Headquarters
Tacoma, WA
Year Founded
1993
Social Media