Job Description
Monday-Friday 8:00AM-4:30PM| Hybrid Position| Weekly Earned Wage Access is an option for this position.
Job Purpose or Goals: The AI Operations and Monitoring Engineer is responsible for ensuring the reliability, performance, documentation, and compliance of production AI/ML systems, keeping them stable and functioning as intended. They play a key role in maintaining system uptime and driving effective incident response to minimize outages and protect overall service quality.
Tasks:
- Monitor AI/ML system health, drift, and anomalies by continuously tracking model performance, identifying unexpected behavior, and ensuring systems operate within expected parameters.
- Build dashboards, alerts, and observability tooling to give teams real-time insight into system performance and enable rapid detection of issues before they impact operations.
- Implement CI/CD pipelines for models to streamline deployments, maintain version control, and ensure that model updates move reliably from development to production.
- Maintain training and deployment environments by keeping infrastructure stable, updated, and optimized so models can be trained and deployed efficiently.
- Develop automation for scaling, retraining and reducing manual workloads and ensuring models automatically adapt to new data or demand.
- Respond to incidents and outages to restore services quickly, conduct rootcause analysis, and prevent future disruptions to AI/ML workloads.
- Collaborate on security, compliance, and governance to ensure AI systems follow organizational standards and regulatory requirements.
- Write runbooks and maintain audit logs to document procedures, support operational consistency, and provide traceability for all system actions.
- Performs other duties as may be assigned.
Job Requirements:
Bachelor's degree in computer science or related field, or 4 years relevant professional experience
3 years' experience in DevOps, SRE, or MLOps Proficiency with tools like Kubernetes, Docker, and cloud services
Experience with observability tools
Familiarity with ML deployment workflows
- Monitor AI/ML system health, drift, and anomalies by continuously tracking model performance, identifying unexpected behavior, and ensuring systems operate within expected parameters.
- Build dashboards, alerts, and observability tooling to give teams real-time insight into system performance and enable rapid detection of issues before they impact operations.
- Implement CI/CD pipelines for models to streamline deployments, maintain version control, and ensure that model updates move reliably from development to production.
- Maintain training and deployment environments by keeping infrastructure stable, updated, and optimized so models can be trained and deployed efficiently.
- Develop automation for scaling, retraining and reducing manual workloads and ensuring models automatically adapt to new data or demand.
- Respond to incidents and outages to restore services quickly, conduct rootcause analysis, and prevent future disruptions to AI/ML workloads.
- Collaborate on security, compliance, and governance to ensure AI systems follow organizational standards and regulatory requirements.
- Write runbooks and maintain audit logs to document procedures, support operational consistency, and provide traceability for all system actions.
- Performs other duties as may be assigned.
Bachelor's degree in computer science or related field, or 4 years relevant professional experience
3 years' experience in DevOps, SRE, or MLOps Proficiency with tools like Kubernetes, Docker, and cloud services
Experience with observability tools
Familiarity with ML deployment workflows