Job Description
๐ง๐ต๐ถ๐ ๐ฟ๐ผ๐น๐ฒ ๐ถ๐ ๐ณ๐ผ๐ฟ ๐ผ๐ป๐ฒ ๐ผ๐ณ ๐๐ต๐ฒ ๐ช๐ฒ๐ฒ๐ธ๐ฑ๐ฎ๐'๐ ๐ฐ๐น๐ถ๐ฒ๐ป๐๐
๐ฆ๐ฎ๐น๐ฎ๐ฟ๐ ๐ฟ๐ฎ๐ป๐ด๐ฒ: ๐ฅ๐ ๐ญ๐ฑ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ - ๐ฅ๐ ๐ฏ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ๐ฌ (๐ถ๐ฒ ๐๐ก๐ฅ ๐ญ๐ฑ-๐ฏ๐ฌ ๐๐ฃ๐)
Experience: 3+ yrs
Location: Bengaluru, Karnataka, India
Job Type: Full-time
We are looking for an experienced MLOps Engineer to build, automate, and operate reliable machine learning infrastructure and production deployment workflows. This role focuses on creating scalable ML platforms, streamlining model lifecycle management, and ensuring the reliability, security, and performance of machine learning workloads across GCP and Azure
The ideal candidate will have strong hands-on expertise in Python, Docker, Kubernetes, Terraform, Airflow, MLflow, Vertex AI, and cloud infrastructure You will work closely with data scientists, ML engineers, software engineers, and cloud architects to take machine learning models from experimentation through production while establishing robust automation and monitoring practices.
Requirements
Key Responsibilities
- Build and automate ML workflows using Apache Airflow / Cloud Composer for data ingestion, preprocessing, training, and deployment.
- Manage MLflow for experiment tracking, model packaging, versioning, and model registry processes.
- Design and optimize training environments for machine learning and LLM workloads.
- Develop scalable model-serving solutions using FastAPI, Flask, API Gateway, and high-performance inference endpoints.
- Manage Docker, Kubernetes, GKE, and AKS infrastructure, including auto-scaling GPU/CUDA workloads.
- Build and maintain CI/CD pipelines to automate ML application and model deployments.
- Monitor model performance, inference latency, data drift, infrastructure health, and production reliability.
- Work with Vertex AI Workbench, Model Garden, Feature Store, Vertex AI Pipelines, and BigQuery ML
- Use Terraform to provision and maintain secure, scalable, and reproducible cloud infrastructure.
- Develop production-quality Python applications using modular design, testing, and engineering best practices.
- Work with CDC, Spark/PySpark, and optimize data movement between BigQuery and ML training environments.
- Implement secure ML infrastructure using IAM, VPC Service Controls, endpoint security, and cloud security best practices.
- Support enterprise-scale ML infrastructure migration and modernization across GCP and Azure
What Makes You a Great Fit
- 3+ years of experience in MLOps, ML Engineering, or a closely related field.
- Strong hands-on expertise in Python, Docker, Kubernetes, GKE/AKS, and Terraform
- Practical experience with Airflow/Cloud Composer and MLflow
- Strong knowledge of GCP, Azure, Vertex AI, BigQuery ML, and Vertex AI Pipelines
- Experience managing Kubernetes operators and resources for ML workloads.
- Hands-on experience with FastAPI/Flask, API Gateway, CI/CD, and model serving
- Strong understanding of CDC, Spark/PySpark, IAM, and VPC Service Controls
- Experience building and operating production ML platforms with a focus on scalability and reliability.
- Exposure to LLMOps, foundation models, prompt versioning, or vector databases is an advantage.
- Experience migrating or managing enterprise-scale ML infrastructure across Azure and GCP is preferred.
- Relevant MLOps/ML Engineering certifications and production ML platform experience are a plus.
- Strong troubleshooting, analytical, communication, and collaboration skills, with a strong ownership mindset.