NextloopTechnologies

Lead / Principal MLOps Engineer

Gurugram, Chennai | 24/Aug/2026

About the job

We are seeking a highly skilled Lead / Principal MLOps Engineer to design, build, and operate secure, scalable, and reliable enterprise ML platform infrastructure. This role will lead the end-to-end operationalization of machine learning workloads—from training and experimentation through deployment, serving, monitoring, and lifecycle management.

Responsibilties

  • Design, build, and maintain scalable, secure, highly available, and reliable MLOps and ML platform infrastructure.
  • Lead end-to-end ML pipelines covering training, validation, deployment, serving, monitoring, and lifecycle management.
  • Build and manage Kubernetes-based model deployment and serving infrastructure.
  • Implement robust CI/CD pipelines for ML applications, models, and platform infrastructure.
  • Manage Infrastructure as Code and automate cloud provisioning, configuration, and environment management.
  • Design and optimize Kubernetes environments for high availability, scalability, disaster recovery, security, and efficient resource utilization.
  • Implement monitoring and observability solutions across ML models, pipelines, applications, infrastructure, and platform services.
  • Monitor and manage model performance, data drift, concept drift, reliability, and operational health.
  • Optimize cloud and GPU infrastructure for performance, scalability, reliability, and cost efficiency.
  • Lead troubleshooting, incident management, root cause analysis, and production issue resolution. Collaborate with Data Science, ML Engineering, Platform Engineering, and Cloud teams to improve ML workflows.
  • Establish MLOps best practices, engineering standards, architecture patterns, governance controls, and documentation.
  • Provide technical leadership, mentoring, architecture guidance, and infrastructure/code reviews..

Required Qualification

  • 8+ years of overall engineering experience, including at least 4 years of hands-on experience in MLOps, ML Platform Engineering, or Cloud Engineering.
  • Strong hands-on experience designing and operating production MLOps or ML platform environments.
  • Strong understanding of ML lifecycle management: experimentation, model versioning, model registry, training, deployment, serving, monitoring, and governance.
  • Hands-on experience with one or more: MLflow, Kubeflow, Apache Airflow, Argo Workflows, Weights & Biases (W&B), or Amazon SageMaker Advanced experience with Docker and Kubernetes, including EKS, GKE, or AKS.
  • Strong practical experience with Helm and KServe for Kubernetes-based application and model deployment/ serving.
  • Strong experience with GitHub Actions, GitLab CI, or Jenkins to automate build, test, deployment, and release processes.
  • Strong cloud-platform experience in AWS, Azure, or GCP; multi-cloud or hybrid-cloud experience is preferred.
  • Strong experience with ML pipeline orchestration, model deployment and serving, model monitoring, data drift and concept drift detection, autoscaling, HA, DR, incident management, and RCA.
  • Experience optimizing cloud and GPU infrastructure for performance, scalability, reliability, and cost efficiency.
  • Strong problem-solving, technical leadership, communication, stakeholder-management, and mentoring abilities.

Skills

MLOps
Kubernetes
Terraform
Python
MLflow

Gurugram, Chennai, On-site | Contract