About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an MLOps / AI Infrastructure Engineer based in India.
Responsibilities
- Architect, build, and maintain scalable model deployment and serving pipelines for LLMs, deep learning models, and predictive analytics workloads.
- Develop production-grade model-serving architectures using technologies such as Triton, Ray, BentoML, or comparable platforms.
- Build and automate CI/CD pipelines covering data ingestion, feature management, model training, validation, deployment, and monitoring.
- Design and maintain reliable ML infrastructure capable of supporting large-scale AI and generative AI workloads.
- Manage containerized workloads and orchestration platforms, including Kubernetes, across cloud environments.
- Monitor GPU and CPU cluster utilization and identify opportunities to improve performance, availability, and resource efficiency.
- Optimize cloud infrastructure and compute expenditure for resource-intensive AI workloads.
- Partner closely with data scientists and AI researchers to transition experimental models into resilient, scalable, low-latency production services.
- Contribute to infrastructure automation, observability, reliability, and continuous improvements across the ML platform.
Requirements
- 4–8 years of professional experience in cloud infrastructure, DevOps, backend engineering, MLOps, or a closely related engineering discipline, with significant exposure to ML systems.
- Strong proficiency in Python and hands-on experience with Docker, Kubernetes, and Terraform.
- Practical experience working with cloud platforms and managed AI/ML services such as AWS Bedrock, Amazon SageMaker, Google Cloud Vertex AI, or equivalent technologies.
- Solid understanding of machine learning infrastructure, model deployment, model serving, and production ML lifecycle management.
- Hands-on familiarity with vector databases such as Pinecone, Milvus, Qdrant, or similar technologies.
- Experience with LLM orchestration frameworks and modern generative AI infrastructure.
- Understanding of CI/CD, infrastructure-as-code, containerization, monitoring, and production reliability practices.
- Strong problem-solving skills and the ability to troubleshoot complex infrastructure and performance challenges.
- Comfortable collaborating with data scientists, AI researchers, software engineers, and other technical stakeholders.
- Bachelor's degree in Computer Science, Engineering, or a related technical field is preferred.
Conditions
- Annual compensation of INR 1,800,000–2,500,000.
- Full-time employment opportunity.
- Remote working arrangement with flexibility to work remotely from India.
- Opportunity to work on large-scale machine learning and generative AI infrastructure.
- Exposure to modern AI/ML technologies, cloud platforms, GPU infrastructure, and model-serving systems.
- Significant technical ownership across deployment, automation, scalability, and infrastructure optimization.
- Collaborative environment working closely with AI researchers, data scientists, and engineering teams.
- Opportunity to develop expertise in rapidly evolving MLOps and AI infrastructure technologies.