About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Platform Engineer based in India.
Responsibilities
- Design, build, deploy, and operate scalable AI platform services with clear API contracts, versioning, service-level objectives, and production-readiness standards.
- Develop production-grade Python services, shared libraries, SDKs, and integration patterns for consumption by multiple product engineering teams.
- Build and evolve a multi-provider model gateway with routing, fallback, retry, rate limiting, cost attribution, and budget enforcement capabilities.
- Develop RAG-as-a-service capabilities covering data ingestion, chunking, embeddings, retrieval, hybrid search, and retrieval-quality measurement.
- Establish prompt management capabilities including templates, versioning, evaluation, rollout controls, and tenant-specific customization.
- Build guardrails and content-safety mechanisms such as input filtering, output validation, PII redaction, and controlled tool execution.
- Develop agent and tool-use patterns suitable for reliable production workflows.
- Build observability capabilities covering prompt and response tracing, latency, errors, cost per request, quality signals, drift detection, and audit logging.
- Design and maintain evaluation infrastructure using benchmark datasets, offline evaluations, LLM-as-judge approaches, human review workflows, and regression testing.
- Build and operate model deployment pipelines, including safe deployment and rollback processes for fine-tuned models where applicable.
- Establish alerting and SLO frameworks for AI services, treating quality regressions as a first-class operational concern.
- Participate in on-call rotations, create incident response runbooks, and continuously improve operational procedures following incidents.
- Collaborate directly with product engineering teams to onboard AI features, provide integration guidance, and troubleshoot platform issues.
- Contribute to architecture decisions, document technical tradeoffs, and constructively challenge approaches when appropriate.
- Evaluate third-party technologies and vendors through structured comparisons, making cost, quality, security, and reliability tradeoffs explicit.
- Strengthen AI security across platform services, including PII protection, tenant isolation, prompt-injection defenses, and auditability.
- Identify opportunities to improve scalability, reliability, developer experience, and operational efficiency across the platform.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
- At least 5 years of professional software engineering experience with a strong production track record.
- Approximately 6+ years of experience building production distributed systems, ideally including internal developer platforms, infrastructure platforms, or API gateways at scale.
- At least 5 years of experience in MLOps, LLMOps, ML platform engineering, or a comparable combination of DevOps and machine-learning engineering at production scale.
- At least 2 years of hands-on production experience with LLM-based systems, including areas such as prompt engineering, RAG, evaluation, or LLM infrastructure.
- Strong Python development skills plus proficiency in Go or Java, with experience writing production-quality code rather than primarily notebooks or scripts.
- Strong understanding of asynchronous programming, backpressure, rate limiting, distributed systems, and production service design.
- Practical experience with LLM observability and evaluation platforms such as LangSmith, Langfuse, Braintrust, Arize, or comparable tools.
- Working knowledge of LLM evaluation methodologies, including benchmark design, LLM-as-judge approaches, regression testing, and human review workflows.
- Familiarity with modern LLM technologies, including OpenAI or Anthropic APIs, an orchestration framework such as LangChain or LlamaIndex, vector databases, and LLM observability tooling.
- Experience designing and operating multi-tenant systems with strong data and service isolation guarantees.
- Strong cloud-native experience with AWS or Azure, including Kubernetes, service mesh technologies, infrastructure-as-code, Terraform, and CI/CD.
- Experience deploying model updates safely using techniques such as canary releases, shadow evaluation, feature controls, and rollback triggers.
- Understanding of the full ML lifecycle, including training pipelines, model serving, monitoring, evaluation, and cost management.
- Strong understanding of AI security fundamentals, particularly PII handling, tenant isolation, prompt-injection risks, and audit logging.
- Strong written and verbal communication skills, with the ability to document technical decisions clearly in an asynchronous, distributed environment.
- Fluent English communication skills.
- Preferred experience in procure-to-pay, ERP integration, accounting.