← Все вакансии/Senior/jobgether
SeniorRemoteIndia

AI Platform Engineer

J
jobgether
Уровень
Senior
Формат
Remote
О роли

Описание вакансии

About the company

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Platform Engineer based in India.

Responsibilities
  • Design, build, deploy, and operate scalable AI platform services with clear API contracts, versioning, service-level objectives, and production-readiness standards.
  • Develop production-grade Python services, shared libraries, SDKs, and integration patterns for consumption by multiple product engineering teams.
  • Build and evolve a multi-provider model gateway with routing, fallback, retry, rate limiting, cost attribution, and budget enforcement capabilities.
  • Develop RAG-as-a-service capabilities covering data ingestion, chunking, embeddings, retrieval, hybrid search, and retrieval-quality measurement.
  • Establish prompt management capabilities including templates, versioning, evaluation, rollout controls, and tenant-specific customization.
  • Build guardrails and content-safety mechanisms such as input filtering, output validation, PII redaction, and controlled tool execution.
  • Develop agent and tool-use patterns suitable for reliable production workflows.
  • Build observability capabilities covering prompt and response tracing, latency, errors, cost per request, quality signals, drift detection, and audit logging.
  • Design and maintain evaluation infrastructure using benchmark datasets, offline evaluations, LLM-as-judge approaches, human review workflows, and regression testing.
  • Build and operate model deployment pipelines, including safe deployment and rollback processes for fine-tuned models where applicable.
  • Establish alerting and SLO frameworks for AI services, treating quality regressions as a first-class operational concern.
  • Participate in on-call rotations, create incident response runbooks, and continuously improve operational procedures following incidents.
  • Collaborate directly with product engineering teams to onboard AI features, provide integration guidance, and troubleshoot platform issues.
  • Contribute to architecture decisions, document technical tradeoffs, and constructively challenge approaches when appropriate.
  • Evaluate third-party technologies and vendors through structured comparisons, making cost, quality, security, and reliability tradeoffs explicit.
  • Strengthen AI security across platform services, including PII protection, tenant isolation, prompt-injection defenses, and auditability.
  • Identify opportunities to improve scalability, reliability, developer experience, and operational efficiency across the platform.
Requirements
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
  • At least 5 years of professional software engineering experience with a strong production track record.
  • Approximately 6+ years of experience building production distributed systems, ideally including internal developer platforms, infrastructure platforms, or API gateways at scale.
  • At least 5 years of experience in MLOps, LLMOps, ML platform engineering, or a comparable combination of DevOps and machine-learning engineering at production scale.
  • At least 2 years of hands-on production experience with LLM-based systems, including areas such as prompt engineering, RAG, evaluation, or LLM infrastructure.
  • Strong Python development skills plus proficiency in Go or Java, with experience writing production-quality code rather than primarily notebooks or scripts.
  • Strong understanding of asynchronous programming, backpressure, rate limiting, distributed systems, and production service design.
  • Practical experience with LLM observability and evaluation platforms such as LangSmith, Langfuse, Braintrust, Arize, or comparable tools.
  • Working knowledge of LLM evaluation methodologies, including benchmark design, LLM-as-judge approaches, regression testing, and human review workflows.
  • Familiarity with modern LLM technologies, including OpenAI or Anthropic APIs, an orchestration framework such as LangChain or LlamaIndex, vector databases, and LLM observability tooling.
  • Experience designing and operating multi-tenant systems with strong data and service isolation guarantees.
  • Strong cloud-native experience with AWS or Azure, including Kubernetes, service mesh technologies, infrastructure-as-code, Terraform, and CI/CD.
  • Experience deploying model updates safely using techniques such as canary releases, shadow evaluation, feature controls, and rollback triggers.
  • Understanding of the full ML lifecycle, including training pipelines, model serving, monitoring, evaluation, and cost management.
  • Strong understanding of AI security fundamentals, particularly PII handling, tenant isolation, prompt-injection risks, and audit logging.
  • Strong written and verbal communication skills, with the ability to document technical decisions clearly in an asynchronous, distributed environment.
  • Fluent English communication skills.
  • Preferred experience in procure-to-pay, ERP integration, accounting.
Стек и навыки

С чем работаем