About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior AI Engineer - Observability based in United States.
This is a hands-on senior engineering opportunity focused on making AI-powered products reliable, measurable, and production-ready.
Responsibilities
- Partner with product engineering teams to design, build, and enhance AI-powered features using LLMs, RAG, semantic search, and agentic workflows.
- Contribute directly to production codebases using Python, C#/.NET, or related technologies, while improving prompts, retrieval, context assembly, ranking, grounding, and response-generation patterns.
- Evaluate architecture and technology trade-offs across AI quality, latency, cost, privacy, maintainability, and operational reliability.
- Design and implement AI evaluation pipelines, scoring methodologies, regression frameworks, and quality gates that establish whether features are ready for production.
- Build and maintain versioned golden datasets covering real-world use cases, edge cases, failure modes, and customer-critical workflows.
- Apply LLM-as-judge, heuristic, human-feedback, and task-specific evaluation approaches, while assessing whether evaluation methods accurately measure meaningful outcomes.
- Instrument LLM interactions, RAG pipelines, tool calls, and agent workflows using observability platforms and establish dashboards, alerts, and production monitoring.
- Track and analyze latency, token usage, cost, retrieval quality, groundedness, safety signals, failure modes, and user feedback to identify regressions and improvement opportunities.
- Develop reusable libraries, SDKs, templates, documentation, and reference implementations for AI instrumentation, evaluation, tracing, cost reporting, and incident response.
- Coach product teams on AI quality, observability, evaluation design, and reliable release practices, helping establish consistent standards across product areas.
- Support responsible AI delivery by ensuring appropriate handling of sensitive data and alignment with security, privacy, SOC 2, ISO 27001, and data-residency requirements.
- Contribute to model and provider evaluations, migrations, fallback strategies, rollout plans, and ongoing optimization of AI-powered systems.
Requirements
- 5+ years of professional software engineering experience building and supporting production systems.
- Hands-on experience developing or operating LLM-powered features, RAG systems, AI workflows, or comparable AI applications.
- Strong programming capabilities in Python, C#/.NET, or both, with the ability to contribute effectively to production engineering environments.
- Practical understanding of prompts, embeddings, vector search, retrieval quality, orchestration patterns, grounding, and modern LLM application architecture.
- Demonstrated experience designing evaluations, quality metrics, regression frameworks, or benchmarking approaches, with strong judgment around whether an evaluation measures what truly matters.
- Strong observability fundamentals covering tracing, logging, metrics, alerting, production debugging, and analysis of system behavior.
- Experience with CI/CD, Git-based workflows, cloud environments, and production release practices; experience with Azure DevOps is advantageous.
- Familiarity with AI-assisted development tools such as Claude Code, PlayerZero, or comparable technologies.
- Experience with AI observability and LLMOps platforms such as Arize, Langfuse, LangSmith, W&B, Humanloop, or Helicone is preferred.
- Knowledge of OpenTelemetry, Azure Monitor, Application Insights, or similar observability technologies is beneficial.
- Experience with vector search technologies such as Azure AI Search, Pinecone, Qdrant, Weaviate, or pgvector is advantageous.
- Familiarity with frameworks such as Semantic Kernel, LangChain, LlamaIndex, AutoGen, or related AI orchestration technologies is a plus.
- Experience with LLM-as-judge evaluation, RAG evaluation, semantic similarity, hallucination detection, groundedness scoring, human-feedback workflows, or dedicated evaluation frameworks is preferred.
- Experience with A/B testing, online experimentation, product analytics, regulated environments, PII handling, or data-residency requirements is beneficial.
- Strong communication, collaboration, decision-making, and stakeholder management skills, with the ability to translate ambiguous AI behavior into concrete engineering actions.
- A strong product mindset, curiosity about real-world AI behavior, adaptability, accountability, and a genuine commitment to quality and continuous improvement.
Conditions
- Fully remote position within the United States.
- Company-provided equipment, including laptop and required software.
- Comprehensive medical and prescription drug plan options, along with dental and vision coverage.
- Employer contribution to a Health Savings Account (HSA) for eligible employees enrolled in a high-deductible healthcare plan.
- Medical and dependent-care Flexible Spending Accounts.
- Basic life insurance valued at $50,000 or one times an