← Все вакансии/Senior/jobgether
SeniorRemoteUS

AI Optimization Engineer

J
jobgether
Зарплата
$8,300
Уровень
Senior
Формат
Remote
О роли

Описание вакансии

About the company

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Optimization Engineer based in United States.

Responsibilities
  • Optimize training and inference workloads to maximize throughput, minimize latency, and improve cost efficiency across large-scale neural network systems.
  • Analyze and improve performance across the full technology stack, including GPU kernels, memory management, communication, distributed systems, and model execution.
  • Profile CPU, GPU, and distributed workloads to identify bottlenecks and use quantitative analysis to guide optimization decisions.
  • Design and implement performance improvements using Python, C++, and relevant AI systems technologies.
  • Optimize distributed training and inference architectures, including model parallelism, communication strategies, and resource utilization.
  • Evaluate and implement model compression techniques while carefully considering their impact on model accuracy and production performance.
  • Investigate complex performance and reliability issues through systematic debugging, instrumentation, benchmarking, and root-cause analysis.
  • Contribute to production-scale optimization of large language model inference and other demanding AI workloads.
  • Develop and improve low-level optimization techniques, including custom GPU kernels where appropriate.
  • Collaborate with product, design, engineering, operations, and business stakeholders to translate ambiguous requirements into scalable, well-engineered technical solutions.
  • Participate in architecture and code reviews, establish engineering best practices, and contribute to long-term technical strategy.
  • Mentor junior and mid-level engineers, helping raise technical quality and strengthen performance engineering capabilities.
  • Identify opportunities to improve the cost structure of AI workloads through infrastructure optimization and FinOps-oriented analysis.
Requirements
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical discipline.
  • 6+ years of professional experience in performance engineering, machine learning systems, high-performance computing, or a closely related field.
  • Strong programming proficiency in Python and C++, with the ability to develop production-quality, maintainable code.
  • Hands-on experience optimizing deep learning workloads on modern GPU architectures.
  • Deep understanding of distributed training and inference techniques, including parallelism strategies and communication primitives.
  • Strong knowledge of memory hierarchies, GPU/CPU performance characteristics, and systems-level optimization.
  • Experience using profiling and instrumentation tools across CPU, GPU, and distributed environments.
  • Familiarity with model compression methods and their implications for accuracy, performance, and production deployment.
  • Excellent measurement, debugging, analytical reasoning, and problem-solving abilities.
  • Strong communication and collaboration skills, with the ability to explain complex technical concepts to cross-functional stakeholders.
  • Demonstrated ability to work independently, make data-driven technical decisions, and deliver meaningful improvements in production environments.
  • Experience with production-scale LLM inference is strongly preferred.
  • Contributions to projects such as vLLM, TensorRT-LLM, DeepSpeed, or comparable AI systems projects are a plus.
  • Experience with custom kernel development using technologies such as Triton or CUTLASS is preferred.
  • Familiarity with FinOps and cost optimization for AI workloads is advantageous.
  • Publications, conference presentations, or technical talks focused on AI systems or performance engineering are a plus.
  • Must be currently based in the United States and authorized to work in the U.S.; U.S. citizens, permanent residents, EAD holders, and candidates eligible for H-1B transfer are encouraged to apply. New H-1B sponsorship is not available.
Conditions
  • $100,000 annual salary for this full-time direct W2 position.
  • 100% remote work within the United States.
  • Opportunity to work on challenging AI optimization and high-performance computing problems.
  • Exposure to large-scale neural networks, modern GPU architectures, distributed systems, and production AI infrastructure.
  • Significant opportunities for technical ownership, mentorship, and career growth.
  • Collaborative environment spanning engineering, product, operations, design, and business teams.
  • Opportunity to contribute to impactful production AI systems and advance performance, scalability, and cost efficiency.
  • Equal employment opportunity and an inclusive workplace committed to fair treatment of employees and applicants.
Стек и навыки

С чем работаем