About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Optimization Engineer based in United States.
Responsibilities
- Optimize training and inference workloads to maximize throughput, minimize latency, and improve cost efficiency across large-scale neural network systems.
- Analyze and improve performance across the full technology stack, including GPU kernels, memory management, communication, distributed systems, and model execution.
- Profile CPU, GPU, and distributed workloads to identify bottlenecks and use quantitative analysis to guide optimization decisions.
- Design and implement performance improvements using Python, C++, and relevant AI systems technologies.
- Optimize distributed training and inference architectures, including model parallelism, communication strategies, and resource utilization.
- Evaluate and implement model compression techniques while carefully considering their impact on model accuracy and production performance.
- Investigate complex performance and reliability issues through systematic debugging, instrumentation, benchmarking, and root-cause analysis.
- Contribute to production-scale optimization of large language model inference and other demanding AI workloads.
- Develop and improve low-level optimization techniques, including custom GPU kernels where appropriate.
- Collaborate with product, design, engineering, operations, and business stakeholders to translate ambiguous requirements into scalable, well-engineered technical solutions.
- Participate in architecture and code reviews, establish engineering best practices, and contribute to long-term technical strategy.
- Mentor junior and mid-level engineers, helping raise technical quality and strengthen performance engineering capabilities.
- Identify opportunities to improve the cost structure of AI workloads through infrastructure optimization and FinOps-oriented analysis.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical discipline.
- 6+ years of professional experience in performance engineering, machine learning systems, high-performance computing, or a closely related field.
- Strong programming proficiency in Python and C++, with the ability to develop production-quality, maintainable code.
- Hands-on experience optimizing deep learning workloads on modern GPU architectures.
- Deep understanding of distributed training and inference techniques, including parallelism strategies and communication primitives.
- Strong knowledge of memory hierarchies, GPU/CPU performance characteristics, and systems-level optimization.
- Experience using profiling and instrumentation tools across CPU, GPU, and distributed environments.
- Familiarity with model compression methods and their implications for accuracy, performance, and production deployment.
- Excellent measurement, debugging, analytical reasoning, and problem-solving abilities.
- Strong communication and collaboration skills, with the ability to explain complex technical concepts to cross-functional stakeholders.
- Demonstrated ability to work independently, make data-driven technical decisions, and deliver meaningful improvements in production environments.
- Experience with production-scale LLM inference is strongly preferred.
- Contributions to projects such as vLLM, TensorRT-LLM, DeepSpeed, or comparable AI systems projects are a plus.
- Experience with custom kernel development using technologies such as Triton or CUTLASS is preferred.
- Familiarity with FinOps and cost optimization for AI workloads is advantageous.
- Publications, conference presentations, or technical talks focused on AI systems or performance engineering are a plus.
- Must be currently based in the United States and authorized to work in the U.S.; U.S. citizens, permanent residents, EAD holders, and candidates eligible for H-1B transfer are encouraged to apply. New H-1B sponsorship is not available.
Conditions
- $100,000 annual salary for this full-time direct W2 position.
- 100% remote work within the United States.
- Opportunity to work on challenging AI optimization and high-performance computing problems.
- Exposure to large-scale neural networks, modern GPU architectures, distributed systems, and production AI infrastructure.
- Significant opportunities for technical ownership, mentorship, and career growth.
- Collaborative environment spanning engineering, product, operations, design, and business teams.
- Opportunity to contribute to impactful production AI systems and advance performance, scalability, and cost efficiency.
- Equal employment opportunity and an inclusive workplace committed to fair treatment of employees and applicants.