← Все вакансии/Senior/Soniox
SeniorRemote

Software Engineer

S
Soniox
Уровень
Senior
Формат
Remote
О роли

Описание вакансии

About the company

  • Soniox is pushing the boundaries of real-time AI, and we're looking for an engineer to help us run large language models with exceptional speed, efficiency, and reliability at production scale.

Responsibilities

  • Build and optimize our vLLM-based inference stack for low latency, high throughput, and maximum GPU utilization.
  • Optimize continuous batching, scheduling, prefill/decode, prefix caching, and KV-cache allocation and reuse.
  • Optimize or implement CUDA and Triton kernels for attention, GEMMs, sampling, normalization, and other critical model operations.
  • Evaluate and integrate technologies such as FlashAttention, FlashInfer, CUDA Graphs, torch.compile, speculative decoding, and quantization.
  • Optimize distributed inference using tensor, data, and expert parallelism, NCCL, NVLink/NVSwitch, and InfiniBand.
  • Work closely with researchers to bring new dense and MoE model architectures into production quickly and efficiently.

Requirements

  • Deep hands-on experience with LLM inference, ideally working inside vLLM, SGLang, TensorRT-LLM, or similar systems, not just deploying them.
  • Understand TTFT, inter-token latency, throughput, continuous batching, PagedAttention, KV caching, and prefill vs. decode performance.
  • Comfortable profiling GPUs and reasoning about compute, memory bandwidth, kernel launches, synchronization, and communication bottlenecks.
  • Experience with CUDA, Triton, PyTorch, NCCL, and modern NVIDIA GPU architectures.
  • Understand Transformer internals including MHA/GQA, RoPE, KV cache, quantization, and MoE.
  • Experience optimizing distributed, performance-critical systems in production.
  • Care deeply about performance, simplicity, and reliability, and take ownership from profiling through production deployment.

Conditions

  • You'll help build one of the most technically advanced voice AI platforms in the world, and push LLM inference performance at every layer of the stack.
  • You'll work directly with a world-class team of engineers and researchers on hard, measurable problems spanning models, GPU kernels, distributed systems, and production infrastructure.
  • You'll have a voice in how our technology evolves, how our company grows, and how AI transforms human communication.
Стек и навыки

С чем работаем