← Все вакансии/Lead/jobgether
LeadRemoteNetherlands

Staff Storage Platform Engineer

J
jobgether
Уровень
Lead
Формат
Remote
О роли

Описание вакансии

About the company

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Storage Platform Engineer (AI Storage) - Radian Arc based in Netherlands.

This is a Staff-level opportunity to shape the storage architecture powering large-scale GPU and AI infrastructure across edge and core environments.

Responsibilities
  • Storage architecture: Design scalable storage architectures for edge and core GPU deployments, covering hyperconverged platforms such as StorPool, local NVMe, and disaggregated systems such as VAST Data and Weka. Define reference architectures, reusable design patterns, fault domains, lifecycle strategies, and scaling approaches while balancing throughput, latency, resilience, data locality, operability, and cost.
  • AI workload optimization: Optimize storage for distributed training, fine-tuning, and inference workloads, including large dataset ingestion, model artifact distribution, checkpointing, and high-concurrency access. Establish realistic performance baselines and ensure storage architecture aligns with actual GPU workload behavior.
  • Distributed inference: Design storage architectures supporting inference platforms such as NVIDIA Dynamo, llm-d, or similar systems. Optimize model distribution, token-generation data paths, and KV-cache persistence and retrieval so infrastructure can scale efficiently across large GPU clusters without storage becoming a throughput or latency bottleneck.
  • High-performance data paths: Engineer efficient storage-to-GPU data paths using technologies such as GPU Direct Storage, RDMA/RoCE, NVMe-oF, and SPDK. Investigate and tune performance across hardware, networking, operating systems, filesystems, storage layers, and distributed workloads.
  • Platform integration: Integrate block, object, and shared file storage into Kubernetes and platform orchestration systems. Implement and maintain CSI integrations, support multi-tenant storage architectures, and define standards for storage integration across different deployment models.
  • Distributed storage: Contribute to large-scale storage platforms, including S3-compatible object storage, distributed file systems, and block storage. Design systems with clear operational boundaries, resilience models, scaling paths, and reusable operating patterns across multi-cluster and multi-site environments.
  • Performance and reliability: Lead storage benchmarking, capacity planning, performance investigations, incident response, and root-cause analysis. Establish measurable standards for throughput, latency consistency, recovery behavior, reliability, and operational maturity.
  • Engineering delivery: Own storage initiatives end to end, from architecture and validation through production rollout. Validate BOMs, topology decisions, node profiles, and deployment assumptions while ensuring changes are introduced safely with minimal customer impact.
  • Operational excellence: Improve storage observability, automation, runbooks, lifecycle management, and day-2 operations. Turn recurring incidents and operational pain points into durable engineering improvements and standardized practices.
  • Technical leadership: Act as the primary storage design authority, influencing platform architecture and roadmap decisions across compute, networking, DevOps, infrastructure, and operations. Communicate technical trade-offs clearly, mentor adjacent engineers, and raise the organization’s expertise in AI storage.
Requirements
  • Distributed storage expertise: Strong hands-on experience designing and operating distributed storage systems in high-performance computing, AI, GPU, or similarly demanding environments.
  • AI infrastructure experience: Proven experience designing storage architectures for large-scale AI training, fine-tuning, or inference, including dataset distribution, model artifacts, checkpointing, and high-concurrency data access.
  • AI storage knowledge: Deep understanding of how AI workload characteristics affect storage throughput, latency, concurrency, data locality, checkpoint recovery, and serving performance.
  • Storage technologies: Hands-on experience with technologies such as Weka, VAST Data, StorPool, local NVMe, distributed filesystems, S3-compatible object storage, block storage, and/or comparable enterprise storage platforms.
  • Linux and systems expertise: Strong knowledge of the Linux storage and I/O stack, storage hardware, NVMe devices, storage fabrics, and high-performance data paths.
  • Kubernetes: Familiarity with Kubernetes storage integrations, particularly CSI, and experience integrating storage into containerized or orchestrated platforms.
  • AI data paths: Practical knowledge of GPU Direct Storage, RDMA/RoCE, NVMe-oF, SPDK, and techniques for minimizing unnecessary data movement between storage and GPU compute.
  • Distributed inference: Experience with storage requirements for inference orchestration and model-serving environments.
Стек и навыки

С чем работаем