About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Storage Platform Engineer (AI Storage) - Radian Arc based in United Kingdom.
This is a Staff-level opportunity to shape the storage architecture powering large-scale GPU and AI infrastructure across edge and core environments.
Responsibilities
- Design scalable storage architectures for edge and core GPU deployments, covering hyperconverged platforms such as StorPool, local NVMe, and disaggregated systems such as VAST Data and Weka.
- Optimize storage for distributed training, fine-tuning, and inference workloads, including large dataset ingestion, model artifact distribution, checkpointing, and high-concurrency access.
- Design storage architectures supporting inference platforms such as NVIDIA Dynamo, llm-d, or similar systems.
- Engineer efficient storage-to-GPU data paths using technologies such as GPU Direct Storage, RDMA/RoCE, NVMe-oF, and SPDK.
- Integrate block, object, and shared file storage into Kubernetes and platform orchestration systems.
- Contribute to large-scale storage platforms, including S3-compatible object storage, distributed file systems, and block storage.
- Lead storage benchmarking, capacity planning, performance investigations, incident response, and root-cause analysis.
- Own storage initiatives end to end, from architecture and validation through production rollout.
- Improve storage observability, automation, runbooks, lifecycle management, and day-2 operations.
- Act as the primary storage design authority, influencing platform architecture and roadmap decisions.
Requirements
- Strong hands-on experience designing and operating distributed storage systems in high-performance computing, AI, GPU, or similarly demanding environments.
- Proven experience designing storage architectures for large-scale AI training, fine-tuning, or inference.
- Deep understanding of how AI workload characteristics affect storage throughput, latency, concurrency, data locality, checkpoint recovery, and serving performance.
- Hands-on experience with technologies such as Weka, VAST Data, StorPool, local NVMe, distributed filesystems, S3-compatible object storage, block storage, and/or comparable enterprise storage platforms.
- Strong knowledge of the Linux storage and I/O stack, storage hardware, NVMe devices, storage fabrics, and high-performance data paths.
- Familiarity with Kubernetes storage integrations, particularly CSI.
- Practical knowledge of GPU Direct Storage, RDMA/RoCE, NVMe-oF, SPDK, and techniques for minimizing unnecessary data movement between storage and GPU compute.
- Experience with storage requirements for inference orchestration and model-serving environments.