About the company
This position is listed on behalf of a partner company, who manages all applications and next steps.
Responsibilities
- Own the long-term technical direction, architecture, and operational strategy for high-performance AI networking and broader infrastructure networking.
- Design, deploy, and operate GPU networking fabrics optimized for distributed AI training and inference workloads.
- Architect large-scale RoCE and Ethernet fabrics, including leaf-spine, fat-tree, rail, and multi-plane topologies.
- Optimize east-west networking for GPU communication patterns and high-throughput distributed workloads, working closely with compute and platform teams.
- Implement and operate technologies including RDMA, RoCE, high-bandwidth Ethernet, Spectrum-X, and related AI networking platforms.
- Define reusable reference architectures, design principles, deployment standards, validation criteria, and operational patterns for future AI fabric deployments.
- Evaluate architectural trade-offs across performance, resilience, scalability, cost, operability, and deployment speed, providing clear recommendations to stakeholders.
- Design and operate Layer 2 and Layer 3 datacenter networks using technologies such as BGP, ECMP, EVPN/VXLAN, VLAN, VRF, OVS/OVN, and Linux networking.
- Build scalable routing, segmentation, overlay, tenant-isolation, and north-south traffic-management architectures.
- Design and maintain edge security and connectivity infrastructure, including WAF, TLS termination, DDoS mitigation, API gateways, and proxy protections.
- Design and operate private inter-datacenter connectivity, including dark fiber, metro fiber rings, DWDM transport, high-capacity WAN links, and redundant backbone architectures.
- Develop scalable principles for inter-site routing, redundancy, failure-domain isolation, and backbone evolution.
- Lead end-to-end delivery of networking initiatives, from architecture and lab validation through production deployment and operational handover.
- Collaborate with procurement and deployment teams on network bills of materials, datacenter layouts, rack elevations, capacity planning, performance modeling, and scaling strategies.
- Establish safe network-change practices, including deployment validation, rollback procedures, acceptance criteria, and minimal-impact production execution.
- Drive automation for provisioning, configuration management, monitoring, validation, and network lifecycle management using software engineering practices.
- Own network reliability and operational performance, defining and tracking SLAs, SLOs, latency, recovery, and other key reliability metrics.
- Serve as the senior escalation point for complex networking incidents, leading deep investigations, root-cause analysis, remediation, and long-term systemic improvements.
- Build network observability and operational tooling that improves visibility, reliability, and day-two operations.
- Act as the primary networking design authority, influencing platform architecture and helping adjacent teams understand how networking capabilities and constraints affect their decisions.
- Mentor engineers across adjacent domains and help establish the standards, practices, and technical foundations for the future networking organization.
Requirements
- Extensive hands-on experience designing, deploying, and operating large-scale datacenter networks in production environments.
- Expert knowledge of modern networking protocols and architectures, including BGP, OSPF, ECMP, and EVPN/VXLAN.
- Proven experience operating high-speed Ethernet networks and networking hardware in production.
- Strong experience with NVIDIA/Mellanox networking platforms and high-performance interconnect technologies.
- Deep expertise designing and operating networking fabrics for large-scale GPU clusters and distributed AI workloads.
- Strong understanding of GPU communication patterns and how distributed training and inference requirements influence network architecture and performance.
- Practical experience with NCCL communication patterns, including all-reduce, all-gather, broadcast, and reduce-scatter, and their impact on network traffic.
- Hands-on experience designing and tuning RoCE/RDMA fabrics for GPU clusters.
- Strong understanding of RDMA transport behavior, failure modes, congestion, and performance characteristics.
- Practical experience implementing and tuning congestion-management technologies such as PFC and ECN.
- Experience designing rail-optimized GPU networking fabrics and diagnosing issues such as NCCL stalls, RDMA congestion, fabric hotspots, and packet loss affecting distributed workloads.
- Understanding of how networking performance affects distributed AI frameworks such as PyTorch.