← Все вакансии/Senior/jobgether
SeniorFlexibleIreland

Staff Network Engineer

J
jobgether
Уровень
Senior
Формат
Flexible
О роли

Описание вакансии

About the company

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Network Engineer (AI Fabric, Datacenter and Edge Networking) based in Ireland.

Responsibilities
  • Own the long-term technical direction, architecture, and operational strategy for high-performance AI networking and broader infrastructure networking.
  • Design, deploy, and operate GPU networking fabrics optimized for distributed AI training and inference workloads.
  • Architect large-scale RoCE and Ethernet fabrics, including leaf-spine, fat-tree, rail, and multi-plane topologies.
  • Optimize east-west networking for GPU communication patterns and high-throughput distributed workloads, working closely with compute and platform teams.
  • Implement and operate technologies including RDMA, RoCE, high-bandwidth Ethernet, Spectrum-X, and related AI networking platforms.
  • Define reusable reference architectures, design principles, deployment standards, validation criteria, and operational patterns for future AI fabric deployments.
  • Evaluate architectural trade-offs across performance, resilience, scalability, cost, operability, and deployment speed, providing clear recommendations to stakeholders.
  • Design and operate Layer 2 and Layer 3 datacenter networks using technologies such as BGP, ECMP, EVPN/VXLAN, VLAN, VRF, OVS/OVN, and Linux networking.
  • Build scalable routing, segmentation, overlay, tenant-isolation, and north-south traffic-management architectures.
  • Design and maintain edge security and connectivity infrastructure, including WAF, TLS termination, DDoS mitigation, API gateways, and proxy protections.
  • Design and operate private inter-datacenter connectivity, including dark fiber, metro fiber rings, DWDM transport, high-capacity WAN links, and redundant backbone architectures.
  • Develop scalable principles for inter-site routing, redundancy, failure-domain isolation, and backbone evolution.
  • Lead end-to-end delivery of networking initiatives, from architecture and lab validation through production deployment and operational handover.
  • Collaborate with procurement and deployment teams on network bills of materials, datacenter layouts, rack elevations, capacity planning, performance modeling, and scaling strategies.
  • Establish safe network-change practices, including deployment validation, rollback procedures, acceptance criteria, and minimal-impact production execution.
  • Drive automation for provisioning, configuration management, monitoring, validation, and network lifecycle management using software engineering practices.
  • Own network reliability and operational performance, defining and tracking SLAs, SLOs, latency, recovery, and other key reliability metrics.
  • Serve as the senior escalation point for complex networking incidents, leading deep investigations, root-cause analysis, remediation, and long-term systemic improvements.
  • Build network observability and operational tooling that improves visibility, reliability, and day-two operations.
  • Act as the primary networking design authority, influencing platform architecture and helping adjacent teams understand how networking capabilities and constraints affect their decisions.
  • Mentor engineers across adjacent domains and help establish the standards, practices, and technical foundations for the future networking organization.
Requirements
  • Extensive hands-on experience designing, deploying, and operating large-scale datacenter networks in production environments.
  • Expert knowledge of modern networking protocols and architectures, including BGP, OSPF, ECMP, and EVPN/VXLAN.
  • Proven experience operating high-speed Ethernet networks and networking hardware in production.
  • Strong experience with NVIDIA/Mellanox networking platforms and high-performance interconnect technologies.
  • Deep expertise designing and operating networking fabrics for large-scale GPU clusters and distributed AI workloads.
  • Strong understanding of GPU communication patterns and how distributed training and inference requirements influence network architecture and performance.
  • Practical experience with NCCL communication patterns, including all-reduce, all-gather, broadcast, and reduce-scatter, and their impact on network traffic.
  • Hands-on experience designing and tuning RoCE/RDMA fabrics for GPU clusters.
  • Strong understanding of RDMA transport behavior, failure modes, congestion, and performance characteristics.
  • Practical experience implementing and tuning congestion-management technologies such as PFC and ECN.
  • Experience designing rail-optimized GPU networking fabrics and diagnosing issues such as NCCL stalls, RDMA congestion, fabric hotspots, and packet loss affecting distributed workloads.
  • Understanding of how networking performance affects distributed AI frameworks such as
Стек и навыки

С чем работаем