← Все вакансии/Senior/jobgether
SeniorOfficeIreland

Technical Operations & Deployment Engineer

J
jobgether
Уровень
Senior
Формат
Office
О роли

Описание вакансии

About the company

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure) based in Ireland.

This is a highly hands-on infrastructure role focused on deploying, commissioning, and operating GPU cloud environments across regional and core datacenters.

Responsibilities
  • Coordinate deployments with datacenter providers, integrators, logistics teams, vendors, and internal engineering.
  • Support rack-and-stack activities for GPU and CPU servers, storage, switches, routers, firewalls, PDUs, serial/OOB systems, and supporting infrastructure.
  • Validate fiber and copper cabling, optics, transceivers, breakout cables, port mappings, link speeds, redundancy, and management, storage, north-south, and east-west connectivity.
  • Commission GPU servers, storage nodes, and platform infrastructure while validating BIOS, BMC, firmware, NICs, DPUs, GPUs, NVMe, RAID/HBA, PCIe topology, NUMA, thermals, power, and hardware health.
  • Execute burn-in, stress, network, storage, and acceptance testing before production handover.
  • Work with network engineering to validate switch configurations, routing, VLAN/VRF segmentation, BGP, ECMP, EVPN/VXLAN, OVS/OVN, VyOS, firewalls, WAF infrastructure, and customer connectivity.
  • Support validation of RoCE/RDMA fabrics for distributed AI workloads and troubleshoot issues such as link flaps, MTU mismatches, route errors, packet loss, PFC/ECN problems, and congestion.
  • Install and validate Ubuntu/Linux environments, NVIDIA drivers, CUDA, OFED or inbox drivers, Docker/containerd, KVM/QEMU, platform agents, and GPU infrastructure components.
  • Support CloudStack, Kubernetes, KubeVirt, GPU Operator, CSI/CNI integrations, GPU passthrough, SR-IOV, BlueField DPUs, VM networking, and container networking.
  • Support integration and validation of StorPool, Weka, local NVMe, and other supported storage platforms.
  • Execute acceptance testing, produce deployment readiness reports, maintain runbooks, and ensure infrastructure is fully operational before customer or production handover.
  • Perform controlled firmware, OS, driver, BIOS, switch, and hardware maintenance while supporting production incidents and infrastructure escalations.
  • Investigate operational failures, perform root-cause analysis, distinguish temporary workarounds from permanent fixes, and work with engineering to eliminate recurring issues.
  • Validate telemetry and monitoring across hosts, GPUs, DPUs, switches, storage, and platform components using tools such as Zabbix, Prometheus, Grafana, Loki, DCGM/NVML, and NVIDIA NetQ or equivalents.
  • Establish baselines for GPU, network, storage, and host performance and support benchmarking and infrastructure validation.
  • Maintain accurate as-built records covering rack elevations, cable maps, port mappings, serial numbers, asset records, IP allocations, changes, and operational procedures.
  • Partner with infrastructure, networking, storage, platform, fleet automation, observability, product engineering, sales engineering, and service delivery teams.
  • Coordinate with datacenter providers, system integrators, server and storage vendors, NVIDIA, and networking suppliers to resolve deployment and infrastructure issues.
  • Feed field experience back into reference architectures, BOMs, rack designs, cabling standards, deployment playbooks, validation processes, and automation.
Requirements
  • Strong hands-on experience deploying and maintaining datacenter infrastructure, ideally within GPU, HPC, AI cloud, private cloud, or high-density compute environments.
  • Proven ability to bring servers from physical installation and bare metal through validation and production readiness.
  • Experience with NVIDIA GPU servers, drivers, firmware, PCIe topology, hardware validation, and high-performance compute environments.
  • Familiarity with NVL72-style rack-scale architectures, NVLink/NVSwitch domains, in-rack networking, high-density power delivery, and OEM/NVIDIA validation requirements.
  • Ability to assess power density, cooling, rack dimensions, floor loading, containment, serviceability, maintenance, and safety.
Стек и навыки

С чем работаем