← Все вакансии/Senior/jobgether
SeniorRemoteUS

Senior Engineer, Platform Tooling

J
jobgether
Уровень
Senior
Формат
Remote
О роли

Описание вакансии

About the company

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Engineer, Platform Tooling based in the United States.

Responsibilities
  • Manage and maintain internal and client-facing monitoring and observability platforms, including software support, threshold adjustments, customizations, and special projects.
  • Troubleshoot monitoring issues involving SNMP, WMI, SYSLOG, APIs, and other monitoring mechanisms.
  • Own assigned incidents and service tickets through resolution, ensuring timely communication and adherence to defined SLAs.
  • Diagnose and resolve configuration, customization, and production issues while coordinating vendor support cases through to resolution.
  • Support enterprise infrastructure monitoring across networking, VMware, storage, Windows, and Unix/Linux environments.
  • Develop, test, and implement new monitoring features and functionality, including validation plans for production deployments.
  • Create and maintain clear technical documentation, runbooks, knowledge-base content, service definitions, and operational procedures.
  • Develop and maintain best-practice policies for supported monitoring and observability products.
  • Use Python, Groovy, Bash, or comparable scripting technologies to automate tasks and enhance platform capabilities.
  • Manage personal and team service queues, prioritize competing requests, and allocate time effectively during periods of high operational demand.
  • Communicate effectively with clients, peers, engineering teams, management, and vendors regarding incidents, changes, and technical issues.
  • Provide emergency on-call support as part of a rotating schedule.
  • Remain accessible during assigned shifts through approved communication channels, including instant messaging, phone, email, and other collaboration tools.
  • Identify opportunities for process improvements and provide constructive feedback to management regarding operational challenges and areas of concern.
  • Support cross-functional engineering teams and develop additional technical expertise across other products and technologies as business needs evolve.
  • Handle and escalate high-impact operational issues to management or third-party vendors when necessary.
Requirements
  • Bachelor’s degree or equivalent professional or military experience.
  • 3–5 years of experience supporting enterprise IT infrastructure in IT Operations, NOC, Managed Services, Infrastructure Operations, Monitoring, or Observability environments.
  • 2+ years of hands-on experience working with LogicMonitor.
  • Experience creating custom LogicModules in LogicMonitor.
  • Solid understanding of networking fundamentals, including TCP, UDP, IP addressing, routing, switching, VLANs, and firewalls.
  • Technical understanding of VMware ESXi from a monitoring and observability perspective.
  • Experience monitoring enterprise networking and storage infrastructure.
  • Working knowledge of Microsoft Windows and Unix/Linux operating systems.
  • Programming or scripting experience with Python, Groovy, Bash, or similar technologies.
  • Experience with monitoring and analytics platforms such as Nagios, NetXMS, Datadog, New Relic, AppDynamics, Prometheus, Grafana, LogicMonitor, Cribl, or comparable tools.
  • Experience with log handling, parsing, enrichment, and related observability workflows.
  • Ability and willingness to continuously develop technical expertise and stay current with emerging technologies and services.
  • Experience working effectively in multi-vendor environments and coordinating complex technical issues across internal and external teams.
  • Strong troubleshooting, analytical, and problem-solving skills with the ability to manage multiple issues simultaneously.
  • Strong interpersonal and relationship-management skills, with the ability to work effectively with technical and management stakeholders at different levels.
  • Excellent written and verbal communication skills, including the ability to explain technical concepts clearly to non-technical audiences.
  • Strong customer-service orientation and client focus.
  • Ability to work independently, take initiative, recognize priorities, and contribute effectively within a collaborative team environment.
  • Strong dependability, attention to detail, adaptability, persistence, integrity, and stress tolerance.
  • Ability to remain calm, professional, and courteous during high-pressure incidents and operational challenges.
  • Cisco or other relevant vendor certifications are preferred.
  • Experience with Cribl Stream and/or Cribl Edge is a plus.
  • Experience handling and escalating high-impact operational incidents is advantageous.
Conditions
  • Remote position available within the Continental United States.
  • Approximately 5% travel associated with the role.
  • Opportunity to work with enterprise monitoring, observability, networking, infrastructure, and managed services technologies.
  • Exposure to a broad range of platforms and vendors in complex technical environments.
  • Opportunities to expand technical expertise across em
Стек и навыки

С чем работаем