About the company
SOFTSWISS is hiring a Monitoring Systems Engineer to join our Team. We are looking for a detail-oriented engineer to design, maintain, and enhance our monitoring and observability ecosystem, ensuring the reliability, performance, and visibility of critical services across our technology landscape.
Responsibilities
- Offering on-duty service coverage, encompassing day and night shifts.
- Addressing incidents by troubleshooting and resolving issues, even seeking assistance from third-party or vendor support when necessary.
- Directing issues or queries to the relevant department as needed.
- Keeping detailed records and documentation of current infrastructure challenges and Root Cause Analyses (RCAs).
- Contribute to safe and effective internal practices for AI usage in monitoring and incident response workflows.
- Collaborating with other teams to understand and define their monitoring needs, then implementing the right solutions.
- Setting up and adjusting the monitoring/observability systems for various teams.
- Designing and tweaking alerts and dashboards to suit specific needs.
- Refining alerts to reduce irrelevant notifications and increase their significance.
- Enhancing dashboards for better clarity, understanding, and a more comprehensive view.
- Building and sustaining connections between the monitoring systems and other platforms like Jira, Opsgenie, etc. when required.
- Establishing and updating a Knowledge Base, covering system configurations, alert processes, troubleshooting guidelines, and user manuals.
- Staying updated with the newest trends and best practices to continuously uplift our organization's monitoring capabilities.
- Identify opportunities to automate repetitive monitoring and support tasks, including with AI-assisted approaches where suitable.
Requirements
- Minimum of 3 years experience as a Systems Engineer, SRE, DevOps, or Monitoring Support Engineer (L2+).
- Good understanding of Linux-like operating systems (Debian-based).
- Experience with containerization, virtualization, and orchestration (LXC/LXD, Docker, Kubernetes).
- Development experience in any scripting language (Bash, Python, Go, etc) and familiarity with REST API.
- Knowledge of basic database concepts (experience with PostgreSQL is preferable), including transactions and WAL.
- English proficiency at an Intermediate (B1) level or higher. It's crucial to understand technical terminology related to our specific tech stack and to be able to interpret technical documentation.
- Russian proficiency at an Upper-Intermediate (B2) level or higher.
- Practical interest in using AI-assisted tools for troubleshooting, automation, documentation, and operational efficiency.
- Ability to critically evaluate AI-generated output and validate it before using it in production environments.
- Understanding of the risks and limitations of AI usage in infrastructure and production operations.
Conditions
- Private health insurance
- Sports benefits
- Comprehensive Mental Health Program
- Free English lessons (online)
- Local language courses
- Paid time off
- Maternity leave support
- Referral program rewards
- Upskilling, internal workshops, and participation in professional conferences and corporate events