Faster chat, better deals — Get the App

Senior DevOps Stability Engineer

MyJob

Company

Job typeFull-time
Workplace typeOnsite
Experience levelNo experience limit
Education levelNo degree limit

Description

Job Summary: Ensure operational continuity, availability, and resilience of the DevOps toolchain and its supporting infrastructure by implementing defined practices to prevent disruptions and degradation in CI/CD pipelines. Key Responsibilities: 1. Operate and maintain the DevOps toolchain (SaaS and PaaS) 2. Participate in incident resolution and high-availability strategies 3. Progressively develop technical autonomy for complex tasks Maintain operational continuity, availability, and resilience of the DevOps toolchain and its supporting infrastructure by executing the team-defined practices, procedures, and automations to prevent disruptions or degradations in integration, testing, and deployment pipelines. Perform monitoring, maintenance, backup, and incident response tasks per guidelines established by Senior roles, participate in incident resolution and implementation of high-availability and recovery strategies, while progressively developing technical autonomy enabling independent handling of increasingly complex tasks. In this role, you will have the opportunity to: - Operate and maintain the DevOps toolchain (SaaS and PaaS) according to team-defined procedures and standards. - Perform monitoring, alerting, and performance verification of infrastructure supporting CI/CD pipelines. - Serve as first-line responder for Toolchain incidents, applying defined runbooks and escalating critical or complex cases to Senior roles. - Support execution of toolchain updates, patches, and migrations under existing plans. - Execute and verify backups and participate in periodic disaster recovery (DR) drills. - Support root cause analysis (RCA) of recurring incidents by providing data and evidence. - Keep operational documentation (runbooks) up to date and record changes and findings. - Collaborate with peer teams within the DevOps Submanagement on operations and support tasks. - Develop and maintain monitoring or repetitive-task automation scripts under defined design specifications. - Report operational status, incidents, and metrics to Senior roles and area management. To succeed in this position, you need: Education: University degree in Computer Science, Industrial Engineering, or equivalent, with a Bachelor's degree in Execution Engineering or Civil Engineering. **Experience:** \+2 years in infrastructure administration, DevOps, platform support, or Systems Engineering roles. Cloud: Azure knowledge (Azure DevOps, infrastructure); AWS or GCP experience is desirable. Operating Systems / Infrastructure: Experience with Linux (RHEL) and/or Windows Server; familiarity with on-premise and hybrid infrastructure administration. Toolchain Tools: Experience using or administering Azure DevOps, GitHub, Jenkins, or Nexus—or equivalent tools (GitLab, TeamCity, Artifactory). (At least one required) Scripting / Automation: Proficiency in PowerShell, Python, or Bash; familiarity with Infrastructure-as-Code (Terraform, Ansible) is desirable. It’s even better if you have: Languages: Basic or intermediate written English (technical documentation reading). Monitoring and Observability: Experience with tools such as Grafana, Prometheus, Dynatrace, Zabbix, or similar. 1\. Infrastructure and Platform Certifications (Fundamentals / Associate Level) Ideally hold or be pursuing at least one of the following: ITIL Foundation – fundamentals of incident, problem, and change management. Microsoft Certified: Azure Fundamentals (AZ\-900\) – Azure cloud fundamentals. Microsoft Certified: Azure Administrator Associate (AZ\-104\) – Azure infrastructure administration (desirable). SRE Foundation – fundamentals of Site Reliability Engineering. DevOps Certification. 2\. Technical / Operational Certifications (Practical Level) Jenkins Certified Engineer – Jenkins administration. HashiCorp Certified: Terraform Associate – Infrastructure-as-Code. Linux Professional Institute (LPIC\-1\) – Linux system administration. 3\. Valued Courses and Training It is beneficial if the candidate has recently completed courses in: Incident Management (ITIL): incident response and logging processes. High Availability and Backup Fundamentals: continuity and recovery concepts. Monitoring and Observability: use of monitoring and alerting tools. Automation and IaC: familiarity with Terraform, Ansible, or task scripting.

Some content was automatically translated

Posted by

Sofía Muñoz

MyJob · HR

Location

Sofía Muñoz

MyJob · HR

Similar jobs

Senior DevOps Stability Engineer job by MyJob in 2026 | ok.com