Description
Job Summary:
We are seeking a proactive SRE Engineer to implement and monitor products, collaborate with teams, improve processes, and ensure platform stability and availability.
Key Highlights:
1. Proactively identify issues and improvement opportunities.
2. Collaborate in building an SRE culture across the organization.
3. Resolve multi-platform issues and manage production incidents.
**Job Description**
----------------------
Safely implement the product and enable monitoring through observability, while collaborating with the team.
1\. Proactively identify issues, determine areas for improvement and performance bottlenecks across platforms; conduct system analysis, configuration management, and develop enhancements to improve software system performance, availability, and reliability. The goal is to enable continuous operational process improvement throughout the product lifecycle, gain deeper insight into how systems operate, and ensure production stability.
2\. Collaborate in building an SRE culture across the organization by sharing best practices, approaches, documentation, and code with other engineering teams, applying observability and security mechanisms. This aims to foster a growing community grounded in shared knowledge and experience.
3\. Resolve complex multi-platform issues—considering operating systems, networking, and databases—in on-premises, SaaS, and cloud-based IaaS environments; manage live production incidents; debug and resolve application and infrastructure issues; adopt and implement SRE best practices leveraging automation and reducing TOIL (toil). This ensures end-to-end understanding and effective resolution of problems.
4\. Document platform knowledge as it is acquired over time, create runbooks, and ensure critical system information is accessible to those who need it. This enables easy access to foundational product information and ensures clear understanding during crisis situations (incidents).
5\. Design and implement mechanisms to define and enforce service-level objectives (SLIs), service-level objectives (SLOs), and service-level agreements (SLAs). The objective is to maintain platform availability per commitments and achieve realistic, effective metrics.
6\. Serve as the initial point of contact in incident management processes, capable of using and improving the process, and applying incident management expertise such as post-mortem analysis. This ensures rapid and effective responses to issues affecting production systems.
7\. Build tools supporting end-to-end product lifecycle management (software and others), optimizing and prioritizing implementation and use of continuous deployment pipelines (CI/CD), and applying automation and scripting to any manual tasks or platform components identified—enabling early problem detection in code. The goal is to ensure quality and speed in releasing new features while reducing TOIL and risks associated with manual work.
**Candidate Requirements**
--------------------------
* 3–5 years of experience in SRE, platform, or infrastructure-related roles.
* Required: Certifications in relevant technology tools.
* DataDog, Grafana, Prometheus, and/or OpenTelemetry.
* Linux system administration and operations.
* Advanced English proficiency.
This opportunity is open to persons with disabilities.
**Compensation**
----------------
3000000**Selection Process**
------------------------
Join our team and become part of Falabella’s transformation:
1\. Apply to this position.
2\. Check your email.
**About Us**
------------------
We are over 88,000 people united by a strong Purpose: "Simplify and Enjoy Life More." Present in 9 countries, we consist of five major brands spanning diverse industries: Falabella Retail, Sodimac, Banco Falabella, Tottus, and Mallplaza. Each brand shapes who we are—and together, as One Team, we strive daily to reinvent ourselves and exceed our customers’ expectations.
A team full of dreams that makes things happen. We dare to launch and innovate, take risks, and create opportunities—keeping us at the forefront and driving us to continuously reinvent ourselves to deliver the best shopping experience at every touchpoint with us.