DevOps / SRE Specialist – Incident Operations & Production Reliability
Key Responsibilities
- Monitor and maintain the reliability and availability of production systems.
- Respond to and coordinate real-time production incidents in SLA-driven environments.
- Participate in on-call rotations and provide timely incident response.
- Investigate incidents, identify root causes, and coordinate remediation activities.
- Work with engineering and technical teams to resolve complex production issues.
- Analyse monitoring, alerts, logs, dashboards, and other observability data to identify potential problems.
- Help improve incident response processes, operational procedures, and reliability practices.
- Contribute to post-incident reviews and identify opportunities for continuous improvement.
- Support system integrations and troubleshoot issues involving APIs and distributed systems.
- Develop or maintain automation and operational tooling using Python, Kotlin, or similar technologies.
- Apply SDLC and production reliability principles throughout the software lifecycle.
- Communicate incident status, technical findings, and resolutions clearly to internal and external stakeholders.
Must-Have Requirements
- 5+ years of experience in Incident Operations, SRE, Technical Operations, DevOps, or similar roles.
- Proven experience working in on-call environments with SLA-driven responsibilities.
- Demonstrated ability to operate effectively during high-pressure, real-time incidents.
- Strong understanding of distributed systems and production environments.
- Hands-on experience with monitoring and alerting platforms, such as:
- Datadog
- Chronosphere
- Similar observability and monitoring solutions
- Experience with incident management and response tools, such as:
- PagerDuty
- Rootly
- Slack workflows
- Similar incident management platforms
- Familiarity with APIs and system integrations.
- Good understanding of observability tooling, monitoring dashboards, and production metrics.
- Hands-on programming experience with Python, Kotlin, or similar languages.
- Strong understanding of the Software Development Lifecycle (SDLC) and production reliability principles.
- Strong operational judgment and the ability to make informed decisions under uncertainty.
- English proficiency: C1 level.
Soft Skills
- Excellent written and verbal communication skills.
- Strong ability to communicate clearly during critical incidents.
- Comfortable working with both technical and non-technical stakeholders.
- Ability to remain calm and structured in high-pressure situations.
- Strong ownership and accountability for production reliability.
- Proactive approach to problem-solving and continuous improvement.
- Excellent collaboration skills in distributed and multidisciplinary teams.
DevOps / SRE Specialist – Incident Operations & Production Reliability
Key Responsibilities
- Monitor and maintain the reliability and availability of production systems.
- Respond to and coordinate real-time production incidents in SLA-driven environments.
- Participate in on-call rotations and provide timely incident response.
- Investigate incidents, identify root causes, and coordinate remediation activities.
- Work with engineering and technical teams to resolve complex production issues.
- Analyse monitoring, alerts, logs, dashboards, and other observability data to identify potential problems.
- Help improve incident response processes, operational procedures, and reliability practices.
- Contribute to post-incident reviews and identify opportunities for continuous improvement.
- Support system integrations and troubleshoot issues involving APIs and distributed systems.
- Develop or maintain automation and operational tooling using Python, Kotlin, or similar technologies.
- Apply SDLC and production reliability principles throughout the software lifecycle.
- Communicate incident status, technical findings, and resolutions clearly to internal and external stakeholders.
Must-Have Requirements
- 5+ years of experience in Incident Operations, SRE, Technical Operations, DevOps, or similar roles.
- Proven experience working in on-call environments with SLA-driven responsibilities.
- Demonstrated ability to operate effectively during high-pressure, real-time incidents.
- Strong understanding of distributed systems and production environments.
- Hands-on experience with monitoring and alerting platforms, such as:
- Datadog
- Chronosphere
- Similar observability and monitoring solutions
- Experience with incident management and response tools, such as:
- PagerDuty
- Rootly
- Slack workflows
- Similar incident management platforms
- Familiarity with APIs and system integrations.
- Good understanding of observability tooling, monitoring dashboards, and production metrics.
- Hands-on programming experience with Python, Kotlin, or similar languages.
- Strong understanding of the Software Development Lifecycle (SDLC) and production reliability principles.
- Strong operational judgment and the ability to make informed decisions under uncertainty.
- English proficiency: C1 level.
Soft Skills
- Excellent written and verbal communication skills.
- Strong ability to communicate clearly during critical incidents.
- Comfortable working with both technical and non-technical stakeholders.
- Ability to remain calm and structured in high-pressure situations.
- Strong ownership and accountability for production reliability.
- Proactive approach to problem-solving and continuous improvement.
- Excellent collaboration skills in distributed and multidisciplinary teams.
- Equipo
- Evaluador
- Manager
- Agencia
- Cliente