Detalles Equipo Calendario Documento FAQ
Challenge

59888-Site Reliability Engineer (SRE) – Incident Operations & Communications

Ranking: 2632

We are looking for an experienced Site Reliability Engineer (SRE) with at least 5 years of experience, combining strong software and platform engineering capabilities with expertise in incident management, production reliability and operational communications.

The role is particularly focused on Incident Operations & Communications, requiring the ability to coordinate critical incidents, troubleshoot complex production environments, assess impact and severity, and communicate effectively with both technical teams and stakeholders.

The ideal candidate has strong experience with distributed systems, observability, automation and reliability engineering, together with the ability to make effective decisions under pressure and ambiguity.

Main Responsibilities

  • Manage and coordinate production incidents and incident response activities.
  • Act as part of the incident command process during critical production events.
  • Troubleshoot and debug complex production issues.
  • Assess incident impact and severity.
  • Coordinate technical teams during incident resolution.
  • Manage external communications related to production incidents.
  • Manage and update status pages during incidents.
  • Work within SLA-driven operational environments.
  • Improve platform availability, reliability and resilience.
  • Develop automation to reduce operational toil.
  • Work with APIs and system integrations.
  • Use observability and monitoring platforms to identify and investigate production issues.
  • Coordinate operational handoffs across Follow-the-Sun teams.
  • Communicate effectively with both technical teams and executive stakeholders.
  • Make decisions in high-pressure situations and environments with incomplete or ambiguous information.

Required Experience

  • Minimum 5 years of professional experience in Site Reliability Engineering or related reliability/platform engineering roles.
  • Strong experience in Site Reliability Engineering (SRE).
  • Experience in Incident Management and Incident Command.
  • Experience in Platform Engineering and/or Software Engineering.
  • Experience working with distributed systems and systems design.
  • Strong production troubleshooting and debugging capabilities.
  • Experience working with high-availability and reliability-critical environments.

Technical Skills

Primary Skill

  • Site Reliability Engineering (SRE)
  • Reliability Engineering

Specialization

  • Incident Operations
  • Incident Management
  • Incident Command
  • Incident Communications
  • Impact Assessment
  • Severity Management
  • External Incident Communications
  • Status Page Management
  • SLA-driven Operations
  • Follow-the-Sun Operations
  • Operational Handoffs

Technical Stack

  • Python and/or Kotlin
  • Datadog
  • Chronosphere
  • PagerDuty
  • Rootly
  • Observability & Monitoring
  • Distributed Systems
  • Systems Design
  • APIs
  • System Integrations
  • Automation
  • Production Troubleshooting & Debugging
  • High Availability
  • Reliability Engineering
  • Toil Reduction

Professional Skills

  • Cross-functional coordination.
  • Executive and technical communication.
  • Decision-making under pressure.
  • Ability to work effectively in ambiguous situations.
  • Coordination across distributed operational teams.

About the projects you will work on

You will work 100% remotely from wherever you decide. Sometimes you may have face-to-face meetings and for that reason you have to reside in Spain or in the European Union.

You will work on projects with a leading company in digital transformation with a passion for technology and innovation in sectors such as banking (35 of the main banks worldwide work with our client), insurance, industrial and automotive in Big Data projects, Blockchain, AI, Cloud, among others.

About the client

  • Global presence in more than 15 markets
  • 8,000+ employees
  • More than 35 years of experience

About the process and your contractual relationship

If you are interested in this offer, we will enroll you in the process and submit your application, blindly, that is, without your contact details, to the technical and human resources department so that they can evaluate your profile and your financial expectations.

If the answer is positive, we organize the meetings so that the client knows you and explains the project in detail.

If after the meeting both parties agree on the conditions, you receive a firm offer to work with us or directly be hired by the client.

59888-Site Reliability Engineer (SRE) – Incident Operations & Communications

Ranking: 2632

We are looking for an experienced Site Reliability Engineer (SRE) with at least 5 years of experience, combining strong software and platform engineering capabilities with expertise in incident management, production reliability and operational communications.

The role is particularly focused on Incident Operations & Communications, requiring the ability to coordinate critical incidents, troubleshoot complex production environments, assess impact and severity, and communicate effectively with both technical teams and stakeholders.

The ideal candidate has strong experience with distributed systems, observability, automation and reliability engineering, together with the ability to make effective decisions under pressure and ambiguity.

Main Responsibilities

  • Manage and coordinate production incidents and incident response activities.
  • Act as part of the incident command process during critical production events.
  • Troubleshoot and debug complex production issues.
  • Assess incident impact and severity.
  • Coordinate technical teams during incident resolution.
  • Manage external communications related to production incidents.
  • Manage and update status pages during incidents.
  • Work within SLA-driven operational environments.
  • Improve platform availability, reliability and resilience.
  • Develop automation to reduce operational toil.
  • Work with APIs and system integrations.
  • Use observability and monitoring platforms to identify and investigate production issues.
  • Coordinate operational handoffs across Follow-the-Sun teams.
  • Communicate effectively with both technical teams and executive stakeholders.
  • Make decisions in high-pressure situations and environments with incomplete or ambiguous information.

Required Experience

  • Minimum 5 years of professional experience in Site Reliability Engineering or related reliability/platform engineering roles.
  • Strong experience in Site Reliability Engineering (SRE).
  • Experience in Incident Management and Incident Command.
  • Experience in Platform Engineering and/or Software Engineering.
  • Experience working with distributed systems and systems design.
  • Strong production troubleshooting and debugging capabilities.
  • Experience working with high-availability and reliability-critical environments.

Technical Skills

Primary Skill

  • Site Reliability Engineering (SRE)
  • Reliability Engineering

Specialization

  • Incident Operations
  • Incident Management
  • Incident Command
  • Incident Communications
  • Impact Assessment
  • Severity Management
  • External Incident Communications
  • Status Page Management
  • SLA-driven Operations
  • Follow-the-Sun Operations
  • Operational Handoffs

Technical Stack

  • Python and/or Kotlin
  • Datadog
  • Chronosphere
  • PagerDuty
  • Rootly
  • Observability & Monitoring
  • Distributed Systems
  • Systems Design
  • APIs
  • System Integrations
  • Automation
  • Production Troubleshooting & Debugging
  • High Availability
  • Reliability Engineering
  • Toil Reduction

Professional Skills

  • Cross-functional coordination.
  • Executive and technical communication.
  • Decision-making under pressure.
  • Ability to work effectively in ambiguous situations.
  • Coordination across distributed operational teams.

About the projects you will work on

You will work 100% remotely from wherever you decide. Sometimes you may have face-to-face meetings and for that reason you have to reside in Spain or in the European Union.

You will work on projects with a leading company in digital transformation with a passion for technology and innovation in sectors such as banking (35 of the main banks worldwide work with our client), insurance, industrial and automotive in Big Data projects, Blockchain, AI, Cloud, among others.

About the client

  • Global presence in more than 15 markets
  • 8,000+ employees
  • More than 35 years of experience

About the process and your contractual relationship

If you are interested in this offer, we will enroll you in the process and submit your application, blindly, that is, without your contact details, to the technical and human resources department so that they can evaluate your profile and your financial expectations.

If the answer is positive, we organize the meetings so that the client knows you and explains the project in detail.

If after the meeting both parties agree on the conditions, you receive a firm offer to work with us or directly be hired by the client.

  • Equipo
  • Evaluador
  • Manager
  • Agencia
  • Cliente

GFT Cliente

Cliente

Comentarios: 0

Teba Gomez-Monche

Agencia

Comentarios: 0

Sandra Lobero

Agencia

Comentarios: 0

Paco Romero

Agencia

Comentarios: 0

claudia herrero

Agencia

Comentarios: 0

Hugo Herrero

Manager

Comentarios: 0

Víctor M. herrero

Evaluador

Comentarios: 3