Qureos

Find The RightJob.

Site Reliability Engineer (SRE)

Role Summary

Ensure high availability, performance, and reliability of production systems through automation, cloud operations, monitoring, and incident management.

Key Responsibilities

- Build and manage scalable AWS cloud infrastructure using Terraform / CloudFormation

- Implement monitoring and observability using Prometheus, Grafana, Splunk, or Datadog

- Handle incidents, RCA, on-call support, SLOs, SLIs, and error budgets

- Automate operations to improve efficiency, reliability, and MTTR

- Manage Kubernetes, Docker, CI/CD pipelines, runbooks, and ITIL processes

Mandatory Skills

- SRE / DevOps / Production Support with AWS cloud experience

- Kubernetes, Docker, Linux, and networking fundamentals

- Scripting using Python, Bash, or Go

- Monitoring tools and incident / RCA management

Preferred Skills

- Terraform / CloudFormation, Jenkins / GitLab CI/CD, ServiceNow / ITSM

- Cloud security tools and enterprise environment experience

Experience

5+ years in SRE / DevOps / Cloud / Production Support

© 2026 Qureos. All rights reserved.