Qureos

Find The RightJob.

Site Reliability Engineer

Evolvice is a nearshore technology services provider that helps businesses scale, innovate, and enhance efficiency. Since 2012, we’ve been developing software solutions and building high-performing remote teams. Today, we focus on integrating AI into business processes and providing IT and security support to drive digital transformation.
Originally based in Germany, we have established development hubs in Egypt, Ukraine, and Portugal, as well as offices in Saudi Arabia. This international presence enables us to deliver high-quality, cost-effective solutions worldwide.
Our Services:
Nearshore Teams – Build and scale remote teams of any size with expert engineers.
AI-Powered Business Productivity – Leverage AI-driven software solutions to boost efficiency.
Cybersecurity – Safeguard your business with advanced security assessments and services.
Managed IT & Application Support – Ensure seamless operations with proactive IT management and support.
We’re proud to work with industry leaders like Bosch, Douglas, WTS, DHL, Tatweer and GOSI, and many others. Combining German precision with nearshore agility, we provide secure, scalable, and cost-effective IT solutions tailored to your business needs.
Currently, we are searching for L3 Specialist / Site Reliability Engineer (SRE) to join the big team of professionals.

Summary:
The L3 Specialist / Site Reliability Engineer (SRE) provides expert-level troubleshooting, root cause analysis, platform optimization, automation, architecture improvements, performance tuning, and reliability engineering for mission-critical environments. The ideal candidate combines deep infrastructure and cloud expertise with strong automation skills to keep large-scale, highly available systems running reliably and efficiently.
_____________________________________________________________________________________________________________________________________________________
Key Responsibilities

Troubleshooting & Root Cause Analysis
  • Provide expert-level troubleshooting for complex, mission-critical infrastructure and application issues.
  • Lead root cause analysis for major incidents and drive permanent remediation to prevent recurrence.
  • Act as an escalation point for L1/L2 teams on complex platform and reliability issues.

Platform Reliability & Performance Engineering
  • Design and implement reliability engineering practices, including SLIs, SLOs, and error budgets.
  • Perform capacity planning and performance tuning across compute, storage, and networking layers.
  • Proactively identify and remediate single points of failure and reliability risks in production systems.

Automation & Infrastructure as Code
  • Build and maintain Infrastructure as Code using Terraform and configuration automation using Ansible.
  • Automate operational tasks, deployments, and remediation workflows using Python, Bash, and PowerShell.
  • Design and maintain containerized workloads using Docker and orchestration with Kubernetes.

CI/CD & Release Engineering
  • Design, maintain, and optimize CI/CD pipelines using Jenkins and GitLab CI/CD.
  • Support release management processes, ensuring safe, repeatable, and well-tested deployments.
  • Collaborate with development teams to embed reliability and automation practices into the software delivery lifecycle.

Monitoring & Observability
  • Implement and maintain monitoring, alerting, and observability solutions using Prometheus and Grafana.
  • Manage centralized logging and analysis using the ELK stack to support troubleshooting and trend analysis.
  • Continuously improve alerting thresholds and dashboards to reduce noise and improve incident response time.

Architecture & Continuous Improvement
  • Contribute to architecture improvements for scalability, resilience, and cost efficiency across cloud environments.
  • Evaluate and recommend new tools, technologies, and practices to improve platform reliability and engineering efficiency.
  • Maintain up-to-date documentation for infrastructure, processes, and runbooks.

Collaboration & Incident Response
  • Participate in on-call rotations and lead incident response for critical production issues.
  • Partner with development, security, and infrastructure teams to align on reliability goals and best practices.
  • Support post-incident reviews and communicate findings and improvement actions to stakeholders.
_____________________________________________________________________________________________________________________________________________________
Required Qualifications

Experience
  • 5+ years of experience in infrastructure engineering, DevOps, or Site Reliability Engineering (SRE) roles.
  • Proven experience supporting mission-critical, highly available production environments.
  • Experience working within Agile/DevOps delivery models and cross-functional engineering teams.

Technical Skills
  • Strong hands-on experience with Kubernetes and Docker for container orchestration and management.
  • Proficiency with Terraform and Ansible for infrastructure automation and configuration management.
  • Experience building and maintaining CI/CD pipelines using Jenkins and GitLab CI/CD.
  • Solid experience with monitoring and observability tools: Prometheus, Grafana, and the ELK stack.
  • Strong Linux systems administration and troubleshooting skills.
  • Hands-on experience with at least one major cloud platform (Azure, AWS, or OCI); multi-cloud experience a plus.
  • Strong scripting skills in Python, Bash, and/or PowerShell.

Certifications (Preferred)
  • Certified Kubernetes Administrator (CKA) or equivalent.
  • Cloud certifications (e.g., Azure Administrator/Solutions Architect, AWS Solutions Architect, OCI Architect) a plus.
  • HashiCorp Certified: Terraform Associate a plus.

Soft Skills
  • Strong analytical and problem-solving skills with a methodical approach to troubleshooting.
  • Excellent communication skills, able to convey technical issues to both technical and non-technical stakeholders.
  • Ability to work independently and take ownership of complex, ambiguous problems.
  • Strong sense of accountability and calm, structured approach during high-pressure incidents.

Education
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
_____________________________________________________________________________________________________________________________________________________
We offer:
  • Financial stability;
  • Interesting and challenging projects within professional self-managed teams;
  • Friendly team and a comfortable working environment;
  • Flexible schedule (8 —10 AM start) with the possibility to work assigned hours and/or adjust work schedule as requested by the manager;
  • Social insurance & Health insurance;
  • Paid sick leave;

Why Work With Us:
We work as a self-driven team without complex management structures. Our teams make independent decisions without recommendations from the client. We nurture an open, transparent environment where we all enjoy our work.

© 2026 Qureos. All rights reserved.