Qureos
الرجوع إلى الوظائف
Ceinsys Tech Ltd.

Site Reliability Engineering Lead

  • غير محدد

نُشرت أول أمس

عن الوظيفة

Location: Chennai
Experience: 5-8+ years
Open Positions: 01

Role Summary

Lead reliability engineering for business-critical manufacturing platforms, owning SLOs, observability, CI/CD, infrastructure automation, incident management and production readiness. Lead AI/Agentic-AI reliability covering model/agent observability, AIOps, drift monitoring, workflow guardrails and GenAI-assisted incident response.

Key Responsibilities

  • Define SLI/SLO/SLA and error budgets; align release velocity with reliability targets.
  • Build observability across metrics, logs and traces using Splunk, Datadog, Prometheus and Grafana.
  • Create actionable alerts and dashboards that reduce alert fatigue.
  • Develop Jenkins CI/CD pipelines and evaluate GitOps/Argo CD.
  • Manage GCP infrastructure with Terraform; operate Docker/Kubernetes workloads.
  • Conduct load/stress testing and production-readiness reviews.
  • Own incident response, on-call, MTTA/MTTR, postmortems, runbooks and improvement.
  • Monitor model-serving, agent orchestration and RAG systems for availability, latency, task success and drift.
  • Define AI SLOs and safeguards including auditing, human-in-the-loop escalation and runaway-agent circuit breakers.
  • Implement AIOps for anomaly detection, predictive alerts, automated triage and GenAI-assisted RCA/runbooks.
  • Reduce operational toil through automation; collaborate across platform, manufacturing, data/AI and IT teams; mentor engineers.
  • Provide weekly reliability, incident and KPI reporting.

Required Qualifications

  • 5–8+ years in SRE, DevOps or platform engineering.
  • Strong SLI/SLO, error-budget, observability and incident-management expertise.
  • Strong GCP experience including networking, compute, managed services and access.
  • Production experience with Kubernetes, Docker and Terraform.
  • Jenkins CI/CD experience; GitOps/Argo CD exposure desirable.
  • Splunk plus Datadog, Prometheus or Grafana experience.
  • Python and/or Java scripting capability.
  • Load/stress testing, capacity analysis and production-readiness experience.
  • PagerDuty, Opsgenie or equivalent on-call tooling knowledge.
  • Exposure to AI/ML reliability, MLOps, AIOps or Agentic-AI operations.

AI & Agentic-AI Reliability Experience

  • Vertex AI: model monitoring, pipelines and agent-building services.
  • LangChain, LangGraph or comparable agent orchestration frameworks.
  • RAG, vector search and retrieval-system observability.
  • Model performance and data-drift monitoring.
  • Inference latency, model availability and agent task-completion observability.
  • GenAI-assisted incident response, runbook retrieval, incident copilots and RCA summarisation.

Behavioural & Leadership Competencies

  • Technical leadership and architectural decision-making.
  • Analytical troubleshooting of distributed and AI-system failures.
  • Clear communication of reliability, risks and progress to technical/executive audiences.
  • Cross-functional collaboration and SRE mentoring.
  • Operational ownership through incident closure, corrective actions and documentation.

Key Deliverables

  • SLO/error-budget framework and integrated observability dashboards.
  • Controlled CI/CD and Terraform-based infrastructure automation.
  • Incident-management process, on-call model, runbooks and postmortems.
  • Load/stress testing evidence for production readiness.
  • Initial AI/Agentic-AI observability and AIOps/GenAI incident-response proof of concept.
  • Modular onboarding architecture and regular reliability/KPI reporting.

Initial Success Measures

  • SLOs operational for at least three critical platforms within three months.
  • Initial model-monitoring or AIOps capability operational within three months.
  • Critical-incident MTTA below 15 minutes with continuous MTTR improvement.
  • Operational toil at or below 50% per sprint, with remaining capacity for automation.

Additional Expectations

  • Flexibility for onsite collaboration, milestones and on-call participation.
  • Approved engineering workstation with GCP and AI/ML tooling access.
  • Strong documentation and modular architecture to support handover and scale-up.

Send your CVs to careers@cstech.ai

وظائف مشابهة