Qureos
Back to jobs
HCLTech

Senior Technical Architect

  • Not specified

Posted 10 hours ago

About the role

Bangalore, Karnataka
Job Summary

AI Observability Principal Architect(18+ Years)

Description

The AI Observability Principal Architect is responsible for defining and delivering end-to-end observability and AIOps capabilities across enterprise applications and platforms. The role focuses on enabling proactive detection, intelligent event correlation, and automated incident response , improving system reliability and operational efficiency.

This role will lead observability strategy, standardization, and implementation across teams, ensuring clear visibility into system health, business workflows, and performance outcomes. The lead will work closely with SRE, application, infrastructure, and service management teams to embed outcome-driven observability (CUJs, SLIs/SLOs) and drive continuous improvement in incident detection and resolution.

Key Responsibilities

Assignment Deliverables

Define and implement observability strategy, standards, and roadmap
Establish telemetry framework (logs, metrics, traces) and instrumentation standards
Define Critical User Journeys (CUJs) and map Service Level Indicators (SLIs) / SLOs
Enable actionable alerting aligned to user/business impact
Implement alert-to-incident automation with correct routing and ownership

Drive AIOps capabilities :
Event correlation and alert noise reduction
Root-cause-based incident generation
Predictive detection and anomaly identification
Build and optimize observability dashboards for operations and leadership visibility
Enable automation and self-healing playbooks for recurring incidents
Lead major incident support and post-incident improvements

Ensure continuous improvement through observability lifecycle management

Skill Requirements

Required Skills

18+ years experience in IT Operations / SRE / Observability / Platform Engineering

Strong expertise in:
Observability (logs, metrics, traces, distributed tracing)
SRE practices (SLIs, SLOs, error budgets)
Incident management and automation

Experience in AIOps / Event Management :
Event correlation, alert deduplication, noise reduction
Hands-on experience with observability platforms and ITSM integrations

Strong understanding of:
Distributed systems and cloud environments
Telemetry pipelines and data integration
Experience in designing automation workflows, runbooks, and self-healing mechanisms
Strong stakeholder management and ability to work across application, infra, and operations teams
Other Requirements

Nice to Have

Experience with Azure observability stack / OpenTelemetry
Experience with ServiceNow ITOM / AIOps
Exposure to AI-driven observability (anomaly detection, predictive analytics)
Experience in building executive dashboards and reporting frameworks
#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-#body.unify div.unify-button-container .unify-apply-now: focus, #body.unify div.unify-button-container .unify-apply-

Similar jobs