Brief Job / Project Description:
Lead onsite L2/L3 operations for Smart Services Solutions to ensure high availability, performance, security, and compliance of platforms and supporting infrastructure. Manage the L1 operations team (8x5) and own end-to-end production support, monitoring, incident/problem management, vendor coordination, change execution, and continuous improvement across Smart Services technologies (EV, IoT, Signage, IBMS integrations, and related components).
Key Responsibilities:
1) Leadership, Operations & SLA Ownership
- Lead and manage the L1 support team (8x5) for monitoring and day-to-day operations.
- Ensure SLAs & KPIs are achieved as per contract (uptime, response time, stability).
- Conduct and govern daily application and platform health checks (apps, services, integrations, jobs, dependencies).
- Maintain and continuously improve runbooks (startup/shutdown, failover, troubleshooting, known errors, escalation matrix).
2) Monitoring, Automation & Proactive Support
- Build/maintain monitoring coverage: dashboards, alerts, and smart rules aligned with business requirements.
- Proactively analyze logs/metrics to detect degradation trends and prevent outages.
- Automate operational routines and housekeeping tasks to reduce manual work and repeat incidents.
3) Incident, Problem & RCA Management (L2/L3)
- Own the incident lifecycle end-to-end: triage, diagnosis, workaround, permanent fix coordination.
- Drive stakeholder communications, status updates, and restoration timelines.
- Lead RCA collection for major incidents and ensure corrective/preventive actions are tracked and closed.
4) Application, Platform & Server Administration
- Manage application services, service accounts, scheduled tasks, configurations, and environment parameters.
- Perform OS-level administration for application servers (Windows/Linux): restarts, capacity checks (CPU/RAM/Disk), performance tuning, and hardening alignment.
- Coordinate and execute platform/application patching and upgrades with validation and rollback readiness.
- Validate backups and restore readiness with relevant infra teams.
5) Change, Release, Testing & Acceptance
- Prepare and manage Change Requests (CRs): impact, risk, rollback, implementation steps, and maintenance windows execution.
- Coordinate testing, deployments, smoke tests, and obtain stakeholder acceptance/sign-off.
- Maintain change calendar and ensure stakeholder notifications.
6) Security, Compliance & Governance
- Ensure platforms comply with Enterprise Architecture, Cybersecurity controls, and government regulations.
- Coordinate vulnerability remediation through patching and configuration fixes.
- Manage and renew SSL/TLS certificates (CSR, PFX/PEM, TLS configuration) and prevent expiry incidents.
- Ensure access controls follow least privilege, approvals, and audit requirements.
7) Stakeholder, Vendor & Business Coordination
- Act as primary onsite interface for vendors: raise tickets, follow up, enforce timelines, and escalate as needed.
- Coordinate with stakeholders for new integrations and deployments.
- Engage business owners to capture requirements and ensure delivery alignment.
8) Documentation, Reporting & Service Performance
- Maintain inventory of applications, servers, versions, certificates, licenses, and integrations.
- Generate agreed performance and business reports (weekly/monthly): availability, SLA adherence, incidents, recurring issues, improvement actions.
- Maintain as-built documentation and service maps (components + dependencies).
9) License & Certificate Management
- Own application licensing tracking, renewals, compliance, and vendor alignment.
- Maintain certificate renewal plan, expiry tracking, and evidence documentation.
10) Continuous Improvement
- Identify key operational improvements (monitoring enhancements, automation, thresholds).
- Reduce repeat incidents through preventive maintenance and root-cause elimination.
Deliverables / Expected Outcomes:
- Daily health check confirmation + exception reporting
- Incident tickets with timelines, actions, and stakeholder updates
- RCA documents and action tracking for major incidents
- Updated runbooks, service maps, and operational documentation
- Certificate renewal plan and expiry tracking
- Monthly service report (availability, incidents, improvements, KPI/SLA performance)
Lead onsite L2/L3 operations for Smart Services Solutions to ensure high availability, performance, security, and compliance of platforms and supporting infrastructure. Manage the L1 operations team (8x5) and own end-to-end production support, monitoring, incident/problem management, vendor coordination, change execution, and continuous improvement across Smart Services technologies (EV, IoT, Signage, IBMS integrations, and related components).
Mandatory Technical Skills:
10+ years in Application Operations / Production Support (L2/L3), with leadership/team management experience.
- Strong hands-on experience in Windows/Linux server and Kubernetes administration.
- Strong troubleshooting and stakeholder communication skills.
- Monitoring & log analysis (APM/log tools, dashboards)
- Web services & integrations (REST APIs, auth, tokens, gateways)
- Implementing(CSR, PFX/PEM, TLS)
- Databases (health checks, queries basics, backup verification)
- Reverse proxy / load balancer concepts, IIS/NGINX basics
Tools / Platforms / Technologies:
- EV Services Platforms (backend apps, integrations, APIs, dashboards)
- IoT Platforms (device onboarding, telemetry/data pipelines, middleware, APIs)
- Digital Signage / Workplace Screens: Appspace (content platform, players, device enrollment, licensing)
- IBMS / BMS integrations (connectors, data exchange, reporting interfaces)
- Windows Server / Linux Server and Kubernetes administration
- Web / Reverse Proxy: IIS, NGINX, Apache
- Load Balancers: F5 / HAProxy / NGINX LB (as applicable)
- Monitoring & APM: Dynatrace / AppDynamics / Zabbix / PRTG (as applicable)
- Log Management / SIEM: Splunk / ELK (Elastic) / Microsoft Sentinel / Graylog (as applicable)
- Certificates & Security: SSL/TLS, CSR, PFX/PEM, hardening alignment, vulnerability remediation coordination
- Databases: Microsoft SQL Server / PostgreSQL / MySQL (health checks, basic queries, backup verification)
- Middleware / Messaging: RabbitMQ / Kafka (as applicable)
- Scheduling / Jobs: Windows Task Scheduler, Cron
- ITSM / Ticketing: ServiceNow / Jira Service Management
- Documentation / Knowledge Base: Confluence
- DevOps / CI-CD: Azure DevOps / Jenkins / GitLab CI, Git version control
- Reporting: SLA/KPI dashboards, operational performance & availability reports
Preferred / Nice-to-Have Skills:
- Smart Services domain experience (EV, IoT, Digital Signage/Appspace, IBMS/BMS integrations)
- Proven team leadership for L1/L2 operations (8x5), including shift handover and task governance
- Strong incident/problem management (Major Incidents, RCA leadership, CAPA tracking/closure)
- Monitoring & observability expertise (APM tuning, alert optimization, dashboarding, SLA/SLO reporting)
- Automation & scripting (PowerShell/Python/Bash) to reduce manual ops and improve reliability
- Integration troubleshooting skills (REST APIs, OAuth/tokens, gateways, JSON, Postman)
- Performance & capacity management (baselining, thresholds, CPU/RAM/Disk planning)
- Security & compliance mindset (TLS hardening, vulnerability remediation, least-privilege access)
- Change & release management excellence (risk/impact, rollback planning, maintenance execution)
- Strong vendor & stakeholder management (ticket follow-up, escalations, service reviews, clear communications)
Certifications (if any):
- Certified Information Systems Security Professional (CISSP)
- AWS Certified SysOps Administrator – Associate
- Linux Foundation Certified System Administrator (LFCS) / RHCSA
- Certified Kubernetes Administrator (CKA)
- ITIL Foundation Certification