Search by job, company or skills

SRE Expert

  • Posted an hour ago
  • Be among the first 10 applicants

Job Description

Key Responsibilities:

- Implement and institutionalize SRE practices including SLIs, SLOs, Error Budgets, reliability reviews, and service health management

- Design and implement end-to-end observability solutions using Splunk Observability Cloud, Splunk ITSI, and Splunk Enterprise

- Define and manage alert quality standards to reduce noise, eliminate alert fatigue, and improve incident response effectiveness

- Configure and maintain business service monitoring, KPI frameworks, service health dashboards, event analytics, and business observability capabilities

- Implement AIOps capabilities including event correlation, anomaly detection, intelligent alerting, and operational analytics

- Drive continuous improvement initiatives to improve MTTD, MTTR, MTBF, service reliability, and operational efficiency

- Integrate monitoring, logging, ITSM, cloud, application, infrastructure, and third-party platforms to enable end-to-end observability

- Support incident management, problem management, RCA, and blameless postmortem activities

- Partner with application, cloud, platform, and operations teams to improve production readiness and operational resilience

- Develop automation solutions for monitoring, alerting, reporting, remediation, and operational workflows

- Participate in Agile ceremonies, reliability reviews, and continuous service improvement programs

Skill Requirements

Must Have Skills:

- 8+ years of experience in Site Reliability Engineering, Production Support, Application Support, Observability, Operations Engineering, or related roles

- Strong hands-on experience implementing SRE practices including SLI, SLO, Error Budgets, reliability governance, and service health management

- Strong expertise with Splunk Observability Cloud (OpenTelemetry, Infrastructure Monitoring, APM, RUM, Synthetic Monitoring)

- Strong expertise with Splunk ITSI (Event Analytics, Service Modeling, KPI Management, Glass Tables, Service Health Monitoring, Business Observability)

- Strong expertise with Splunk Enterprise (Logging, SPL, Log Analytics, Dashboarding)

- Experience implementing alert quality management, alert rationalization, and noise reduction initiatives

- Experience with Incident Management, Problem Management, RCA, and Blameless Postmortems

- Experience implementing AIOps and intelligent operations capabilities

- Hands-on experience integrating observability platforms with applications, cloud services, infrastructure, ITSM platforms, and enterprise tools

- Experience supporting business applications including custom applications (Java, .NET and other technologies running on VM and container platforms), SAP, Salesforce, and SaaS/COTS platforms

- Experience with automation technologies to reduce the operational toil (Ansible, Python, RPA and other automation platforms)

- Strong analytical, problem-solving, communication, and stakeholder management skills

Other Requirements

Good to Have Skills:

- Knowledge of DevOps practices, CI/CD pipelines, GitOps, and release automation

- Experience with Platform Engineering and Internal Developer Platforms (IDP)

- Experience with Resilience Engineering and Chaos Engineering practices

- Exposure to Agentic AI, AI-driven Operations, and AI-assisted observability solutions

- Experience with cloud platforms including Azure, AWS, and GCP

- Splunk, SRE, Cloud, or Observability-related certifications

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152384333

Beware of Scammers

We don’t charge money for job offers