Site Reliability Engineer

Overview

The Site Reliability Engineer (SRE) will play a vital role in supporting critical project work within a heavily AWS-oriented environment. This position involves collaborating with a team to ensure the reliability and availability of production services, with a focus on optimizing cloud infrastructure and application performance. The SRE will utilize their experience in incident management and monitoring to effectively manage and resolve issues, all while working remotely.

Responsibilities

  • Develop and implement observability solutions using tools like Splunk and Grafana.
  • Monitor, support, and optimize live production services in high-availability environments.
  • Investigate and resolve application, infrastructure, and performance-related issues.
  • Analyze logs, metrics, and traces to identify service risks and performance degradation.
  • Collaborate with stakeholders for effective incident resolution and communication.
  • Work with service level objectives (SLOs) and reliability engineering best practices.
  • Utilize automation tools to enhance service stability and reduce manual intervention.

Requirements

  • Strong experience with AWS services including EC2, ECS, RDS, and CloudWatch.
  • Proven ability to monitor, troubleshoot, and optimize live services in production environments.
  • Experience in incident management and performing root cause analysis (RCA).
  • Familiarity with at least one programming language such as Java, C#, Python, or JavaScript.
  • Understanding of SLOs, SLIs, and alerting strategies.
  • Experience in supporting stakeholder communications and issue investigation.
  • Desirable skills include familiarity with Jenkins, CI/CD pipelines, and Splunk administration.