Overview
The Site Reliability Engineer (SRE) will play a vital role in supporting critical project work within a heavily AWS-oriented environment. This position involves collaborating with a team to ensure the reliability and availability of production services, with a focus on optimizing cloud infrastructure and application performance. The SRE will utilize their experience in incident management and monitoring to effectively manage and resolve issues, all while working remotely.
Responsibilities
- Develop and implement observability solutions using tools like Splunk and Grafana.
- Monitor, support, and optimize live production services in high-availability environments.
- Investigate and resolve application, infrastructure, and performance-related issues.
- Analyze logs, metrics, and traces to identify service risks and performance degradation.
- Collaborate with stakeholders for effective incident resolution and communication.
- Work with service level objectives (SLOs) and reliability engineering best practices.
- Utilize automation tools to enhance service stability and reduce manual intervention.
Requirements
- Strong experience with AWS services including EC2, ECS, RDS, and CloudWatch.
- Proven ability to monitor, troubleshoot, and optimize live services in production environments.
- Experience in incident management and performing root cause analysis (RCA).
- Familiarity with at least one programming language such as Java, C#, Python, or JavaScript.
- Understanding of SLOs, SLIs, and alerting strategies.
- Experience in supporting stakeholder communications and issue investigation.
- Desirable skills include familiarity with Jenkins, CI/CD pipelines, and Splunk administration.