Overview
We are seeking a skilled Site Reliability Engineer to support high-profile government projects. The successful candidate will collaborate with multiple teams to enhance the reliability, performance, and cost-effectiveness of cloud and on-premises services. This hybrid role involves both working on-site and remotely, with a focus on ensuring system availability and introducing automated solutions to improve operational efficiency.
Responsibilities
- Design and maintain reliable, scalable physical and virtual infrastructure.
- Monitor system performance and proactively resolve issues.
- Automate processes using tools such as Ansible to improve efficiency and consistency.
- Collaborate with engineers and stakeholders across the business.
- Support continuous improvement of systems, tools, and practices.
- Operate across the full infrastructure stack, from bare metal to virtualized deployments.
Requirements
- Proven experience with configuration management tools (e.g., Ansible, Chef).
- Familiarity with Terraform and container orchestration tools (e.g., Kubernetes, OpenShift).
- Experience with CI/CD tools (e.g., Jenkins) and monitoring tools (e.g., InfluxDB, Prometheus).
- Understanding of relational databases and SQL, along with Linux command line skills.
- Knowledge of network security protocols and cloud hosting services (preferably AWS).
- Experience in writing well-tested code in programming languages like Java, Go, or Python is a plus.