Overview
We are looking for an experienced DevOps Engineer to join a financial services consultancy for a long-term contract role focused on supporting a high-throughput, enterprise API platform. This position combines engineering and operations responsibilities, including maintaining system health across various cloud environments while contributing to automation and performance improvements. Collaborating with cross-functional teams, the contractor will play a key role in enhancing platform stability and performance through their engineering expertise.
Responsibilities
- Maintain and manage a Kubernetes-based platform across public and private cloud environments.
- Drive end-to-end incident response, including root-cause analysis and post-incident reviews.
- Build automation, Helm charts, and custom diagnostic tooling to eliminate manual processes.
- Ensure monitoring alerts are actionable through clear, up-to-date runbooks and observability tools.
- Proactively manage capacity scaling and oversee software and infrastructure upgrades.
- Deliver practical engineering improvements to ensure platform stability and performance.
Requirements
- Proven hands-on experience operating production-grade Kubernetes, preferably with Service Mesh experience.
- Strong Linux administration and command-line troubleshooting skills.
- Practical experience with major cloud providers such as AWS or Azure.
- Confidence in debugging complex systems across all application and infrastructure layers.
- Familiarity with observability tools like Grafana, Prometheus, Loki, or Tempo is preferred.
- Hands-on experience with Infrastructure as Code (Terraform, Helm), CI/CD pipelines, or scripting in Python/Java is a plus.