Overview
We are seeking a Cloud Site Reliability Engineer (SRE) to enhance and maintain our client's Azure infrastructure. This role involves collaboration in planning and supporting activities across virtual machines, networking, databases, and storage. The position focuses on ensuring service health, resolving incidents, implementing changes, and enhancing operational processes through automation in a fully remote capacity.
Responsibilities
- Monitor service health and resolve incidents in a timely manner.
- Implement and execute controlled changes, including application deployments and configuration adjustments.
- Conduct infrastructure discovery and maintain an inventory of servers and applications.
- Automate repetitive operational tasks using scripting languages, contributing to AIOps initiatives.
- Administer Azure compute, storage, and OS-layer services to support applications.
- Create and maintain documentation, including runbooks and standard operating procedures.
- Mentor junior engineers to enhance team capabilities and knowledge.
- Improve monitoring and observability processes to enhance service reliability.
Requirements
- Strong hands-on experience with Microsoft Azure infrastructure.
- Proven ability to support workloads migrating from VMware/ESX to Azure.
- Solid experience in Windows Server administration across various versions.
- Deep understanding of Active Directory, DNS, and IIS operational dependencies.
- Proficient in SQL Server infrastructure management and performance tuning.
- Strong scripting skills in PowerShell and familiarity with Terraform or Azure Bicep.
- Knowledge of Azure networking concepts and configurations.
- Experience in ITSM/change-controlled environments and troubleshooting across various layers.