Work Type: 3 days onsite per week
We are seeking a technically strong Azure Infrastructure L2 Engineer based in TX to support, optimize, automate, and improve enterprise cloud operations. The ideal candidate should bring hands-on experience across Azure technologies, Azure Automation, Terraform, Ansible, Azure Update Manager, Logic Apps, Azure Monitor, Application Insights, and ServiceNow. This role requires a solid DevOps mindset, strong infrastructure-as-code capability, vulnerability remediation experience, and the ability to operate within ITIL-aligned incident, change, problem, and agile delivery processes. The candidate must be comfortable driving major incidents, preparing RCA documentation, facilitating PIR reviews, communicating with stakeholders, and coordinating with vendors to deliver stable, secure, and well-governed cloud operations.
Roles and responsibilities:
Provide L2 support for Azure infrastructure, applications, and platform services across production and non-production environments.
Monitor, troubleshoot, and resolve incidents related to Azure compute, storage, networking, identity, automation, patching, and application availability.
Design, maintain, and enhance automation using Azure Automation runbooks, PowerShell, Azure CLI, Logic Apps, Terraform, and Ansible.
Implement and support infrastructure as code using Terraform, ensuring repeatable, scalable, secure, and compliant Azure deployments.
Use Ansible for configuration management, patch orchestration, application configuration, and operational automation across Windows and Linux workloads.
Drive vulnerability remediation activities, including assessment, prioritization, patch planning, execution tracking, exception handling, and closure reporting.
Support Azure Update Manager based patching and maintenance activities, including compliance review, scheduling, execution validation, and post-maintenance checks.
Manage incidents, service requests, changes, problems, and agile stories using ServiceNow and other ITSM or DevOps platforms.
Lead or actively participate in major incident bridges, coordinate technical resolver groups, provide timely updates, and drive restoration with minimum business impact.
Create high-quality RCA documents, support corrective and preventive action tracking, and facilitate Post-Incident Review discussions with stakeholders.
Maintain observability using New Relic, Azure Monitor, Azure Application Insights, Log Analytics, alerts, dashboards, and ServiceNow event or incident workflows.
Collaborate with application, security, network, database, DevOps, and vendor teams to resolve issues, implement improvements, and reduce recurring incidents.
Identify operational gaps, recommend best-practice improvements, update SOPs/runbooks, and contribute to continuous service improvement.
Required Mandatory experience
Minimum 5 8 years of IT infrastructure experience with strong hands-on exposure to Azure cloud operations and L2 production support.
Strong practical experience with Azure services including virtual machines, storage, networking, Azure Active Directory, RBAC, Azure Monitor, Log Analytics, Application Insights, Logic Apps, and Azure Automation.
Strong expertise in Terraform for Azure infrastructure provisioning, module development, state management, reusable templates, and infrastructure lifecycle management.
Strong experience with Ansible for configuration management, patch automation, operational tasks, and repeatable server/application configuration.
Hands-on experience with Azure Update Manager, patch governance, maintenance windows, compliance tracking, and remediation reporting.
Experience in vulnerability remediation across Windows and Linux environments, including coordination with security teams, patch validation, risk-based prioritization, and closure governance.
Good understanding of DevOps concepts, CI/CD pipelines, Git-base