Site Reliability Engineer
Job Summary
The Site Reliability Engineer will design, build, automate, and support highly available, secure, scalable, and resilient production systems. The role combines cloud infrastructure, DevOps, automation, monitoring, incident management, and reliability engineering to ensure critical LexisNexis applications meet availability and performance objectives.
Key Responsibilities
- Design, implement, and maintain highly available cloud infrastructure.
- Develop and maintain Infrastructure as Code (IaC) using tools such as Terraform, Ansible, or CloudFormation.
- Build and maintain CI/CD pipelines and automated deployment processes.
- Support cloud migration initiatives from on-premises environments to public cloud.
- Manage containerized environments using Docker and Kubernetes.
- Implement monitoring, logging, alerting, and observability solutions.
- Establish service-level baselines and reliability metrics.
- Participate in 24/7 on-call support and production incident response.
- Troubleshoot complex infrastructure and application issues.
- Lead or participate in Root Cause Analysis (RCA), post-incident reviews, and problem management.
- Develop automation to reduce manual operational work and eliminate repetitive tasks.
- Participate in disaster recovery testing and improve system resilience.
- Support blue-green and other highly reliable deployment strategies.
- Identify and remediate security vulnerabilities and infrastructure risks.
- Work closely with Development, QA, IT Operations, Customer Operations, and Project Management teams.
- Maintain technical documentation, operational procedures, and SRE knowledge bases.
- Contribute to infrastructure cost optimization and continuous improvement.
Required Technical Skills
- 5+ years of experience in SRE, DevOps, Cloud Engineering, Platform Engineering, or Systems Engineeringfor senior-level roles.
- Strong experience with Azure.
- Terraform and Infrastructure as Code.
- Kubernetes and Docker.
- CI/CD technologies such as GitHub Actions or equivalent.
- Monitoring and observability tools such as Grafana, Splunk, Elastic Stack, Pingdom, or Uptrends.
- Strong Linux/system administration and troubleshooting skills.
- Experience with databases such as MySQL or Microsoft SQL Server.
- Understanding of networking, security, availability, scalability, and disaster recovery.
- Strong scripting/programming skills, such as Python, Bash, or similar.
specifically emphasize Terraform, cloud platforms, Kubernetes/containerized workloads, monitoring, troubleshooting, 24/7 support, resilient application stacks, and incident/RCA processes.
Education
- Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent practical experience.
- Relevant cloud/DevOps certifications are a plus.
Thanks & Regards,
Mansi Jain