Title : Site Reliability Engineer (SRE)
Location : Dallas, TX - Hybrid
Duration : 6 months
Position : W2 Position
Rate :$51/hr on W2
Relevant Experience (in Yrs.): 8+
Detailed Job Description:
· 10+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, Linux administration, production operations, or platform engineering.
· Strong hands-on experience designing, operating, and supporting cloud infrastructure across AWS, Azure, and Google Cloud Platform.
· Deep experience with Kubernetes platforms such as EKS, AKS, GKE, and containerization using Docker.
· Experience building and maintaining CI/CD pipelines using Jenkins, GitHub Actions.
· Strong observability experience with Prometheus, Grafana, ELK Stack, OpenSearch, Log Analytics, Application Insights, and Google Cloud Platform Cloud Monitoring.
· Experience with disaster recovery, high availability, backup automation, multi-region failover, and recovery validation.
· Hands-on scripting and automation experience using Python, Bash, PowerShell, and Ansible.
· Linux systems administration experience across enterprise production environments.
Key Responsibilities:
· Site Reliability Engineering and Reliability Governance (SLIs, SLOs, error budgets, reliability reviews, and blameless postmortem practices, MTTD, MTTR, recurring incidents)
· Kubernetes, Containers, and Platform Engineering (Kubernetes, EKS, AKS, GKE, ECS, and Docker-based platforms)
· Infrastructure as Code and Automation (Terraform)
· CI/CD and Release Reliability (GitHub Actions, blue-green, canary, rolling deployments, automated rollback, deployment validation, and automated testing)
· Observability, Monitoring, and Logging (Prometheus, Grafana)
· Disaster Recovery, High Availability, and Resilience
· Security, Compliance, and Cloud Governance
· Linux Systems Administration and Production Support
Have Skills
· Google Cloud Platform (Google Cloud Platform)
· SRE
· Kubernetes & Docker
· Terraform & Infrastructure Automation
· CI/CD & Release Engineering
· Monitoring