GovCIO is currently hiring for a Senior Site Reliability Engineer (SRE) to design, implement, and maintain highly available, scalable, and resilient infrastructure systems. The ideal candidate will bridge the gap between development and operations, focusing on automation, reliability, and performance optimization across multi-cloud environments. This position is located in Arlington, VA and is a hybrid remote/onsite position.
Responsibilities
Key Responsibilities:
Infrastructure & Automation
Design, deploy, and manage cloud infrastructure using Infrastructure as Code (IaC) principles
Develop and maintain Terraform modules for AWS and Azure environments
Create and manage Ansible playbooks for configuration management and application deployment
Implement CI/CD pipelines using GitHub Actions to automate build, test, and deployment processes
Implement GitOps workflows for declarative infrastructure and application delivery
Build self-service tools and platforms to enable development teams
Reliability & Performance
Establish and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
Implement comprehensive monitoring, logging, and alerting solutions
Conduct capacity planning and performance tuning
Perform root cause analysis and implement preventive measures
Design and execute chaos engineering experiments to validate system resilience
Disaster Recovery & Business Continuity
Design and implement disaster recovery strategies across multi-cloud environments
Develop and maintain backup and restore procedures
Create and test business continuity plans
Implement automated failover mechanisms
Document recovery time objectives (RTO) and recovery point objectives (RPO)
Cloud Operations
Manage and optimize AWS services (EC2, S3, RDS, Lambda, ECS, EKS, CloudWatch, etc.)
Manage and optimize Azure services (VMs, Storage, SQL Database, AKS, Monitor, etc.)
Implement cost optimization strategies and resource tagging
Ensure security best practices and compliance requirements
Manage identity and access management (IAM) policies
Collaboration & Leadership
Participate in on-call rotation and incident response
Collaborate with development teams on architecture and design decisions
Mentor team members on SRE practices and tools
Document systems, processes, and runbooks
Drive continuous improvement initiatives
Qualifications
Required Education and Experience
Bachelor's Degree with 12+ yrs experience
Clearance Level: Active Secret with the ability to obtain and hold DEA suitability
Technical Skills
Cloud Platforms: 3+ years of hands-on experience with AWS and Azure
Infrastructure as Code: Expert-level proficiency with Terraform
Configuration Management: Strong experience with Ansible
Scripting: Proficiency in Python, Bash, or PowerShell
Containerization: Experience with Docker and Kubernetes
Version Control: Strong Git and GitHub workflow knowledge
GitOps: Experience implementing GitOps practices and workflows
Monitoring Tools: Experience with Prometheus, Grafana, ELK Stack, or similar
CI/CD: Hands-on experience with GitHub Actions, Jenkins, GitLab CI, or Azure DevOps
Core Competencies
Deep understanding of Microsoft/Linux systems administration
Strong networking knowledge (TCP/IP, DNS, load balancing, VPN)
Experience with database administration (PostgreSQL, MySQL, SQL Server)
Knowledge of security best practices and compliance frameworks
Understanding of microservices architecture and distributed systems
Experience with disaster recovery planning and execution
Soft Skills
Excellent problem-solving and analytical abilities
Strong communication skills, both written and verbal
Ability to work independently and in team environments
Customer-focused mindset with emphasis on reliability
Adaptability to rapidly changing technologies and requirements
Preferred Qualifications
AWS Certified Solutions Architect or SysOps Administrator
Azure Administrator or Solutions Architect certification
Certified Kubernetes Administrator (CKA)
HashiCorp Certified: Terraform Associate
GitHub Certified or demonstrated expertise with GitHub Enterprise
Experience with service mesh technologies (Istio, Linkerd)
Knowledge of observability platforms (Datadog, New Relic, Dynatrace)
Experience with GitOps tools and practices (ArgoCD, Flux, GitHub Actions for GitOps)
Familiarity with compliance frameworks (SOC 2, HIPAA, FedRAMP)
Previous experience in a DevOps or Platform Engineering role
Posted Salary Range
USD $210,000.00 - USD $230,000.00 /Yr.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: 10443217
- Position Id: 8668
- Posted 4 hours ago