IT Applications Solutions Architect Mid / SRE Engineer
Location: Remote with Quarterly PI Planning Travel
Duration: 12 Months+
Client is looking for a DevOps/Site Reliability Engineer (SRE) with strong AWS cloud, automation, and production support experience. This is a hands-on engineering role.The ideal candidate will have experience building CI/CD pipelines (GitHub Actions, Jenkins, AWS CodePipeline), Infrastructure as Code (Terraform, CloudFormation, or AWS CDK), scripting with Python, and supporting cloud infrastructure in AWS (Azure exposure is a plus). Candidates should also have experience with Dynatrace observability, incident response, root cause analysis, ServiceNow, Linux, containers (Docker/Kubernetes/ECS), and SRE concepts such as SLIs, SLOs, error budgets, resiliency, and performance monitoring. Strong troubleshooting skills, automation experience, and the ability to support production environments and participate in on-call rotations are key to success.
| Labor Category / Position Title | IT Applications Solutions Architect Mid / SRE Engineer |
| Skill Set | - Deployment & Automation
- Implement CI/CD pipelines using tools such as GitHub Actions, AWS Code Pipeline, and Jenkins - Automate infrastructure provisioning through Infrastructure-as-Code (IaC) using Terraform, CloudFormation, or AWS CDK. - Develop automation scripts and self-service tools to enhance operational efficiency. - Leveraging Dynatrace Observability Platform
Demonstrated expertise in - Standardized installation through automation - Integrating with CI/CD pipelines - Enforcing Tagging and Metadata standards - Use of environment-aware configuration - Implement distributed tracing with appropriate context propagation. - Optimize alerts, create dashboards, alerts, and anomaly detectors - Incident Management & Response
- Proficient in ITIL framework and ITSM tools such as ServiceNow. - Production on-call responder with strong troubleshooting capabilities. - Develop RCA documentation, and Knowledge articles - Apply SRE principles, including SLIs, SLOs, and error budgets. - Capacity Planning & Performance
- Implement operational cost optimization initiatives. - Configure and maintain auto-scaling policies and thresholds. - Develop Resiliency Test plans and support Performance testing. - Security & Compliance Implementation
- Manage service accounts and access permissions - Create, deploy, and manage digital certificates. - Respond to security incidents and execute remediation tasks effectively. - Education & Experience
- Bachelor s degree in computer science, Engineering, or related field - 2 to 4 years of experience in DevOps, SRE, or infrastructure roles - Mid-level proficiency in Python or other scripting languages. - Mid-level proficiency in Configuration management tool including Ansible. - Practical experience with cloud platforms AWS and Azure. - Knowledge of containerization (Docker, Kubernetes/ECS). - Knowledge of Linux systems and networking. - Knowledge of relational, cloud, and NoSQL databases. - Excellent written and verbal communication skills. - Demonstrated ability to work independently and manage priorities. - Availability to work outside of standard business hours as required. |