Job Summary
We are seeking a motivated Site Reliability Engineer (Contractor) with 3 to 5 years of experience in automation, cloud infrastructure, and production operations. The ideal candidate has a strong automation-first mindset and enjoys solving complex operational challenges. This role focuses on improving reliability, observability, operational efficiency, and automation across on-premises and cloud platforms.
Key Responsibilities
Automation & Platform Engineering
• Develop Python-based automation solutions to reduce manual operational effort.
• Automate infrastructure management across Linux, Windows, Kubernetes, Google Cloud Platform, and cloud-native environments.
• Integrate tools and platforms through APIs and client libraries.
• Assist in implementing infrastructure automation using Ansible, Terraform, or similar technologies.
• Support CI/CD automation and deployment reliability initiatives.
Reliability & Operations
• Monitor and maintain production systems to meet reliability and availability objectives.
• Participate in incident response, troubleshooting, and root cause analysis activities.
• Develop automation and operational improvements to prevent recurring issues.
• Support disaster recovery, failover testing, and operational readiness activities.
• Perform performance analysis and system health reviews.
Observability & Monitoring
• Build and maintain dashboards, alerts, and monitoring solutions using Splunk, Grafana, Prometheus, Google Cloud Platform Operations Suite, or similar tools.
• Improve visibility into application and infrastructure health through metrics, logs, and traces.
• Investigate alerts and identify opportunities to reduce noise and improve detection.
AIOps & Intelligence
• Explore AI/ML-driven operational improvements such as anomaly detection, intelligent alerting, and log analytics.
• Assist in developing automation solutions that leverage AI to improve operational efficiency.
• Participate in evaluating emerging AIOps capabilities and observability technologies.
Required Qualifications
• Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience.
• 3 to 5 years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or Platform Engineering.
• Strong programming skills in Python for automation and tooling development.
• Experience supporting Kubernetes and cloud platforms (Google Cloud Platform, AWS, or Azure).
• Familiarity with infrastructure automation and configuration management tools.
• Experience with monitoring and observability platforms such as Splunk, Grafana, Prometheus, Datadog, or similar.
• Understanding of Linux systems, networking, and distributed applications.
• Strong analytical, troubleshooting, and problem-solving skills.
• Ability to work effectively in fast-paced, mission-critical environments.