Site Reliability Engineer (SRE)
Location: Ann Arbor, MI (4 days onsite, Friday remote)
Employment Type: Contract-to-Hire (W2 Only)
Pay Rate: Up to $60/hour W2
About the Role
We are seeking a Site Reliability Engineer (SRE) to join a team responsible for maintaining and enhancing the reliability, performance, and scalability of critical Linux-based infrastructure supporting high-availability enterprise applications.
This role is focused on reliability engineering, observability, system performance, and continuous improvement. Candidates with a strong Linux engineering background and experience in monitoring and cloud environments will be highly successful in this position.
Key Responsibilities
- Maintain and support enterprise Linux environments in virtualized infrastructure.
- Monitor system health, performance, and availability across production environments.
- Implement and manage observability solutions, including logging, monitoring, alerting, and incident response.
- Troubleshoot complex infrastructure and application issues to improve system reliability.
- Identify opportunities to reduce operational toil through automation and process improvements.
- Collaborate with engineering and operations teams to optimize platform stability and performance.
- Support cloud-based infrastructure and reliability initiatives.
Required Skills & Experience
- Strong experience with Enterprise Linux administration and engineering (RHEL preferred).
- Experience working within virtualized environments such as VMware or similar platforms.
- Hands-on experience with at least one observability or monitoring platform, including:
- Datadog
- Splunk
- Grafana
- New Relic
- Similar tools
- Experience creating and managing monitoring dashboards, alerts, and log analysis.
- Public cloud experience, preferably Azure (AWS or Google Cloud Platform experience also considered).
- Strong troubleshooting and problem-solving skills focused on system reliability and performance.
Preferred Qualifications
- Experience supporting large-scale production environments.
- Background in Site Reliability Engineering, Infrastructure Engineering, or Systems Engineering.
- Familiarity with automation and operational efficiency initiatives.
- Understanding of reliability, monitoring, incident management, and performance optimization best practices.
What We're Looking For
This position is ideal for candidates who are passionate about system reliability, observability, and platform engineering. Rather than focusing primarily on CI/CD pipelines and deployments, the successful candidate will have a strong interest in understanding system behavior, reducing operational complexity, and improving overall service reliability.
Interview Process
- Virtual Interview
- Onsite Interview
Work Schedule: Monday through Thursday onsite, Friday remote.