Job Title: Lead Site Reliability Engineer
Work Location: Remote
Remote F2F Interview or Laptop Pick up Required San Francisco, Arlington, VA, Denver, CO, Chicago, Boston, NYC, Houston, Miami, Los Angeles, Seattle, Dallas, Minneapolis, Minn, Atlanta, GA, Birmingham, MI or Irvine, CA
Long Term Contract
If you find this opportunity suitable, kindly share your updated resume. Also, please let me know a good time to connect with you for a quick discussion.
Looking forward to hearing from you!
Job Description
The Lead Site Reliability Engineer is a member of the SRE team, where you'll play a pivotal role in ensuring the reliability, security, and efficiency of our cloud environments and the innovative products within our Enterprise Imaging suite of solutions. In this role, you will enjoy the flexibility to work remotely1 from anywhere within the U.S. as you take on some tough challenges.
Your Impact:
- Drive our SRE Practice: Contribute to the improvement of our SRE practices; Lead incident response; Ensure high availability, scalability, and performance of cloud environments.
- Develop and maintain software: Design, code, test, and operate software using Python & Node.js.
- Automate Infrastructure & Operations: Develop Infrastructure as Code using Terraform & GitHub Actions while adhering to best practices.
- Collaborate, Influence, & Mentor: Work closely with development and SRE teams to improve system design, advocate for reliability, and mentor fellow engineers.
Your Skills & Experience:
- 8+ years of experience in automation and administration of Public Cloud systems (Google Cloud Platform preferred), including compute, network, storage, and IAM.
- 3+ years of experience in Site Reliability Engineering, including system design, error budgets, monitoring, alerting, observability, incident response, and software development (Python or Node.js preferred)
- 3+ years of CI/CD (GitHub Actions preferred) & Infrastructure as Code experience with Terraform
- 2+ years of experience managing Kubernetes environments (EKS, AKS, GKE, or self-hosted)
- You are available to work rotating 24x7 primary and secondary on-call shifts.
Preferred Qualifications:
- Bachelor s degree in information systems, Computer Science, Engineering, or related field or equivalent certification
- 3+ years of experience in public cloud networking and security, including architecture, implementation, and best practices across major cloud platforms
- 3+ years of hands-on experience in designing, developing, and maintaining distributed systems at scale