Remote or New York, New York
•
Today
Description WHO WE ARE Come join the company at the center of how the world adopts AI securely. Cyera's mission is to give enterprises the confidence to embrace AI safely - deciding exactly what it can see and do as it reaches deeper into the business. We started by solving the hardest problem in data security: finding and securing data faster and more precisely than anyone thought possible. That foundation is now the essential AI trust infrastructure for the Fortune 1000. We're hiring mission
Full-time
USD 160,000.00 - 200,000.00 per year





Location: New York- Hybrid
Employment Type: Contract
Experience: 15+ Years
Job Summary
We are looking for an experienced Site Reliability Engineer (SRE) to build, maintain, and improve highly reliable, scalable, and secure production systems. The ideal candidate will have strong experience in cloud infrastructure, automation, monitoring, incident management, and DevOps practices.
Key Responsibilities
Design, implement, and maintain highly available and scalable production environments.
Develop automation to reduce manual operational tasks and improve system reliability.
Monitor system performance, availability, capacity, and overall health.
Define and maintain SLIs, SLOs, and SLAs for critical applications and services.
Participate in incident response, troubleshooting, root-cause analysis, and post-incident reviews.
Build and maintain CI/CD pipelines for reliable and automated software deployments.
Implement infrastructure as code using tools such as Terraform, CloudFormation, or Ansible.
Manage and optimize cloud infrastructure across AWS, Azure, or Google Cloud Platform.
Develop monitoring, logging, and alerting solutions using tools such as Prometheus, Grafana, Datadog, Splunk, or ELK.
Improve system performance, scalability, resilience, and disaster recovery capabilities.
Collaborate with software engineering, DevOps, security, and infrastructure teams.
Establish best practices around reliability, observability, deployment, and operational readiness.
Automate infrastructure provisioning, configuration management, and operational workflows.
Required Skills
5+ years of experience in SRE, DevOps, Cloud Engineering, or Infrastructure Engineering.
Strong Linux/Unix administration and troubleshooting skills.
Strong experience with at least one major cloud platform: AWS, Azure, or Google Cloud Platform.
Hands-on experience with Kubernetes and Docker.
Strong scripting/programming skills in Python, Bash, Go, or similar.
Experience with Terraform or other Infrastructure-as-Code tools.
Strong knowledge of CI/CD pipelines and tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, or ELK.
Understanding of networking, DNS, HTTP/HTTPS, TCP/IP, load balancing, and security fundamentals.
Experience with incident management, troubleshooting, and root-cause analysis.
Strong understanding of distributed systems, scalability, availability, and fault tolerance.
Preferred Skills-
Experience with AWS EKS, ECS, EC2, Lambda, CloudWatch, IAM, S3, or equivalent cloud services.
SRE OR "Site Reliability Engineer" OR "Site Reliability Engineering" OR DevOps OR "Cloud Engineer" AND AWS OR Azure OR Google Cloud Platform AND Kubernetes AND Terraform AND Docker AND Linux AND Python AND "CI/CD" AND Prometheus OR Grafana OR Datadog.