About the Team
The SRE Platform Engineering team builds and operates the infrastructure that
powers our cloud. We focus on delivering reliable, scalable, and simple
platforms that enable product teams to move quickly while meeting the
requirements of regulated environments such as FedRAMP High and DoD
IL5.
About the Role
We’re seeking a Senior Site Reliability Engineer to support the development and
operation of our Kubernetes-based platform in regulated environments. In this
role, you’ll collaborate with engineers across the stack to ensure the platform
remains reliable, observable, compliant, and easy for developers to use—while
maintaining team velocity.
This is a hands-on position centered on strengthening reliability, advancing
automation, and driving operational excellence within environments that have
rigorous security and compliance standards.
What You Will Do
Own and operate components of the Kubernetes platform, including
deployment, upgrades, and maintenance
Contribute to the design and implementation of scalable and reliable
platform features
Build and improve automation, tooling, and CI/CD workflows to reduce
operational overhead
Monitor system health and respond to issues; participate in on-call
rotations and incident response
Contribute to defining and tracking SLIs, SLOs, and error budgets
Support FedRAMP High / IL5 compliance efforts, including system
hardening, documentation, and audit readiness
Collaborate with senior engineers, technical leaders, and cross-
functional teams to deliver platform improvements
Participate in on-call rotations supporting customer requests and paging
alerts
Participate in post-incident reviews and implement follow-up
improvements
What You Bring
8+ years of experience in SRE, DevOps, or platform engineering
Hands-on experience with Kubernetes and containerized workloads in
production
Hands-on experience with cloud platforms (AWS, Azure, or similar;
GovCloud experience a plus)
Strong working knowledge of Linux systems, networking, and
distributed systems fundamentals
Experience with Infrastructure as Code (e.g., Terraform)
Ability to write and maintain scripts or services (e.g., Python, Go, Bash)
Experience with monitoring and observability tools (Prometheus,
Grafana, logging systems)
Basic understanding of security and compliance concepts (e.g., NIST
800-53, STIGs, RMF)
Nice to Have
Exposure to FedRAMP High or DoD IL5 environments
Experience with CI/CD systems and deployment automation (e.g.,
ArgoCD)
Familiarity with container security and vulnerability management
Experience working in regulated or audited environments
How You Work
You take ownership of well-scoped problems and drive them to
completion
You are comfortable working independently, while seeking input on more
complex decisions
You collaborate effectively with senior engineers and technical leaders
You prioritize simplicity, reliability, and maintainability in your work
You are eager to learn and grow into broader technical ownership over