Title: Mid-Senior Site Reliability Engineer Kubernetes Platform
Location: San Jose, CA/Remote
Duration: Full Time
Client: Confidential
Salary Range: $120k-$130K/year
Visa: ,
Experience required: 8 - 10 years of experience
Interview Mode: Video
Job Description
Must Have Technical/Functional Skills:
8+ years of experience in SRE, DevOps, or platform engineering
Hands-on experience with Kubernetes and containerized workloads in production
Hands-on experience with cloud platforms (AWS, Azure, or similar; GovCloud experience a plus)
Strong working knowledge of Linux systems, networking, and distributed systems fundamentals
Experience with Infrastructure as Code (e.g., Terraform)
Ability to write and maintain scripts or services (e.g., Python, Go, Bash)
Experience with monitoring and observability tools (Prometheus, Grafana, logging systems)
Basic understanding of security and compliance concepts (e.g., NIST 800-53, STIGs, RMF)
Roles & Responsibilities:
Own and operate components of the Kubernetes platform, including deployment, upgrades, and maintenance
Contribute to the design and implementation of scalable and reliable platform features
Build and improve automation, tooling, and CI/CD workflows to reduce operational overhead
Monitor system health and respond to issues; participate in on-call rotations and incident response
Contribute to defining and tracking SLIs, SLOs, and error budgets
Support FedRAMP High / IL5 compliance efforts, including system hardening, documentation, and audit readiness
Collaborate with senior engineers, technical leaders, and cross-functional teams to deliver platform improvements
Participate in on-call rotations supporting customer requests and paging alerts
Participate in post-incident reviews and implement follow-up improvements Nice to Have
Experience with service mesh technologies (Istio, Linkerd)
Familiarity with policy-as-code (OPA/Gatekeeper, Kyverno)
Experience with GitOps workflows
Exposure to multi-cluster or hybrid cloud architectures
Knowledge of FIPS-compliant systems or DoD Cloud SRG
Relevant certifications (CKA, CKS, cloud provider certs, Security+)
Nice to have skills:
Exposure to FedRAMP High or DoD IL5 environments
Experience with CI/CD systems and deployment automation (e.g., ArgoCD)
Familiarity with container security and vulnerability management
Experience working in regulated or audited environments
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: 91171926
- Position Id: OOJ - 1264-268-1785438517
- Posted 18 hours ago