Incredible opportunity to join one of the fastest growing AI Security startups in the world // VC Backed Startup experience is required // Fully Remote
This Jobot Job is hosted by: Craig Rosecrans
Are you a fit? Easy Apply now by clicking the "Apply Now" button and sending us your resume.
Salary: $220,000 - $260,000 per year
A bit about us:
We are partnering with a rapidly growing, venture-backed technology company operating at the intersection of Artificial Intelligence, Cybersecurity, and Cloud Infrastructure to hire a Director of Engineering focused on Site Reliability Engineering.
The company's technology protects mission-critical AI systems and is deployed across sophisticated enterprise and government environments. As adoption continues to accelerate, the organization is looking for an exceptional engineering leader to own and evolve the reliability, scalability, resiliency, and operational excellence of its platform.
This is not a traditional people-management-only Director position.
We are looking for someone who combines exceptional leadership ability with deep, current technical expertise across SRE, cloud infrastructure, distributed systems, platform engineering, automation, observability, and production operations.
You will inherit a highly technical engineering team, and credibility matters.
The engineers reporting to this person need to trust that their leader understands the technology at their level, can challenge their thinking, can make difficult architectural decisions, and-when necessary-can sit beside them during a complex production incident and help solve the problem.
You don't need to write production code every day.
But you absolutely need to be capable of doing it.
Why join us?
This is an opportunity to join a rapidly scaling company tackling one of the most important emerging challenges in technology: protecting the AI systems enterprises and government organizations increasingly depend upon.
You will have significant ownership over the infrastructure and reliability strategy supporting a sophisticated cybersecurity platform while leading an experienced technical team.
For the right engineering leader, this represents a rare combination of:
AI + Cybersecurity + Cloud Infrastructure + Distributed Systems + Engineering Leadership + Mission-Critical Reliability.
Compensation: $230,000-$260,000 Base Salary + 10% Annual Bonus + Stock Options
Location: Fully Remote - United States
Job Details
What You'll Own
Lead a Highly Technical SRE Organization
Lead, mentor, develop, and grow a team responsible for the reliability and operational excellence of a sophisticated AI security platform.
You will:
Develop and mentor senior-level SRE, Infrastructure, and Platform engineers.
Establish clear expectations around ownership, execution, technical quality, and accountability.
Build an engineering culture centered around reliability, automation, continuous improvement, and operational excellence.
Recruit and retain exceptional engineering talent as the organization scales.
Provide meaningful technical mentorship rather than simply managing projects and people.
Make thoughtful, decisive engineering decisions in a fast-moving startup environment where perfect information isn't always available.
Balance short-term operational requirements with long-term platform and infrastructure investments.
This organization values leaders who can move quickly, make difficult decisions, and create clarity in ambiguous environments.
Remain Deeply Technical
This Director will remain close to the technology and serve as a senior technical authority across SRE and infrastructure.
You should be capable of contributing meaningfully to conversations involving:
Cloud architecture
AWS
Kubernetes and container orchestration
Linux
Networking
Distributed systems
Infrastructure as Code
CI/CD
Observability
Production troubleshooting
Reliability engineering
Infrastructure automation
Security
Performance
Scalability
Resiliency and disaster recovery
You will partner with senior engineers on architectural decisions, identify systemic risks, challenge assumptions, and help troubleshoot particularly complex production problems.
Your engineers should consider you one of the strongest technical resources in the organization-not simply the person managing the strongest technical resources.
Site Reliability & Production Engineering
Own and continuously improve how production systems are designed, deployed, monitored, and operated.
Responsibilities will include:
Establishing and evolving SLIs, SLOs, availability objectives, and error budgets.
Improving system availability, fault tolerance, scalability, and performance.
Building mature observability practices across metrics, logs, traces, dashboards, and alerting.
Leading improvements to incident response, escalation, root-cause analysis, and postmortems.
Reducing Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR).
Improving capacity planning and infrastructure forecasting.
Identifying systemic reliability risks before they become customer-impacting incidents.
Developing production-readiness standards across engineering.
Reducing operational toil through automation.
The goal is not simply to respond effectively when systems fail.
It is to engineer systems so failures become less frequent, less severe, easier to detect, and faster to recover from.
Cloud, Platform & Infrastructure Engineering
Help evolve the infrastructure supporting a rapidly scaling, security-focused technology platform.
Relevant experience may include:
AWS and cloud-native infrastructure
Kubernetes
Docker / containerized workloads
Infrastructure as Code
Terraform or comparable technologies
CI/CD and software delivery automation
Linux systems engineering
Networking and distributed systems
Cloud security and IAM
Secrets and key management
Configuration management
Observability and monitoring platforms
Automated provisioning
Performance optimization
Capacity management
High-availability architecture
Experience operating complex distributed SaaS platforms at scale is strongly preferred.
Air-Gapped & Disconnected Environments
Experience supporting air-gapped, disconnected, restricted, or highly regulated environments will be particularly valuable.
Some customers operate environments where traditional cloud assumptions simply don't apply.
We are especially interested in engineering leaders who understand the complexities involved in deploying and maintaining sophisticated platforms when infrastructure may have:
No direct internet connectivity
Restricted ingress and egress
Limited access to external APIs and SaaS services
Private container and artifact registries
Offline software and dependency management
Controlled software-update processes
Restricted remote access
Strict security boundaries
Customer-managed infrastructure
On-premise or hybrid deployments
Specialized observability and monitoring requirements
Experience solving challenges involving deployment, upgrades, dependency management, observability, troubleshooting, security, and lifecycle management in disconnected environments would be highly relevant.
Experience supporting government, defense, intelligence, financial services, critical infrastructure, or other highly regulated environments is also valuable.
Security & Resilience
Security and reliability are inseparable within this environment.
You will partner closely with Security, Engineering, Product, and Architecture teams to maintain strong operational practices around:
Infrastructure security
Production access
Identity and permissions
Secrets management
Encryption
Vulnerability management
Auditability
Disaster recovery
Business continuity
Availability commitments
RPO/RTO objectives
Enterprise security and compliance requirements
Experience operating within SOC 2, ISO 27001, FedRAMP, NIST, or similarly controlled environments is valuable.
Developer Experience & Automation
A great SRE organization should make software engineers faster-not create another layer of process.
This leader will help create the platforms, automation, tooling, and standards that allow engineering teams to safely build, deploy, and operate software.
You will drive improvements around:
Self-service infrastructure
Infrastructure automation
Automated deployments
CI/CD
Production readiness
Service ownership
Observability
Developer tooling
Operational documentation
Reduction of repetitive manual work
You should naturally look at repetitive operational processes and ask:
"Why are humans still doing this?"
What We're Looking For
The strongest candidates will bring:
Significant experience across Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, DevOps, Cloud Engineering, or Production Engineering.
Proven leadership of highly technical engineering teams.
Deep hands-on experience operating complex production systems.
Strong cloud infrastructure expertise, particularly AWS.
Strong Kubernetes/container orchestration experience.
Excellent Linux and networking fundamentals.
Experience with Infrastructure as Code and infrastructure automation.
Deep understanding of distributed systems.
Experience designing and operating highly available, fault-tolerant systems.
Strong observability and production-monitoring experience.
Significant incident-response and production-troubleshooting experience.
Experience building or improving SLO/SLI practices.
Strong CI/CD and deployment-automation knowledge.
Experience balancing reliability, security, scalability, performance, and cost.
Strong understanding of security within cloud and production environments.
Experience scaling engineering practices in a rapidly growing technology organization.
Experience with air-gapped, disconnected, government, defense, highly regulated, or on-premise environments is a significant plus.
The Leadership Profile
Titles are less important than technical depth and leadership ability.
Potential backgrounds could include:
Director of SRE | Director of Platform Engineering | Director of Infrastructure Engineering | Head of SRE | Head of Platform | Head of Infrastructure | Senior Engineering Manager - SRE | Senior Engineering Manager - Platform | Principal SRE / Platform Engineer with significant leadership experience
The ideal candidate has reached engineering leadership without losing their engineering instincts.
You have managed and developed exceptional engineers, but you can still dive deeply into an architecture discussion.
You understand organizational strategy, but you also understand what is happening inside the infrastructure.
You can communicate with executives, but you can also earn the respect of a Principal Engineer.
And when a critical production system is failing, your team wants you in the room.
Interested in hearing more? Easy Apply now by clicking the "Apply Now" button.
Jobot is an Equal Opportunity Employer. We provide an inclusive work environment that celebrates diversity and all qualified candidates receive consideration for employment without regard to race, color, sex, sexual orientation, gender identity, religion, national origin, age (40 and over), disability, military status, genetic information or any other basis protected by applicable federal, state, or local laws. Jobot also prohibits harassment of applicants or employees based on any of these protected categories. It is Jobot's policy to comply with all applicable federal, state and local laws respecting consideration of unemployment status in making hiring decisions.
Sometimes Jobot is required to perform background checks with your authorization. Jobot will consider qualified candidates with criminal histories in a manner consistent with any applicable federal, state, or local law regarding criminal backgrounds, including but not limited to the Los Angeles Fair Chance Initiative for Hiring and the San Francisco Fair Chance Ordinance.
Information collected and processed as part of your Jobot candidate profile, and any job applications, resumes, or other information you choose to submit is subject to Jobot's Privacy Policy, as well as the Jobot California Worker Privacy Notice and Jobot Notice Regarding Automated Employment Decision Tools which are available at jobot.com/legal.
By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Jobot, and/or its agents and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy here: jobot.com/privacy-policy
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
- Dice Id: 91113390
- Position Id: 857737951
- Posted 2 hours ago