Director of Engineering, SRE (AI Security / Startup)

Boston, MA, US • Posted 2 hours ago • Updated 2 hours ago
Full Time
On-site
USD $220,000.00 - 260,000.00 per year
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Venture Capital
  • Videoconferencing
  • People Management
  • Reporting
  • System Integration Testing
  • Cyber Security
  • Accountability
  • Continuous Improvement
  • Operational Excellence
  • Mentorship
  • Startups
  • Investments
  • Clarity
  • Cloud Architecture
  • Budget
  • Dashboard
  • Root Cause Analysis
  • MEAN Stack
  • Recovery
  • Forecasting
  • Docker
  • Terraform
  • Systems Engineering
  • Cloud Security
  • Configuration Management
  • Provisioning
  • Performance Tuning
  • Capacity Management
  • High Availability
  • Internet
  • SaaS
  • Remote Access
  • Lifecycle Management
  • Financial Services
  • Security Engineering
  • Management
  • Encryption
  • Vulnerability Management
  • Disaster Recovery
  • Business Continuity Planning
  • RPO
  • Regulatory Compliance
  • System On A Chip
  • ISO/IEC 27001:2005
  • FedRAMP
  • Documentation
  • Reliability Engineering
  • DevOps
  • Production Engineering
  • IaaS
  • Amazon Web Services
  • Kubernetes
  • Orchestration
  • Linux
  • Computer Networking
  • Incident Management
  • Continuous Integration
  • Continuous Delivery
  • Scalability
  • Cloud Computing
  • Adobe AIR
  • Leadership
  • Business Strategy
  • Military
  • SAP BASIS
  • Authorization
  • Law
  • LOS
  • Recruiting
  • Legal
  • Artificial Intelligence
  • Privacy

Summary

Incredible opportunity to join one of the fastest growing AI Security startups in the world // VC Backed Startup experience is required // Fully Remote

This Jobot Job is hosted by: Craig Rosecrans
Are you a fit? Easy Apply now by clicking the "Apply Now" button and sending us your resume.
Salary: $220,000 - $260,000 per year

A bit about us:

We are partnering with a rapidly growing, venture-backed technology company operating at the intersection of Artificial Intelligence, Cybersecurity, and Cloud Infrastructure to hire a Director of Engineering focused on Site Reliability Engineering.

The company's technology protects mission-critical AI systems and is deployed across sophisticated enterprise and government environments. As adoption continues to accelerate, the organization is looking for an exceptional engineering leader to own and evolve the reliability, scalability, resiliency, and operational excellence of its platform.

This is not a traditional people-management-only Director position.

We are looking for someone who combines exceptional leadership ability with deep, current technical expertise across SRE, cloud infrastructure, distributed systems, platform engineering, automation, observability, and production operations.

You will inherit a highly technical engineering team, and credibility matters.

The engineers reporting to this person need to trust that their leader understands the technology at their level, can challenge their thinking, can make difficult architectural decisions, and-when necessary-can sit beside them during a complex production incident and help solve the problem.

You don't need to write production code every day.

But you absolutely need to be capable of doing it.

Why join us?

This is an opportunity to join a rapidly scaling company tackling one of the most important emerging challenges in technology: protecting the AI systems enterprises and government organizations increasingly depend upon.

You will have significant ownership over the infrastructure and reliability strategy supporting a sophisticated cybersecurity platform while leading an experienced technical team.

For the right engineering leader, this represents a rare combination of:

AI + Cybersecurity + Cloud Infrastructure + Distributed Systems + Engineering Leadership + Mission-Critical Reliability.

Compensation: $230,000-$260,000 Base Salary + 10% Annual Bonus + Stock Options

Location: Fully Remote - United States

Job Details

What You'll Own

Lead a Highly Technical SRE Organization

Lead, mentor, develop, and grow a team responsible for the reliability and operational excellence of a sophisticated AI security platform.

You will:

Develop and mentor senior-level SRE, Infrastructure, and Platform engineers.

Establish clear expectations around ownership, execution, technical quality, and accountability.

Build an engineering culture centered around reliability, automation, continuous improvement, and operational excellence.

Recruit and retain exceptional engineering talent as the organization scales.

Provide meaningful technical mentorship rather than simply managing projects and people.

Make thoughtful, decisive engineering decisions in a fast-moving startup environment where perfect information isn't always available.

Balance short-term operational requirements with long-term platform and infrastructure investments.

This organization values leaders who can move quickly, make difficult decisions, and create clarity in ambiguous environments.

Remain Deeply Technical

This Director will remain close to the technology and serve as a senior technical authority across SRE and infrastructure.

You should be capable of contributing meaningfully to conversations involving:

Cloud architecture

AWS

Kubernetes and container orchestration

Linux

Networking

Distributed systems

Infrastructure as Code

CI/CD

Observability

Production troubleshooting

Reliability engineering

Infrastructure automation

Security

Performance

Scalability

Resiliency and disaster recovery

You will partner with senior engineers on architectural decisions, identify systemic risks, challenge assumptions, and help troubleshoot particularly complex production problems.

Your engineers should consider you one of the strongest technical resources in the organization-not simply the person managing the strongest technical resources.

Site Reliability & Production Engineering

Own and continuously improve how production systems are designed, deployed, monitored, and operated.

Responsibilities will include:

Establishing and evolving SLIs, SLOs, availability objectives, and error budgets.

Improving system availability, fault tolerance, scalability, and performance.

Building mature observability practices across metrics, logs, traces, dashboards, and alerting.

Leading improvements to incident response, escalation, root-cause analysis, and postmortems.

Reducing Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR).

Improving capacity planning and infrastructure forecasting.

Identifying systemic reliability risks before they become customer-impacting incidents.

Developing production-readiness standards across engineering.

Reducing operational toil through automation.

The goal is not simply to respond effectively when systems fail.

It is to engineer systems so failures become less frequent, less severe, easier to detect, and faster to recover from.

Cloud, Platform & Infrastructure Engineering

Help evolve the infrastructure supporting a rapidly scaling, security-focused technology platform.

Relevant experience may include:

AWS and cloud-native infrastructure

Kubernetes

Docker / containerized workloads

Infrastructure as Code

Terraform or comparable technologies

CI/CD and software delivery automation

Linux systems engineering

Networking and distributed systems

Cloud security and IAM

Secrets and key management

Configuration management

Observability and monitoring platforms

Automated provisioning

Performance optimization

Capacity management

High-availability architecture

Experience operating complex distributed SaaS platforms at scale is strongly preferred.

Air-Gapped & Disconnected Environments

Experience supporting air-gapped, disconnected, restricted, or highly regulated environments will be particularly valuable.

Some customers operate environments where traditional cloud assumptions simply don't apply.

We are especially interested in engineering leaders who understand the complexities involved in deploying and maintaining sophisticated platforms when infrastructure may have:

No direct internet connectivity

Restricted ingress and egress

Limited access to external APIs and SaaS services

Private container and artifact registries

Offline software and dependency management

Controlled software-update processes

Restricted remote access

Strict security boundaries

Customer-managed infrastructure

On-premise or hybrid deployments

Specialized observability and monitoring requirements

Experience solving challenges involving deployment, upgrades, dependency management, observability, troubleshooting, security, and lifecycle management in disconnected environments would be highly relevant.

Experience supporting government, defense, intelligence, financial services, critical infrastructure, or other highly regulated environments is also valuable.

Security & Resilience

Security and reliability are inseparable within this environment.

You will partner closely with Security, Engineering, Product, and Architecture teams to maintain strong operational practices around:

Infrastructure security

Production access

Identity and permissions

Secrets management

Encryption

Vulnerability management

Auditability

Disaster recovery

Business continuity

Availability commitments

RPO/RTO objectives

Enterprise security and compliance requirements

Experience operating within SOC 2, ISO 27001, FedRAMP, NIST, or similarly controlled environments is valuable.

Developer Experience & Automation

A great SRE organization should make software engineers faster-not create another layer of process.

This leader will help create the platforms, automation, tooling, and standards that allow engineering teams to safely build, deploy, and operate software.

You will drive improvements around:

Self-service infrastructure

Infrastructure automation

Automated deployments

CI/CD

Production readiness

Service ownership

Observability

Developer tooling

Operational documentation

Reduction of repetitive manual work

You should naturally look at repetitive operational processes and ask:

"Why are humans still doing this?"

What We're Looking For

The strongest candidates will bring:

Significant experience across Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, DevOps, Cloud Engineering, or Production Engineering.

Proven leadership of highly technical engineering teams.

Deep hands-on experience operating complex production systems.

Strong cloud infrastructure expertise, particularly AWS.

Strong Kubernetes/container orchestration experience.

Excellent Linux and networking fundamentals.

Experience with Infrastructure as Code and infrastructure automation.

Deep understanding of distributed systems.

Experience designing and operating highly available, fault-tolerant systems.

Strong observability and production-monitoring experience.

Significant incident-response and production-troubleshooting experience.

Experience building or improving SLO/SLI practices.

Strong CI/CD and deployment-automation knowledge.

Experience balancing reliability, security, scalability, performance, and cost.

Strong understanding of security within cloud and production environments.

Experience scaling engineering practices in a rapidly growing technology organization.

Experience with air-gapped, disconnected, government, defense, highly regulated, or on-premise environments is a significant plus.

The Leadership Profile

Titles are less important than technical depth and leadership ability.

Potential backgrounds could include:

Director of SRE | Director of Platform Engineering | Director of Infrastructure Engineering | Head of SRE | Head of Platform | Head of Infrastructure | Senior Engineering Manager - SRE | Senior Engineering Manager - Platform | Principal SRE / Platform Engineer with significant leadership experience

The ideal candidate has reached engineering leadership without losing their engineering instincts.

You have managed and developed exceptional engineers, but you can still dive deeply into an architecture discussion.

You understand organizational strategy, but you also understand what is happening inside the infrastructure.

You can communicate with executives, but you can also earn the respect of a Principal Engineer.

And when a critical production system is failing, your team wants you in the room.

Interested in hearing more? Easy Apply now by clicking the "Apply Now" button.

Jobot is an Equal Opportunity Employer. We provide an inclusive work environment that celebrates diversity and all qualified candidates receive consideration for employment without regard to race, color, sex, sexual orientation, gender identity, religion, national origin, age (40 and over), disability, military status, genetic information or any other basis protected by applicable federal, state, or local laws. Jobot also prohibits harassment of applicants or employees based on any of these protected categories. It is Jobot's policy to comply with all applicable federal, state and local laws respecting consideration of unemployment status in making hiring decisions.

Sometimes Jobot is required to perform background checks with your authorization. Jobot will consider qualified candidates with criminal histories in a manner consistent with any applicable federal, state, or local law regarding criminal backgrounds, including but not limited to the Los Angeles Fair Chance Initiative for Hiring and the San Francisco Fair Chance Ordinance.

Information collected and processed as part of your Jobot candidate profile, and any job applications, resumes, or other information you choose to submit is subject to Jobot's Privacy Policy, as well as the Jobot California Worker Privacy Notice and Jobot Notice Regarding Automated Employment Decision Tools which are available at jobot.com/legal.

By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails from Jobot, and/or its agents and contracted partners. Frequency varies for text messages. Message and data rates may apply. Carriers are not liable for delayed or undelivered messages. You can reply STOP to cancel and HELP for help. You can access our privacy policy here: jobot.com/privacy-policy
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91113390
  • Position Id: 857737951
  • Posted 2 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

•

Today

Full-time

USD 180,240.00 - 300,360.00 per year

Remote

•

12d ago

Easy Apply

Full-time

Depends on Experience

Remote

•

Today

Full-time

USD 235,000.00 - 275,000.00 per year

Remote

•

Today

Full-time

USD 130,000.00 - 160,000.00 per year

Search all similar jobs