Lead DevOps Platform Engineer
Location: Hybrid (Princeton, NJ)
Duration : Long Term
Job Summary
We are seeking a highly skilled Senior DevOps Platform Engineer to design, build, and support scalable cloud-native platforms, CI/CD pipelines, infrastructure automation, and developer enablement solutions. The ideal candidate will have deep expertise in Platform Engineering, Kubernetes, Infrastructure as Code, DevSecOps practices, and cloud technologies while partnering closely with Development, Security, and SRE teams.
Key Responsibilities
Platform Engineering & Infrastructure
- Design, implement, and maintain scalable cloud-native platform solutions.
- Build and manage Kubernetes environments across development, staging, and production.
- Develop reusable Infrastructure as Code (IaC) modules using Terraform/OpenTofu.
- Implement platform standards, golden-path templates, and automation frameworks.
- Improve platform reliability, scalability, and operational efficiency.
CI/CD & DevOps Automation
- Design and maintain enterprise-grade CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, ArgoCD, or similar tools.
- Automate application deployment, testing, release management, and rollback processes.
- Implement GitOps practices and deployment automation.
- Manage artifact repositories and software package lifecycle processes.
DevSecOps & Security
- Integrate security controls into CI/CD pipelines.
- Implement SAST, SCA, container scanning, secrets detection, and policy-as-code solutions.
- Support software supply chain security initiatives including SBOM generation, artifact signing, and dependency management.
- Collaborate with Security teams to ensure compliance with enterprise security standards.
Cloud & Container Platforms
- Manage and optimize cloud environments across AWS, Azure, or Google Cloud Platform.
- Deploy and maintain Kubernetes clusters and containerized workloads.
- Implement monitoring, logging, observability, and alerting solutions.
- Support high availability, disaster recovery, and platform resilience initiatives.
Developer Experience & Enablement
- Build self-service developer capabilities and platform automation.
- Support Internal Developer Portal initiatives such as Backstage or similar platforms.
- Create technical documentation, runbooks, standards, and onboarding guides.
- Improve developer productivity through automation and platform enhancements.
Monitoring & Reliability
- Implement observability solutions using Prometheus, Grafana, ELK, Datadog, Splunk, or similar tools.
- Monitor platform health, availability, and performance metrics.
- Participate in incident response, root cause analysis, and continuous improvement activities.
Required Skills & Experience
Core Technical Skills
- 10+ years of experience in DevOps, Platform Engineering, Cloud Engineering, or Site Reliability Engineering.
- Strong expertise in Kubernetes administration and container orchestration.
- Hands-on experience with Terraform/OpenTofu and Infrastructure as Code.
- Experience designing and supporting CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, ArgoCD, or Tekton.
- Strong Linux system administration and scripting skills (Python, Bash, PowerShell).
- Experience with AWS, Azure, or Google Cloud Platform.
- Knowledge of GitOps practices and deployment automation.
DevSecOps Experience
- Experience integrating security tools such as Semgrep, Snyk, Checkmarx, Trivy, Prisma Cloud, Gitleaks, or HashiCorp Vault.
- Working knowledge of Policy-as-Code using OPA/Rego or Kyverno.
- Familiarity with SBOM, SLSA, software supply chain security, and artifact signing.
Observability & Reliability
- Experience implementing monitoring and observability solutions.
- Knowledge of DORA metrics and platform performance measurement.
- Strong troubleshooting and incident management skills.
Nice to Have
- Experience with Backstage or Internal Developer Portals.
- Exposure to AI-assisted development tools such as GitHub Copilot, Cursor, or Agentic workflows.
- Experience in regulated industries such as Financial Services, Healthcare, or Government.
- Knowledge of Service Mesh technologies (Istio, Linkerd).
- Experience with eBPF security tools such as Falco or Tetragon.
- Cloud certifications (AWS, Azure, Google Cloud Platform, Kubernetes).
Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent practical experience.
- Relevant certifications in Cloud, Kubernetes, DevOps, or Security are highly preferred.
Education:
Bachelor's degree in Business, Computer Science, Mathematics, Engineering, or related discipline, or equivalent.