Site Reliability Engineer Cloud / DevOps

New York, NY, US • Posted 8 hours ago • Updated 8 hours ago
Full Time
No Travel Required
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Python
  • OpenTelemetry
  • Linux
  • Docker

Summary

NMK Global Inc. is a global IT Services & Solutions company specializing in IT Staffing, Technology Consulting, Workforce Solutions, and Application Development Services. We help organizations build high-performing technology teams by connecting them with exceptional talent across a wide range of industries. Our expertise, industry knowledge, and consultative approach enable clients to achieve successful business outcomes while helping professionals advance their careers. 

Role: Site Reliability Engineer — Cloud / DevOps

Location: New York City, NY – Hybrid
Onsite: 3–4 Days/Week
Employment Type: Full-Time, W2 Direct Hire


Job Summary

Our client is a fast-growing financial technology firm based in New York City, transforming how institutional markets operate. They are seeking a highly experienced Site Reliability Engineer (SRE) / Cloud DevOps Engineer to own and drive infrastructure, cloud architecture, CI/CD, containerization, observability, and production reliability.

This is not a ticket-based support role. The ideal candidate is a hands-on infrastructure builder and technical leader who understands systems at a deep level and has experience owning production environments from code to cloud.

The successful candidate will have strong architectural judgment and the ability to establish and enforce engineering standards across the organization.

Key Responsibilities

  • Own and evolve highly available, scalable, and secure cloud infrastructure.
  • Design and maintain AWS cloud architecture based on performance, security, reliability, cost, and ROI considerations.
  • Build reusable and scalable CI/CD pipelines supporting end-to-end code-to-cloud delivery.
  • Implement and enforce engineering standards through:
    • Linters
    • Security scanning
    • Policy-as-code
    • Automated quality gates
  • Design, deploy, and manage containerized workloads using Docker and related technologies.
  • Demonstrate a strong understanding of Docker isolation, including Linux namespaces and cgroups.
  • Establish and enforce organization-wide observability standards.
  • Implement and maintain metrics, logs, traces, and alerting for distributed production systems.
  • Work with OpenTelemetry and modern observability practices.
  • Manage infrastructure using Terraform and Infrastructure as Code principles.
  • Develop automation and infrastructure tooling using Python.
  • Troubleshoot complex production issues across Linux, cloud infrastructure, networking, and distributed systems.
  • Establish reliability, scalability, security, and operational best practices across engineering teams.
  • Take ownership of production environments and drive improvements rather than simply supporting existing systems.
  • Evaluate new technologies and architectural approaches based on business value and technical requirements.
  • Partner closely with software engineering, security, and platform teams.
  • Contribute to the infrastructure roadmap for future AI workload infrastructure.

Required Qualifications

  • 5+ years of experience in Infrastructure, DevOps, Cloud Engineering, or Site Reliability Engineering.
  • Strong hands-on experience with AWS.
  • Strong experience with Terraform / Infrastructure as Code.
  • Strong Python scripting and automation experience.
  • Advanced Linux knowledge.
  • Production experience supporting and owning distributed systems.
  • Demonstrated experience owning production environments end to end.
  • Strong understanding of cloud architecture, reliability, scalability, security, and cost optimization.
  • Experience building reusable CI/CD systems and pipelines.
  • Strong understanding of containerization and Docker.
  • Experience implementing observability across production systems.
  • Strong troubleshooting and problem-solving skills.
  • Ability to make and defend technical architecture decisions.

Preferred / Strong-Signal Qualifications

  • Production experience with OPA (Open Policy Agent) and policy-as-code.
  • Experience enforcing security and engineering standards directly within CI/CD pipelines.
  • Strong understanding of OpenTelemetry and observability specifications.
  • Experience with distributed systems and large-scale production environments.
  • Experience implementing organization-wide observability practices.
  • Experience with container orchestration and cloud-native infrastructure.
  • Experience with AI/ML workload infrastructure is a plus.

 

NMK Global Inc. is an Equal Opportunity Employer. We are committed to creating an inclusive workplace and consider all qualified applicants without regard to age, race, color, religion, sex, national origin, disability, protected veteran status, sexual orientation, gender identity, or any other characteristic protected by applicable law. 

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90773860
  • Position Id: 3871-4981-1789076612
  • Posted 8 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote or New York, New York

Today

Full-time

USD 160,000.00 - 200,000.00 per year

New York, New York

Today

Full-time

USD 145,000.00 - 195,000.00 per year

New York, New York

Today

Easy Apply

Full-time

USD250,000 - USD300,000

Newark, New Jersey

13d ago

Easy Apply

Contract

Depends on Experience

Search all similar jobs