Data Infrastructure Site Reliability Engineer

Remote • Posted 6 hours ago • Updated 6 hours ago
Contract W2
Contract Independent
Remote
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🎯 Assessing qualifications...

Job Details

Skills

  • AWS
  • Hadoop
  • on-premises
  • AI

Summary

Job Description Data Infrastructure Site Reliability Engineer (SRE) AWS & Big Data Platforms
Dallas, Texas (Remote)
Skills : (AWS, Hadoop, on-premises, AI)

Skills - Site Reliability Engineering, Big Data, AWS, AWS EMR, AWS EKS, AWS MSK, AWS Athena, AWS Glue, Spark, Iceberg, Python, Terraform, Cloudera Hadoop, Agentic Ai, Monitoring Tools


Role Overview
We are seeking a highly skilled L2+/Senior Data Infrastructure Site Reliability Engineer (SRE) to operate, optimize, and automate large-scale data infrastructure environments across AWS and on-premises Hadoop platforms. This role requires a strong reliability engineering mindset focused on platform stability, performance, observability, automation, and incident response.
As a key member of the Data Infrastructure SRE team, you will ensure the availability, scalability, security, and operational excellence of mission-critical data platforms while driving continuous improvements through Infrastructure as Code (IaC), AI-enabled automation, and modern SRE practices.
Key Responsibilities

  • Maintain and support highly available, scalable, and secure data infrastructure platforms across AWS and on-premises environments.
  • Drive operational excellence through automation of repetitive tasks, incident reduction, and proactive reliability improvements.
  • Monitor platform health, troubleshoot complex issues, and lead root cause analysis efforts to minimize downtime and improve system resiliency.
  • Collaborate closely with engineering, platform, and data teams to support deployments, integrations, upgrades, and performance optimization initiatives.
  • Participate in on-call rotations and partner with teams across the US and India to provide 24x7 operational support. US support is aligned to Pacific Time zone.
  • Continuously improve observability, alerting, and incident response processes to enhance platform reliability and reduce operational noise.


Required Skills & Experience
Site Reliability Engineering (SRE)
Strong SRE mindset with a proven focus on reliability, availability, performance optimization, incident management, and operational excellence.
Experience delivering services within defined SLAs and ensuring timely resolution of production issues.
Expertise in troubleshooting complex distributed systems and identifying root causes quickly and effectively.
AWS & Cloud Infrastructure
Deep hands-on experience with AWS services, including:
EMR
EKS
MSK
Athena
Glue
IAM
Amazon S3
VPC
AWS networking and security services
Strong understanding of cloud-native architectures, scalability, and infrastructure resilience.

Big Data Platforms

  • Extensive operational experience managing Hadoop clusters, with a strong focus on administration, platform maintenance, and automation of day-to-day operational activities.
  • Experience supporting both AWS-based data platforms and on-premises Cloudera CDH/CDP environments.
  • Solid understanding of Kerberos authentication and security implementation within Hadoop ecosystems.

Hands-on experience with:

  • Apache Spark
  • Apache Iceberg
  • Big Data platform architecture
  • Performance tuning and optimization


Linux & System Administration
Strong Linux administration and operational support experience.
Expertise in user and access management, system configuration, customization, and platform administration.

Observability & Incident Management
Hands-on experience with monitoring and observability platforms such as:
AWS CloudWatch
Datadog
PagerDuty
Similar enterprise monitoring solutions
Proven success improving alert quality, reducing false positives, and minimizing alert fatigue.
Excellent debugging and troubleshooting skills across infrastructure, applications, Spark workloads, and Iceberg environments.

Java Platform Operations
Strong understanding of Java application administration, including:
JVM tuning
Thread dump analysis
Heap dump analysis
JVM parameters
Application log analysis and troubleshooting

Automation & Infrastructure as Code
Expert-level Python scripting for operational automation, monitoring, and reliability engineering use cases.
Strong experience with:
Terraform
Infrastructure as Code (IaC)
CI/CD pipeline implementation and automation
Platform engineering best practices

AI-Driven Operations
Current hands-on experience applying AI technologies to Data Infrastructure and SRE operations.
Demonstrated ability to design and implement:
Agentic AI solutions
AI-assisted operational workflows
Intelligent automation for routine SRE activities
Strong creativity and problem-solving skills in leveraging AI to improve operational efficiency, reliability, and productivity.

Preferred Candidate Profile
The ideal candidate combines deep expertise in AWS cloud platforms, Hadoop ecosystems, SRE practices, observability, automation, and AI-driven operations, with a passion for supporting resilient data platforms at scale and eliminating operational toil through engineering excellence.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91138214
  • Position Id: 9051000
  • Posted 6 hours ago

Company Info

About Marici Solutions

MARICI Solutions is a global business consulting organization, where inspired & visionary people with a shared passion for innovation come together to make a difference. We take a holistic approach from various perspectives to deliver measurable results along the entire value chain. We work closely with our clients and provide distinct advantages to sustainably transform their business processes, which eventually grow their businesses.

Driven by the passion of offering state of the art, customized and high-quality information technology services, our experts from Germany & India came together in 2017 and formed MARICI Solutions GmbH. The company is registered in Germany & India and well funded for long term sustenance. The team at MARCI comprises of experts from various industry sectors who have more than a decade of diverse & international consulting experience. MARICI is expanding its business across multiple continents and has secured various projects from some of the eminent market giants.

We believe that Excellence, Collaboration, Commitment and Innovation are foundations for success.

By Collaborating with our stakeholders, we identify opportunities for growth and Innovation. Ongoing innovation results from a combination of strategy, processes, systems, and culture. Through our Commitment to Excellence, we accomplish complex assignments and inspire people to do great things together. 

Contact the job poster
AT

Anish Tarte

Recruiter @ Marici Solutions
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs