Druid Engineer

Remote • Posted 1 day ago • Updated 10 hours ago
Full Time
No Travel Required
Remote
$50 - $55/hr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Ansible
  • Root Cause Analysis
  • Scalability
  • Query Optimization
  • Optimization
  • Production Support
  • PB
  • Meta-data Management
  • Management
  • MySQL
  • Disaster Recovery
  • Apache HTTP Server
  • Big Data
  • Analytical Skill
  • Capacity Management
  • Computer Cluster Management
  • Configuration Management
  • Dashboard
  • Data Analysis
  • Backup
  • Data Engineering
  • DevOps
  • Database
  • Collaboration
  • Incident Management
  • Grafana
  • High Availability
  • Operational Excellence
  • Distribution
  • Performance Engineering
  • Recovery
  • Performance Tuning
  • IaaS

Summary

Druid Engineer

Location: Dallas, TX
Rate: $55/hr C2C
Employment Type: Contract

Job Description

We are seeking an experienced Druid Engineer to design, deploy, administer, optimize, and support large-scale Apache Druid environments supporting PB-scale datasets. The ideal candidate will have strong expertise in Druid cluster administration, performance tuning, ingestion, monitoring, automation, and production support.

Key Responsibilities

  • Design, deploy, configure, administer, and optimize large-scale Apache Druid clusters supporting PB-scale analytical workloads.
  • Manage, monitor, and troubleshoot Apache Druid services, clusters, and distributed components.
  • Perform Druid cluster upgrades, patching, capacity planning, scaling, and platform modernization.
  • Monitor cluster health, ingestion performance, query latency, segment distribution, resource utilization, and overall platform performance.
  • Troubleshoot ingestion failures, stuck tasks, compaction issues, retention policies, segment management, and indexing problems.
  • Optimize Druid queries, partitioning strategies, indexing specifications, segment allocation, and compaction configurations.
  • Configure and maintain Druid metadata stores using MySQL, including metadata management and database connectivity.
  • Develop operational automation using Ansible, Infrastructure as Code (IaC), and configuration management practices.
  • Build and maintain Grafana dashboards, Prometheus monitoring, alerts, metrics, and centralized logging solutions.
  • Implement and support High Availability (HA), Disaster Recovery (DR), backup/recovery, security, and compliance requirements.
  • Participate in production support, incident management, root cause analysis (RCA), troubleshooting, and performance tuning.
  • Collaborate with DevOps, SRE, platform engineering, architects, data engineering, and business teams to translate analytical requirements into scalable Druid solutions.
  • Support Druid performance optimization, scalability, reliability, availability, and operational excellence across enterprise environments.

Required Skills

  • Strong hands-on experience with Apache Druid / Druid Engineering.
  • Experience administering large-scale distributed Druid clusters and PB-scale datasets.
  • Strong knowledge of Druid ingestion, indexing, segments, partitioning, compaction, retention, and query optimization.
  • Experience with Druid MiddleManager/Indexer, Historical, Broker, Coordinator, Overlord, Router, and Metadata Store components.
  • Strong experience with MySQL for Druid metadata management.
  • Experience with Ansible and Infrastructure as Code (IaC).
  • Hands-on experience with Grafana, Prometheus, monitoring, alerting, metrics, and logging.
  • Strong troubleshooting, performance tuning, capacity planning, and production support experience.
  • Knowledge of High Availability, Disaster Recovery, backup/recovery, security, and compliance.

 

Apache Druid, Druid Engineer, Druid Administrator, Druid Cluster Administration, Druid Cluster Management, Druid Performance Tuning, Druid Optimization, Druid Ingestion, Druid Indexing, Druid Segments, Druid Compaction, Druid Partitioning, Druid Query Optimization, Druid Metadata Store, MySQL, PB-Scale Data, Distributed Systems, Big Data, Data Analytics, Ansible, Infrastructure as Code, IaC, Grafana, Prometheus, Monitoring, Alerting, Logging, High Availability, Disaster Recovery, Backup & Recovery, Capacity Planning, Cluster Scaling, Production Support, Incident Management, Root Cause Analysis, Performance Engineering, DevOps, SRE, Cloud Infrastructure.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10451163
  • Position Id: 9067933
  • Posted 1 day ago
Contact the job poster
JC

Johnson Carter

Recruiter @ Innovative IT Solutions Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Easy Apply

Contract, Third Party

Depends on Experience

Remote or Burlingame, California

Today

Full-time

USD 185,000.00 - 225,000.00 per year

Remote or Dallas, Texas

11d ago

Full-time

USD 90,700.00 - 153,925.00 per year

Remote

Yesterday

Easy Apply

Full-time

Depends on Experience

Search all similar jobs