We are specifically looking for an Iceberg DBA / Big Data Administrator / Lakehouse Operations Engineer with strong production administration and operational support experience.
This is NOT primarily a Big Data Developer or Data Engineer role.
Candidates coming from a Hadoop Administrator, Cloudera Administrator, Big Data DBA, Data Platform Administrator, or Lakehouse Operations background who have hands-on Apache Iceberg administration and table operations are strongly preferred.
Mid-level candidates with solid hands-on administration and production support experience will be considered.
Position Summary
We are seeking a Senior Iceberg DBA / Lakehouse Operations Engineer to support the reliability, performance, availability, and operational integrity of an enterprise-scale Apache Iceberg Lakehouse environment.
The ideal candidate will have 4 6 years of Big Data / Data Operations / DBA experience, including 4+ years with the Cloudera ecosystem (CDP) and at least 1+ year of hands-on Apache Iceberg experience.
The role focuses on Iceberg table administration, maintenance, troubleshooting, performance optimization, metadata management, production support, and incident resolution across Spark, Hive, and Impala environments.
Required Skills
- 4 6 years of experience in Big Data Administration, Data Operations, DBA, Hadoop Administration, or Data Platform Operations
- 1+ year hands-on Apache Iceberg experience
- 4+ years of experience with Cloudera / CDP ecosystem
- Strong experience with Iceberg table administration and maintenance
- Hands-on experience with:
- Iceberg table operations
- Table maintenance
- Compaction
- Small-file management
- Snapshot expiration
- Metadata management
- Orphan file cleanup / vacuum
- Partition management and optimization
- Strong experience with Spark SQL, Hive and/or Impala
- Production support and L2/L3 incident management
- Strong troubleshooting and root-cause-analysis skills
- Experience supporting large-scale data platforms
- Knowledge of TB/PB-scale data environments
- Understanding of data lake / Lakehouse architecture
- Strong knowledge of:
- Partitioning strategies
- Parquet / ORC
- Distributed query processing
- Table-level performance optimization
- Data lifecycle management
- Experience with monitoring, alerting, troubleshooting and operational support
- Experience with data validation, reconciliation and data consistency
- Knowledge of Ranger, RBAC and data access controls
- Strong scripting skills using Python and/or Shell
Preferred Skills
- Cloudera CDP
- CDE / CDW
- Hadoop Administration
- Hive Administration
- Apache Iceberg
- Trino
- NiFi
- AWS or Azure
- Hive-to-Iceberg migration
- Teradata-to-Iceberg migration
- Experience supporting Bronze / Silver / Gold / Medallion architecture
- Experience with Iceberg across multiple query engines
Key Responsibilities
- Own day-to-day operational administration of Apache Iceberg tables
- Maintain reliability, availability, and consistency of enterprise Lakehouse datasets
- Perform Iceberg table maintenance and optimization
- Manage:
- Compaction
- Small-file mitigation
- Snapshot expiration
- Metadata cleanup
- Orphan-file cleanup
- Partition evolution
- File-size optimization
- Troubleshoot Iceberg table and metadata issues across Spark, Hive and Impala
- Monitor production workloads and resolve performance issues
- Troubleshoot query failures, inefficient scans and execution-plan issues
- Support large-scale Lakehouse environments at multi-TB/PB scale
- Perform data validation and reconciliation between source and Iceberg datasets
- Support Hive/Teradata modernization to Iceberg
- Assist with schema and data-type alignment during migration
- Provide L2/L3 production support
- Participate in on-call rotation
- Handle P1/P2 incidents and meet defined SLAs
- Conduct RCA and implement preventive actions
- Collaborate with Data Engineering, Platform, Application and Data Governance teams
- Support Ranger policies, RBAC and secure data access
- Maintain data lifecycle, retention and archival policies
Ideal Candidate Background
The strongest candidates may have previously worked as:
Hadoop Administrator OR Cloudera Administrator OR Big Data Administrator OR Big Data DBA OR Data Platform Administrator OR Data Operations Engineer OR Lakehouse Operations Engineer
and have subsequently gained hands-on experience with:
Apache Iceberg + Spark SQL + Hive/Impala + Cloudera CDP
NOT a Development-Heavy Role
Candidates whose experience is primarily:
- Big Data Development
- ETL Development
- Spark Development
- Python Data Engineering
- Data Pipeline Development
with only limited exposure to Iceberg administration may not be a good fit.
We are looking for candidates who have owned, administered, maintained, monitored, troubleshot and supported data platforms in production.