Data Quality Engineer (Databricks, Kafka, AWS)

Hybrid in Dallas, TX, US • Posted 12 hours ago • Updated 12 hours ago
Full Time
No Travel Required
Able to Sponsor
Hybrid
Depends on Experience
Fitment

Dice Job Match Score™

👾 Reticulating splines...

Job Details

Skills

  • Python
  • ETL Pipeline Building
  • AWS Services
  • Data Quality Checks
  • Amazon S3
  • Amazon Redshift
  • Amazon Web Services
  • Data Quality
  • Apache Kafka
  • Apache Spark
  • Data Validation
  • PySpark
  • Continuous Integration
  • Data Engineering
  • Data Integrity
  • SQL
  • Streaming
  • ELT

Summary

We are looking for a Data Quality Engineer to own validation across batch and streaming data pipelines. This role focuses on ensuring data correctness, reliability, and performance across platforms built on Databricks, Kafka, AWS, SQL, and Python.
This is a hands-on role focused on building scalable data validation frameworks and ensuring production-grade data systems.

Key Responsibilities
End-to-End Data Validation
* Validate data pipelines for accuracy, completeness, consistency, and timeliness
* Build SQL-based validations for business rules and transformations
* Implement reconciliation between source and downstream systems
* Ensure data lineage and traceability

ETL / ELT & Spark Testing
* Test pipelines built on AWS (Glue, Lambda, EMR, Step Functions)
* Validate transformations using SQL and Python
* Test ingestion, transformation, aggregation, and serving layers
* Handle backfills, reprocessing, and historical data loads
* Validate Spark pipelines (PySpark/Scala) on Databricks

Streaming (Kafka)
* Validate data integrity, ordering, and delivery guarantees
* Test producer and consumer logic and serialization formats (Avro, JSON, Protobuf)
* Validate topics, partitions, offsets, retention, and schema evolution
* Simulate late events, duplicates, and failure scenarios

Automation & Frameworks
* Build Python-based data testing frameworks
* Develop reusable validation utilities and synthetic datasets
* Integrate data tests into CI/CD pipelines
* Enable automated alerts for data quality issues

Performance & Reliability
* Validate throughput, latency, and concurrency at scale
* Test retry logic, idempotency, and recovery mechanisms
* Perform regression, soak, and failover testing

Observability
* Validate logs, metrics, and alerts using tools such as CloudWatch, Prometheus, and Grafana
* Define and monitor data SLAs and SLOs
* Support incident response, root cause analysis, and postmortems

Required Qualifications & Experience
* 7+ years of total experience in QA, SDET, or Data Quality Engineering
* Minimum 4–6 years of hands-on experience working with data platforms, data pipelines, or data engineering ecosystems
* 3+ years of hands-on experience with Databricks and Apache Spark
* Strong SQL skills for data validation, reconciliation, and complex analysis
* Proficiency in Python for automation and data validation
* Experience testing ETL/ELT pipelines (batch and streaming)
* Hands-on experience with Kafka or similar streaming platforms
* Strong understanding of AWS data services (S3, Glue, Lambda, Redshift, Athena)
* Experience working with large-scale distributed data systems
* Strong debugging, analytical, and problem-solving skills

Nice to Have
* Experience with data quality or observability tools such as Great Expectations or Monte Carlo
* Knowledge of schema registry and data contracts
* Experience with CI/CD tools such as GitHub Actions or Jenkins
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91174352
  • Position Id: 8986768
  • Posted 12 hours ago
Contact the job poster
VD

Venkat Dharmik

Recruiter @ Plugins Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Hybrid in Dallas, Texas

7d ago

Easy Apply

Full-time

90,000 - 100,000

Hybrid in Dallas, Texas

11d ago

Easy Apply

Full-time

60 - 70

Hybrid in Dallas, Texas

Today

Easy Apply

Contract

50+

Hybrid in Dallas, Texas

11d ago

Easy Apply

Contract, Third Party

Depends on Experience

Search all similar jobs