Data Quality Engineer (Data QE)
Data & AI | Wealth Management Division
Fort Mill , SC or New York City , NY or Florham Park, NJ
Department: Data Engineering & Analytics
Employment Type: Full-Time
About the Role
We are seeking an experienced Data Quality Engineer (Data QE) to ensure the quality, reliability, and integrity of data across our Wealth Management data platform. This role will focus on testing and validating data across modern cloud-native architectures, including AWS data pipelines, ELT/ETL workflows, Medallion Lakehouse architecture (Bronze, Silver, Gold), Kafka/event streams, APIs, and analytical data products.
The ideal candidate combines strong expertise in Data QA, test automation, SQL, Python, cloud data platforms, and streaming architectures with domain knowledge in wealth management data such as client accounts, positions, transactions, securities, advisors, and portfolio performance.
Key Responsibilities
- Define and execute the Data Quality Engineering strategy across enterprise data platforms, APIs, batch pipelines, and streaming architectures.
- Design and develop automated test frameworks covering:
- Data integrity and completeness
- Transformation and business rule validation
- Data reconciliation
- Schema and contract testing
- Data freshness and latency
- Data lineage and traceability
- Validate data across AWS-based ELT/ETL pipelines built on services such as S3, Glue, Lambda, Athena, EMR, MSK/Kafka, Redshift, and Iceberg-based data lakes.
- Test data movement and transformations across Medallion Architecture (Bronze → Silver → Gold) layers and ensure consistency between source and curated datasets.
- Build SQL and Python-based automation frameworks for large-scale data validation and regression testing.
- Develop and execute API test automation for REST and event-driven services, including payload validation, schema enforcement, error handling, and downstream data verification.
- Validate Kafka and streaming pipelines, including:
- Producer and consumer testing
- Event schema validation
- Message ordering and duplication controls
- Replay and recovery testing
- Throughput and latency validation
- Support testing of modern Lakehouse platforms including Apache Iceberg, Delta Lake, Spark/PySpark, Databricks, and AWS analytics platforms.
- Design reconciliation testing across custodians, brokerage platforms, internal systems, market data providers, and downstream reporting systems.
- Establish automated data quality controls, monitoring dashboards, anomaly detection, and alerting mechanisms.
- Investigate production data issues, perform root-cause analysis, and collaborate with engineering teams to drive permanent remediation.
- Integrate automated testing into CI/CD pipelines and enable quality gates for data and streaming releases.
- Leverage AI-assisted engineering tools (GitHub Copilot, ChatGPT, Claude, etc.) to accelerate test development, coverage expansion, anomaly detection, and documentation.
- Promote a shift-left, automation-first quality culture across Data Engineering and Analytics teams.
Required Qualifications
- 6-10 years of experience in Data QA, Data Quality Engineering, QA Automation, or Software Quality Engineering.
- Strong hands-on expertise with SQL and Python for data validation, automation, and analytical testing.
- Experience validating complex ELT/ETL pipelines and large-scale data transformation workflows.
- Strong understanding of modern data architectures, including:
- Data Lake
- Lakehouse
- Medallion Architecture
- Event-Driven Architectures
- Hands-on experience with AWS data services including:
- S3
- Glue
- Lambda
- Athena
- RDS
- Redshift
- EMR
- MSK/Kafka
- Experience testing data stored in Apache Iceberg, Delta Lake, Parquet, and other analytical formats.
- Strong experience with API testing and automation.
- Hands-on experience validating Kafka, streaming platforms, and real-time data processing systems.
- Experience building automated data quality, reconciliation, and regression testing frameworks.
- Familiarity with CI/CD pipelines and test automation integration.
- Strong understanding of data governance, data lineage, metadata, and data observability principles.
- Excellent analytical, debugging, and root-cause analysis skills.
- Wealth Management, Asset Management, Capital Markets, or Financial Services experience preferred.
Preferred Qualifications
- AWS Certified Data Analytics, AWS Certified Developer, or related certifications.
- Experience with:
- Apache Iceberg
- Databricks
- PySpark
- Delta Lake
- Apache Kafka
- Schema Registry
- EventBridge
- Kinesis
- Data Quality & Observability tools:
- Great Expectations
- Soda
- Monte Carlo
- Deequ
- Collibra DQ
- Testing frameworks:
- PyTest
- Postman
- REST Assured
- Karate
- CI/CD Tools:
- GitHub Actions
- Jenkins
- Azure DevOps
- Experience with DBT-based ELT testing and validation.
Nice-to-Have Skills
- Testing data products built on AWS Glue + S3 + Apache Iceberg architecture.
- Experience validating data across Bronze/Silver/Gold Medallion layers.
- Data mesh, data contracts, and schema evolution testing.
- AI-assisted test generation and intelligent anomaly detection.
- Wealth Management domains:
- Portfolio Management
- Performance Reporting
- Investment Advisory
- Client Onboarding
- Security Master
- Positions & Transactions
What Success Looks Like
You build scalable and automated quality controls that proactively identify data, API, and streaming defects before production; ensure trusted financial data across Medallion layers and AWS data platforms; reduce manual validation efforts; and improve confidence in analytics, reporting, and advisor-facing applications.
Hiring Manager Notes (I'd specifically add for TBAR)
Must Have
- AWS Data Lake / Data Pipeline testing
- S3, Glue, Athena, RDS
- Strong SQL + Python
- Kafka/Event Streaming validation
- Data Reconciliation
- Wealth Management data domain
Preferred
- Apache Iceberg
- Medallion Architecture
- Databricks/PySpark
- Data Quality Frameworks (Great Expectations/Soda)
- AI-assisted QE