Hello
Description:
Job Title: Senior Technology Consultant Data / AI Native Engineer
Location: Arlington, VA, New York, NY and St. Louis, MO; Hybrid work model | Local candidates preferred.
Overall Experience: 7+ Years
Role Summary:
We are seeking a hands-on Senior Data / AI Native Engineer with strong expertise in Databricks, AWS, PySpark, and modern Lakehouse architecture. The ideal candidate combines deep Data Engineering expertise with strong Software Engineering practices and has experience leading complex enterprise data initiatives.
The role requires daily hands-on use of GitHub Copilot and/or Claude Code CLI to accelerate software and data pipeline development. The candidate should be comfortable rapidly building POCs and MVPs and applying AI directly within data engineering workflows-not just using AI for application development.
Day to Day Job Duties
- Design, develop, and optimize high-volume enterprise data pipelines using PySpark, Spark, and Databricks.
- Build scalable Lakehouse solutions using Delta Lake and Medallion Architecture (Bronze/Silver/Gold).
- Implement data governance, access controls, and cataloging using Databricks Unity Catalog.
- Design and develop cloud-native data solutions using AWS S3, EMR, Glue, Lambda, and Redshift.
- Lead complex Data Engineering initiatives and provide technical guidance to engineering teams.
- Develop batch and real-time data processing pipelines for large-scale enterprise datasets.
- Use GitHub Copilot and/or Claude Code CLI daily to accelerate coding, pipeline development, testing, troubleshooting, and documentation.
- Rapidly develop POCs and MVPs using AI-assisted engineering practices.
- Apply AI within data pipelines for use cases such as schema inference, automated data-quality rule generation, anomaly detection, and PySpark transformation generation.
- Build AI-ready data platforms supporting RAG, embeddings, vector stores, and AI/ML applications.
- Develop streaming pipelines using Structured Streaming, Kafka, Kinesis, or Databricks Auto Loader.
- Implement modern data engineering patterns including CDC, SCD Type 2, schema evolution, idempotent processing, and data contracts.
- Implement data quality, validation, monitoring, lineage, and observability across data pipelines.
- Conduct code/design reviews and establish reusable Data Engineering patterns and standards.
- Collaborate with Data Architects, AI/ML Engineers, Software Engineers, and business stakeholders to deliver enterprise data products.
Basic Qualifications Must Have:
- 7+ years of experience in Data Engineering and Software Engineering, building production-grade enterprise data solutions.
- 4+ years of hands-on experience with Databricks, Delta Lake, Medallion Architecture, and PySpark/Spark.
- 4+ years of experience with the AWS data ecosystem, including S3, EMR, Glue, Lambda, and/or Redshift.
- Proven experience leading complex Data Engineering initiatives or providing technical leadership to Data Engineering teams.
- Strong hands-on experience processing high-volume datasets using PySpark, rather than Scala-only Spark development.
- Demonstrated daily use of GitHub Copilot and/or Claude Code CLI, with the ability to explain specific examples of how these tools improve Data Engineering productivity.
- Proven experience rapidly developing POCs and MVPs using AI-assisted development practices.
Technical Skills
Data Platform: Databricks, Delta Lake, Unity Catalog
Data Processing: PySpark, Apache Spark
Cloud: AWS S3, EMR, Glue, Lambda, Redshift
Architecture: Lakehouse, Medallion Architecture, Data Lakes
Streaming: Spark Structured Streaming, Kafka, Kinesis, Auto Loader