Job Overview
We are seeking a Senior Data Engineer to join a large-scale banking data migration initiative supporting the integration of Discover and Capital One platforms.
The team is responsible for migrating and processing high volumes of customer and transactional data across banking products, including checking, savings, money market accounts, IRAs, and CDs.
This is a data engineering and data pipeline-focused role, with a strong emphasis on Java, Spring Boot, Scala, Spark/PySpark, and AWS data services. The ideal candidate will have experience building and optimizing high-volume batch data pipelines and working with large-scale data migration workloads.
Candidates should be able to make an immediate impact with minimal ramp-up time.
Key Responsibilities
- Design, develop, enhance, and support large-scale data pipelines using Java, Spring Boot, Scala, and AWS.
- Build and maintain batch-oriented data processing applications supporting large-scale migration activities.
- Develop and customize Spring Boot-based applications used within data processing pipelines.
- Work with Spark/PySpark to process and transform large volumes of data.
- Develop and manage AWS-based data pipelines using services such as AWS Glue, S3, Lambda, and Step Functions.
- Support high-volume data migration and scheduled migration executions.
- Perform performance tuning and optimization of existing pipelines and data-processing applications.
- Troubleshoot data processing issues and improve pipeline reliability and scalability.
- Integrate external APIs into data pipelines where required.
- Collaborate with engineering teams to deliver enhancements and customizations to existing migration applications.
- Support testing, deployment, and production readiness of data pipelines.
- Work with distributed processing technologies to efficiently process large datasets.
- Contribute to CI/CD processes and engineering practices within the development environment.
- Leverage AI-assisted development tools effectively as part of the engineering workflow.
Required Skills
- Strong hands-on experience with Java.
- Strong experience with Spring Boot.
- Strong data engineering and data pipeline development experience.
- Hands-on experience with Scala β Must Have.
- Hands-on experience developing AWS data pipelines β Must Have.
- Experience with AWS Glue.
- Experience with Apache Spark and/or PySpark.
- Experience with AWS Lambda.
- Experience with AWS Step Functions.
- Experience with Amazon S3.
- Strong understanding of ETL/data processing concepts.
- Experience working with batch processing and large-scale data workloads.
- Experience with performance tuning and optimization of data pipelines.
- Strong problem-solving and troubleshooting skills.
Preferred Skills
- Previous Capital One experience.
- Experience with Capital One OnePipeline.
- Financial services or banking industry experience.
- Experience with large-scale customer or transactional data migrations.
- Experience with Kafka.
- Experience with Amazon EventBridge.
- Experience with Amazon SNS.
- API integration experience.
- Experience with CI/CD pipelines.
- Experience working with AI-assisted software development tools.
Data Engineering Focus
This position is primarily focused on data engineering rather than traditional backend/API development.
The successful candidate should be able to demonstrate experience with:
- Large-scale data pipelines
- Batch processing
- ETL/data transformation
- Distributed data processing
- High-volume data migration
- Spark/PySpark
- AWS Glue
- Pipeline performance optimization
- Scheduled data processing
- Production-scale data workloads
API development experience is valuable but is secondary to hands-on data pipeline experience.
Project Environment
The team is supporting a major financial data migration involving extremely high data volumes, including workloads of up to approximately 500 million transactions.
The project is already significantly established, with multiple engineering teams working across the broader migration initiative. The selected engineer will contribute to existing pipelines and applications through:
- Enhancements
- Customizations
- Performance tuning
- Pipeline optimization
- Migration preparation
- High-volume batch execution support
The role is primarily focused on batch migration workloads rather than real-time streaming systems.
Ideal Candidate
The ideal candidate is a Senior Data Engineer / Senior Java Data Engineer with strong hands-on experience across:
Java + Spring Boot + Scala + Spark/PySpark + AWS Glue + AWS Data Pipelines
The candidate should be comfortable working independently, troubleshooting complex data-processing problems, optimizing existing applications, and contributing to high-volume migration workloads with minimal ramp-up time.
Previous Capital One experience is highly desirable because familiarity with Capital One's engineering environment, processes, pipelines, and OnePipeline can enable faster productivity.
Locations
- Wilmington, DE
- McLean, VA
- Richmond, VA
- New York City, NY