Location: NY / NJ / Pittsburgh Onsite
Employment Type: Full-Time
Experience: 8+ Years
Job Summary
We are looking for a hands-on Technology Engineering Lead to design and deliver enterprise-scale data engineering solutions, including a batch feature store supporting analytics and AI use cases.
This is a technical leadership + individual contributor role. The ideal candidate must be comfortable working directly with clients while remaining deeply hands-on with PySpark, SQL, database architecture, and ETL/ELT pipeline development.
Key Responsibilities
-
Lead architecture and development of enterprise-scale data pipelines.
-
Design and build scalable PySpark and Impala SQL solutions.
-
Develop ETL/ELT pipelines across Hive and Oracle data platforms.
-
Own data architecture, feature logic, data quality, and reconciliation.
-
Drive pipeline performance tuning, monitoring, and failure recovery.
-
Manage batch scheduling and dependencies using CA7 or equivalent.
-
Conduct design reviews, code reviews, and hands-on troubleshooting.
-
Establish Git/Bitbucket, CI/CD, and testing practices.
-
Lead production releases and SDLC/change-management activities.
-
Work directly with business, analytics, and technology stakeholders.
-
Mentor engineers while remaining actively involved in technical delivery.
Required Qualifications
-
8+ years of Data Engineering / Data Platform experience.
-
5+ years of hands-on PySpark and enterprise ETL/ELT development.
-
Advanced SQL and strong pipeline performance-tuning skills.
-
Strong database/data-platform architecture experience.
-
Hands-on experience with Hive and Oracle.
-
Technical leadership experience across complex data-engineering programs.
-
Experience with Linux, Git/Bitbucket, and CI/CD.
-
Strong client-facing communication skills.
-
Ability to work effectively in a fast-paced environment.
Preferred Skills
-
Banking / Financial Services experience.
-
CA7 or equivalent batch scheduling tools.
-
Jira and Agile/Scrum experience.
-
Experience with AI-assisted development tools.
Primary Skills
PySpark, SQL, ETL/ELT, Data Engineering, Data Architecture, Hive, Oracle, Impala SQL, Database Architecture, Data Pipelines, Performance Tuning, CI/CD, Git/Bitbucket, Linux, CA7