Hadoop Hive Python Developer

Charlotte, NC, US • Posted 9 hours ago • Updated 9 hours ago
Full Time
No Travel Required
On-site
$110,000 - $125,000/yr
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • hadoop
  • python
  • pyspark
  • apache
  • kafka
  • Hive
  • ETL

Summary

Job Description
Must Have Technical/Functional Skills
 
Primary skills: Hadoop, Hive, Python, PySpark, Apache Kafka, Hadoop Ecosystem, Hive, Databricks Lakehouse Architecture, Delta Lake, Bronze/Silver/Gold Data Modeling, Big Data ETL Pipeline Development, SQL, Real-time Data Ingestion Frameworks, Data Governance & Cataloging, CI/CD Tools – Git, Jenkins, Bitbucket, Workflow Orchestration, and Cloud & On-Prem Big Data Platforms.
 
Experience: Minimum 9+ years
 
Roles & Responsibilities
 
Seeking a Senior Big Data Engineer with 9-14 years of experience specializing in Hadoop, Python, Hive PySpark, Kafka, and strong experience designing data solutions for large-scale financial systems.
In addition, the candidate must possess advanced expertise in Databricks Lakehouse architecture, particularly around Bronze/Silver/Gold layer data modeling, Delta Lake optimizations, and building reliable, scalable pipelines for regulatory, risk, trading, and analytics workloads.
This role focuses on delivering highly performant, well-governed data platforms that support the bank’s mission-critical global markets functions.
 
Key Responsibilities:
 
Big Data Platform Engineering
· Design, develop, and optimize PySpark-based ETL pipelines running on on‑prem Hadoop clusters and cloud environments.
· Build high‑volume ingestion frameworks using Kafka for real-time and near-real-time trading and market data. 
· Develop, tune, and manage Hadoop ecosystem components—HDFS, YARN, MapReduce, Tez, Oozie/Airflow. 
· Build high-performance, optimized Hive data models for regulatory reporting, trade lifecycle, and market risk processing. 
 
Databricks Lakehouse & Delta Framework 
· Architect and implement Bronze/Silver/Gold layer modeling patterns within the Databricks Lakehouse. 
· Apply Delta Lake best practices including: 
o optimized file management 
o Z-Ordering 
o Delta Change Data Feed (CDF) 
o schema evolution & enforcement 
o ACID transaction handling 
· Build reusable frameworks for ingestion, cleansing, transformation, and consumption of data across Lakehouse layers. 
· Enable governance, lineage, and auditability using Unity Catalog or equivalent cataloging tools. 
 
Collaboration, Leadership & Delivery 
· Collaborate closely with quants, product owners, architects, risk tech, and business users. 
· Participate in agile ceremonies — sprint planning, refinement, design reviews. 
· Mentor junior engineers and contribute to building strong engineering practices across tech teams. 
 
Required Skills & Experience 
· 9-14 years of hands-on experience in Big Data engineering. 
· Expert skills in: 
o PySpark — dataframe optimizations, partitioning, broadcast strategies, distributed computing. 
o Kafka — producer/consumer design, schema registry, streaming ETLs. 
o Hadoop ecosystem — HDFS, YARN, MapReduce/Tez, Oozie/Airflow. 
o Hive — advanced query tuning, TEZ optimization, partition/bucket management. 
· Extensive hands-on experience with Databricks Lakehouse, including: 
o Bronze/Silver/Gold layer modeling 
o Delta Lake optimizations 
o Data quality frameworks on Lakehouse 
o Structured & unstructured data handling 
· Experience in Global Markets, Risk, Treasury, Trade Surveillance, or Regulatory Reporting. 
· Strong SQL knowledge with experience working on massive datasets (TB/PB scale). 
· Experience with CI/CD practices — Git, Jenkins, Bitbucket, build pipelines.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91172467
  • Position Id: 9071911
  • Posted 9 hours ago

Company Info

About TATA Consultancy Services Limited

Tata Consultancy Services (TCS) is the technology partner of choice for industry leading organizations worldwide. Since its inception in 1968, TCS has upheld the highest standards of innovation, engineering excellence, and customer service.

It has set an aspiration to become the world's largest AI-led technology services company and is enabling its clients to transform themselves across the full AI stack, from infrastructure to intelligence. Rooted in the heritage of the Tata Group, TCS is focused on creating long term value for its clients, its investors, its employees, and the community at large.

With a highly skilled workforce spread across 55 countries and 202 service delivery centers across the world, the company has been recognized as a top employer in six continents. With the ability to rapidly apply and scale new technologies, the company has built long term partnerships with its clients – helping them emerge as perpetually adaptive enterprises.

Many of these relationships have endured into decades and navigated every technology cycle, from mainframes in the 1970s to artificial intelligence today.

TCS sponsors 14 of the world’s most prestigious marathons and endurance events, including the TCS New York City Marathon, TCS London Marathon and TCS Sydney Marathon with a focus on promoting health, sustainability, and community empowerment.

About_Company_OneAbout_Company_Two
Contact the job poster
VS

Vigneshwaran Srinivasan

Recruiter @ TATA Consultancy Services Limited
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs