Job Title: Senior DataStage Developer
Location: Remote,
Across USA any Location
Full Time
Objective: Drive the end-to-end rationalization, reverse engineering, and automated validation of legacy DataStage environments migrating to modern Databricks architectures. This role is critical to eliminating legacy code redundancies, mapping complex data lineage, and implementing automated testing frameworks to guarantee zero data loss and business disruption during system cutovers.
Key Responsibilities:
Legacy Rationalization: Analyze DataStage .dsx/.isx exports and metadata to map end to end data lineage.
Code Elimination: Identify and isolate, dead jobs, redundant code, and duplicate logic to streamline migration waves.
Technical Translation: Provide functional logic explanations of complex parallel/server jobs to the PySpark development team.
Test Automation: Deploy automated test frameworks to execute large-scale data reconciliation between DataStage and Databricks.
Data Validation: Build automated scripts to validate data schemas, row counts, and complex transformations across dual run environments.
Regression Testing: Execute regression testing on newly refactored PySpark code against historical legacy data.
Migration Sign off: Document validation execution KPIs and formally sign off on data accuracy before live migration cutovers.
Technical Skills
Competencies
Legacy ETL: IBM InfoSphere DataStage (Parallel/Server jobs, Sequences) and operational metadata analysis.
Data Quality
Testing: Automated ETL Testing tools, PyTest, and Great Expectations.
Languages
Querying: Advanced SQL, Python, and XML/JSON parsing.
Target Platforms: Familiarity with Databricks, PySpark, and modern cloud data warehouses.