Role: Data Engineer
Location: Cincinnati, OH (Local Preferred; Open to Strong Remote Candidates)
Role: Data Engineer
Location: Cincinnati, OH (Local Preferred; Open to Strong Remote Candidates)
Duration: 12-month contract, with a high likelihood of extension/conversion
Requirement Description:
Overview
Isoft, a large-scale retail organization, is seeking a talented Data Engineer to join a high-performing data organization responsible for building and maintaining scalable data products that power analytics, machine learning, personalization, and real-time business operations.
This position partners closely with Engineering, Data Science, Machine Learning, and Product teams to design, develop, and optimize modern data solutions across both batch and streaming environments. The ideal candidate combines strong data engineering expertise with solid software engineering fundamentals and a passion for building reliable, high-quality data platforms at scale.
Success in this role requires hands-on experience with Spark-based processing, cloud-native technologies, performance optimization, automated data quality frameworks, and production-grade engineering practices.
What You'll Do:
Data Pipeline Development
Design, build, and maintain scalable data pipelines for ingestion, transformation, and integration across diverse data sources.
Develop reliable batch processing solutions using PySpark, SQL, and distributed data processing technologies.
Build data products that support analytics, machine learning, personalization, and operational reporting initiatives.
Data Processing & Performance Optimization
Optimize pipeline performance, scalability, reliability, and cost efficiency across large-scale datasets.
Leverage deep knowledge of Spark architecture, including partitioning, caching, shuffles, join strategies, and cluster tuning.
Troubleshoot and resolve production data processing bottlenecks.
Software Engineering Excellence
Apply strong Python and software engineering fundamentals to create maintainable, reusable, and scalable solutions.
Develop clean, testable code using object-oriented programming principles and engineering best practices.
Participate in code reviews and contribute to engineering standards across the team.
Data Quality & Reliability
Design and implement automated data quality frameworks and validation processes.
Develop testing strategies including unit, integration, regression, and end-to-end testing.
Ensure data accuracy, reliability, and compliance with organizational standards.
Analytics & Machine Learning Enablement
Partner with Data Scientists and Machine Learning Engineers to support feature engineering and machine learning workflows.
Help modernize and optimize core data assets supporting advanced analytics initiatives.
Collaboration & Documentation
Work closely with Engineering, Product, Data Science, and Machine Learning teams to deliver high-quality data solutions.
Create and maintain technical documentation, architectural diagrams, and operational processes.
Participate in Agile ceremonies, planning sessions, estimation activities, and sprint execution.
Required Qualifications :
3-5+ years of experience in Data Engineering, Software Engineering, or a related field.
Strong hands-on experience building and supporting production data pipelines using PySpark and SQL.
Deep understanding of Spark architecture and performance optimization techniques, including partitions, shuffles, caching, joins, and cluster tuning.
Strong Python development experience.
Experience working with distributed data systems and large-scale datasets.
Experience implementing automated testing, validation, and data quality frameworks.
Strong understanding of data modeling concepts and distributed system fundamentals.
Experience with Git, GitHub, CI/CD pipelines, and modern software development practices.
Experience working within Agile/Scrum development environments.
Strong communication, collaboration, and problem-solving skills.
Preferred Qualifications:
Experience with Google Cloud Platform (Google Cloud Platform) and BigQuery.
Experience with Databricks and cloud-native data platforms.
Experience with Microsoft Azure or other cloud environments.
Experience with Kafka or other streaming technologies.
Familiarity with machine learning workflows and feature engineering pipelines.
Exposure to MLOps tools and practices.
Experience supporting enterprise-scale analytics and machine learning platforms.
Technical Skills:
Data Processing
Programming
Python (Required)
Java (Preferred)
Cloud & Data Platforms
Streaming Technologies
DevOps & CI/CD
Git
GitHub
GitHub Actions
CI/CD Best Practices
Collaboration Tools
JIRA
Confluence
Microsoft Teams
Ideal Candidate Profile :
The ideal candidate understands not only how to use modern data processing frameworks such as Spark, but also the engineering principles that enable scalable and maintainable systems. This includes experience with:
Object-oriented programming
Data structures and algorithms
Memory management concepts
Variable scoping and application design
Reusable module development
Production-grade testing and reliability practices
You are passionate about building performant, scalable solutions while maintaining high standards for code quality, testing, and operational excellence.
Project Overview :
Initiative
Modernizing and optimizing core data assets that support advanced analytics, machine learning, and personalization initiatives.
Business Impact
You'll help build and maintain scalable data products that enable critical business capabilities, including customer analytics, machine learning, personalization, and real-time operational decision-making.
Day-to-Day Responsibilities:
Build and maintain scalable data pipelines using PySpark and SQL.
Optimize Spark workloads through partitioning, caching, shuffling, and join tuning.
Develop and support data products used across analytics and machine learning environments.
Implement automated testing and data quality frameworks.
Partner with Data Science, Machine Learning, Product, and Engineering teams.
Participate in Agile ceremonies, code reviews, and technical planning.
Create and maintain technical documentation.
Team & Culture
Collaborative, high-performing Agile environment.
Close partnership with Data Science, Product, Engineering, and Machine Learning teams.
Strong focus on engineering excellence, scalability, automation, and continuous improvement.
Opportunity to contribute to impactful analytics and machine learning initiatives within a major retail organization.
Why Apply?
This is an opportunity to join a modern data engineering team that is building scalable data platforms supporting advanced analytics, machine learning, and personalization at enterprise scale. You'll work with leading technologies, collaborate with cross-functional teams, and directly influence the future of data-driven decision-making.
Duration: 12-month contract, with a high likelihood of extension/conversion
Requirement Description:
Overview
Isoft, a large-scale retail organization, is seeking a talented Data Engineer to join a high-performing data organization responsible for building and maintaining scalable data products that power analytics, machine learning, personalization, and real-time business operations.
This position partners closely with Engineering, Data Science, Machine Learning, and Product teams to design, develop, and optimize modern data solutions across both batch and streaming environments. The ideal candidate combines strong data engineering expertise with solid software engineering fundamentals and a passion for building reliable, high-quality data platforms at scale.
Success in this role requires hands-on experience with Spark-based processing, cloud-native technologies, performance optimization, automated data quality frameworks, and production-grade engineering practices.
What You'll Do:
Data Pipeline Development
Design, build, and maintain scalable data pipelines for ingestion, transformation, and integration across diverse data sources.
Develop reliable batch processing solutions using PySpark, SQL, and distributed data processing technologies.
Build data products that support analytics, machine learning, personalization, and operational reporting initiatives.
Data Processing & Performance Optimization
Optimize pipeline performance, scalability, reliability, and cost efficiency across large-scale datasets.
Leverage deep knowledge of Spark architecture, including partitioning, caching, shuffles, join strategies, and cluster tuning.
Troubleshoot and resolve production data processing bottlenecks.
Software Engineering Excellence
Apply strong Python and software engineering fundamentals to create maintainable, reusable, and scalable solutions.
Develop clean, testable code using object-oriented programming principles and engineering best practices.
Participate in code reviews and contribute to engineering standards across the team.
Data Quality & Reliability
Design and implement automated data quality frameworks and validation processes.
Develop testing strategies including unit, integration, regression, and end-to-end testing.
Ensure data accuracy, reliability, and compliance with organizational standards.
Analytics & Machine Learning Enablement
Partner with Data Scientists and Machine Learning Engineers to support feature engineering and machine learning workflows.
Help modernize and optimize core data assets supporting advanced analytics initiatives.
Collaboration & Documentation
Work closely with Engineering, Product, Data Science, and Machine Learning teams to deliver high-quality data solutions.
Create and maintain technical documentation, architectural diagrams, and operational processes.
Participate in Agile ceremonies, planning sessions, estimation activities, and sprint execution.
Required Qualifications :
3-5+ years of experience in Data Engineering, Software Engineering, or a related field.
Strong hands-on experience building and supporting production data pipelines using PySpark and SQL.
Deep understanding of Spark architecture and performance optimization techniques, including partitions, shuffles, caching, joins, and cluster tuning.
Strong Python development experience.
Experience working with distributed data systems and large-scale datasets.
Experience implementing automated testing, validation, and data quality frameworks.
Strong understanding of data modeling concepts and distributed system fundamentals.
Experience with Git, GitHub, CI/CD pipelines, and modern software development practices.
Experience working within Agile/Scrum development environments.
Strong communication, collaboration, and problem-solving skills.
Preferred Qualifications:
Experience with Google Cloud Platform (Google Cloud Platform) and BigQuery.
Experience with Databricks and cloud-native data platforms.
Experience with Microsoft Azure or other cloud environments.
Experience with Kafka or other streaming technologies.
Familiarity with machine learning workflows and feature engineering pipelines.
Exposure to MLOps tools and practices.
Experience supporting enterprise-scale analytics and machine learning platforms.
Technical Skills:
Data Processing
Programming
Python (Required)
Java (Preferred)
Cloud & Data Platforms
Streaming Technologies
DevOps & CI/CD
Git
GitHub
GitHub Actions
CI/CD Best Practices
Collaboration Tools
JIRA
Confluence
Microsoft Teams
Ideal Candidate Profile :
The ideal candidate understands not only how to use modern data processing frameworks such as Spark, but also the engineering principles that enable scalable and maintainable systems. This includes experience with:
Object-oriented programming
Data structures and algorithms
Memory management concepts
Variable scoping and application design
Reusable module development
Production-grade testing and reliability practices
You are passionate about building performant, scalable solutions while maintaining high standards for code quality, testing, and operational excellence.
Project Overview :
Initiative
Modernizing and optimizing core data assets that support advanced analytics, machine learning, and personalization initiatives.
Business Impact
You'll help build and maintain scalable data products that enable critical business capabilities, including customer analytics, machine learning, personalization, and real-time operational decision-making.
Day-to-Day Responsibilities:
Build and maintain scalable data pipelines using PySpark and SQL.
Optimize Spark workloads through partitioning, caching, shuffling, and join tuning.
Develop and support data products used across analytics and machine learning environments.
Implement automated testing and data quality frameworks.
Partner with Data Science, Machine Learning, Product, and Engineering teams.
Participate in Agile ceremonies, code reviews, and technical planning.
Create and maintain technical documentation.
Team & Culture
Collaborative, high-performing Agile environment.
Close partnership with Data Science, Product, Engineering, and Machine Learning teams.
Strong focus on engineering excellence, scalability, automation, and continuous improvement.
Opportunity to contribute to impactful analytics and machine learning initiatives within a major retail organization.
Why Apply?
This is an opportunity to join a modern data engineering team that is building scalable data platforms supporting advanced analytics, machine learning, and personalization at enterprise scale. You'll work with leading technologies, collaborate with cross-functional teams, and directly influence the future of data-driven decision-making.