We are looking for a versatile Data Scientist / Data Engineer to join our data team. This hybrid role combines strong data engineering capabilities with advanced data science and machine learning expertise. The ideal candidate will design, build, and maintain robust data pipelines while developing predictive models, performing advanced analytics, and delivering actionable insights that drive business value. You will work at the intersection of data infrastructure and analytics, turning raw data into production-ready solutions and impactful machine learning applications. This position requires a strong technical foundation, problem-solving skills, and the ability to collaborate effectively with cross-functional teams.
Key Responsibilities
• Design, build, and maintain scalable, reliable data ingestion, processing, and transformation pipelines using tools such as Python, Spark, Databricks, Airflow, or cloud-native services (Azure, AWS).
• Develop and optimize data models, data warehouses, and lakehouse architectures to support both analytical and machine learning workloads.
• Build, train, deploy, and monitor machine learning and deep learning models for various use cases including prediction, classification, recommendation, and generative AI.
• Perform feature engineering, exploratory data analysis (EDA), and statistical analysis to support model development and business insights.
• Implement MLOps practices including model versioning, CI/CD for ML, containerization, and monitoring of production models.
• Conduct statistical analysis, A/B testing, causal inference, and predictive analytics to solve complex business problems.
• Ensure high data quality, implement data validation, monitoring, and governance processes across the data lifecycle.
• Work closely with business teams, analysts, and leadership to understand requirements, translate them into technical solutions, and communicate findings effectively.
• Optimize data pipelines and ML models for scalability, cost-efficiency, and performance at scale.
• Leverage modern data stack technologies including SQL, Python (Pandas, Scikit-learn, TensorFlow/PyTorch), Spark, Databricks, MLflow, and cloud platforms.
• Maintain clear documentation of data pipelines, models, and processes. Mentor junior team members and promote best practices.
• Stay up-to-date with emerging technologies in data engineering and data science and evaluate their potential value for the organization.
Qualifications
• Bachelors degree in Computer Science, Data Science, Engineering or related quantitative field
• 5+ years’ experience working in a hybrid Data Scientist and Data Engineer capacity
• Advanced proficiency in Python and strong command of SQL for complex data querying and manipulation
• Hands on experience with big data processing frameworks like Apache Spark or Databricks, and workflow orchestration tools like Apache Airflow
• Solid understanding of data warehousing, lakehouse architectures, and designing scalable ETL/ELT pipelines
• Practical experience building and deploying models using libraries
• Experience with cloud-native data services on Azure and AWS
• Knowledge of data quality standards, data validation techniques, and data security governance practices
• Strong analytical mindset with ability to turn raw data into structured actionable business solutions
• Excellent written and verbal communication skills, including the ability to clearly explain technical concepts to non-technical audiences