Data Scientist II - Big Data Engineer

Remote • Posted 7 hours ago • Updated 7 hours ago
Contract Independent
12 Months
No Travel Required
Remote
$65/hr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • ETL/ELT
  • Structured & Unstructured Data
  • Apache Spark
  • Databricks
  • Python
  • R
  • SQL
  • Microsoft Azure
  • Azure Data Lake Storage (ADLS)
  • Azure Data Factory (ADF)
  • Data Warehousing
  • Data Modeling
  • Data Schemas
  • Database Structures
  • Data Pipelines
  • Spark Architecture
  • RDDs
  • DataFrames
  • Spark SQL
  • Spark Job Optimization
  • Performance Optimization
  • Cost Optimization
  • Databricks Notebooks
  • Databricks Clusters
  • Databricks Jobs
  • Delta Lake
  • Unity Catalog
  • Data Validation
  • Data Quality
  • Data Governance
  • Metadata Management
  • Data Lineage
  • Data Cataloging
  • Data Security
  • Encryption
  • Access Controls
  • Auditing
  • CI/CD
  • DevOps
  • Version Control
  • Troubleshooting
  • Debugging
  • Agile
  • Cross-Functional Collaboration
  • Multicultural Environment
  • MLflow
  • Scikit-learn
  • TensorFlow

Summary

ONLY LOCAL TO TEXAS

Data Scientist (Big Data Engineer) II – Databricks / Azure 
Position: Data Scientist (Big Data Engineer) II
Openings: 2
Location: 100% Remote
Duration: 12 Months ( Up to 3 years extension )

Rate : $65/Hr on C2C
Key Responsibilities
  • Design, develop, and maintain scalable data pipelines using Apache Spark on Databricks.
  • Implement ETL/ELT workflows for structured and unstructured data.
  • Develop and optimize Spark jobs for performance and cost efficiency.
  • Build and maintain data models, schemas, and database structures supporting analytical and operational use cases.
  • Integrate Databricks solutions with Azure Data Factory and other Azure cloud services.
  • Work with Azure Data Lake Storage and data warehouse solutions.
  • Implement data validation and quality checks to ensure data accuracy, consistency, and reliability.
  • Contribute to data governance initiatives, including metadata management, data lineage, and data cataloging.
  • Implement data security measures, including encryption, access controls, and auditing.
  • Support compliance with applicable regulations, security requirements, and industry best practices.
  • Automate deployments using CI/CD pipelines, DevOps practices, and version control systems.
  • Work with Databricks notebooks, clusters, jobs, and Delta Lake.
  • Utilize Unity Catalog and/or Delta Lake to support data quality, governance, and security.
  • Troubleshoot and debug data pipelines, Spark applications, and related technical issues.
  • Collaborate with data scientists, data analysts, stakeholders, and cross-functional teams.
  • Work effectively within Agile and multicultural environments.
Required Qualifications
  • 4+ years of experience implementing ETL/ELT workflows for structured and unstructured data.
  • 4+ years of experience automating deployments using CI/CD tools.
  • 4+ years collaborating with data scientists, analysts, stakeholders, and cross-functional teams.
  • 4+ years designing and maintaining data models, schemas, and database structures.
  • 4+ years working with data storage solutions, including Azure Data Lake Storage and data warehouses.
  • 4+ years implementing data validation and data quality checks.
  • 4+ years contributing to data governance, metadata management, data lineage, and data cataloging.
  • 4+ years implementing data security measures, including encryption, access controls, and auditing.
  • 4+ years of proficiency in Python and R programming languages.
  • 4+ years of strong SQL querying and data manipulation experience.
  • 4+ years of experience with the Microsoft Azure cloud platform.
  • 4+ years of experience with DevOps, CI/CD pipelines, and version control systems.
  • 4+ years working in Agile and multicultural environments.
  • 4+ years of strong troubleshooting and debugging capabilities.
  • 3+ years designing and developing scalable data pipelines using Apache Spark on Databricks.
  • 3+ years optimizing Spark jobs for performance and cost efficiency.
  • 3+ years integrating Databricks with Azure Data Factory.
  • 3+ years ensuring data quality, governance, and security using Unity Catalog or Delta Lake.
  • 3+ years of strong understanding of Apache Spark architecture, RDDs, DataFrames, and Spark SQL.
  • 3+ years of hands-on experience with Databricks notebooks, clusters, jobs, and Delta Lake.
Preferred Qualifications
  • Knowledge of machine learning libraries such as:
    • MLflow
    • Scikit-learn
    • TensorFlow
  • Databricks Certified Associate Developer for Apache Spark certification.
  • Microsoft Certified: Azure Data Engineer Associate certification.
Role Overview:
The position involves designing, developing, and optimizing scalable data pipelines and big data solutions using Azure, Databricks, and Apache Spark. The role also involves data quality, governance, security, CI/CD, and collaboration with data engineering, analytics, and business teams.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10211499
  • Position Id: 9100334
  • Posted 7 hours ago
Contact the job poster
Srija Nakka

Srija Nakka

Recruiter @ Cogent IBS, Inc
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

•

Today

Easy Apply

Third Party, Contract

$50+

Remote or Hybrid in San Francisco, California

•

Yesterday

Easy Apply

Contract

Depends on Experience

Remote or Chantilly, Virginia

•

Today

Easy Apply

Contract

$$50/hr

Remote or Texas

•

Today

Full-time

USD 118,450.00 - 236,900.00 per year

Search all similar jobs