Data Engineer

Austin, TX, US • Posted 1 day ago • Updated 2 hours ago
Full Time
On-site
Compensation information provided in the description
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Brand
  • Research
  • Manufacturing Operations
  • Scratch
  • ROOT
  • Computer Science
  • Computer Engineering
  • Statistics
  • Mechanical Engineering
  • Manufacturing Engineering
  • Operations Research
  • Applied Mathematics
  • Data Science
  • Modeling
  • Data Analysis
  • Python
  • SQL
  • Regression Analysis
  • Clustering
  • Time Series
  • IoT
  • Semiconductors
  • Aerospace
  • Energy
  • Telecommunications
  • Manufacturing
  • Sensors
  • Quality Management
  • Apache Spark
  • PySpark
  • Cloud Computing
  • Analytics
  • Statistical Process Control
  • Artificial Intelligence
  • Data Engineering
  • Data Modeling
  • Star Schema
  • Unity
  • ADF
  • DS
  • DirectShow
  • Machine Learning Operations (ML Ops)
  • Machine Learning (ML)
  • Amazon SageMaker
  • Continuous Integration
  • Continuous Delivery
  • Generative Artificial Intelligence (AI)
  • LangChain
  • Business Intelligence
  • Reporting
  • Microsoft Power BI
  • Tableau
  • KPI
  • MES
  • SCADA
  • Programmable Logic Controller
  • Databricks
  • eXist

Summary

  • Location: Austin, Texas
  • Type: Contract
  • Job #107180


Type: W2, 4 months (HIGH likelihood of extension to 12 months)
Pay: $97.00-$107.40/hour
Location: Austin, TX | Onsite, Monday-Friday
Start Date: ASAP
Job Summary:

The team is building a brand-new AI application/system from the ground up and needs an experienced Data Scientist who can dig into manufacturing, quality, and operational data, build data models, and help identify new ways AI can support the business.
Most of the job is in the data: understanding how a process behaves, cleaning noisy and incomplete signals, defining "normal" vs. "abnormal" with people who run the line, creating features, validating against real outcomes, and explaining limits when the data cannot support a model.
You will not own the data platform. This is not a research lab role and not a platform-engineering role. This is an opportunity to help build an AI application from the ground up, work with extensive manufacturing and quality data, and contribute to solutions that can have a significant impact across manufacturing operations.
Job Responsibilities:
  • Frame manufacturing problems with plant and engineering partners; push back when labels, ground truth, or "accuracy" expectations are not real.
  • Explore, clean, and join fragmented operational data, including machines, sensors, quality, maintenance, production, and MES/historian extracts.
  • Dig into large amounts of manufacturing and quality data and help build the system from scratch.
  • Build and validate statistical and machine-learning models for anomaly, quality, health, and process monitoring.
  • Report false positives/negatives and business cost, not only a leaderboard metric.
  • Create features and define "normal" vs. "abnormal" behavior based on real operational data.
  • Hand usable outputs to engineers and operators, including thresholds, explanations, and "what to do when this fires."
  • Support models after they are in use.
  • Work with data engineering and software on pipelines, Databricks, and production. You are the customer of the platform, not the person hired to build it.
  • Review the available data, bring creative ideas to the team, and identify additional opportunities to help the business with AI.
  • Work on problems such as process drift, abnormal machine behavior, quality prediction, equipment health, bottlenecks, downtime, scrap/rework, and root-cause support.
Required Qualifications:
  • Bachelor's or master's degree in a quantitative, technical, or engineering field such as Data Science, Computer Science, Computer Engineering, Statistics, Industrial/Mechanical/Manufacturing Engineering, Operations Research, Applied Mathematics, or a related field.
  • 8-10+ years of applied data science experience, including analysis, feature work, statistical or ML modeling on real operational or business datasets.
  • Strong experience with data modeling, data pipelines, and data analytics, particularly building something useful out of messy data.
  • Strong Python and SQL skills with evidence of working with large, messy tables, not only notebooks on clean extracts.
  • Experience producing models you can defend, including classification, regression, clustering, anomaly detection, or time series, with a clear target and validation approach.
  • Experience creating features from machine, sensor, process, quality, maintenance, or other operational data.
  • Ability to explore, clean, analyze, and derive meaningful findings from incomplete or fragmented datasets.
  • Ability to explain how a model was validated and determine when it is wrong or right.
  • Comfort telling stakeholders when a model should not ship.
  • Ability to learn an unfamiliar plant process quickly.
Preferred Qualifications
  • Experience in manufacturing, industrial IoT, semiconductor, automotive, aerospace, energy, telecom/operations, or equipment-heavy environments.
  • Manufacturing experience working with machine, sensor, process, quality, maintenance, or MES/historian data.
  • Experience with Databricks, Spark/PySpark, or similar cloud analytics.
  • Familiarity with MLOps concepts such as model tracking, monitoring, and drift while partnering with platform teams.
  • Experience with SPC, model explainability, historians, or MES data.
  • AI application/platform experience is a plus but not required.
What This Role Is NOT:

This position is NOT primarily focused on:
  • Data engineering/data modeling centered on lakehouse, medallion, star schema, Unity Catalog, ADF, or building "pipelines for the DS team"
  • MLOps/ML platform work such as Airflow, SageMaker plumbing, FastAPI services, or CI/CD with little hands-on dataset analysis
  • GenAI product work such as RAG chatbots, LangChain agents, or Copilot applications as the primary experience
  • BI/reporting work centered on Power BI/Tableau KPI applications without model development
  • Simply listing MES, SCADA, PLC, JPH, Databricks, or anomaly detection without being able to explain the dataset, finding, model, and validation result

Those skills exist on the team or in partner teams. We need the person who works the data.


#LI-LS1
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10111081
  • Position Id: d529de25fbdf0db07516555908118cf6
  • Posted 1 day ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Austin, Texas

•

2d ago

Easy Apply

Contract

$65 - $75

Taylor, Texas

•

Today

Full-time

USD 90,000.00 - 174,500.00 per year

No location provided

•

Today

Full-time

USD 135,000.00 - 160,000.00 per year

Remote

•

Today

Full-time

USD 112,000.00 - 179,000.00 per year

Search all similar jobs