AI Data Engineer

Remote • Posted 44 minutes ago • Updated 44 minutes ago
Full Time
Occasional Travel Required
Remote
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • A/B Testing
  • Amazon Neptune
  • Amazon SageMaker
  • Amazon Web Services
  • Analytical Skill
  • Apache Spark
  • Artificial Intelligence
  • Collaboration
  • Computer Science
  • Conflict Resolution
  • Data Engineering
  • Data Governance
  • Data Quality
  • Data Science
  • Databricks
  • Distributed Computing
  • Distribution
  • Documentation
  • Evaluation
  • Generative Artificial Intelligence (AI)
  • Good Clinical Practice
  • Google Cloud Platform
  • scikit-learn
  • Vector Databases
  • TensorFlow
  • Real-time
  • Unstructured Data

Summary

Role Overview:

The AI Data Engineer will specialize in building and optimizing machine learning data pipelines, focusing on AI model tracking, lifecycle management, and integration with AI governance systems. This role combines data engineering expertise with AI/ML knowledge to support the organization's broader data and AI infrastructure initiatives.

 

Key Responsibilities:

• Design and implement specialized data pipelines for AI model metadata, training data lineage, and model performance metrics tracking.

• Build data infrastructure on Databricks leveraging Spark for large-scale distributed dataset processing.

• Develop MCP servers and enable AI data distribution via MCP.

• Develop feature engineering pipelines and data preprocessing workflows for AI model training and inference.

• Implement model versioning, experiment tracking, and model registry integration using MLflow or similar tools.

• Create automated workflows for AI agent discovery, classification, and inventory management across the enterprise.

• Design and maintain knowledge graph structures for representing AI model relationships, dependencies, and data lineage.

• Build real-time data pipelines for AI model monitoring, drift detection, and performance tracking.

• Develop data quality frameworks specific to AI training datasets and validation data.

• Collaborate with data scientists to optimize data access patterns and feature store implementations.

• Implement security and compliance controls for sensitive AI training data and model artifacts.

• Create comprehensive documentation for AI data architectures, schemas, and integration patterns.

Required Skills and Qualifications:

• Bachelor's or Master's degree in Computer Science, Data Science, Machine Learning, or related field.

• 5-7 years of hands-on experience in data engineering, with at least 2 years focused on AI/ML workloads.

• Expert proficiency in Python and experience with ML frameworks such as PyTorch, TensorFlow, or scikit-learn.

• Strong experience with Databricks, Apache Spark, and distributed computing for ML workflows.

• Deep understanding of the machine learning lifecycle, including model training, deployment, and monitoring processes.

• Experience with feature engineering, data preprocessing techniques, and ML data pipelines.

• Knowledge of vector databases, embeddings, and similarity search for AI applications.

• Proficiency in SQL for structured and unstructured data management.

• Understanding of data governance, model governance, and AI ethics principles.

• Strong analytical and problem-solving capabilities with attention to data quality.

• Excellent collaboration skills for working with data scientists, ML engineers, and architects.

Preferred/Nice-to-Have Skills:

• Experience with generative AI applications, including RAG (Retrieval-Augmented Generation) and fine-tuning.

• Knowledge of LangChain, HuggingFace, or other GenAI frameworks.

• Familiarity with Azure ML, AWS SageMaker, or Google Vertex AI platforms.

• Experience with graph databases (Neo4j, Amazon Neptune) for knowledge graph implementation.

• Understanding of AI model explainability and interpretability techniques.

• Experience with A/B testing frameworks for ML model evaluation.

• Certification in Databricks, AWS, Azure, or Google Cloud Platform AI/ML services.

• Publications or contributions to open-source ML projects

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90888109
  • Position Id: suhael.ahamed@vegaintellisoft.com
  • Posted 44 minutes ago

Company Info

About Vega Intellisoft Inc.

Founded in 2004, Vega IntelliSoft excels in innovative products , IT services & staffing solutions. The company delivers cutting-edge technology and skilled professionals, ensuring exceptional service and customer satisfaction across various industries.

About_Company_OneAbout_Company_Two
Contact the job poster
YA

Yuvaraj Anbu

Recruiter @ Vega Intellisoft Inc.
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs