AI Data Engineer- Databricks & Snowflake
Location: Tallahassee, FL, USA
Duration: 12 Months + Extension
Bill Rate: $90/hr on C2C
Job Type: C2C/1099 Contract
Client: To Be Discussed Later
Work Authorization: US-Citizen, H-1B, OPT-EAD, GC-EAD
Job Description:
We are seeking an experienced AI Data Engineer Databricks & Snowflake to design, build, and optimize modern cloud-based data platforms that support enterprise analytics and AI initiatives. The ideal candidate will have strong expertise in Databricks, Snowflake, Apache Spark, Python, SQL, cloud platforms (AWS/Azure/Google Cloud Platform), and Generative AI technologies. This role involves developing scalable data pipelines, integrating AI/ML capabilities, and enabling data-driven decision-making through modern data architecture and AI-powered solutions.
Key Responsibilities:
Data Engineering & Platform Development
- Design, develop, and maintain scalable data pipelines using Databricks, Apache Spark, and Snowflake.
- Build batch and real-time ETL/ELT workflows for enterprise data processing.
- Develop data ingestion frameworks from structured, semi-structured, and unstructured data sources.
- Optimize data pipelines for scalability, reliability, and performance.
- Implement Delta Lake architecture and data lakehouse best practices.
Snowflake Data Warehouse
- Design and implement enterprise data warehouse solutions using Snowflake.
- Develop data models including star schema, snowflake schema, and dimensional modeling.
- Optimize Snowflake performance through clustering, partitioning, caching, and warehouse tuning.
- Implement secure data sharing, governance, and role-based access control (RBAC).
- Develop SQL-based transformations, stored procedures, streams, and tasks.
Databricks Engineering
- Develop notebooks, workflows, and jobs using Databricks.
- Implement Spark applications using PySpark and Spark SQL.
- Build Delta Live Tables (DLT) and Auto Loader pipelines.
- Optimize Spark jobs for high-performance distributed data processing.
- Implement data quality validation and monitoring frameworks.
AI & Generative AI Integration
- Develop AI-enabled data platforms leveraging Generative AI and Large Language Models (LLMs).
- Build Retrieval-Augmented Generation (RAG) pipelines using enterprise data stored in Databricks and Snowflake.
- Integrate vector databases and embedding models for semantic search.
- Develop AI-powered analytics, document intelligence, and conversational AI solutions.
- Implement prompt engineering techniques and LLM integrations using OpenAI, Azure OpenAI, or Google Vertex AI.
- Build agentic AI workflows and intelligent automation solutions.
Cloud & DevOps
- Design cloud-native data solutions on AWS, Azure, or Google Cloud Platform (Google Cloud Platform).
- Develop CI/CD pipelines for automated deployment of data engineering solutions.
- Implement Infrastructure as Code (IaC) using Terraform or similar tools.
- Monitor cloud infrastructure, data pipelines, and platform performance.
- Ensure cloud security, governance, and compliance.
Data Integration
- Integrate enterprise applications using APIs, streaming platforms, and messaging services.
- Develop data ingestion pipelines from ERP, CRM, SaaS, and third-party systems.
- Implement Change Data Capture (CDC) solutions.
- Build event-driven architectures using Kafka, Event Hubs, or Pub/Sub.
Performance Optimization
- Tune Spark workloads and optimize distributed processing performance.
- Optimize Snowflake queries and warehouse utilization.
- Implement partitioning, caching, indexing, and workload management.
- Improve overall platform scalability, reliability, and cost efficiency.
Data Governance & Security
- Implement enterprise data governance and metadata management.
- Ensure data quality, lineage, cataloging, and compliance.
- Develop secure data access models and encryption strategies.
- Implement role-based security and auditing.