Role: AI Data Engineer
Location: Remote
Must have: Agentic AI, Databricks, Datalake.
If someone has Genie and Agent Bricks experience will be great
Job requirements
We are seeking a highly skilled and hands-on Databricks Specialist to lead the design, development, deployment, and optimization of enterprise-scale Artificial Intelligence solutions built on the Databricks Lakehouse Platform. This role is focused on the development of next-generation AI applications leveraging Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), Agentic AI architectures, Vector Search, Machine Learning Operations (MLOps), and modern cloud-native engineering practices. The ideal candidate will possess deep expertise in Databricks, Python, PySpark, SQL, MLflow, and modern AI frameworks, along with practical experience designing production-ready solutions that integrate advanced AI capabilities into real-world business processes. You will work closely with business stakeholders, data scientists, software engineers, enterprise architects, and client leadership teams to transform complex business challenges into scalable, secure, and high-performing AI-driven solutions. This role requires a strong combination of technical expertise, solution architecture capabilities, client engagement skills, and a passion for emerging AI technologies. Key Responsibilities Generative AI Solution Development Design and implement enterprise-grade Generative AI solutions using Databricks. Develop advanced Retrieval-Augmented Generation (RAG) systems for enterprise knowledge management. Build Agentic AI solutions utilizing orchestration frameworks and autonomous workflows. Develop intelligent assistants, copilots, AI agents, and domain-specific virtual experts. Design scalable prompt engineering frameworks and reusable AI components. Optimize retrieval strategies to improve answer quality, accuracy, and relevance. Develop evaluation frameworks for measuring AI system effectiveness. Databricks Platform Engineering Build and maintain robust data pipelines on Databricks. Develop scalable ETL and ELT processes using PySpark and SQL. Implement lakehouse architectures using Delta Lake. Design efficient data ingestion pipelines for structured, semi-structured, and unstructured data. Create reusable Databricks workflows and job orchestration frameworks. Optimize Databricks workloads for performance and cost efficiency. Implement Unity Catalog-based governance and access management. Retrieval-Augmented Generation (RAG) Design end-to-end RAG systems using enterprise data sources. Develop document ingestion and indexing pipelines. Implement chunking and document segmentation strategies. Build embedding generation workflows. Configure vector databases and retrieval mechanisms. Perform semantic search optimization. Implement hybrid search approaches using vector and keyword search. Improve context retrieval precision and relevance. Agentic AI Development Design autonomous AI agents capable of reasoning and executing tasks. Build multi-agent architectures for complex enterprise use cases. Implement workflow orchestration across multiple systems. Design tool-calling mechanisms. Create memory-based agent systems. Develop planning and reasoning capabilities. Build AI agents that integrate with enterprise services and APIs. Monitor agent performance and continuously improve operational effectiveness. Model Lifecycle Management Build, train, fine-tune, evaluate, and deploy machine learning models. Implement model versioning using MLflow. Establish model governance and auditability practices. Develop model performance monitoring solutions. Implement automated retraining pipelines. Ensure reproducibility of AI and machine learning experiments. Manage end-to-end model lifecycle activities. Databricks Mosaic AI Build enterprise applications using Mosaic AI. Leverage Mosaic AI Gateway for secure model access. Configure model serving infrastructure. Implement LLM evaluation frameworks. Design AI governance and monitoring solutions. Work with foundation models available through Databricks. Build production-grade applications utilizing Mosa