We are seeking a Data Engineer with strong AI/GenAI expertise to design, build, and productionize scalable data platforms and AI-driven solutions. The ideal candidate will have strong hands-on experience in data engineering, cloud technologies, Python, and modern AI/LLM technologies.
Key Responsibilities
Design, develop, and maintain scalable data pipelines for structured and unstructured data.
Build production-grade ETL/ELT pipelines using Python, SQL, Spark/PySpark, and cloud technologies.
Develop and optimize data processing solutions using Databricks, Snowflake, or similar platforms.
Integrate data engineering pipelines with AI/ML and Generative AI applications.
Build and support RAG pipelines, including data ingestion, chunking, embeddings, retrieval, and vector search.
Work with LLMs, LangChain, LlamaIndex, Azure OpenAI/OpenAI, and related AI frameworks.
Develop data pipelines supporting AI agents and Agentic AI applications.
Implement data quality, governance, security, monitoring, and performance optimization.
Collaborate with Data Scientists, ML Engineers, AI Engineers, and application teams to productionize AI solutions.
Deploy and monitor AI/data workloads using CI/CD and MLOps practices.
Troubleshoot data pipeline and production issues and continuously improve reliability and scalability.
Required Skills
7+ years of experience in Data Engineering.
Strong Python and SQL skills.
Hands-on experience with Apache Spark / PySpark.
Experience building ETL/ELT and data pipelines.
Strong experience with AWS, Azure, or Google Cloud Platform.
Experience with Databricks and/or Snowflake.
Experience with Generative AI, LLMs, or AI/ML applications.
Hands-on experience with RAG, embeddings, vector databases, or semantic search.
Experience with LangChain, LlamaIndex, Azure OpenAI/OpenAI, or similar frameworks.
Understanding of MLOps, CI/CD, Docker, and Kubernetes is preferred.