Role: AI & Data SME
Location: Remote
Contract: 6+ Months with Possible Extension
Experience: 10+ years
Mandatory Skills: Google Cloud data stack: BigQuery, Vertex AI Search, AlloyDB/Cloud SQL, Dataflow/Dataproc.
Job Description:
Role Summary
Serves as the senior subject-matter expert bridging enterprise data engineering and AI grounding for the Agent Factory. Owns the strategy and design for extracting, normalizing, and grounding enterprise data (Databricks, ServiceNow, Snowflake, OneDrive, Adobe, PDFs) for high-quality, reliable AI agent consumption, and advises on data architecture, RAG quality, and AI data governance.
Key Responsibilities
• Define the enterprise data-to-AI strategy: ingestion, normalization, embeddings, vector stores, and RAG grounding for agents.
• Architect and oversee data pipelines from source systems (Databricks, ServiceNow, Snowflake, OneDrive, Adobe, PDFs) into Google Cloud (BigQuery, Vertex AI Search, AlloyDB, Vector Search).
• Own retrieval quality: chunking strategies, embedding models, hybrid search, and grounding accuracy for Finance/Marketing/HR agents.
• Establish AI data governance: data privacy, PII controls, lineage, quality, and access management.
• Advise architects and developers on data readiness, feature/knowledge design, and evaluation of grounded responses.
• Guide model selection and cost/performance trade-offs for data-heavy agent workloads.
• Act as a senior technical authority in stakeholder discussions with Google on data-track dependencies.
• Mentor Data & AI Engineers and set standards for reusable data assets in the Agent Library.
Mandatory (Must-Have) Skills
• Deep enterprise data engineering: pipelines, ETL/ELT, data modelling, and data quality at scale.
• Hands-on with Google Cloud data stack: BigQuery, Vertex AI Search, AlloyDB/Cloud SQL, Dataflow/Dataproc.
• Strong RAG design: embeddings, vector databases, chunking, hybrid/semantic search, grounding evaluation.
• Experience integrating enterprise sources (Databricks, ServiceNow, Snowflake, SharePoint/OneDrive).
• Proficiency in Python and SQL; solid grasp of GenAI/LLM data patterns.
• Strong data governance, privacy (PII/DLP), and security knowledge.
Preferred (Good-to-Have) Skills
• Knowledge graphs and advanced retrieval (GraphRAG, corrective/self-RAG).
• Snowflake / Databricks certifications; Google Professional Data Engineer.
• Experience with AI evaluation frameworks and observability for data quality.
• Prior high-tech / semiconductor domain exposure.
Experience & Certifications
• 9–12+ years in data engineering / data architecture with recent GenAI/RAG delivery (Tier 5–6).
• Google Professional Data Engineer / relevant cloud data certification preferred.
• Onsite (USA), aligned to hours and data-track stakeholder collaboration.