• Create high-quality, customer-specific synthetic data.
• Own RAG / knowledge pipelines for each deployment of CCAI, voice, and chat.
• Configure, ground, demonstrate, and validate deployments without using real customer PII.
• Design generation and ingestion pipelines and load data into correct Google Cloud Platform and AWS services.
• Document synthetic dataset for each customer engagement, covering channels in scope.
• Ensure each in-scope customer has a working RAG / knowledge pipeline: corpus prepared, indexed, retrievable, and evaluated.
• Validate data and retrieval quality for configuration, evaluation, and stakeholder demos.
• Ensure safe for isolation and compliance expectations.
• Parameterize and repeat generation and indexing, not a one-off manual copy-paste per customer.
• Analyze each customer’s domain: intents, entities, knowledge topics, document types, languages, tone, and edge cases.
• Generate synthetic conversation transcripts for voice and chat, plus CCAI training/evaluation dialogues.
• Generate supporting content: customer/agent profiles, knowledge-base articles, FAQs, and structured entity values.
• Schedule and document index refresh processes when customer knowledge changes.
• Use appropriate techniques while documenting parameters and limitations.
• Validate realism, coverage, diversity, and absence of residual real-world PII in synthetic data and source corpora.
• Maintain reusable generators, ingestion jobs, and quality checklists that can be parameterized per customer.
• Partner
• The Conversational Platform Specialist is responsible for loading data and indexes to drive the deployed experience.
• They must work with DevOps to automate and isolate pipeline jobs, stores, and secrets per customer.
• Required qualifications include 4+ years of experience in data engineering, conversation design operations, applied NLP data work, or knowledge-pipeline engineering.
• They should have a working knowledge of how conversational platforms consume training, FAQ, transcript, and retrieval-grounded knowledge data.
• Strong judgment on synthetic-data quality, retrieval quality, and privacy safety is necessary.
• Preferred qualifications include experience with LLM-assisted synthetic data generation in a production or implementation setting, familiarity with BigQuery, S3, and document stores used as knowledge sources, and multilingual data generation or evaluation experience.
• Work experience of 7-10 years is required.