Data Engineer with Conversational AI

Hybrid in New York, NY, US • Posted 1 day ago • Updated 1 day ago
Contract W2
Hybrid
$60 - $70/hr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • conversational AI
  • LLM
  • Gen AI
  • RAG
  • Natural Language Processing

Summary

Job Title: Data Engineer with Conversational AI

Location: Remote(USA)

Experience: 15+ Years (Must)

Job Description:

Create high-quality, customer-specific synthetic data and own RAG / knowledge pipelines so each deployment ofCCAI, voice, and chat can be configured, grounded, demonstrated, and validated without using real customer PII.You design generation and ingestion pipelines and load data into the correct Google Cloud Platform and AWS services.What success looks likeEach customer engagement has a documented synthetic dataset covering the channels in scope Each in-scope customer has a working RAG / knowledge pipeline: corpus prepared, indexed, retrievable,and evaluated.Data and retrieval quality are good enough for configuration, evaluation, and stakeholder demos, andsafe enough for isolation and compliance expectations.Generation and indexing are parameterized and repeatable, not a one-off manual copy-paste percustomer.Key responsibilitiesAnalyze each customer s domain: intents, entities, knowledge topics, document types, languages, tone,and edge cases.Generate synthetic conversation transcripts for voice and chat, plus CCAI training/evaluation dialogues.Generate supporting content: customer/agent profiles, knowledge-base articles, FAQs, and structuredentity values.Schedule and document index refresh processes when customer knowledge changes.Use appropriate techniques whiledocumenting parameters and limitations.Validate realism, coverage, diversity, and absence of residual real-world PII in synthetic data and sourcecorpora.Maintain reusable generators, ingestion jobs, and quality checklists that can be parameterized percustomer.Partner with the Conversational Platform Specialist so loaded data and indexes actually drive thedeployed experience.Partner with DevOps so pipeline jobs, stores, and secrets are automated and isolated per customer.Required qualifications4+ years in data engineering, conversation design operations, applied NLP data work, or knowledge-pipeline engineering.Working knowledge of how conversational platforms consume training, FAQ, transcript, and retrieval-grounded knowledge data.Strong judgment on synthetic-data quality, retrieval quality, and privacy safety.Preferred qualificationsLLM-assisted synthetic data generation in a production or implementation setting.Familiarity with BigQuery, S3, and document stores used as knowledge sources.Multilingual data generation or evaluation experience.

Thanks

Thamiz

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10217051
  • Position Id: 9094771
  • Posted 1 day ago
Contact the job poster
VK

Vinoth Kumar

Recruiter @ ITBMS Inc.
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

New York, New York

•

9d ago

Easy Apply

Contract

Depends on Experience

Hybrid in New York, New York

•

Today

Easy Apply

Full-time

140,000 - 150,000

Hybrid in New York, New York

•

Today

Easy Apply

Full-time

140,000 - 150,000

New York, New York

•

Today

Full-time

USD 137,300.00 - 166,000.00 per year

Search all similar jobs