Data Analyst – AI & Synthetic Data- 5+ yrs- New York, United States- Onsite

New York, NY, US • Posted 3 hours ago • Updated 2 hours ago
Contract W2
Contract Corp To Corp
Contract Independent
6 Months
No Travel Required
On-site
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • detail-oriented and technical Data Analyst to lead synthetic data curation
  • relational modeling
  • and dataset validation for an AI-driven
  • multi-industry demo platform. In thisrole
  • you will design synthetic data models across four distinct domains (Financial
  • Healthcare
  • Manufacturing
  • and Retail). You will write programmatic generation scripts to produce realistic
  • non-PII datasets
  • load them into an Oracle Autonomous Database (ADB)
  • and validate datasetcoverage to ensure accurate Natural-Language-to-SQL (NL?SQL) performance.Key Responsibilities1. Relational Data Modeling? Design realistic
  • normalized relational data models (dimension and fact tables) for fourindustry verticals: Financial
  • and Retail.? Ensure schema design supports representative industry use cases and natural-languagequerying patterns without exposing PHI/PII or sensitive data.? Map data requirements to align with the Oracle AI Database Agent metadata indexingand query execution engine.2. Synthetic Data Generation & Loading? Develop and execute reusable Python/SQL scripts to programmatically synthesize dataat scale.? Generate target volume thresholds (~50–500 rows per dimension table;~10
  • 000–100
  • 000 rows per fact table per industry).? Perform one-time seed data loads into designated schemas within the shared OracleAutonomous Database.3. Query Validation & Golden Dataset Development? Validate data quality
  • primary/foreign key integrity
  • and representative query coverageacross all four schemas.? Collaborate with the AI/Gemini engineering team to build prompt libraries
  • goldendatasets
  • and sample Critical User Journeys (CUJs).? Perform end-to-end testing (Question ? SQL ? Result Execution) to confirm theaccuracy of generated SQL queries.? Participate in cross-industry isolation testing to ensure dataset security across schemaboundaries.4. Post-Deployment & UAT Support? Assist with User Acceptance Testing (UAT) bug fixing and schema refinements based ondemo feedback

Summary

Job Description :

Job Requirements Job SummaryWe are seeking a detail-oriented and technical Data Analyst to lead synthetic data curation,relational modeling, and dataset validation for an AI-driven, multi-industry demo platform. In thisrole, you will design synthetic data models across four distinct domains (Financial, Healthcare,Manufacturing, and Retail). You will write programmatic generation scripts to produce realistic,non-PII datasets, load them into an Oracle Autonomous Database (ADB), and validate datasetcoverage to ensure accurate Natural-Language-to-SQL (NL?SQL) performance.Key Responsibilities1. Relational Data Modeling? Design realistic, normalized relational data models (dimension and fact tables) for fourindustry verticals: Financial, Healthcare, Manufacturing, and Retail.? Ensure schema design supports representative industry use cases and natural-languagequerying patterns without exposing PHI/PII or sensitive data.? Map data requirements to align with the Oracle AI Database Agent metadata indexingand query execution engine.2. Synthetic Data Generation & Loading? Develop and execute reusable Python/SQL scripts to programmatically synthesize dataat scale.? Generate target volume thresholds (~50–500 rows per dimension table;~10,000–100,000 rows per fact table per industry).? Perform one-time seed data loads into designated schemas within the shared OracleAutonomous Database.3. Query Validation & Golden Dataset Development? Validate data quality, primary/foreign key integrity, and representative query coverageacross all four schemas.? Collaborate with the AI/Gemini engineering team to build prompt libraries, goldendatasets, and sample Critical User Journeys (CUJs).? Perform end-to-end testing (Question ? SQL ? Result Execution) to confirm theaccuracy of generated SQL queries.? Participate in cross-industry isolation testing to ensure dataset security across schemaboundaries.4. Post-Deployment & UAT Support? Assist with User Acceptance Testing (UAT) bug fixing and schema refinements based ondemo feedback.Work Experience

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91143549
  • Position Id: 4793-31204-1787669979
  • Posted 3 hours ago

Company Info

About iMedhas Consulting Services

Welcome to iMedhas Consulting Services. We are an IT consulting and services enterprises with precision expertise in Digital Transformations, Big data and Analytics. Through our expert team, we provide greater adaptabilities to disseminate with new technologies to increase the overall productivity, operational efficiency, and productivity of a particular business process of an organization.

The goal is the achievement of higher customer satisfaction and providing lucrative returns on investment to the business. On the other hand, we deliver content management expertise, ERP (SAP) systems integration, EAI, and Information Management services.

About_Company_One
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

New York, New York

Today

Easy Apply

Contract, Third Party

Depends on Experience

New York, New York

Today

Easy Apply

Third Party, Contract

Depends on Experience

Columbus, Ohio

Today

Easy Apply

Contract, Third Party

Depends on Experience

Search all similar jobs