AI Architect - Santa Clara, CA - Fulltime

Santa Clara, CA, US • Posted 23 hours ago • Updated 23 hours ago
Full Time
Part Time
On-site
Company Branding Image
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

Summary

Job Title: AI Architect

Location: Santa Clara, CA

Duration: Fulltime

Key Requirements: Looking for a strong customer-facing resource with excellent communication and stakeholder management skills. Bay Area local candidates are highly preferred.

  • Will work on the intelligence layer for multiple programs owns all model quality, RAG accuracy, prompt engineering, and AI safety across applications
  • Socratic tutor persona, adaptive learning recommendation engine, multi-modal AI (text and voice), RAG evaluation framework, and feedback loop into retrieval
  • 6-LLM call chain orchestration (NeMoGuardrails intent classification query rewriting RAG synthesis), , and compatibility check logic
  • Production-grade AI quality from launch this is not a research or prototyping role; accuracy thresholds, latency requirements, and safety guardrails must pass InfoSec adversarial testing before Release 1

Required Skills

  • Total IT 15+ Years
  • 4-7 years of software engineering with at least 2 years focused on LLM application development in production not research, not demos, not internal tools with 10 users
  • Has shipped an LLM-powered feature or product to production where real users depend on the accuracy and the engineer owns the quality metrics
  • Has owned an AI safety or guardrails implementation for a customer-facing product not just added an off-the-shelf filter; designed and tested the safety layer
  • Has built RAG evaluation pipelines and used them to make go/no-go release decisions accuracy gating is part of the workflow.
  • Has profiled and optimized a multi-step LLM call chain for latency

LLM Application Development

  • LLM prompt engineering system prompts, few-shot examples, chain-of-thought, instruction following Expert Must-have
  • Multi-step LLM chain orchestration LangChain, LlamaIndex, or custom orchestration Expert Must-have
  • Multi-turn conversation design context window management, conversation summarization, session memory Advanced Must-have
  • Streaming LLM response handling token-by-token streaming, partial response rendering Advanced Must-have
  • Model selection and benchmarking matching model size to task; balancing latency, cost, and accuracy Advanced Must-have

RAG Pipeline Design & Quality

  • RAG pipeline design chunking strategy, embedding model selection, retrieval configuration Expert Must-have
  • Vector similarity search tuning index parameters, similarity thresholds, retrieval depth Advanced Must-have
  • Reranking cross-encoder rerankers, relevance scoring Advanced Must-have
  • RAG evaluation frameworks RAGAS, TruLens, or equivalent; automated eval pipelines Advanced Must-have
  • Hybrid search combining dense vector retrieval with BM25 or keyword search Proficient Nice to have

AI Safety & Guardrails

  • Prompt injection detection and mitigation Advanced Must-have
  • Jailbreak testing and red-teaming LLM systems Advanced Must-have
  • Content safety classifier integration Advanced Must-have
  • Hallucination detection and mitigation strategies Advanced Must-have
  • Topical control enforcing scope boundaries on LLM responses Advanced Must-have

Evaluation & Production Quality

  • Automated evaluation pipeline design test set curation, metric selection, regression detection Advanced Must-have
  • A/B evaluation methodology for prompt and model changes Proficient Must-have
  • Latency profiling for LLM call chains identifying bottlenecks across multi-step pipelines Proficient Must-have
  • Feedback loop design user signal collection, signal-to-retrieval-weight integration Proficient Must-have
  • Production model monitoring accuracy drift detection, quality degradation alerting Proficient Must-have

Development

  • Python ML/AI application development, async programming Expert Must-have
  • API design for AI services streaming endpoints, error handling, timeout management Advanced Must-have
  • Embedding model operations model selection, batch embedding, index updates Advanced Must-have

Nice to Have

  • Adaptive learning systems or personalization engine experience
  • Knowledge graph integration with RAG
  • Multi-agent orchestration patterns
  • ServiceNow API integration
  • Prior experience building AI products on NVIDIA infrastructure

Regards

Rajesh

Arrowminds Inc

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91166603
  • Position Id: OOJ - 1857-858-1789426885
  • Posted 23 hours ago

Company Info

About Arrowminds inc

Arrowminds staffing practice delivers high-quality staffing services built on industry best practices. We work with our clients to recruit and retain the best information technology talent possible. Our team manages the acquisition and deployment of professionals for temporary staffing needs. Our flexible recruiting process provides client with consistent, quick access to skilled professionals.

Business managers need a knowledgeable technology partner to help them select the best-fit technology platform & business applications to effectively capture the maximum ROI benefits. Arrowminds can offer objective advice on choosing the right technology solutions for your business needs. We also deliver cost-effective software customizations, infrastructure support using global delivery model.

Arrowminds provide onsite/offshore  resources for any IT technology, Healthcare, financial and manufacturing industries.

About_Company_OneAbout_Company_Two
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs