Job: AI Engineer – Generative
& Agentic Workflows
Location: Sunnyvale,
CA/Austin, TX (Hybrid)
VISA: USC
Generative
AI Core Development
- Model Optimization & Fine-Tuning: Fine-tune, distill, and optimize open-source and
proprietary foundation models (e.g., Llama, Claude, GPT-4, Mistral) using
techniques like PEFT/LoRA and RLHF/DPO.
- Retrieval-Augmented Generation (RAG): Build scalable, low-latency Hybrid RAG architectures
integrating vector databases, graph databases, and semantic routing for
complex domain contexts.
Design and deploy multimodal pipelines (text, vision, audio, structured
data) to extract insights and generate rich artifacts.
Agentic
AI & Autonomous Systems
- Multi-Agent Architectures: Design and deploy multi-agent orchestration frameworks
(e.g., LangGraph, AutoGen, CrewAI, Semantic Kernel) with dedicated roles,
shared memory, and cross-agent negotiation strategies.
- Planning & Reasoning Protocols: Implement advanced reasoning paradigms—such as
Chain/Tree/Graph-of-Thought, ReAct loops, self-reflection, and
reflection-based error correction.
- Tool Augmentation & Function Calling: Connect agents to external APIs, databases, software
environments, and browser tools, enabling reliable function calling,
schema validation, and tool execution.
- Human-in-the-Loop (HITL) Workflows: Build human-in-the-loop safety checkpoints, approval
triggers, and oversight interfaces into autonomous agent loops.
Platform
Performance, Guardrails & MLOps
- Agentic Observability & Evaluation: Set up rigorous evaluation frameworks (LLM-as-a-judge,
trajectory tracing, latency profiling) using tools like LangSmith,
Phoenix, or Arize.
- Guardrails & Alignment: Implement strict safety, hallucination mitigation,
context-window optimization, and prompt injection defenses using guardrail
frameworks (e.g., NeMo Guardrails, Guardrails AI).
- Production Deployment: Scale agent workflows on cloud infrastructure
(AWS/Google Cloud Platform/Azure) with async task queues, durable execution state, and
low-latency API integration.
Key Skills & Technologies:
Languages
Python (Expert), TypeScript / Node.js (Plus)
Gen AI Frameworks
PyTorch, Hugging Face Transformers, vLLM, Ollama,
LangChain, LlamaIndex
Agentic Frameworks
LangGraph, AutoGen, CrewAI, Semantic Kernel, Temporal
Databases & Search
Qdrant, Pinecone, Milvus, Weaviate, Pgvector, Neo4j
Infrastructure & MLOps
Docker, Kubernetes, Ray, FastAPI, LangSmith, MLflow,
AWS/Google Cloud Platform/Azure
Experience
& Qualifications
- Experience: 6+ years of professional software engineering experience, with 2+ years
dedicated to building and deploying Gen AI and/or LLM applications in
production.
Experience building stateful LLM applications, custom RAG systems, or
autonomous agentic workflows deployed to real users.
- Strong Algorithmic Foundation: Solid understanding of transformer architectures,
attention mechanisms, vector embeddings, and non-deterministic state
machine design.
- Problem-Solving Mindset: Comfort dealing with model non-determinism, edge cases
in tool calling, and designing robust fallback mechanisms.