Duration: 6+ Months
Location: Dallas, TX / New York, NY / Louisville, KY / Washington, DC
The LLM Engineer will design, build, optimize, deploy, and operate Large Language Model (LLM) and Small Language Model (SLM) capabilities supporting an enterprise Agent Factory. This role focuses on production-grade AI applications, agentic workflows, Retrieval-Augmented Generation (RAG), prompt and context engineering, model evaluation, fine-tuning, model serving, private AI hosting, and LLMOps.
The engineer will collaborate with AI architects, engineering leads, platform teams, security, enterprise architecture, product owners, and domain teams to develop secure, reliable, reusable, and enterprise-ready AI capabilities.
Required Skillset:
- 5+ years of software engineering, AI engineering, machine learning engineering, or platform engineering experience.
- 2+ years of hands-on LLM, Generative AI, or agentic AI development experience.
- Production experience building LLM applications using RAG, prompts, APIs, and cloud or on-premise platforms.
- Experience with model evaluation, prompt testing, and AI quality measurement.
- Experience integrating AI capabilities into enterprise applications, workflows, or developer platforms.
Responsibilities:
- Build enterprise-grade LLM applications and intelligent agent capabilities.
- Develop reusable prompts, context, retrieval, memory, evaluation, and agent components.
- Design and implement enterprise RAG architectures and retrieval pipelines.
- Optimize chunking, embeddings, indexing, ranking, reranking, and retrieval strategies.
- Evaluate, fine-tune, deploy, and optimize LLMs and SLMs.
- Apply LoRA, QLoRA, quantization, distillation, and model compression techniques.
- Build private, hybrid, cloud, and on-premise model hosting solutions.
- Develop scalable model-serving and inference architectures.
- Optimize latency, throughput, concurrency, GPU utilization, and cost.
- Implement model routing, load balancing, caching, and fallback strategies.
- Build LLMOps, ModelOps, and AgentOps capabilities.
- Develop observability for prompts, retrieval, responses, latency, cost, and failures.
- Build automated evaluation and regression-testing frameworks.
- Implement Responsible AI, security, governance, and data-protection controls.
Preferred Skillset:
- Experience with open-source models such as Llama, Mistral, Mixtral, Phi, Gemma, Qwen, DeepSeek, Granite, or Falcon.
- Experience with GPU infrastructure such as NVIDIA H100, H200, B200, B300, A100, L40S, GH200, or AMD MI300X.
- Experience with private AI, hybrid AI, or air-gapped AI environments.
- Experience with MCP, tool registries, agent runtimes, or enterprise integration patterns.
- Experience in healthcare, financial services, insurance, or other regulated industries.
- Experience with Responsible AI, model governance, model risk management, or AI compliance.