Job Title: Senior AI/ML Developer (LLMOps & Model Deployment)
Location: USA – Remote (Work in Pacific Time)
Duration:- 6-12 Months+
Job Overview
We are seeking a highly skilled AI/ML Engineer to lead the deployment, optimization, and scaling of open-weights foundation models within our Azure cloud ecosystem. In this role, you will bridge the gap between machine learning and core software engineering, turning raw models into highly available, low-latency APIs that power our cloud applications.
The ideal candidate has a deep understanding of LLMOps, hands-on experience optimizing model inference (including knowledge distillation), and a proven track record of architecting production-grade infrastructure on Azure.
Key Responsibilities
Model Deployment & API Development
⚬ Host & Scale Open-Weights Models: Deploy and maintain state-of-the-art open-weights models (e.g., Llama, Mistral, Phi) on Azure infrastructure.
⚬ API Engineering: Design, build, and secure high-performance, low-latency REST/gRPC APIs (using frameworks like FastAPI) to serve models to downstream cloud applications.
⚬ Inference Optimization: Implement advanced optimization techniques (e.g., vLLM, TensorRT-LLM, DeepSpeed) to maximize throughput and minimize time-to-first-token (TTFT).
LLMOps & Infrastructure
⚬ Pipeline Automation: Build and manage end-to-end LLMOps pipelines for continuous integration, deployment, and monitoring of models.
⚬ Azure Cloud Architecture: Leverage Azure AI Studio, Azure Machine Learning (Azure ML), Azure Kubernetes Service (AKS), and managed compute (GPUs like A100/H100) efficiently.
⚬ Monitoring & Observability: Implement comprehensive logging, tracing, and evaluations for LLM outputs (tracking drift, latency, costs, and hallucination rates).
Model Efficiency & Distillation
⚬ Knowledge Distillation: Train smaller, task-specific student models from larger, high-performing teacher models to reduce operational costs and latency without sacrificing accuracy.
⚬ Quantization & Fine-Tuning: Apply quantization techniques (AWQ, GPTQ, GGUF) and parameter-efficient fine-tuning (PEFT/LoRA) to adapt models to specific business domains.
Required Skills & Qualifications
Technical Requirements
⚬ Experience: 4+ years of professional experience as an ML Engineer, Data Scientist, or Backend Engineer with a heavy focus on AI deployment.
⚬ Cloud Proficiency: Strong hands-on experience with Microsoft Azure (Azure ML, AKS, Azure Container Apps, Key Vault).
⚬ AI/ML Frameworks: Deep proficiency with PyTorch, Hugging Face ecosystem (Transformers, Accelerate), and LangChain or LlamaIndex.
⚬ API & Backend: Strong Python programming skills and experience with containerization (Docker, Kubernetes) and API development (FastAPI).
⚬ Model Optimization: Demonstrated experience with model distillation, pruning, quantization, and utilizing inference engines like vLLM or TGI.
Soft Skills & Culture Fit
⚬ Problem Solver: Ability to triage infrastructure bottlenecks, memory constraints (OOM errors), and CUDA-related issues independently.
⚬ Collaborator: Comfortable working cross-functionally with backend engineers, product managers, and security teams.
⚬ Cost-Conscious Mindset: A sharp focus on balancing model accuracy with cloud spend and compute efficiency.
Preferred Qualifications
⚬ Certifications such as Azure AI Engineer Associate or Azure Solutions Architect.
⚬ Experience implementing secure guardrails (e.g., NeMo Guardrails, Azure AI Content Safety).
⚬ Contributions to open-source ML/LLMOps projects.