Senior Lead AI Engineer (Foundation Model Hosting & LLM Inference)/Remote

Remote • Posted 2 hours ago • Updated 2 hours ago
Contract Corp To Corp
Contract Independent
Contract W2
1 Year
No Travel Required
Remote
Depends on Experience
Fitment

Dice Job Match Score™

🔗 Matching skills to job...

Job Details

Skills

  • Senior Lead AI Engineer (Foundation Model Hosting & LLM Inference)

Summary

Job Description: Senior Lead AI Engineer (Foundation Model Hosting & LLM Inference)

Location-Remote

Position Overview

We are seeking a Senior Lead AI Engineer to lead the design, deployment, and optimization of infrastructure for hosting foundation models and serving large language model (LLM) inference workloads. The ideal candidate has strong expertise in AI infrastructure, distributed systems, cloud platforms, and production-scale model serving.

Key Responsibilities

  • Lead the design and implementation of scalable infrastructure for hosting foundation models and LLM inference services.
  • Build and optimize high-performance inference pipelines for latency, throughput, reliability, and cost efficiency.
  • Deploy, monitor, and manage AI workloads across cloud and on-premises environments.
  • Collaborate with AI researchers and ML engineers to productionize machine learning models.
  • Design APIs and backend services for model serving and inference.
  • Optimize GPU utilization, model parallelism, batching, and resource scheduling.
  • Implement monitoring, logging, security, and observability for AI platforms.
  • Drive architecture decisions, engineering best practices, and technical standards.
  • Mentor engineers and lead technical initiatives across cross-functional teams.
  • Stay current with advancements in AI infrastructure, model serving, and inference optimization.

Qualifications

  • Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Software Engineering, or a related field.
  • Strong programming skills in Python and/or C++, with experience in backend systems.
  • Experience deploying and managing LLMs or other foundation models in production.
  • Hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.
  • Strong understanding of distributed systems, microservices, and container orchestration using Docker and Kubernetes.
  • Experience with GPU computing and model serving frameworks.
  • Familiarity with REST APIs, networking, and scalable backend architectures.

Preferred Skills

  • Experience with inference frameworks such as vLLM, TensorRT-LLM, Triton Inference Server, or similar technologies.
  • Knowledge of model optimization techniques, including quantization, batching, and caching.
  • Experience with distributed inference, autoscaling, and GPU scheduling.
  • Familiarity with MLOps tools, CI/CD pipelines, and Infrastructure as Code.
  • Experience with monitoring and observability tools for production AI systems.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10513292
  • Position Id: 73381-12895-1786047354
  • Posted 2 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Full-time

USD 200,000.00 - 270,000.00 per year

Remote

Today

Full-time

USD 114,600.00 - 234,600.00 per year

Remote

Today

Full-time

USD 114,600.00 - 234,600.00 per year

Remote

Today

Full-time

USD 185,000.00 - 217,400.00 per year

Search all similar jobs