Apply Now

LLM Inference & GPU Systems Consultant

Charlotte, NC, US • Posted 2 hours ago • Updated 2 hours ago

Contract Corp To Corp

Contract Independent

Contract W2

6 Months

No Travel Required

Able to Sponsor

On-site

Depends on Experience

Fitment

Dice Job Match Score™

🔗 Matching skills to job...

Job Details

Skills

LLM Systems Engineer
NVIDIA H200
KV Cache
OpenShift AI
GPU
RunAI
vLLM
TensorRT-LLM

Summary

Role : LLM Inference & GPU Systems Consultant

Location : Charlotte , NC ( Locals only)

We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involves no model training infrastructure or fine-tuning pipelines.

Key Responsibilities
NVIDIA GPU Runtime Optimization: Drive extreme runtime efficiency and optimization for the token generation pipeline. Specifically manage prefill/decode optimization and KV cache management.
Inference Serving: Deploy and manage inference engines including vLLM and TensorRT-LLM.
Hardware Utilization: Optimize GPU throughput tuning, batching strategies, and latency optimization. Manage workload orchestration using RunAI and Kubernetes GPU orchestration.
Model Lifecycle Management: Oversee the complete Hugging Face model lifecycle, including model onboarding, deployment, and retirement.
Platform Operations: Operate and maintain the OpenShift AI ecosystem as the primary container platform for GenAI workloads.

Required Qualifications
8+ years experience working as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.
8+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques (KV Cache, prefill/decode).
Proficiency in OpenShift AI and GPU orchestration tools like RunAI.
Strong experience with modern inference frameworks, specifically vLLM and TensorRT-LLM.
Proven track record managing the Hugging Face deployment lifecycle.
Must be onsite at client in Charlotte, NC at least 3 days/week

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.

Dice Id: 90970970
Position Id: 8973788
Posted 2 hours ago

Create job alert

Never miss an opportunity! Create an alert based on the job you applied for.

Charlotte, North Carolina

•

Today

TECHNOGEN, Inc. is a Proven Leader in providing full IT Services, Software Development and Solutions for 15 years. TECHNOGEN is a Small & Woman Owned Minority Business with GSA Advantage Certification. We have offices in VA; MD & Offshore development centers in India. We have successfully executed 100+ projects for clients ranging from small business and non-profits to Fortune 50 companies and federal, state and local agencies. Description: Local candidates preferred. Role Overview: We are se

Easy Apply

Contract, Third Party

$0,00/-

AI Architect

Hybrid in Charlotte, North Carolina

•

10d ago

Locations : - Charlotte NC, Dallas, Iselin, NJ Hybrid - 2/3 days onsite 12 months contract with possible extension Company Overview: NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now. We are currently seeking a AI Architect to join our team in Charlotte, North Carolina (US-NC), United States (US). Job Description: Job Duties: Role Overview: We are se

Easy Apply

Third Party, Contract

$133

AI Architect

Hybrid in Charlotte, North Carolina

•

10d ago

Company Overview: Req ID: 371108 NTT DATA strives to hire exceptional, innovative and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now. We are currently seeking a AI Architect to join our team in Charlotte, North Carolina (US-NC), United States (US). Job Description: Job Duties: Role Overview: We are seeking a Principal GenAI Architect to serve as a hands-on practitioner and core technical visionary.

Easy Apply

Contract

$133

AI Program Manager

Charlotte, North Carolina

•

3d ago

Client: Photon Position: Program Manager/Delivery Lead Location: Newark, DE (Onsite) Duration: Full time Description for Internal Candidates Serve as the named Delivery Lead responsible for day-to-day delivery management, staffing alignment, sprint execution, and issue identification across the AI pod. Own pod-level impediment removal and stakeholder coordination onsite in Newark/Wilmington chasing access provisioning, unblocking SLM dependencies, and escalating to engagement leads when need

Easy Apply

Contract, Third Party

Depends on Experience

Search all similar jobs

LLM Inference & GPU Systems Consultant

Dice Job Match Score™

Job Details

Skills

Summary

Similar Jobs