LLM Inference & GPU Systems Consultant || Charlotte, NC (Hyrbid-3days onsite in a week)

Charlotte, NC, US • Posted 3 hours ago • Updated 42 minutes ago
Contract W2
Contract Corp To Corp
Contract Independent
75% Travel Required
On-site
Fitment

Dice Job Match Score™

🫥 Flibbertigibetting...

Job Details

Skills

  • Source Code Management (SCM) and DevOps-Containerization-Openshift

Summary

TECHNOGEN, Inc. is a Proven Leader in providing full IT Services, Software Development and Solutions for 15 years.

TECHNOGEN is a Small & Woman Owned Minority Business with GSA Advantage Certification. We have offices in VA; MD & Offshore development centers in India. We have successfully executed 100+ projects for clients ranging from small business and non-profits to Fortune 50 companies and federal, state and local agencies.


Hi,

Greetings of the day!

We are looking to Hire a Talented Professional for the below Job opportunity with one of our clients,
If you're interested, please share your updated resume at your earliest convenience, and I'll be happy to provide more details about the role.

Position: LLM Inference & GPU Systems Consultant

Location: Charlotte, NC (Hyrbid-3days onsite in a week)

Duration: Long Term

Local candidates preferred. This is 3 days onsite every week.

Job Description:

We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involves no model training infrastructure or fine-tuning pipelines.

Key Responsibilities

NVIDIA GPU Runtime Optimization: Drive extreme runtime efficiency and optimization for the token generation pipeline. Specifically manage prefill/decode optimization and KV cache management.

Inference Serving: Deploy and manage inference engines including vLLM and TensorRT-LLM.

Hardware Utilization: Optimize GPU throughput tuning, batching strategies, and latency optimization. Manage workload orchestration using RunAI and Kubernetes GPU orchestration.

Model Lifecycle Management: Oversee the complete Hugging Face model lifecycle, including model onboarding, deployment, and retirement.

Platform Operations: Operate and maintain the OpenShift AI ecosystem as the primary container platform for GenAI workloads.

Required Qualifications

8+ years experience working as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.

8+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques (KV Cache, prefill/decode).

Proficiency in OpenShift AI and GPU orchestration tools like RunAI.

Strong experience with modern inference frameworks, specifically vLLM and TensorRT-LLM.

Proven track record managing the Hugging Face deployment lifecycle.

Must be onsite at client in Charlotte, NC at least 3 days/week

Ranjitha P | Sr. IT Recruiter

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10217412
  • Position Id: 2026-42837
  • Posted 3 hours ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Hybrid in Charlotte, North Carolina

Today

Easy Apply

Contract, Third Party

Depends on Experience

Charlotte, North Carolina

22d ago

Easy Apply

Full-time, Third Party

130000 - 140000

Charlotte, North Carolina

Today

Full-time

Charlotte, North Carolina

Today

Full-time

Search all similar jobs