AI Performance Engineer

Hybrid in Phoenix, AZ, US • Posted 22 hours ago • Updated 22 hours ago
Contract W2
Contract Independent
12 Months
No Travel Required
Able to Sponsor
Hybrid
$55 - $60/hr
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • GPU Computing
  • Computer Engineering
  • Computer Hardware
  • CUDA
  • Caching
  • Computer Architecture
  • Artificial Intelligence
  • Benchmarking
  • Optimization
  • PyTorch
  • Python
  • Regression Testing
  • Linux
  • Machine Learning (ML)
  • Mechanics
  • Fusion
  • GPU
  • InfiniBand
  • Leadership
  • C
  • C++
  • Data Compression
  • Deep Learning
  • Research
  • Rust
  • SAN
  • Training

Summary

AI Performance Engineer

Client: Infosys/Microsoft

Pheonix, AZ// Seattle, WA

Rate: $55-60/hr C2C

 

AI Model tuning:
AI Performance Engineer — Model Optimization & Systems
About the role You will own the question: "How well do our AI models run on our hardware, and how do we make them run better?" You'll benchmark and profile training and inference workloads, identify compute, memory, and I/O bottlenecks, and recommend (and implement) optimizations — from model-level techniques like quantization and batching to system-level tuning of GPU utilization, memory bandwidth, and data pipelines.
Responsibilities
Benchmark AI models (LLMs, vision, multimodal) across hardware configurations; measure latency, throughput, utilization, memory behavior, and scaling efficiency
Profile workloads end-to-end using tools such as Nsight Systems/Compute, PyTorch Profiler, and system telemetry (nvidia-smi, DCGM) to isolate bottlenecks
Build roofline/performance models to quantify achieved vs. theoretical performance and prioritize the highest-impact optimizations
Apply and evaluate optimizations: quantization, pruning, distillation, operator/kernel fusion, graph compilation, KV-cache management, batching strategies, speculative decoding
Recommend hardware/system configurations (GPU selection, memory sizing, interconnect, storage/network I/O) for given model workloads
Establish performance baselines, SLAs, and regression testing so models stay fast as they evolve
Write clear analyses and recommendations for engineering and leadership audiences
Required qualifications
BS/MS in CS, Computer Engineering, EE, or equivalent practical experience
Strong Python; working proficiency in at least one systems language (C++/Rust/C)
Hands-on experience with a deep-learning framework (PyTorch preferred), including model execution, export, and profiling
Demonstrated experience delivering measurable performance improvements in DL training or inference
Solid grounding in computer architecture: memory hierarchy, bandwidth vs. compute limits, parallelism
Ability to reason quantitatively about latency, throughput, batching, memory footprint, and utilization under real workloads
Fluency with Linux and GPU computing environments
Preferred qualifications
GPU programming (CUDA, Triton, ROCm/HIP) and low-level libraries (cuBLAS, cuDNN, CUTLASS)
Inference runtimes/serving engines: TensorRT(-LLM), ONNX Runtime, vLLM, SGLang, Triton Inference Server
LLM inference mechanics: attention, KV caching, prefill vs. decode, continuous batching, speculative decoding
Distributed training/inference: data/tensor/pipeline parallelism, NCCL, InfiniBand/RoCE
Model compression research or MLPerf-style benchmarking experience
Edge/on-device deployment (Jetson, NPUs, Core ML) if your systems include e

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 90769335A
  • Position Id: 9076124
  • Posted 22 hours ago
Contact the job poster
HC

Hema Chandiran

Recruiter @ Info Way Solutions
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

Today

Full-time

Remote

Today

Easy Apply

Full-time, Part-time, Contract, Third Party

USD 45-45

No location provided

Today

Full-time

USD 142,800.00 - 274,800.00 per year

Remote

Yesterday

Easy Apply

Part-time, Third Party

45 - 55

Search all similar jobs