Software Development Engineer - GPU Fleet Management & AI Infrastructure

San Jose, CA, US • Posted 1 day ago • Updated 3 hours ago
Full Time
On-site
USD 150,500.00 per year
Fitment

Dice Job Match Score™

🔢 Crunching numbers...

Job Details

Skills

  • Fleet Management
  • Health Care
  • Innovation
  • Health Care Administration
  • Distributed Computing
  • Lifecycle Management
  • Data Storage
  • Orchestration
  • Strong Authentication
  • Auditing
  • Computer Hardware
  • Machine Learning (ML)
  • IT Management
  • Design Review
  • Mentorship
  • Conflict Resolution
  • Problem Solving
  • Software Development
  • C++
  • Rust
  • IaaS
  • Recovery
  • WebSocket
  • Resource Management
  • Scheduling
  • PostgreSQL
  • Concurrent Computing
  • Command-line Interface
  • Authentication
  • Authorization
  • Testing
  • Version Control
  • Continuous Integration and Development
  • Continuous Integration
  • Automated Testing
  • Debugging
  • Collaboration
  • Training
  • Communication
  • PyTorch
  • Capacity Management
  • Performance Analysis
  • Hardware Troubleshooting
  • Kubernetes
  • Computer Networking
  • Routing
  • Streaming
  • Storage
  • CheckPoint
  • Management
  • Cloud Computing
  • GPU
  • Threat Modeling
  • Computer Science
  • Computer Engineering
  • Electrical Engineering
  • Military
  • Law
  • Recruiting
  • Artificial Intelligence

Summary

ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology has the power to solve the world's most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future.

Whether you're designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger - technology that moves the world forward. Join us and, together, we'll advance your career.

THE ROLE:

AMD is looking for an experienced software engineer to help build Fleet Manager, a secure control plane for operating large-scale AMD GPU infrastructure.

Fleet Manager provides scheduling, workload orchestration, hardware health management, interactive development environments, and model-serving capabilities across GPU clusters. You will design and build production systems spanning distributed control planes, Kubernetes, GPU scheduling, inference infrastructure, and developer-facing APIs and tools.

You will join a team working at the intersection of systems software, cloud infrastructure, and accelerated computing. Your work will directly influence how engineers and customers train, serve, debug, and operate workloads on current and future AMD GPU platforms.

THE PERSON:

The ideal candidate is a hands-on systems engineer who enjoys solving complex infrastructure problems and turning them into reliable, easy-to-use products.

You have strong technical judgment, can reason about distributed-system failure modes, and are comfortable working across service, cluster, and hardware boundaries. You communicate clearly, collaborate effectively across organizations, and can lead substantial projects from architecture through production deployment.

You care deeply about correctness, security, operability, and the experience of both end users and platform operators.

KEY RESPONSIBILITIES:
  • Design and develop Fleet Manager's distributed control-plane services, APIs, schedulers, inference gateway, and command-line tools.
  • Build reliable orchestration for GPU training, inference, custom jobs, and interactive development workloads.
  • Develop scalable scheduling and admission-control capabilities, including priority, fairness, topology-aware placement, quotas, backfilling, and multi-node workload coordination.
  • Implement durable reconciliation, lifecycle management, retries, idempotency, and recovery across PostgreSQL and external execution systems.
  • Integrate Fleet Manager with Kubernetes and technologies such as Kueue, JobSet, container runtimes, storage systems, and observability platforms.
  • Help evolve Fleet Manager into a portable orchestration layer capable of supporting Kubernetes, Slurm, Spur, and future execution environments.
  • Develop GPU health, diagnostics, quarantine, and controlled-remediation capabilities using ROCm and AMD hardware telemetry.
  • Build secure multi-tenant infrastructure with strong authentication, authorization, workload isolation, auditing, rate limiting, and least-privilege defaults.
  • Improve the reliability and performance of AI inference services, including routing, streaming, load shedding, health detection, and usage metering.
  • Define and maintain stable APIs, data models, compatibility contracts, and operational procedures.
  • Diagnose complex failures across distributed services, Kubernetes, networking, storage, GPU runtimes, drivers, and hardware.
  • Develop automated unit, integration, failure-injection, and production-readiness tests.
  • Work with AMD architecture, driver, platform, security, and machine-learning software teams to enable current and future GPU products.
  • Participate in new GPU, system, cluster, and software-stack bring-up.
  • Provide technical leadership through design reviews, code reviews, mentoring, and cross-functional problem solving.

PREFERRED EXPERIENCE:
  • Strong systems-software development experience in Rust, C++, Go, or a comparable language. Production Rust experience is highly desirable.
  • Experience designing and operating distributed systems, control planes, schedulers, or cloud infrastructure.
  • Strong understanding of concurrency, asynchronous programming, state machines, and failure recovery.
  • Experience building reliable services using REST, streaming, WebSocket, or gRPC APIs.
  • Experience with Kubernetes internals, controllers, operators, scheduling, resource management, or custom resources.
  • Familiarity with workload scheduling technologies such as Kueue, JobSet, Slurm, or other batch and cluster schedulers.
  • Experience with PostgreSQL-backed services, schema evolution, transactions, leader election, and optimistic concurrency.
  • Experience developing command-line tools and stable, user-focused APIs.
  • Understanding of container security, multi-tenant isolation, authentication, authorization, and secrets management.
  • Experience with production observability, including metrics, structured logging, tracing, alerting, and incident diagnosis.
  • Ability to write high-quality, maintainable code with careful attention to correctness, testing, and operational behavior.
  • Experience with source control, continuous integration, automated testing, profiling, and debugging tools.
  • Demonstrated ability to lead technically challenging projects and collaborate across organizational boundaries.
  • Effective written and verbal communication skills.

Experience in one or more of the following areas would be beneficial but is not required:
  • AMD GPU architecture, ROCm, HIP, amd-smi, RCCL, or GPU device plugins
  • Distributed AI training and multi-node collective communication
  • Large-model inference using platforms such as vLLM, SGLang, PyTorch, or similar runtimes
  • GPU topology, capacity management, performance analysis, and hardware diagnostics
  • Kubernetes networking, storage, admission control, and workload isolation
  • OpenAI-compatible inference APIs, request routing, streaming, and rate limiting
  • High-performance shared storage and large model or checkpoint management
  • Bare-metal, virtualized, and cloud GPU infrastructure
  • Production security and threat modeling for multi-tenant compute platforms

PREFERRED ACADEMIC CREDENTIALS :
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.

This role is not eligible for visa sponsorship.

#LI-G11

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants' needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's "Responsible AI Policy" is available here.

This posting is for an existing vacancy.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10127278
  • Position Id: 721a15ca8b87743a0043d9e2a0716c3
  • Posted 1 day ago
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

San Jose, California

Today

Full-time

USD 201,000.00 - 402,000.00 per year

Palo Alto, California

Today

Full-time

USD 135,000.00 - 175,000.00 per year

Remote or Sunnyvale, California

Today

Full-time

San Jose, California

Today

Full-time

USD 180,000.00 - 215,000.00 per year

Search all similar jobs