Devops Enigneer/Devops运维

Hybrid in Los Angeles, CA, US • Posted 12 hours ago • Updated 12 hours ago
Full Time
No Travel Required
Hybrid
$70,000 - $100,000/yr
Fitment

Dice Job Match Score™

👾 Reticulating splines...

Job Details

Skills

  • mandarin
  • aws
  • CI/CD
  • Kubernetes
  • GPU Infrastructure

Summary

Core Responsibilities

  • Design, build, and maintain AWS cloud infrastructure and private data center environments using Infrastructure as Code (IaC).

  • Manage large-scale Kubernetes (EKS) clusters, including cluster deployment, upgrades, scaling, networking (CNI), storage management, and operational maintenance.

  • Develop internal DevOps platforms, automation tools, and command-line utilities using Python or Go to improve engineering productivity and operational efficiency.

  • Build and maintain end-to-end monitoring and observability platforms based on Prometheus, Grafana, and ELK Stack to ensure system reliability and rapid troubleshooting.

  • Manage AI infrastructure, including GPU servers (NVIDIA A100, T4, etc.), CUDA environments, driver versions, and GPU resource allocation.

  • Deploy, maintain, and optimize containerized AI applications such as ComfyUI and other Generative AI services for high concurrency and production environments.

Requirements

Education

  • Bachelor''s degree or above in Computer Science, Information Technology, or related disciplines.

Experience

  • Minimum 3 years of experience in DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering.

Languages

  • Fluent in Mandarin (spoken and written).

Technical Skills

Cloud & Containers

  • Strong experience with AWS services, including EC2, EKS, S3, VPC, IAM, and related cloud infrastructure.

  • Deep understanding of Kubernetes architecture, scheduling, networking, storage, and container orchestration.

Programming

  • Strong programming skills in Python or Go.

  • Experience developing backend services, automation tools, or internal DevOps platforms.

System Administration

  • Strong knowledge of Linux operating systems.

  • Familiarity with TCP/IP, HTTP, DNS, Shell scripting, and system troubleshooting.

CI/CD

  • Hands-on experience with Jenkins, GitLab CI/CD, GitHub Actions, or similar continuous integration and deployment platforms.

Preferred Qualifications

GPU Infrastructure

  • Experience managing large-scale GPU clusters.

  • Knowledge of GPU monitoring, resource scheduling, memory optimization, and Spot Instance cost optimization.

Generative AI Infrastructure

  • Hands-on experience deploying and maintaining ComfyUI, Stable Diffusion WebUI, or similar AI inference platforms.

  • Experience with dependency management, multi-user concurrency optimization, and AI model loading acceleration.

MLOps

  • Familiarity with Kubeflow, MLflow, Triton Inference Server, or similar MLOps platforms.

High Performance Computing

  • Experience with RDMA networking, distributed computing, and large-scale parallel processing environments.

Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: PTP5FkzgL34bW7u
  • Position Id: 9057210
  • Posted 12 hours ago
Contact the job poster
JB

Julian Black

Recruiter @ luminarytech
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Remote

10d ago

Easy Apply

Full-time, Third Party

Depends on Experience

Wisconsin

Today

Full-time

Remote or Pennsylvania

Today

Full-time

USD 155,000.00 - 170,000.00 per year

Remote

Today

Full-time

USD 170,000.00 - 220,000.00 per year

Search all similar jobs