AI-Enabled Platform/SRE Engineer

Phoenix, AZ, US • Posted 8 hours ago • Updated 8 hours ago
Contract Corp To Corp
Contract W2
Contract Independent
12 Months
No Travel Required
On-site
Depends on Experience
Fitment

Dice Job Match Score™

🔗 Matching skills to job...

Job Details

Skills

  • Artificial Intelligence
  • SRE Engineer
  • Site Reliability Engineering
  • GKE
  • GCP
  • Terraform
  • Helm
  • GitHub
  • CI/CD
  • Python
  • Splunk
  • Grafana
  • Datadog
  • AppDynamics
  • alerting
  • Apigee
  • Apigee X
  • GraphQL
  • AIOps
  • LLMs
  • Gemini
  • Llama
  • Mistral
  • Rancher RKE2
  • Generative AI
  • Leverage AI
  • Google Cloud Platform

Summary

Role: AI-Enabled Platform/SRE Engineer

Location:  Dallas, TX / Scottsdale, AZ (Hybrid)

Duration: Long Term Contract

Pay Rate: $55/hr on C2C

Note: 2nd round inperson 

Key Responsibilities

  • Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
  • Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
  • Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
  • Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
  • Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacentre Kubernetes environments.
  • Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.
  • Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.

 

Core Technical Skills

  • Site Reliability Engineering (SRE) – Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
  • Kubernetes Platform Engineering – 5+ years of Strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.
  • Cloud & Infrastructure Automation – Strong experience in Google Cloud Platform, Terraform, Helm, GitHub, CI/CD, and production-grade automation.
  • Software Development – 5+ years of Advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
  • Observability & Monitoring – Splunk, Grafana, Datadog, AppDynamics, alerting, and platform health monitoring.
  • API & Microservices Engineering – Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
  • AI-Driven Operations (AIOps) – Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91163673
  • Position Id: 9039721
  • Posted 8 hours ago
Contact the job poster
PK

Prudhvi Krishna

Team Lead @ Trinite Consulting Group LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Phoenix, Arizona

Today

Full-time

USD 144,250.00 - 256,250.00 per year

Phoenix, Arizona

6d ago

Easy Apply

Contract

Depends on Experience

Remote

6d ago

Easy Apply

Contract

Depends on Experience

Remote

29d ago

Easy Apply

Contract

Depends on Experience

Search all similar jobs