SRE Platform Engineer AI Agents
Job Title: SRE Platform Engineer AI Agents
Experience: 8+ Years
Location: Atlanta, GA (Inperson Interview)
Job Summary
We are looking for an experienced SRE Platform Engineer with 8+ years of experience in Site Reliability Engineering, platform engineering, cloud infrastructure, Kubernetes, and automation. The ideal candidate will also have hands-on exposure to AI Agents / Agentic AI and be comfortable working on modern AI-driven platform solutions.
The candidate should have strong knowledge of SRE principles, SLA/SLO, Kubernetes, cloud platforms, automation, monitoring, and production support, along with solid programming and problem-solving skills.
Key Responsibilities
- Design, build, maintain, and improve highly reliable and scalable platform infrastructure.
- Implement Site Reliability Engineering (SRE) practices across production environments.
- Define and monitor SLAs, SLOs, SLIs, error budgets, and service reliability metrics.
- Develop and maintain Kubernetes-based applications and platform infrastructure.
- Troubleshoot production issues, perform root-cause analysis, and implement long-term corrective actions.
- Build automation for deployment, monitoring, infrastructure management, and operational processes.
- Work with AI Agents / Agentic AI solutions and integrate AI capabilities into platform and operational workflows.
- Support CI/CD pipelines and DevOps automation.
- Monitor application and infrastructure health using logging, metrics, and alerting tools.
- Collaborate with development, DevOps, cloud, and AI engineering teams.
- Participate in technical design discussions and contribute to platform architecture.
- Write clean, efficient code/scripts and demonstrate strong problem-solving abilities.
Required Skills
- 8+ years of experience in SRE, Platform Engineering, DevOps, or related infrastructure roles.
- Strong understanding of SRE concepts and practices.
- Strong knowledge of SLA, SLO, SLI, and Error Budgets.
- Hands-on experience with Kubernetes and containerized environments.
- Strong understanding of cloud infrastructure and production systems.
- Experience with CI/CD, automation, monitoring, logging, and incident management.
- Strong programming/scripting skills in languages such as Python, Java, Go, or similar.
- Experience with AI Agents / Agentic AI or AI-driven automation.
- Strong troubleshooting, debugging, and analytical skills.
- Excellent communication and collaboration skills.
Interview Focus Areas
Candidates should be prepared to discuss:
- Current/recent project architecture, responsibilities, and technical contributions.
- SRE concepts: SLA, SLO, SLI, error budgets, reliability, and incident management.
- Kubernetes fundamentals: pods, deployments, services, namespaces, scaling, configuration, and troubleshooting.
- Programming/coding problems, including exercises such as palindrome-number logic.
- Practical problem-solving and debugging scenarios.
- Experience designing or implementing AI Agents / Agentic AI solutions.