Tittle – SRE Engineer with ML Ops and Arize product
Location – Malvern, PA - 100% ONSITE FROM DAY 1
ROLE_DESCRIPTION -
AWS Cloud Platform Expertise - EC2, EKS, ECS, Lambda, CloudWatch, SNS, SQS, Event Bridge for highly available and scalable services.
Arize AI Observability & Monitoring - Model performance monitoring, drift detection, evaluation analytics, and AI/LLM observability.
Site Reliability Engineering (SRE) - SLI/SLO definition, Error Budgets, reliability improvement, service resilience, and uptime management.
Incident Management & RCA - Major Incident Management (MIM), outage tracking, root cause analysis, problem management, and MTTR reduction.
Failure Analysis & Risk Assessment - FMEA, risk quantification, reliability assessments, and proactive mitigation of platform failures.
Testing & Validation Engineering - Scenario testing, regression testing, impact analysis, release validation, and upstream change testing.
Monitoring, Alerting & Automation - CloudWatch, Grafana, Prometheus, PagerDuty, automated notifications, dashboards, and operational metrics.
DevOps & MLOps Practices - Kubernetes, Terraform, CI/CD pipelines, Python scripting, AI/ML platform operations, and LLM reliability optimization.
Preferred Technologies: AWS, Arize AI, Kubernetes (EKS), CloudFormation, Python, CloudWatch, Grafana, Prometheus, PagerDuty, GitHub Actions/Jenkins.
SYSMIND LLC is an Equal Employment Opportunity employer. All qualified applicants will receive consideration for employment without any discrimination. We promote and support a diverse workforce at all levels in the company. All job offers are contingent upon completion of a satisfactory background check and reference checks. Additionally passing the drug test may also be required. All contractors intending to work on SYSMIND's W2 are "at will" employees.