Principle Architect || Onsite in Phoenix, AZ || W2

Phoenix, AZ, US • Posted 11 days ago • Updated 11 days ago
Contract W2
6 Months
Travel Required
On-site
$65 - $70/yr
Company Branding Image
Fitment

Dice Job Match Score™

🛠️ Calibrating flux capacitors...

Job Details

Skills

  • Apache HTTP Server
  • Artificial Intelligence
  • Cloud Architecture
  • Database
  • Finance
  • High Availability
  • Kubernetes
  • LangChain
  • Large Language Models (LLMs)
  • MEAN Stack
  • Machine Learning (ML)
  • Orchestration
  • Performance Monitoring
  • Productivity
  • Roadmaps
  • Root Cause Analysis
  • Scalability
  • Software Performance Management
  • Vertex
  • Workflow

Summary

Position Summary

We are seeking an experienced Principal Architect to lead the design and implementation of an enterprise-grade Centralized Observability, AIOps, and Self-Healing Platform. The platform will unify metrics, logs, traces, events, and infrastructure telemetry across cloud and on-premises environments while leveraging AI/ML for anomaly detection, intelligent alerting, root cause analysis, and automated remediation.

The ideal candidate has deep expertise in Observability, Platform Engineering, Cloud Architecture, SRE, Kubernetes, Event-Driven Architecture, AI/ML integration, and Automation.


Key Responsibilities

Platform Architecture

  • Design enterprise-wide centralized observability architecture.
  • Define platform standards and reference architectures.
  • Build a multi-tenant observability platform supporting multiple customers and business units.
  • Design for high availability, scalability, and resilience.
  • Establish governance, onboarding standards, and platform lifecycle management.

Observability Architecture

Design and standardize:

  • Metrics collection
  • Distributed tracing
  • Centralized logging
  • Event correlation
  • Synthetic monitoring
  • Real User Monitoring (RUM)
  • Infrastructure monitoring
  • Application Performance Monitoring (APM)
  • Database observability
  • Network observability

Implement observability using tools such as:

  • Prometheus
  • Grafana
  • OpenTelemetry
  • Datadog
  • Splunk
  • CloudWatch
  • Loki
  • Tempo
  • Jaeger
  • Elasticsearch

AIOps

Design AI-driven capabilities including:

  • Intelligent alert correlation
  • Event deduplication
  • Dynamic thresholding
  • Anomaly detection
  • Predictive analytics
  • Capacity forecasting
  • Root cause analysis
  • Incident prioritization
  • Service dependency mapping
  • Change impact analysis

Self-Healing Platform

Design automated remediation workflows for:

  • Kubernetes pod failures
  • Container restarts
  • Node failures
  • Database connectivity issues
  • Memory leaks
  • Disk space issues
  • High CPU utilization
  • Service failures
  • Network issues
  • Certificate expiry
  • Auto-scaling
  • Rollback automation

Integrate with:

  • Rundeck
  • StackStorm
  • Ansible
  • Terraform
  • Kubernetes Operators
  • Argo Workflows
  • GitHub Actions
  • Jenkins

Cloud & Platform Engineering

Architect solutions across:

  • AWS
  • Azure
  • Google Cloud Platform
  • Kubernetes
  • OpenShift
  • Docker
  • Service Mesh (Istio, Linkerd)

Design:

  • Multi-cluster observability
  • Multi-region deployment
  • Hybrid cloud observability
  • Disaster recovery

AI Integration

Build AI capabilities using:

  • Large Language Models (LLMs)
  • Retrieval-Augmented Generation (RAG)
  • Vector databases
  • AI agents
  • Knowledge graphs
  • Model orchestration
  • AI-assisted runbooks
  • Automated incident summarization
  • Conversational operations assistants

Experience with:

  • OpenAI-compatible APIs
  • Amazon Bedrock
  • Azure OpenAI
  • Google Vertex AI
  • LangGraph, LangChain, or similar orchestration frameworks

Event-Driven Architecture

Design integrations using:

  • Apache Kafka
  • IBM MQ
  • RabbitMQ
  • Amazon EventBridge
  • Event-driven microservices

SRE Practices

Implement:

  • SLIs
  • SLOs
  • Error budgets
  • Incident management
  • Chaos engineering
  • Reliability engineering
  • Capacity planning
  • Production readiness reviews

Security

Implement:

  • RBAC
  • OAuth2 / OIDC
  • mTLS
  • Secrets management
  • Audit logging
  • Zero Trust principles
  • Compliance controls

Required Technical Skills

Observability

  • Prometheus
  • Grafana
  • OpenTelemetry
  • Datadog
  • Splunk
  • Loki
  • Tempo
  • Jaeger
  • Elasticsearch

Cloud

  • AWS
  • Azure
  • Google Cloud Platform

Kubernetes

  • Kubernetes
  • OpenShift
  • Helm
  • Argo CD

Automation

  • Terraform
  • Ansible
  • Python
  • Bash
  • PowerShell

Programming

  • Python
  • Go
  • Java
  • Node.js

AI/ML

  • LLM integration
  • RAG
  • Vector databases
  • ML-based anomaly detection
  • AI agents

Messaging

  • Kafka
  • IBM MQ
  • RabbitMQ

Databases

  • PostgreSQL
  • MySQL
  • MongoDB
  • Redis

Preferred Experience

  • Banking / Financial Services
  • Insurance
  • Healthcare
  • Enterprise SaaS
  • Multi-cloud environments
  • Large-scale production operations

Responsibilities

The successful candidate will:

  • Define the enterprise observability strategy.
  • Build a centralized telemetry platform.
  • Design AI-driven incident detection and correlation.
  • Implement automated self-healing workflows.
  • Standardize dashboards, alerts, and telemetry collection.
  • Establish SRE best practices.
  • Create reusable platform components.
  • Mentor engineering teams.
  • Lead architecture reviews.
  • Drive cloud modernization initiatives.

Nice-to-Have Skills

  • OpenTelemetry Collector customization
  • eBPF-based observability
  • FinOps
  • ServiceNow integration
  • PagerDuty integration
  • Opsgenie integration
  • Knowledge graph technologies
  • Digital twins for IT operations
  • MLOps experience

Leadership Skills

  • Enterprise architecture
  • Executive communication
  • Technical mentoring
  • Cross-functional leadership
  • Vendor evaluation
  • Strategic roadmap planning

Success Metrics

The architect will be expected to deliver:

  • A single-pane-of-glass observability platform.
  • Reduction in Mean Time to Detect (MTTD).
  • Reduction in Mean Time to Resolve (MTTR).
  • Intelligent alert noise reduction.
  • Automated remediation for common operational issues.
  • Enterprise-wide telemetry standards.
  • High platform availability and scalability.
  • Improved developer and operations productivity.
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 91165686
  • Position Id: 9021289
  • Posted 11 days ago

Company Info

About Value Spectrum Technologies LLC

Step into a future defined by empowerment at Value Spectrum Technologies. With leading-edge software solutions and strategic consulting, were dedicated to shaping and elevating your digital tomorrow. Experience the synergy of innovation and collaboration as we unlock unparalleled opportunities for growth in the dynamic landscape of technology. Welcome to empowerment.

Join us in navigating the ever-evolving digital landscape with confidence, as we work together to unlock unprecedented opportunities and build a tomorrow that is truly empowered by the limitless possibilities of technology. Your digital future starts here.

About_Company_OneAbout_Company_Two
Contact the job poster
LP

Laxminath Pabboju

Recruiter @ Value Spectrum Technologies LLC
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

Phoenix, Arizona

11d ago

Easy Apply

Contract

60 - 65

Phoenix, Arizona

7d ago

Easy Apply

Contract

60 - 65

Remote

7d ago

Easy Apply

Contract

Depends on Experience

Remote

11d ago

Easy Apply

Contract

65 - 70

Search all similar jobs