![]()
Job Title: Observability Operations Engineer
Location: Phoenix, AZ (Hybrid)
Duration: Temp - 12 months
Pay Range: $50/hr to $60/hr (W2)
Job ID: 407523
About BCforward
BCforward is a leading global IT consulting and workforce solutions firm providing services and support to Fortune 500 and government clients. Founded in 1998, BCforward has grown with our customers needs into a full-service business solutions provider. With delivery centers and offices across North America and India, we take pride in building long-term relationships and delivering excellence through innovation, collaboration, and integrity.
Job Description
We are seeking a Senior Observability Operations Engineer to manage and enhance our enterprise observability platform. The ideal candidate will have deep expertise in Dynatrace, Splunk, OpenSearch/Elasticsearch, Kubernetes, Linux, and cloud-native observability solutions. Experience leveraging AI/ML and Generative AI to improve observability, automate operations, and accelerate incident resolution is desirable. The role will operate at approximately 20% automation and 80% operations and will ensure availability, scalability, operational excellence, and continuous improvement of monitoring and logging platforms that support mission-critical applications.
Work Schedule & Coverage:
- Onsite presence 3 days per week.
- Standard shifts: 9:00 AM-6:00 PM or 10:00 AM-6:30 PM.
- Work 1 weekend day every 2-3 weeks to provide coverage.
Responsibilities:
- Administer and optimize Dynatrace, Splunk, and OpenSearch/Elasticsearch platforms.
- Design, deploy, configure, and maintain monitoring, logging, tracing, and alerting solutions.
- Manage large-scale OpenSearch/Elasticsearch clusters, including indexing strategies, performance tuning, shard optimization, backups, and capacity planning.
- Configure Dynatrace OneAgent, ActiveGate, Synthetic Monitoring, RUM, DEM, Davis AI, and APM.
- Administer Splunk Enterprise, Universal Forwarders, Indexers, Search Heads, Cluster Manager, Deployment Server, and Splunk ITSI.
- Develop dashboards, alerts, reports, and executive operational metrics.
- Support Linux infrastructure and Kubernetes environments, including Docker, OpenShift, or Rancher.
- Implement observability best practices using OpenTelemetry for distributed tracing, metrics, logs, and events.
- Perform root cause analysis for production incidents using observability platforms.
- Collaborate with Platform Engineering, SRE, DevOps, Infrastructure, and Application teams.
- Automate operational tasks using Python, Shell scripting, REST APIs, Terraform, or Ansible, including AI-assisted automation.
- Participate in incident, problem, change, and release management processes.
- Drive platform upgrades, patching, security compliance, and operational governance.
- Improve platform reliability through automation, self-healing, and AI-assisted operations.
Required Skills & Qualifications:
- Dynatrace, Splunk Enterprise, OpenSearch, and Elasticsearch administration.
- Grafana, Prometheus, Kibana, Jaeger, and OpenTelemetry.
- Kubernetes and Linux administration with Docker, OpenShift, or Rancher.
- Networking fundamentals including TCP/IP, DNS, load balancers, and firewalls.
- AWS, Azure, or Google Cloud Platform with CI/CD pipelines, Git, Terraform, Ansible, and REST APIs.
- Scripting with Python and Bash/Shell. PowerShell preferred.
- 6-10+ years in IT infrastructure or observability operations with 4+ years administering Dynatrace, Splunk, OpenSearch, or Elasticsearch.
- Strong Linux system administration and production support experience.
- Excellent troubleshooting, analytical, communication, and stakeholder management skills.
Preferred Skills:
- Kafka exposure. Experience with AI/ML for observability and AIOps.
- Experience with Dynatrace Davis AI, predictive monitoring, and intelligent alerting.
- Generative AI tools such as ChatGPT, GitHub Copilot, Amazon Q, or Microsoft Copilot to improve operational efficiency.
- AI-assisted runbooks, incident summarization, log analysis, and automated ticket enrichment.
- Knowledge of RAG, vector databases, embeddings, and AI-powered knowledge search.
- Python with AI frameworks such as LangChain, LangGraph, or OpenAI APIs.
Education & Certifications:
- Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
- Preferred: Dynatrace Associate or Professional, Splunk Enterprise Certified Administrator, Elastic Certified Engineer, Kubernetes CKA/CKAD, AWS/Azure/Google Cloud Platform, ITIL Foundation, and AI/ML or Generative AI certification.
Must Haves:
- Dynatrace, Splunk, OpenSearch/Elasticsearch, and OpenTelemetry.
- Kubernetes, AI/ML, Grafana, and Sahara.
- Automation using AI with a focus on operations.
Why BCforward?
At BCforward, we believe in advancing lives and careers. When you join our team, you gain access to:
- Competitive compensation and benefits.
- Opportunities for growth with global clients.
- A supportive, inclusive culture that values innovation and people.
- Exposure to cutting-edge technologies and projects.
About Our Commitment
BCforward is an equal opportunity employer. We value diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, or veteran status.
Interested? Apply Now!
If this sounds like the right opportunity for you, please apply with your most recent resume.