Site Reliability Engineer

Remote • Posted 12 hours ago • Updated 9 hours ago
Contract Independent
Contract W2
12 Months
Remote
Depends on Experience
Company Branding Image
Fitment

Dice Job Match Score™

✨ Finding the perfect fit...

Job Details

Skills

  • Git
  • Concourse
  • Linux
  • SUSE
  • Kafka
  • Zookeeper
  • OpenSearch

Summary

Hello,

Role: Site Reliability Engineer

Location: 100% Remote

Job Description:

Site Reliability Engineer OpenSearch to help ensure the highest levels of availability, performance, scalability, and Quality of Service (QoS) for mission-critical cloud services. This role will focus on the reliability, operations, automation, and continuous improvement of distributed search and analytics platforms built on OpenSearch, while working in a diverse, globally distributed team environment.

The ideal candidate brings deep experience in site reliability engineering, DevOps, cloud operations, automation, observability, and distributed systems, with proven hands-on expertise architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments.

General Responsibilities

  • Provision, build, deploy, monitor, operate, and support cloud services in a globally distributed team environment
  • Architect, build, deploy, and maintain high-performance OpenSearch clusters and platforms from the ground up
  • Administer and optimize OpenSearch environments for high availability, resiliency, scalability, security, and performance
  • Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation, replication, and storage utilization
  • Analyze and resolve operational issues, platform instability, and production incidents across infrastructure, platform, and application layers
  • Conduct incident response, root cause analysis, and post-incident remediation to drive continuous improvement
  • Maintain the integrity and security of servers, systems, and OpenSearch platform infrastructure
  • Support platform lifecycle activities including installation, configuration, upgrades, patching, hotfixes, backup, restore, and disaster recovery
  • Develop and maintain monitoring policies, alerting standards, operational runbooks, and support procedures
  • Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services
  • Ensure proper resource allocation and capacity planning across compute, memory, storage, and network resources
  • Partner with product development and engineering teams to design and enhance service reliability and operational readiness
  • Develop and implement testing strategies and document results for platform changes and operational improvements
  • Support log ingestion, index management, retention policies, lifecycle management, and search performance tuning
  • Work in a diverse environment and cross-train with other global team members
  • Participate in an on-call rotation and support weekend or after-hours operational needs as required

Requirements

  • Expert with Kubernetes, including troubleshooting, operations, management, and configuration of complex Kubernetes services.
  • Proven hands-on expertise designing, building, deploying, supporting, and maintaining OpenSearch clusters and platforms from scratch in production environments
  • Strong experience with OpenSearch administration, cluster architecture, performance tuning, scaling, upgrades, and troubleshooting
  • Experience with index design, shard and replica strategy, cluster sizing, node management, snapshot/restore, backup, and disaster recovery
  • Strong understanding of distributed systems, search platforms, indexing pipelines, query optimization, and high-availability architectures
  • Expertise with Git
  • Expertise with Concourse, including setup, management, and troubleshooting of new pipelines
  • Expertise with Linux, specifically SUSE and Ubuntu
  • Expertise with Kafka, Zookeeper, and Big Data technologies
  • Expert in development of automation for testing, deployment, scalability, and management of cloud services
  • Expertise with building, implementing, and/or supporting cloud monitoring tools
  • Expert knowledge of cloud computing, infrastructure operations, and databases
  • Expert understanding of web services, networking, virtualization, and internet protocols
  • Ability to multitask and handle various projects, deadlines, and changing priorities
  • Excellent communication and prioritization skills
  • Expertise with security fundamentals as they pertain to SaaS multi-tenant application systems
  • Strong interpersonal, presentation, and customer service skills

Desired Qualifications

  • Experience with AWS services including Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, and VPC
  • Experience deploying and operating OpenSearch in AWS-based environments
  • Experience with Cloud Foundry-based environments
  • Experience with Jenkins, Chef, and/or Terraform
  • Exposure to and understanding of troubleshooting IP networks and application stacks
  • Experience with observability tools such as Prometheus and Grafana
  • Experience with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls
  • Familiarity with capacity forecasting, performance benchmarking, and resilience testing for distributed search platforms

Education

  • BS/BA degree in Computer Science, Management Information Systems, or related IT discipline preferred
  • Allowable substitution: An additional four (4) years of experience may be substituted for a BS/BA degree
  • 8+ years of experience

Additional Requirements

  • Participation in an on-call rotation for handling P1 incidents is required
  • Flexible schedule which may include weekend or after-hours work
  • Ability to work effectively in a diverse, collaborative, and globally distributed team environment
Employers have access to artificial intelligence language tools (“AI”) that help generate and enhance job descriptions and AI may have been used to create this description. The position description has been reviewed for accuracy and Dice believes it to correctly reflect the job opportunity.
  • Dice Id: 10432128
  • Position Id: 9050473
  • Posted 12 hours ago

Company Info

About TechWish

TechWish is a global technology services and digital transformation partner, Headquartered in Virginia, Washington, DC, helping enterprises drive innovation, efficiency, and growth through modern technology solutions. With more than 20 years of enterprise IT experience, TechWish specializes in cloud transformation, application modernization, data engineering, and artificial intelligence solutions, including generative AI and AIOps.

The company helps organizations unlock insights from data, modernize legacy systems, and implement intelligent automation to improve business outcomes. With over 750+ professionals worldwide and offices across the United States and India, TechWish works with Fortune 500 organizations across industries to deliver scalable, future‑ready digital solutions.


TechWish partners with leading technology providers, including Microsoft, AWS, Salesforce, Snowflake, Databricks, and Google Cloud.

About_Company_OneAbout_Company_Two
Contact the job poster
MA

Manisha Arige

Recruiter @ TechWish
Create job alert
Set job alertNever miss an opportunity! Create an alert based on the job you applied for.

Similar Jobs

It looks like there aren't any Similar Jobs for this job yet.

Search all similar jobs