Location: Charlotte, NC
Salary: $53.00 USD Hourly - $57.00 USD Hourly
Description: Site Reliability Engineer (SRE) / Production EngineerWe are not accepting C2C or 1099 arrangements.Location: Charlotte, NC
Employment Type: Contract-to-Hire (Conversion Desired)
Sponsorship: Not Available
About the RoleJoin a team responsible for supporting and modernizing Small Business Lending and Deposit platforms that power critical digital banking experiences across the branch network. This team is focused on improving platform reliability, accelerating cloud adoption, enhancing observability, and driving operational excellence through automation and engineering best practices.
As a Site Reliability Engineer (SRE), you will play a key role in maintaining production stability while helping modernize applications and onboard workloads to OpenShift Container Platform (OCP). This position offers the opportunity to influence reliability strategies, implement automation solutions, and drive continuous improvement initiatives across a large enterprise environment.
What You'll Do Drive Reliability & Production Excellence
- Monitor and support mission-critical banking applications and services.
- Participate in daily production support activities, including reviewing offshore handoffs and operational health metrics.
- Investigate, troubleshoot, and resolve incidents, alerts, and service disruptions.
- Perform root cause analysis and implement preventative solutions.
- Support production deployments, infrastructure changes, and release activities.
Enhance Observability & Monitoring
- Design and maintain monitoring dashboards and alerting strategies.
- Improve platform observability across applications and infrastructure.
- Leverage tools such as Grafana, Splunk, Dynatrace, Prometheus, AppDynamics, and ThousandEyes to proactively identify performance issues.
- Establish metrics, SLIs, SLOs, and operational best practices.
Automate & Modernize Platforms
- Develop automation solutions that reduce manual operational effort.
- Build scripting solutions using Python, Shell, and Linux-based technologies.
- Implement Infrastructure as Code (IaC) using tools such as Terraform and Ansible.
- Support migration and onboarding of applications to OpenShift Container Platform (OCP) and cloud-native environments.
- Drive efficiency improvements across production support and platform engineering functions.
Collaborate Across Teams
- Partner with engineering, infrastructure, product, and offshore support teams.
- Share knowledge, contribute to operational processes, and promote reliability engineering best practices.
- Identify opportunities for innovation and help shape future-state platform capabilities.
Minimum Qualifications- Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent practical experience.
- 5+ years of experience supporting enterprise production environments.
- Experience with Site Reliability Engineering (SRE), Production Engineering, or DevOps practices.
- Strong expertise in observability, monitoring, alerting, and incident response.
- Experience with Linux/Unix administration and scripting using Python, Shell, or other automation languages.
- Experience supporting containerized environments and OpenShift or Kubernetes platforms.
- Experience using monitoring and performance management platforms such as Grafana, Splunk, Dynatrace, AppDynamics, or Prometheus.
- Experience conducting root cause analysis and implementing long-term reliability improvements.
Preferred Qualifications- Experience with cloud platforms and cloud-native architectures.
- Hands-on experience with Terraform, Ansible, or Infrastructure as Code methodologies.
- Knowledge of Salesforce and Salesforce platform support.
- Experience supporting nCino applications.
- Experience leveraging AI-powered tools such as Microsoft Copilot or AI-driven observability solutions.
- Experience with ServiceNow, incident management, and operational workflow automation.
- Financial services or banking industry experience.
- Experience working in large-scale enterprise environments.
Preferred Technical Skills- Site Reliability Engineering (SRE)
- Production Engineering
- Observability & Monitoring
- Incident Management & Incident Response
- Dynatrace
- Grafana
- Splunk
- Prometheus
- Kubernetes
- OpenShift Container Platform (OCP)
- Terraform
- Ansible
- Python
- Linux / Unix
- Shell Scripting
- Salesforce
- ServiceNow
- AI & Automation Technologies
- Cloud Technologies
What Makes You SuccessfulWe're looking for someone who goes beyond technical execution and demonstrates:
- Strong ownership and accountability.
- Curiosity and a continuous learning mindset.
- Ability to influence and improve operational processes.
- Innovative thinking and passion for modernization.
- Excellent collaboration and communication skills.
- Ability to work independently and execute with minimal direction.
- Positive, team-first attitude with a focus on knowledge sharing.
Why Join Us?You'll have the opportunity to support critical banking platforms while helping shape the future of reliability engineering, cloud adoption, automation, and AI-driven operations within a highly collaborative team environment. This role offers strong conversion potential for candidates seeking long-term career growth.
By providing your phone number, you consent to: (1) receive automated text messages and calls from the Judge Group, Inc. and its affiliates (collectively "Judge") to such phone number regarding job opportunities, your job application, and for other related purposes. Message & data rates apply and message frequency may vary. Consistent with Judge's Privacy Policy, information obtained from your consent will not be shared with third parties for marketing/promotional purposes. Reply STOP to opt out of receiving telephone calls and text messages from Judge and HELP for help.
Contact: This job and many more are available through The Judge Group. Please apply with us today!