job summary:
Platform Engineering & Operations
Lead the administration, monitoring, and performance tuning of Oracle Enterprise Linux (OEL) environments in a large-scale enterprise ecosystem.
Oversee the design, build, and lifecycle management of Linux servers, including storage, virtualization, and associated infrastructure.
Manage high availability (HA) configurations, clustering, and load-balanced environments to ensure minimal downtime.
Drive capacity planning, performance optimization, and system scalability initiatives.
Reliability & Automation (SRE Practices)
Define and implement SRE principles, including SLIs, SLOs, and error budgets.
Lead initiatives for infrastructure automation (provisioning, configuration, patching) using tools such as Ansible.
Build and maintain self-healing systems, reducing manual intervention and improving system resilience.
Develop automation for system installation, configuration, and deployment pipelines.
System Administration & Infrastructure Management
Install, configure, and maintain Oracle Enterprise Linux (OEL) operating systems and related software stacks.
Manage Logical Volume Manager (LVM) configurations, including volume groups and filesystem expansion.
Administer distributed file systems, NFS servers/clients, and automount configurations.
Maintain network services such as DNS, NTP, LDAP/Kerberos, SMTP (sendmail/postfix), and OpenSSH.
Troubleshoot and support network protocols (TCP/IP, HTTP, HTTPS, RPC).
location: Chandler, Arizona
job type: Contract
salary: $51.66 - 61.66 per hour
work hours: 8am to 5pm
education: Bachelors
responsibilities:
Platform Engineering & Operations
- Lead the administration, monitoring, and performance tuning of Oracle Enterprise Linux (OEL) environments in a large-scale enterprise ecosystem.
- Oversee the design, build, and lifecycle management of Linux servers, including storage, virtualization, and associated infrastructure.
- Manage high availability (HA) configurations , clustering, and load-balanced environments to ensure minimal downtime.
- Drive capacity planning, performance optimization, and system scalability initiatives.
Reliability & Automation (SRE Practices)- Define and implement SRE principles , including SLIs, SLOs, and error budgets.
- Lead initiatives for infrastructure automation (provisioning, configuration, patching) using tools such as Ansible .
- Build and maintain self-healing systems , reducing manual intervention and improving system resilience.
- Develop automation for system installation, configuration, and deployment pipelines .
System Administration & Infrastructure Management- Install, configure, and maintain Oracle Enterprise Linux (OEL) operating systems and related software stacks.
- Manage Logical Volume Manager (LVM) configurations, including volume groups and filesystem expansion.
- Administer distributed file systems , NFS servers/clients, and automount configurations.
- Maintain network services such as DNS, NTP, LDAP/Kerberos, SMTP (sendmail/postfix), and OpenSSH.
- Troubleshoot and support network protocols (TCP/IP, HTTP, HTTPS, RPC).
qualifications:
Monitoring, Incident Management & Support
Implement and enhance monitoring, alerting, and observability frameworks for proactive issue detection.
Lead incident response, root cause analysis (RCA), and postmortem reviews.
Drive continuous improvement by identifying systemic issues and implementing preventive solutions.
Oversee break/fix operations, ensuring timely resolution and minimal business impact.
Security & Compliance
Ensure systems are secure, hardened, and compliant with enterprise security standards.
Manage patching, vulnerability remediation, and OS upgrades.
Partner with security teams to implement best practices for access control, auditing, and encryption.
Leadership & Collaboration
Provide technical leadership and mentorship to SRE and infrastructure teams.
Collaborate with application, DevOps, and platform teams to improve system reliability and deployment processes.
Define and enforce operational standards, runbooks, and best practices.
Drive cross-functional initiatives to enhance platform stability and efficiency.
Documentation & Governance
Maintain comprehensive documentation for architecture, processes, and operational procedures.
Ensure adherence to change management and incident governance frameworks.
Standardize operational workflows across environments.
5+ years of experience in Linux system administration in enterprise environments.
Strong expertise in Oracle Enterprise Linux (OEL) systems and FPP.
Proven experience in high availability systems, virtualization, and storage management.
Hands-on experience with automation and configuration management tools (Ansible preferred).
Proficiency in at least one scripting/programming language (Bash, Python preferred).
Strong experience with infrastructure troubleshooting, performance tuning, and incident management.
Solid understanding of enterprise infrastructure (compute, storage, network).
Excellent analytical, problem-solving, and organizational skills.
Strong communication and collaboration skills in a global team environment.
skills:
Ansible,Bash,storage management,deployment pipelines,DevOps,distributed file systems,DNS,infrastructure automation,configuration management tools,Kerberos,LDAP,Linux servers,Linux system administration,Logical Volume Manager,LVM,network protocols,network services,NTP,OpenSSH,Oracle Enterprise Linux,performance tuning,performance optimization,postfix,Python,system reliability,scripting,sendmail,SMTP,SRE Practices,vulnerability remediation,high availability,TCP/IP,analytical,communication,organizational skills,Leadership,Troubleshoot,troubleshooting,problem-solving,Reliability,proactive,business impact,Collaboration,access control,System Administration,administration,architecture,Install,auditing,Automation,budgets,continuous improvement,capacity planning,change management,enterprise security,encryption,Governance,governance frameworks,Incident Management,incident response,infrastructure,Infrastructure Management,enterprise infrastructure,comprehensive documentation,lifecycle management,mentorship,operating systems,issue detection,collaboration skills,root cause analysis,Security,system scalability,SRE principles,Standardize,system installation,deployment processes,technical leadership,operational workflows
Equal Opportunity Employer: Race, Color, Religion, Sex, Sexual Orientation, Gender Identity, National Origin, Age, Genetic Information, Disability, Protected Veteran Status, or any other legally protected group status.
At Randstad Digital, we welcome people of all abilities and want to ensure that our hiring and interview process meets the needs of all applicants. If you require a reasonable accommodation to make your application or interview experience a great one, please contact
Pay offered to a successful candidate will be based on several factors including the candidate's education, work experience, work location, specific job duties, certifications, etc. In addition, Randstad Digital offers a comprehensive benefits package, including: medical, prescription, dental, vision, AD&D, and life insurance offerings, short-term disability, and a 401K plan (all benefits are based on eligibility).
This posting is open for thirty (30) days.
![]()