Job Title : Technical Manager/SRE Senior Lead
Location: - Phoenix,AZ
Duration: 6 months
Role Description
"Lead and manage the SRE/Production Support team to ensure client expectations and service level objectives are consistently met.
(Core) Possess a strong understanding of ITSM and Site Reliability Engineering (SRE) processes to effectively manage production support operations.
(Core) Prepare and maintain operational reports and documentation, including bridge updates, RCA tracking, incident trends, service availability, platform health metrics, monthly operational deliverables, and trend analysis.
(Core) Facilitate technical discussions and bridge calls for critical and escalated production issues, ensuring effective communication and stakeholder alignment throughout the incident lifecycle.
Participate in Incident Management bridge calls, driving timely resolution by coordinating with engineering teams and escalating issues to the appropriate Subject Matter Experts (SMEs) as required.
Monitor platform health and identify opportunities to improve system reliability, availability, observability, and operational efficiency.
Drive continuous service improvement initiatives by analyzing recurring incidents, identifying root causes, and recommending preventive actions.
Plan, coordinate, and facilitate Disaster Recovery (DR) exercises, ensuring readiness, documentation, and post-exercise review of outcomes."
Required Skills:
"Lead and manage the SRE/Production Support team to ensure client expectations and service level objectives are consistently met.
(Core) Possess a strong understanding of ITSM and Site Reliability Engineering (SRE) processes to effectively manage production support operations.
(Core) Prepare and maintain operational reports and documentation, including bridge updates, RCA tracking, incident trends, service availability, platform health metrics, monthly operational deliverables, and trend analysis.
(Core) Facilitate technical discussions and bridge calls for critical and escalated production issues, ensuring effective communication and stakeholder alignment throughout the incident lifecycle.
Participate in Incident Management bridge calls, driving timely resolution by coordinating with engineering teams and escalating issues to the appropriate Subject Matter Experts (SMEs) as required.
Monitor platform health and identify opportunities to improve system reliability, availability, observability, and operational efficiency.
Drive continuous service improvement initiatives by analyzing recurring incidents, identifying root causes, and recommending preventive actions.
Plan, coordinate, and facilitate Disaster Recovery (DR) exercises, ensuring readiness, documentation, and post-exercise review of outcomes."
ESSENTIAL_SKILLS
"Lead and manage the SRE/Production Support team to ensure client expectations and service level objectives are consistently met.
(Core) Possess a strong understanding of ITSM and Site Reliability Engineering (SRE) processes to effectively manage production support operations.
(Core) Prepare and maintain operational reports and documentation, including bridge updates, RCA tracking, incident trends, service availability, platform health metrics, monthly operational deliverables, and trend analysis.
(Core) Facilitate technical discussions and bridge calls for critical and escalated production issues, ensuring effective communication and stakeholder alignment throughout the incident lifecycle.
Participate in Incident Management bridge calls, driving timely resolution by coordinating with engineering teams and escalating issues to the appropriate Subject Matter Experts (SMEs) as required.
Monitor platform health and identify opportunities to improve system reliability, availability, observability, and operational efficiency.
Drive continuous service improvement initiatives by analyzing recurring incidents, identifying root causes, and recommending preventive actions.
Plan, coordinate, and facilitate Disaster Recovery (DR) exercises, ensuring readiness, documentation, and post-exercise review of outcomes."
Technical Manager/SRE Senior Lead
EXPERIENCE_RANGE_IN_REQUIRED_SKILLS : 10+ Years
Role Descriptions: Lead and manage the SRE/Production Support team to ensure client expectations and service level objectives are consistently met.(Core) Possess a strong understanding of ITSM and Site Reliability Engineering (SRE) processes to effectively manage production support operations.(Core) Prepare and maintain operational reports and documentation| including bridge updates| RCA tracking| incident trends| service availability| platform health metrics| monthly operational deliverables| and trend analysis.(Core) Facilitate technical discussions and bridge calls for critical and escalated production issues| ensuring effective communication and stakeholder alignment throughout the incident lifecycle.Participate in Incident Management bridge calls| driving timely resolution by coordinating with engineering teams and escalating issues to the appropriate Subject Matter Experts (SMEs) as required.Monitor platform health and identify opportunities to improve system reliability| availability| observability| and operational efficiency.Drive continuous service improvement initiatives by analyzing recurring incidents| identifying root causes| and recommending preventive actions.Plan| coordinate| and facilitate Disaster Recovery (DR) exercises| ensuring readiness| documentation| and post-exercise review of outcomes.
Essential Skills:
Lead and manage the SRE/Production Support team to ensure client expectations and service level objectives are consistently met.(Core) Possess a strong understanding of ITSM and Site Reliability Engineering (SRE) processes to effectively manage production support operations.(Core) Prepare and maintain operational reports and documentation| including bridge updates| RCA tracking| incident trends| service availability| platform health metrics| monthly operational deliverables| and trend analysis.(Core) Facilitate technical discussions and bridge calls for critical and escalated production issues| ensuring effective communication and stakeholder alignment throughout the incident lifecycle.Participate in Incident Management bridge calls| driving timely resolution by coordinating with engineering teams and escalating issues to the appropriate Subject Matter Experts (SMEs) as required.Monitor platform health and identify opportunities to improve system reliability| availability| observability| and operational efficiency.Drive continuous service improvement initiatives by analyzing recurring incidents| identifying root causes| and recommending preventive actions.Plan| coordinate| and facilitate Disaster Recovery (DR) exercises| ensuring readiness| documentation| and post-exercise review of outcomes.
Keyword:
Skills: Digital : Site Reliability Engineering (SRE)
SYSMIND LLC is an Equal Employment Opportunity employer. All qualified applicants will receive consideration for employment without any discrimination. We promote and support a diverse workforce at all levels in the company. All job offers are contingent upon completion of a satisfactory background check and reference checks. Additionally passing the drug test may also be required. All contractors intending to work on SYSMIND's W2 are "at will" employees.