Senior Manager, Failure Analysis Engineering
Role Overview
The Senior Manager, Failure Analysis Engineering serves as the technical authority for system-level failure analysis and product reliability across the full manufacturing lifecycle, from New Product Introduction (NPI) through High Volume Manufacturing (HVM).
This role leads the identification of complex failure mechanisms, defines structured root cause methodologies, and drives cross-functional resolution to improve product quality, manufacturing yield, and long-term reliability.
The position requires deep expertise in server hardware architectures, failure physics, and data-driven analysis, combined with the ability to influence engineering, quality, and manufacturing organizations.
Scope
End-to-End Failure Analysis Ownership
NPI, including DVT / PVT readiness
Production, including L6, L10, and system-level testing
Field and customer returns, including RMA and DOA
System-Level Technical Scope
CPU
Memory
Storage
Power
Networking
Thermal systems
Cross-Functional Engagement
Additional Scope
Data-driven reliability and failure trend analysis
Influence on product design
Test coverage improvements
Manufacturing process improvements
Key Responsibilities
Lead complex failure analysis (FA) and root cause analysis (RCA) to identify system-level failure mechanisms across server platforms.
Define and standardize failure analysis methodologies, tools, processes, and best practices.
Drive reliability strategy and influence NPI readiness, including DFR, DFT, and test coverage.
Establish failure trend analysis across yield, escapes, and field returns to enable data-driven decision-making.
Serve as an escalation point for critical quality issues and lead cross-functional technical problem solving.
Drive corrective and preventive actions across design, test, and manufacturing to eliminate repeat failures.
Improve test effectiveness, reduce NTF (No Trouble Found) loops, and strengthen feedback loops into engineering.
Mentor engineers and elevate failure analysis capabilities across the organization.
Identify systemic failure drivers and develop technical strategies to improve product reliability and manufacturing performance.
Required Skills
Deep knowledge of server hardware architectures, including:
CPU
Memory
Storage
Power
Networking
Strong expertise in failure analysis methodologies and root cause analysis techniques.
Experience with system-level debugging, including:
Strong statistical analysis and data interpretation skills, including:
Manufacturing yield
Reliability
Failure trends
Understanding of manufacturing test flows, including:
L6
L10
System-level testing
Experience with technical and data analysis tools, including:
Preferred Skills
Experience with GPU systems and liquid cooling.
Experience in hyperscale manufacturing environments.
Automation experience.
Six Sigma certification.
Experience supporting hyperscale or data center server environments.
Knowledge of reliability modeling, including:
Weibull analysis
MTBF
HALT
HASS
Exposure to DFX methodologies, including:
Experience automating failure analysis workflows and data pipelines.
Experience working with suppliers and supporting component-level failure analysis.
Experience & Education
Bachelor's or Master's degree in:
10+ years of experience in:
Proven experience supporting NPI through HVM transitions in complex hardware systems.
Demonstrated track record of solving complex, cross-domain technical problems.
Strong ability to influence engineering, quality, manufacturing, and supplier organizations.
Success Criteria β First 6 Months
Establish a structured failure analysis framework and RCA methodology across programs.
Identify the top systemic failure drivers and implement corrective actions.
Improve failure containment and reduce repeat issues and NTF rates.
Build strong cross-functional alignment across Test Engineering, Quality, Product Engineering, and Manufacturing.
Enable data-driven visibility into failure trends and reliability risks.
Strengthen feedback loops between failure analysis findings and product design, testing, and manufacturing processes.
Organizational Fit
Reporting Structure
Key Interfaces
Test Engineering
Product Engineering
Manufacturing
Infrastructure
Works Closely With
Leadership Expectations
Acts as a technical leader and escalation point for complex failure analysis and reliability issues.
Provides technical direction and mentorship.
Direct people management is not required.
Position Summary
Owns the system-level failure analysis strategy, driving root cause identification, reliability improvements, and cross-functional resolution across NPI and manufacturing.
About INSPYR SolutionsTechnology is our focus and quality is our commitment. As a national expert in delivering flexible technology and talent solutions, we strategically align industry and technical expertise with our clients' business objectives and cultural needs. Our solutions are tailored to each client and include a wide variety of professional services, project, and talent solutions. By always striving for excellence and focusing on the human aspect of our business, we work seamlessly with our talent and clients to match the right solutions to the right opportunities. Learn more about us at inspyrsolutions.com.