Skip to content
All jobs

This job comes from OpenAI careers page, not from an Interstack member. You apply on their site, so Interstack can't track your application or tell you when it has been seen.

O

Data Center Hardware Quality & Reliability Engineer

OpenAI · San Francisco

Full time $226K – $285K • Offers Equity Scaling Posted 1 week ago

Skills this job asks for

Technical leadership Procurement Physics Statistics SQL Python

About the role

About The Role Own the end-to-end data-center hardware quality and reliability loop for OpenAI’s 3P infrastructure and 1P current and next-gen platforms. Turn field failures into quantified risk, fast containment, verified root cause, improved MQE/NPI and manufacturing-test coverage, accurate spares forecasts, and upstream changes that prevent recurrence. The first hire must combine practical hardware/system understanding, reliability engineering, data fluency, and cross-functional technical leadership at data-center scale. Key Responsibilities • Build and govern the field-quality data model across telemetry, tickets, RMA/repair, FA, firmware, configuration, supplier, and manufacturing genealogy. • Define AFR, ASR, DPPM, MTBF/MTTR, repeat-repair, NTF, repair-cycle-time, and forecast-versus-actual metrics with explicit denominators and uncertainty. • Provide fleet-level macro views and unit/FRU/cohort-level micro views; detect shifts and bound affected populations. • Lead systemic field-failure triage, containment, failure analysis, 8D/CAPA, risk assessment, corrective-action verification, and recurrence monitoring. • Develop cohort, life-data, reliability-growth, and spare-demand projections by product, FRU, supplier, configuration, geography, and age. • Partner with MQE and NPI to convert field mechanisms into manufacturing-test coverage, screening/stress profiles, diagnostics, control plans, DFR/DFS requirements, FMEA/FTA, mission profiles, FRU strategy, and qualification gates. • Close the loop by verifying whether upstream changes reduce field recurrence. • Define supplier/CM FA standards, field-data contracts, scorecards, escalation paths, and closure evidence. • Provide serviceability, TCO, and spares inputs without owning inventory execution or procurement. • Create concise executive decision packages: population at risk, exposure, confidence, options, cost/risk, and recommendation. • Run the cross-functional reliability council and, as the team grows, mentor the 1P and 3P Field Quality Engineers. Qualifications • BS in electrical, mechanical, computer, materials, reliability engineering, physics, or equivalent experience; MS preferred. • 8+ years in hardware quality/reliability, server/rack systems, or mission-critical infrastructure; 3+ years owning field-failure, RMA, or CAPA outcomes. • Solid working understanding of hardware and system architecture across board, tray, rack, firmware, telemetry, manufacturing test, and fleet behavior; deep expertise in every subsystem is not required. • Reliability statistics: censored life data, Weibull/Poisson/binomial methods, confidence bounds, MTBF/MTTR, and reliability growth. • Hands-on FMEA/FTA, accelerated or reliability-demonstration testing, 8D/CAPA, FA, and corrective-action verification. • Working proficiency with SQL and Python/R or equivalent analytics tools. • Ability to influence design, validation, operations, suppliers/CMs, and senior leaders without direct authority. Preferred Skill...

A summary from the original listing. Read the full details on their site.

Apply on OpenAI's site

Opens in a new tab.

Similar jobs