Skip to content
All jobs

This job comes from OpenAI careers page, not from an Interstack member. You apply on their site, so Interstack can't track your application or tell you when it has been seen.

O

Rack Power Engineer

OpenAI · San Francisco

Full time $287K – $485K • Offers Equity Scaling Posted 2 weeks ago

Skills this job asks for

Performance

About the role

About the Team OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads. Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models. By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform. About the Role We’re looking for a Rack Power Engineer with deep expertise in high-power conversion and distribution to design, qualify, and support power systems for AI supercomputers. You will own rack power solutions—including power shelves, AC/DC rectifiers, power supply units (PSUs), power management controllers (PMCs), and high-current distribution—from requirements and supplier development through deployment. You will also monitor fleet rack power health, lead debugging and root-cause investigations, and drive improvements into hardware, firmware, and qualification coverage. This role is based in San Francisco, CA. We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role, you will: - Own rack power architecture and requirements for high-power AI supercomputing systems, including power budgets, AC input interfaces, DC distribution, redundancy, efficiency, serviceability, and integration with data center infrastructure. - Drive the design and supplier development of power shelves, rectifiers, PSUs, PMCs, busbars, connectors, and protection circuits. Review electrical designs and control behavior, and evaluate performance, cost, reliability, and availability trade-offs. - Define and execute component, shelf, and rack qualification plans covering load transients, current sharing, hot-swap, startup and shutdown, redundancy failover, fault protection and recovery, thermal limits, and AC disturbances and ride-through. - Partner with firmware and software teams to define and validate PMC functionality, telemetry, alarms, power sequencing, power capping, and fault handling. Verify communication interfaces, firmware updates, interoperability, and regression coverage. - Characterize rack power under representative AI workloads. Translate steady-state demand and fast power excursions into design margins, protection settings, energy-storage needs, and deployment requirements with system and data center teams. - Establish fleet rack power health metrics with operations teams. Analyze telemetry, event logs, and failure trends to detect degraded power capacity, redundancy loss, current imbalance, and recurring power faults. - Lead lab and fleet debugging across power hardware, firmware, and system interfaces. Reproduce failures, identify root causes with suppliers, validate corrective actions, and feed lessons into designs, qualification tests, and fleet remediation. - Work with electrical, mechanica...

A summary from the original listing. Read the full details on their site.

Apply on OpenAI's site

Opens in a new tab.

Similar jobs