OpenAI · Infrastructure · Unspecified · Posted 2026-09-16
Software Engineer, Compute Foundations
OpenAI · San Francisco · $255k–490k base
This range's midpoint is above 95% of posted infrastructure ranges at AI companies right now. See the salary index.
Apply on OpenAI's site Watch OpenAI for new roles
ABOUT THE TEAM
Compute Foundations builds the software that manages OpenAI’s GPU compute infrastructure across sites, data centers, and infrastructure providers, supporting model training and inference. Our systems turn large, heterogeneous fleets of machines into dependable compute for research and products.
We build Kubernetes-based control planes, controllers, services, and APIs that coordinate the lifecycle of machines and clusters. We connect global infrastructure management with the realities of bare-metal systems, giving clients consistent interfaces across differences in hardware, topology, and provider behavior.
ABOUT THE ROLE
You will build distributed systems that provision, configure, and manage compute throughout its lifecycle. Your work will connect global services and Kubernetes controllers with the systems that bring machines online, update them safely, and recover them when something goes wrong.
This role combines software architecture with an understanding of how machines and data centers work. You might design a lifecycle API, improve controller performance under high concurrency and provider rate limits, or trace a provisioning failure from an API through reconciliation to network boot or host configuration. You will help these systems remain reliable as the fleet expands across sites and generations of GPU hardware.
We value depth in relevant systems and the ability to connect layers. You do not need to arrive as an expert in every component of the stack.
IN THIS ROLE, YOU WILL:
- Design, build, and operate Kubernetes-based controllers and distributed services that coordinate infrastructure across sites, isolate failures, and scale as GPU capacity grows.
- Define APIs and resource models that let clients request and track lifecycle operations through consistent interfaces across hardware platforms and providers.
- Build provisioning and configuration services that coordinate network boot, hardware management interfaces, and the deployment of firmware, operating-system images, drivers, and host configuration.
- Develop lifecycle management for discovery, allocation, provisioning, upgrades, maintenance, recovery, and decommissioning, integrating with health and validation systems.
- Design reliable reconciliation and recovery through concurrent changes, interrupted operations, and partial failures, with staged rollouts that limit disruption across nodes, racks, and clusters.
- Improve control-plane throughput, API latency, and the time infrastructure takes to reach its desired state, while respecting the limits of site systems and provider APIs.
- Build the software integrations that bring new sites and GPU hardware generations into the platform, partnering with hardware, networking, data-center, and other infrastructure teams.
YOU MIGHT THRIVE IN THIS ROLE IF YOU:
- Have strong software engineering fundamentals and experience designing, implementing, and owning production distributed systems or infrastructure services.
- Have experience developing infrastructure systems that use Kubernetes APIs and reconciliation to manage resources.
- Understand how a bare-metal node moves from power-on to a configured, workload-ready system, with depth in one or more areas such as PXE, DHCP/DNS, baseboard management controllers (BMCs), firmware, Linux, drivers, images, or configuration management.
- Can design reliable APIs and asynchronous workflows, reasoning about concurrency, consistency, idempotency, and failures across service and provider boundaries.
- Can diagnose reliability and performance problems across service, operating-system, and machine boundaries, turning production evidence into lasting software improvements.
- Work effectively across engineering specialties and communicate system behavior and technical tradeoffs clearly.
BONUS POINTS IF YOU:
- Have built infrastructure control planes that coordinate operations across multiple sites or regions.
- Have worked with G …
More infrastructure roles at OpenAI
-
Platform Engineering Manager, Forward Deployed Engineering (FDE)
InfrastructureLead / ManagerRemote US$302k–335ktoday
-
Data Center Hardware Quality & Reliability Engineer
InfrastructureStaff+Remote US$226k–285ktoday
-
Software Engineer, DevOps
InfrastructureRemote US$177k–327ktoday
-
Technical Program Manager, Hardware Systems
InfrastructureUS$207k–242ktoday
-
Software Engineer, Search Infrastructure
InfrastructureRemote US$266k–445k3d
See also: Software Engineer jobs · AI jobs in San Francisco Bay Area · OpenAI salaries · Kubernetes jobs.
This listing is reproduced from OpenAI's public careers feed and links to the original. AI Hiring Index is not the employer and does not accept applications. All OpenAI roles · AI salaries.