Primary record

Software Engineer - Data Aquisition (systems)

OpenAI Indexed employerSan Francisco
Source-hosted applyChecked 3h ago$255K–$405K/yrFull-Time
Apply at OpenAI

OpenAI receives this application through Ashby. Babu Careers does not claim delivery.

Workplace

On-site

Employment

Full-Time

Published

Jul 27, 2026

Closes

No date supplied

The role

ABOUT THE TEAM This team builds and operates the systems that enable OpenAI researchers to run reliable, scalable, and efficient research workflows. The team sits close to research and works across infrastructure, systems, and automation to make sure researchers have the tools and environments they need to move quickly. The work spans software engineering, infrastructure, systems administration, cluster operations, and reliability engineering. As OpenAI’s infrastructure evolves from bespoke bare-metal systems toward more standard, scalable platforms, the team needs engineers who can understand how systems work end-to-end and build the right abstractions without reinventing the wheel. ABOUT THE ROLE As a Software Engineer on this team, you will build and operate the infrastructure that supports frontier research and critical research-facing systems. You will work on systems that sit close to the metal, but the role is not limited to classic operations or sysadmin work. We are looking for someone who can reason about networking, bootstrapping, Kubernetes, scalability, automation, and reliability - while also writing software to make these systems better over time. This role is a strong fit for an independent, high-ownership engineer who enjoys reliability-heavy infrastructure work but still wants to build. You do not need to come in as a kernel expert or highly algorithmic optimization engineer, but you should be deeply curious about infrastructure, comfortable debugging complex systems, and excited to support researchers doing novel work. WE EXPECT YOU TO: - Build and operate reliable infrastructure for research workloads and research-facing services. - Support and improve systems across data infrastructure, processing, crawl and ingest, caching, search, observability, and clusterwide services. - Improve cluster bootstrapping, provisioning, automation, and deployment workflows. - Debug issues across networking, compute, storage, orchestration, and service reliability

Requirements

Department: Research; Team: Foundations