Primary record

Software Engineer, Model Deployment- ChatGPT Engineering

OpenAI Indexed employerLondon, UK
Source-hosted applyChecked 3h agoFull-Time
Apply at OpenAI

OpenAI receives this application through Ashby. Babu Careers does not claim delivery.

Workplace

hybrid

Employment

Full-Time

Published

Aug 17, 2026

Closes

No date supplied

The role

ABOUT THE TEAM ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. We develop the systems and tools that make it possible to introduce new models, manage production deployments, respond to operational issues, and use infrastructure effectively at scale. Our work spans distributed systems, platform engineering, infrastructure automation, and developer experience. We partner closely with research, infrastructure, and product teams to make model deployment more reliable, more efficient, and easier to manage. ABOUT THE ROLE We are looking for a software engineer with experience building or operating large-scale production systems. You will design and develop systems that support the model lifecycle in production, including deployment orchestration, configuration management, operational automation, reliability, and capacity management. You will help transform complex operational processes into scalable platform capabilities that enable teams across OpenAI to deploy and manage models with greater confidence and less manual effort. This role is a good fit for engineers who enjoy solving complex operational problems and building software that makes production infrastructure easier to run at scale. IN THIS ROLE, YOU WILL - Build and evolve the platform used to deploy, configure, and manage models across ChatGPT. - Develop systems for deployment orchestration, model rollouts, operational visibility, and production readiness. - Create abstractions and tooling that simplify complex infrastructure and improve the developer experience. - Automate operational workflows, including incident detection, diagnosis, mitigation, and recovery. - Improve the reliability, scalability, and efficiency of model deployments and the infrastructure that supports them. - Build systems that support capacity planning, resource allocation, and infrastructure utilization. - Partner with research, infrastructure, and product engineering teams to identi

Requirements

Department: Applied AI; Team: Applied AI Engineering