Primary record

Engineering Manager, Safeguards Interventions

Anthropic Indexed employerSan Francisco, CA · San Francisco, California, United States
Source-hosted applyChecked 3h agoInternship
Apply at Anthropic

Anthropic receives this application through Greenhouse. Babu Careers does not claim delivery.

Workplace

hybrid

Employment

Internship

Published

Aug 11, 2026

Closes

No date supplied

The role

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role

The Safeguards team is responsible for ensuring our models and products are developed and deployed safely. We're looking for an Engineering Manager to lead the Interventions team: the group responsible for what happens when a safety system fires. It owns the composable arsenal of systems that sit between our detection stack (classifiers and probes) and the user, across every Anthropic surface: 1P products, the API, and third-party clouds. This includes inline interventions for areas like bio, cyber, and acceptable usage as well as downstream areas like child safety and copyright. This team is responsible for ensuring that we evolve and drive the quality of our interventions to enable our products to grow safely.

Key responsibilities

• Hands-on lead and grow a team of engineers; own roadmap, OKRs, and execution.

• Drive cross-functional work with ML Infra, Research, Product, Policy, and Legal - and with cloud partners for 3P deployment.

• Set the bar for when an intervention is good enough to ship - backed by measurement - and represent safety and product tradeoffs to leadership and external stakeholders.

• Own production reliability for intervention and compliance systems: incident response, postmortems, SLOs, and the verification processes that prevent repeat incidents.

Minimum qualifications

• Have managed engineering teams shipping production ML or safety-enforcement systems where the system's decisions directly affected users.

• Have run high-stakes, compliance-adjac

Requirements

Department: Safeguards (Trust & Safety)