Primary record

Staff+ Software Engineer, Safeguards

Anthropic Indexed employerSan Francisco, CA | New York City, NY · New York, New York, United States · San Francisco, California, United States
Source-hosted applyChecked 3h agoInternship
Apply at Anthropic

Anthropic receives this application through Greenhouse. Babu Careers does not claim delivery.

Workplace

hybrid

Employment

Internship

Published

Aug 11, 2026

Closes

No date supplied

The role

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role

We are looking for software engineers to help build safety and oversight mechanisms for our AI systems. As a software engineer on the Safeguards team, you will work to monitor models, prevent misuse, and ensure user well-being. This role will focus on building systems to detect unwanted model behaviors and prevent disallowed use of models. You will apply your technical skills to uphold our principles of safety, transparency, and oversight while enforcing our terms of service and acceptable use policies.

We have multiple teams that are currently hiring within Safeguards. Team placement occurs after the interview process, taking into account your interests and experience alongside organizational needs. This flexible approach allows us to match talented engineers where they'll have the greatest impact and growth potential.

Safeguards Acceleration: Builds the agentic systems that let Anthropic's trust & safety teams work at the speed of the models they're protecting. Our flagship project extends Claude Tag, Anthropic's agentic AI collaborator, so it can safely operate on the sensitive data at the heart of Safeguards work: investigating abuse, calibrating detection systems, and closing the loop from signal to enforcement. The hard part isn't making the agent capable, it's making it safe. We develop sandboxed agent architectures, brokered and audited data access, and provenance guarantees that keep humans firmly in control as model capabilities grow. You'll work at the intersection of

Requirements

Department: Safeguards (Trust & Safety)