Anthropic
Anthropic

Staff+ Software Engineer, Safeguards Evals

Full-timeSan Francisco, CA | New York City, NYSafeguards (Trust & Safety)1mo ago

Job Overview

About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role How do we know whether a model is safe - and how do we know whether the systems we built to catch misuse actually catch it?

Anthropic answers both questions with evaluations. We measure model behavior across misuse, prompt injection, and user well-being to inform training and deployment decisions. We also use AI to investigate potential misuse of Claude, analyzing real-world traffic to surface bad actors and emerging threats that drive enforcement actions. Neither is worth much unless the evaluations behind them are representative, robust, and trustworthy.

This role builds the methods and infrastructure that make them so. Sitting at the intersection of applied ML research and engineering, you'll design experiments to improve how we evaluate both model behavior and the agentic systems that govern it, build datasets that repres

Core Requirements

Safeguards (Trust & Safety)