HybridSan Francisco, CA | New York City, NY | Washington, DC
Safeguards Enforcement Analyst
A
Anthropic
Зарплата
$20,400–$23,800
Формат
Hybrid
О роли
Описание вакансии
About the company
Anthropic's mission is to create reliable, interpretable, and steerable AI systems.
A quickly growing team of researchers, engineers, policy experts, and business leaders building beneficial AI systems.
About the role
As a Safeguards Analyst on the User Well-being team, you will support the design and deployment of mental health guardrails – iterating on detection systems, managing review queues, evaluating new interventions, and monitoring existing ones.
Interventions range from steering how Claude responds in the conversation to in-product features that connect users to other resources.
The team covers suicide, self-harm, disordered eating, AI sycophancy, and emotional dependence on AI.
Note: may be exposed to explicit content of a sexual, violent, or psychologically disturbing nature.
Responsibilities
Support the design and execution of interventions, defining key metrics, and curating evaluation datasets
Partner with Engineering and Data Science to build, tune, and validate detection models, including threshold-setting and precision/recall tradeoffs
Monitor how interventions and detection systems perform over time
Review flagged content to drive enforcement and policy improvements
Support development of in-product features that connect users to crisis resources
Support the Safeguards Policy Design team with detailed feedback on policy gaps
Keep up to date with emerging AI policy and external research on AI's relationship to mental health
Requirements
Experience in trust & safety, product policy, content moderation, or a related field
Experience designing or running experiments, evaluations, or measurement studies
Experience translating policy definitions into measurable form (rubrics, review guidelines, classification criteria)
Experience managing or coordinating content review operations, including quality assurance and workflow management
Proficiency in SQL and/or other data analysis tools
Experience working with generative AI products, including writing effective prompts
Experience turning open questions and data into concise analysis
Experience identifying emerging risks and communicating findings to cross-functional stakeholders
Understanding of implementing product policies at scale in content moderation
Sound judgment in ambiguous, high-consequence cases
Preferred qualifications
Subject matter expertise in mental health
Experience building or evaluating LLM-based classification systems
Experience using agentic tools (e.g. Claude Code)
Experience working within crisis support
Conditions
Annual salary: $245,000 – $285,000 USD
Location-based hybrid policy: at least 25% of the time in one of the offices
Visa sponsorship available
Minimum education: Bachelor's degree or equivalent combination of education, training, and/or experience