← Все вакансии/Anthropic
HybridSan Francisco, CA | New York City, NY | Washington, DC

Safeguards Enforcement Analyst

A
Anthropic
Зарплата
$20,400–$23,800
Формат
Hybrid
О роли

Описание вакансии

About the company

  • Anthropic's mission is to create reliable, interpretable, and steerable AI systems.
  • A quickly growing team of researchers, engineers, policy experts, and business leaders building beneficial AI systems.

About the role

  • As a Safeguards Analyst on the User Well-being team, you will support the design and deployment of mental health guardrails – iterating on detection systems, managing review queues, evaluating new interventions, and monitoring existing ones.
  • Interventions range from steering how Claude responds in the conversation to in-product features that connect users to other resources.
  • The team covers suicide, self-harm, disordered eating, AI sycophancy, and emotional dependence on AI.
  • Note: may be exposed to explicit content of a sexual, violent, or psychologically disturbing nature.

Responsibilities

  • Support the design and execution of interventions, defining key metrics, and curating evaluation datasets
  • Partner with Engineering and Data Science to build, tune, and validate detection models, including threshold-setting and precision/recall tradeoffs
  • Monitor how interventions and detection systems perform over time
  • Review flagged content to drive enforcement and policy improvements
  • Support development of in-product features that connect users to crisis resources
  • Support the Safeguards Policy Design team with detailed feedback on policy gaps
  • Keep up to date with emerging AI policy and external research on AI's relationship to mental health

Requirements

  • Experience in trust & safety, product policy, content moderation, or a related field
  • Experience designing or running experiments, evaluations, or measurement studies
  • Experience translating policy definitions into measurable form (rubrics, review guidelines, classification criteria)
  • Experience managing or coordinating content review operations, including quality assurance and workflow management
  • Proficiency in SQL and/or other data analysis tools
  • Experience working with generative AI products, including writing effective prompts
  • Experience turning open questions and data into concise analysis
  • Experience identifying emerging risks and communicating findings to cross-functional stakeholders
  • Understanding of implementing product policies at scale in content moderation
  • Sound judgment in ambiguous, high-consequence cases

Preferred qualifications

  • Subject matter expertise in mental health
  • Experience building or evaluating LLM-based classification systems
  • Experience using agentic tools (e.g. Claude Code)
  • Experience working within crisis support

Conditions

  • Annual salary: $245,000 – $285,000 USD
  • Location-based hybrid policy: at least 25% of the time in one of the offices
  • Visa sponsorship available
  • Minimum education: Bachelor's degree or equivalent combination of education, training, and/or experience
Стек и навыки

С чем работаем