HybridSan Francisco, CA | New York City, NY | Washington, DC
Safeguards Enforcement Analyst
A
Anthropic
Зарплата
$23,800–$27,500
Формат
Hybrid
О роли
Описание вакансии
About the company
Anthropic's mission is to create reliable, interpretable, and steerable AI systems.
A quickly growing team of researchers, engineers, policy experts, and business leaders building beneficial AI systems.
About the role
As a Safeguards Enforcement Analyst focused on Violence & Extremism, you will build and execute operational workflows to assess model behavior, drive enforcement decisions, and develop evals across a technically demanding range of policy areas.
Work spans detecting and mitigating misuse of AI systems to facilitate real-world harm, including weapons and dangerous technology, critical infrastructure attacks, violent extremism, and threats of violence.
Note: may be exposed to explicit content of a violent, graphic, hateful, or psychologically disturbing nature.
Responsibilities
Design and architect automated enforcement systems and review workflows that scale while maintaining high accuracy
Develop and maintain evals that measure model performance, surface regressions, and inform policy and model improvements
Partner with Engineering and Data Science to optimize detection and automated enforcement systems
Review flagged content to drive enforcement decisions and surface policy gaps
Support the Safeguards policy design team with structured feedback on policy gaps and enforcement ambiguities
Develop and maintain enforcement guidelines and reviewer documentation
Keep up to date with emerging threats, extremist movements, regulatory changes, and AI policy enforcement best practices
Identify and escalate emerging misuse patterns, novel attack vectors, and coordinated violent extremist activity
Requirements
Experience in policy enforcement, threat intelligence, counterterrorism, government, or a closely related field
Experience standing up and scaling policy enforcement or content review workflows
Proficiency in SQL and/or other data analysis tools
Experience identifying emerging risks and threat actors and communicating findings to cross-functional stakeholders
Experience working with generative AI products, including writing effective prompts
Understanding of implementing product policies at scale, including content moderation
Preferred qualifications
Subject matter expertise in weapons and dangerous technology, violent extremism, terrorism, autonomous systems, or critical infrastructure protection
Familiarity with legal and regulatory frameworks governing dangerous technology, critical infrastructure, or terrorism
Experience developing evals or red-teaming AI systems
Experience with threat actor profiling and threat intelligence frameworks (e.g., MITRE ATT&CK)
Experience tracking threat actors across surface, deep, and dark web environments
Experience with large language models
Proficiency in Python for data analysis and workflow automation
Background in law enforcement, national security, defense, counterterrorism, or regulatory environment
Familiarity with cross-platform threat analysis and OSINT techniques
Conditions
Annual salary: $285,000 – $330,000 USD
Location-based hybrid policy: at least 25% of the time in one of the offices
Visa sponsorship available
Minimum education: Bachelor's degree or equivalent combination of education, training, and/or experience