HybridSan Francisco, CA | New York City, NY | Washington, DC
Safeguards Enforcement Analyst
A
Anthropic
Зарплата
$23,800–$27,500
Формат
Hybrid
О роли
Описание вакансии
About the company
Anthropic's mission is to create reliable, interpretable, and steerable AI systems.
A quickly growing team of researchers, engineers, policy experts, and business leaders building beneficial AI systems.
About the role
As an Enforcement Analyst, you will review content and execute enforcement actions across products and services, focusing on detecting and mitigating misuse of AI systems for malicious cyber operations.
Initial focus on reviewing flagged activity related to cyberattacks, malware development, and offensive exploitation; may later expand to broader areas of enforcement.
Note: may be exposed to explicit content of a violent, technical, or psychologically disturbing nature. May require responding to escalations during weekends and holidays.
Responsibilities
Review flagged content and accounts to make accurate, well-documented enforcement decisions
Detect and mitigate misuse of AI systems to facilitate cyberattacks, malware creation, exploitation tooling, and related harmful cyber operations
Triage and escalate novel, ambiguous, or high-severity cases
Provide detailed feedback to the Safeguards policy design team on policy gaps
Partner with Engineering and Data Science by surfacing detection model errors and quality signals
Maintain high accuracy and consistency standards across review queues
Keep up to date with AI policy enforcement best practices, threat actor tactics, and the evolving cyber threat landscape
Requirements
Experience in cybersecurity, including knowledge of offensive techniques, exploit development, malware analysis, or vulnerability research
Experience performing content review, abuse investigations, or policy enforcement at volume
Proficiency in SQL and/or Python for data analysis and threat detection
Experience identifying emerging risks and communicating findings to cross-functional stakeholders
Experience working with generative AI products, including writing effective prompts
Preferred qualifications
Experience in trust & safety, abuse investigations, cybersecurity investigations, or threat intelligence
Experience with large language models
Experience operating within abuse monitoring programs or enforcement review systems
Understanding of implementing product policies at scale, including content moderation
Experience working with government agencies, regulated environments, or information sharing communities
Conditions
Annual salary: $285,000 – $330,000 USD
Location-based hybrid policy: at least 25% of the time in one of the offices
Visa sponsorship available
Minimum education: Bachelor's degree or equivalent combination of education, training, and/or experience