← Все вакансии/jobgether
RemoteUS

AI Evaluator

J
jobgether
Зарплата
$100
Формат
Remote
О роли

Описание вакансии

About the company

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for AI Evaluators: Assessing A Shopping Assistant based in United States.

Responsibilities
  • Review real user interaction traces with an AI-powered shopping assistant, carefully assessing conversations and responses within a dedicated evaluation platform.
  • Identify logical failures, factual inaccuracies, irrelevant or unhelpful responses, and poor product recommendations, including subtle issues that may negatively affect the shopping experience.
  • Analyze the quality of AI-generated responses from an e-commerce perspective, considering whether recommendations appropriately address user queries and real-world shopping needs.
  • Create structured evaluation rubrics that establish clear, repeatable criteria for judging response accuracy, helpfulness, reasoning quality, and overall usefulness.
  • Develop verifiers and other structured evaluation mechanisms that can consistently assess future responses and help surface recurring model weaknesses.
  • Contribute insights from individual evaluations to broader efforts to improve AI model behavior, response quality, and performance on real-world e-commerce scenarios.
  • Maintain a sustained evaluation workload of at least 20 hours per week while working independently and maintaining a high level of accuracy and consistency.
Requirements
  • Experience in data evaluation, quality assurance, AI training, data annotation, software testing, prompt engineering, or a closely related analytical discipline.
  • Strong analytical and critical-thinking skills, with the ability to identify subtle logical errors, inaccuracies, inconsistencies, and quality issues within written AI interactions.
  • Familiarity with e-commerce search, online shopping journeys, product discovery, recommendations, and digital shopping experiences.
  • Ability to analyze complex text interactions in depth and distinguish between technically correct responses and responses that are genuinely useful to the user.
  • Experience creating structured evaluation criteria, annotation frameworks, testing methodologies, rubrics, or similar quality-assurance systems is valuable.
  • Strong attention to detail and consistency, with the ability to apply evaluation standards objectively across a high volume of interactions.
  • Ability to work independently in a remote environment, learn new evaluation tools and processes, and communicate findings clearly.
  • Availability to commit to a sustained workload of 20+ hours per week.
Conditions
  • Compensation of $50 USD per hour.
  • Fully remote work, providing flexibility to complete evaluation activities from within the United States.
  • Sustained part-time engagement requiring 20+ hours per week, allowing for a consistent workload.
  • Opportunity to contribute directly to the evaluation and improvement of AI-powered shopping technology.
  • Hands-on exposure to AI evaluation, model quality assessment, e-commerce interactions, structured rubrics, and verification frameworks.
  • Opportunity to apply expertise in quality assurance, data evaluation, e-commerce, software testing, or AI training to real-world AI development.
Стек и навыки

С чем работаем