About the company
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for AI Evaluators: Assessing A Shopping Assistant based in United States.
Responsibilities
- Review real user interaction traces with an AI-powered shopping assistant, carefully assessing conversations and responses within a dedicated evaluation platform.
- Identify logical failures, factual inaccuracies, irrelevant or unhelpful responses, and poor product recommendations, including subtle issues that may negatively affect the shopping experience.
- Analyze the quality of AI-generated responses from an e-commerce perspective, considering whether recommendations appropriately address user queries and real-world shopping needs.
- Create structured evaluation rubrics that establish clear, repeatable criteria for judging response accuracy, helpfulness, reasoning quality, and overall usefulness.
- Develop verifiers and other structured evaluation mechanisms that can consistently assess future responses and help surface recurring model weaknesses.
- Contribute insights from individual evaluations to broader efforts to improve AI model behavior, response quality, and performance on real-world e-commerce scenarios.
- Maintain a sustained evaluation workload of at least 20 hours per week while working independently and maintaining a high level of accuracy and consistency.
Requirements
- Experience in data evaluation, quality assurance, AI training, data annotation, software testing, prompt engineering, or a closely related analytical discipline.
- Strong analytical and critical-thinking skills, with the ability to identify subtle logical errors, inaccuracies, inconsistencies, and quality issues within written AI interactions.
- Familiarity with e-commerce search, online shopping journeys, product discovery, recommendations, and digital shopping experiences.
- Ability to analyze complex text interactions in depth and distinguish between technically correct responses and responses that are genuinely useful to the user.
- Experience creating structured evaluation criteria, annotation frameworks, testing methodologies, rubrics, or similar quality-assurance systems is valuable.
- Strong attention to detail and consistency, with the ability to apply evaluation standards objectively across a high volume of interactions.
- Ability to work independently in a remote environment, learn new evaluation tools and processes, and communicate findings clearly.
- Availability to commit to a sustained workload of 20+ hours per week.
Conditions
- Compensation of $50 USD per hour.
- Fully remote work, providing flexibility to complete evaluation activities from within the United States.
- Sustained part-time engagement requiring 20+ hours per week, allowing for a consistent workload.
- Opportunity to contribute directly to the evaluation and improvement of AI-powered shopping technology.
- Hands-on exposure to AI evaluation, model quality assessment, e-commerce interactions, structured rubrics, and verification frameworks.
- Opportunity to apply expertise in quality assurance, data evaluation, e-commerce, software testing, or AI training to real-world AI development.