Side by Side Quality Evaluator – English
We’re looking for English-fluent experts to engage on a project about the overall quality of AI-generated responses.
As an independent contractor, you’ll interact directly with large language models (LLMs) and compare their responses side by side to decide which one better meets the user’s needs. Your evaluations and written rationales will help build AI systems that are more helpful, accurate, and reliable.
Experts may be asked to:
Interact directly with different AI models and tools
Compare two or more model responses side by side (SxS)
Judge responses on several quality dimensions, such as accuracy, helpfulness, instruction-following, clarity, and overall quality
Follow multi-step evaluation workflows and detailed rating guidelines
Write clear, well-reasoned explanations for each evaluation decision
Record consistent judgments across a wide range of prompts and topics
This project may be a strong match for experts who:
Are fluent in English, with strong reading comprehension and writing skills
Have extensive hands-on experience using LLMs and evaluating their outputs
Published about 2 hours ago • Expires November 10, 2026 05:57