Overview
Side by Side runs a platform for evaluating AI-generated responses. This contract role involves comparing LLM outputs side by side, judging them on quality dimensions like accuracy and helpfulness, and writing reasoned explanations for your evaluations to help improve AI systems.
What You’ll Do
- Interact directly with different AI models and tools
- Compare two or more model responses side by side (SxS)
- Judge responses on several quality dimensions, such as accuracy, helpfulness, instruction-following, clarity, and overall quality
- Follow multi-step evaluation workflows and detailed rating guidelines
- Write clear, well-reasoned explanations for each evaluation decision
- Record consistent judgments across a wide range of prompts and topics
Requirements
- Fluent in English, with strong reading comprehension and writing skills
- Extensive hands-on experience using LLMs and evaluating their outputs
- Strong critical-thinking skills and ability to compare responses on several quality dimensions
- Pay close attention to detail and follow instructions carefully
- Ability to explain evaluation decisions clearly and convincingly
- Comfortable following multi-step workflows and switching between different models and tools
Who Should Apply
Apply if you have real hands-on experience evaluating LLM outputs—not just using ChatGPT casually—and you enjoy detailed, analytical work where precision matters. This suits methodical people comfortable with repetitive workflows and switching contexts. It will frustrate anyone looking for high pay, fast earnings, or work that feels varied day to day; this is piecemeal evaluation work at a modest hourly rate.
Salary Insight
$6/hour is below minimum wage in most developed countries and well below market rates for skilled evaluation work. This is positioned as flexible contract work, likely attracting international applicants where the rate may be more competitive. Clarify upfront whether this scales with complexity, and research typical rates for AI evaluation ($15–25/hr for comparable work) before committing significant time.
About Meridial
A contractor marketplace connecting specialists with project-based AI training and evaluation work.
Remote policy: Remote, contractor-based, project assignments vary by expertise and geographic eligibility. No visa sponsorship. Work available across 120+ countries but individual projects have location restrictions.
How hiring works at Meridial →Required Skills
Compensation
$6.00 /hr
- Location
- Worldwide
- Engagement
- Contract
- Posted
- 2d ago
Opens Meridial’s listing on MyRemoteJobs — we don’t collect applications ourselves.
About Meridial
A contractor marketplace connecting specialists with project-based AI training and evaluation work.
Remote policy: Remote, contractor-based, project assignments vary by expertise and geographic eligibility. No visa sponsorship. Work available across 120+ countries but individual projects have location restrictions.
How hiring works at Meridial →Will your CV get past the filter?
Most applications are rejected by software before a person ever reads them. Scan your CV using our ATS Checker tool, see exactly what's blocking it, then fix it in the CV Builder and apply knowing it'll get through.
Scan my CVGet roles like this by email
Tell us what you do and we'll email you when a similar remote job is posted. No account needed.
See our jobs more often in your search
Add MyRemoteJobs as a Preferred Source on Google. Free, takes seconds, and you stay on this page.
Application Tip
Lead with a concrete example of LLM evaluation work you've done—a project where you assessed model outputs, the dimensions you used to judge them, and how you explained your reasoning. The posting emphasizes "clear, well-reasoned explanations" repeatedly; show in your application that you can write one. Generic assertions of critical thinking won't land; specific evidence that you understand what makes one AI response better than another will.
Sourced from MyRemoteJobs · verified against the original posting
Similar open positions
Language Specialist – AI Language Evaluation
Meridial
VerifiedEveryday AI Agent Tester
micro1
VerifiedAI Agent Evaluator
micro1
Verified