Overview
Mercor partners with frontier AI labs to train and evaluate models using human expertise. This role involves auditing software-engineering benchmark tasks (repository-level tasks, reference patches, test harnesses) used to train and grade a frontier AI lab's models, ensuring quality, correctness, and reproducibility.
What You’ll Do
- Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks
- Assess repository-level tasks, reference patches, and test harnesses
- Audit grading integrity and detect answer leakage or reward hacking
- Provide clear, rubric-based written feedback
Requirements
- 3+ years professional software engineering experience
- Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles)
- Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking
- Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++)
- Familiarity with SWE-Bench (Verified) or similar repository benchmarks (preferred)
- Maintainer history on major Python OSS such as Django, Flask, scikit-learn, sympy, pytest (preferred)
- Prior code-review or task-grading experience (preferred)
Who Should Apply
This suits an experienced software engineer with genuine open-source maintainer or heavy-contributor credibility who enjoys meticulous, detail-oriented review work over building new features; engineers who prefer fast-paced greenfield coding or dislike writing detailed critique feedback will likely find this tedious.
Salary Insight
The $70-$90/hour rate is solid for contract-based code review/audit work and reflects the specialized skill of vetting benchmark integrity rather than standard development, though as a contractor role there's no benefits or guaranteed hours, so treat the hourly rate as the full picture rather than a supplement to a salary.
About Mercor
A marketplace placing specialists into paid remote contract work with AI companies. One profile and one AI interview, reused across every role you apply to.
Remote policy: Fully remote, worldwide, on hourly contractor agreements rather than employment. But payment eligibility is the thing to check before you invest any time, because it is decided by country and there is no way around it. Mercor pays through Stripe Connect in…
How hiring works at Mercor →Required Skills
Compensation
$70-$90 per hour
- Location
- United States
- Engagement
- Remote
- Posted
- Sep 1
Opens Mercor’s listing on MyRemoteJobs Team — we don’t collect applications ourselves.
About Mercor
A marketplace placing specialists into paid remote contract work with AI companies. One profile and one AI interview, reused across every role you apply to.
Remote policy: Fully remote, worldwide, on hourly contractor agreements rather than employment. But payment eligibility is the thing to check before you invest any time, because it is decided by country and there is no way around it. Mercor pays through Stripe Connect in…
How hiring works at Mercor →Will your CV get past the filter?
Most applications are rejected by software before a person ever reads them. Scan your CV using our ATS Checker tool, see exactly what's blocking it, then fix it in the CV Builder and apply knowing it'll get through.
Scan my CVGet roles like this by email
Tell us what you do and we'll email you when a similar remote job is posted. No account needed.
See our jobs more often in your search
Add MyRemoteJobs as a Preferred Source on Google. Free, takes seconds, and you stay on this page.
Application Tip
Lead with concrete evidence of your open-source maintainer or heavy-contributor history — links to merged PRs, committer status, or repos you've reviewed for others — since this is the credibility signal the posting explicitly screens for. If you've caught subtle bugs in test harnesses, flagged flaky tests, or identified data/answer leakage in an evaluation pipeline before, describe that incident specifically; it demonstrates the exact adversarial-auditing mindset this role needs far better than a general resume of coding projects.
Sourced from MyRemoteJobs Team · verified against the original posting
Similar open positions
Language Specialist – AI Language Evaluation
Meridial
VerifiedEveryday AI Agent Tester
micro1
VerifiedAI Agent Evaluator
micro1
Verified