Overview
A leading AI lab is seeking behavioral health experts to evaluate and improve how its AI models handle sensitive conversations around relationships, emotional wellbeing, and personal beliefs. You'll review interactions, assess model responses for appropriate judgment and balance, and help researchers define what safe, neutral guidance looks like in contexts where people seek support.
What You’ll Do
- Evaluate conversations between users and AI models on topics including relationship advice, emotional wellbeing, spiritual questions, and unconventional beliefs, assessing whether responses are neutral, appropriate and safe
- Identify patterns of concern such as excessive agreement, taking sides, reinforcing unfounded beliefs, moralizing, or overstepping into clinical advice, and document with clear written rationale
- Develop rubrics, guidelines and reference responses that define balanced, supportive and appropriately bounded replies grounded in established counseling and behavioral health practice
- Design test scenarios and conversations that probe how models handle sensitive but non-crisis topics
- Collaborate with researchers and fellow experts to maintain consistent, calibrated and well-documented evaluation standards
Requirements
- Degree in psychology, counseling, social work, behavioral health, behavioral science, human services or closely related field, or equivalent professional experience in mental health
- 3+ years professional experience supporting people in mental health, counseling or social services setting (therapist, counselor, clinical social worker, psychologist, psychiatric nurse, case manager, crisis counselor, peer support specialist or mental health advocate)
- Demonstrated ability to remain neutral and nonjudgmental across diverse perspectives, relationships, belief systems and worldviews, and explain professional judgment
- Working familiarity with concepts such as sycophancy, cognitive distortions, healthy boundaries and client-centered approaches like motivational interviewing
- Ability to engage reliably for at least 20 hours per week during weekdays
- Strong written communication skills and ability to deliver precise, well-structured written feedback
Who Should Apply
This role suits licensed therapists, counselors, social workers or experienced mental health practitioners who have genuine expertise in reading interpersonal dynamics and can write clearly about why one response is better than another. It fits people who understand the difference between supporting someone and colluding with their worldview. Skip it if you expect clinical decision-making authority or if you're uncomfortable rating AI responses against an external rubric rather than your own professional judgment.
Salary Insight
$45–$70/hour is wide, likely reflecting the span from entry-level (3 years experience, bachelor's degree) to more senior practitioners (licensed clinical expertise). At 20 hours per week minimum, expect $3,600–$5,600/month before tax; at full 40 hours, $7,200–$11,200/month. This is contingent work, not salaried employment, so there's no mention of health insurance or benefits in the main role—however, note that the employer-of-record structure (Cincinnatus) may offer some protections.
About Mercor
A marketplace placing specialists into paid remote contract work with AI companies. One profile and one AI interview, reused across every role you apply to.
Remote policy: Fully remote, worldwide, on hourly contractor agreements rather than employment. But payment eligibility is the thing to check before you invest any time, because it is decided by country and there is no way around it. Mercor pays through Stripe Connect in…
How hiring works at Mercor →Required Skills
Compensation
$45–$70 per hour
- Location
- United States
- Engagement
- Part-time
- Posted
- Oct 4
Opens Mercor’s listing on MyRemoteJobs — we don’t collect applications ourselves.
About Mercor
A marketplace placing specialists into paid remote contract work with AI companies. One profile and one AI interview, reused across every role you apply to.
Remote policy: Fully remote, worldwide, on hourly contractor agreements rather than employment. But payment eligibility is the thing to check before you invest any time, because it is decided by country and there is no way around it. Mercor pays through Stripe Connect in…
How hiring works at Mercor →Will your CV get past the filter?
Most applications are rejected by software before a person ever reads them. Scan your CV using our ATS Checker tool, see exactly what's blocking it, then fix it in the CV Builder and apply knowing it'll get through.
Scan my CVGet roles like this by email
Tell us what you do and we'll email you when a similar remote job is posted. No account needed.
See our jobs more often in your search
Add MyRemoteJobs as a Preferred Source on Google. Free, takes seconds, and you stay on this page.
Application Tip
Bring a specific example of a difficult conversation you've held where neutrality mattered—where you resisted the urge to steer someone toward your own belief or preference, and explain what you did instead and why. This team is training AI to do exactly that, so evidence you can articulate your own reasoning about judgment in ambiguous cases will set you apart far more than listing credentials. A concise account of a case where you caught yourself (or a peer) straying into moralizing or sycophancy will land harder than a generic description of your clinical training.
Sourced from MyRemoteJobs · verified against the original posting
Similar open positions
Language Specialist – AI Language Evaluation
Meridial
VerifiedEveryday AI Agent Tester
micro1
VerifiedAI Agent Evaluator
micro1
Verified