Overview
Google is evaluating a new personalization feature for Gemini that uses information from past conversations, Gmail, Google Search, and YouTube activity to make responses more relevant. As an AI Quality Analyst, you will design creative prompts from your personal experiences and assess how well the model applies personalization across dimensions like Grounding, Integration, and Helpfulness.
What You’ll Do
- Design and execute multi-turn conversational prompts (1-5 turns) that test the AI's use of personal information
- Evaluate model responses for appropriate application of personalization based on intent
- Analyze responses for Grounding issues to ensure claims are supported by evidence and free from hallucinations
- Assess Integration quality to ensure personal data is woven naturally without robotic overnarrating
- Stack-rank two model responses side-by-side and write clear, defensible rationales for comparisons
- Extract and verify Debug Info from the model to confirm chat summaries and data sources were utilized
- Maintain data hygiene by deleting evaluation conversations to prevent pollution of chat history
Requirements
- English proficiency with ability to read and write at a high degree of competence
- Willingness to use personal Google account and enable personal data sources
- Full-time availability in local time zone with 4 hours of overlap with PST
- Exceptional analytical thinking to evaluate nuanced and ambiguous AI responses
- Creative prompt engineering skills and experience designing multi-turn prompts
- Strong understanding of personalization concepts and ability to identify poor inferences
- Meticulous attention to detail when reviewing and comparing model responses
- Superior written communication skills to write clear rationales referencing specific turn numbers
- Ability to provide constructive feedback and detailed annotations
- Self-motivated and able to work independently in remote setting
- Desktop/Laptop with good internet connection
- BS/BA degree or equivalent experience in Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or related analytical field
- Preferred: Experience in data annotation, AI quality evaluation, or content moderation
Who Should Apply
This role suits someone with strong analytical mindset and prior experience in AI quality work (annotation, moderation, evaluation), who can think critically about nuance and spot subtle problems in language and reasoning. You should be comfortable with ambiguity—there are no right answers, just defensible judgments. This will frustrate you if you prefer structured tasks with clear criteria, dislike writing detailed rationales, or are uncomfortable linking your personal data (Gmail, search history, YouTube activity) to work evaluation.
Salary Insight
At $30/hr for a 3-month contract role requiring a degree and prior AI evaluation experience, this sits at the lower end for specialized contract work, particularly for roles requiring technical acumen and analytical expertise. The pay does not account for regional variation, though the 4-hour PST overlap requirement suggests US-focused hiring may command different rates. Compare this rate against local freelance data annotation and AI evaluation roles in your region before committing.
Required Skills
Compensation
$30/hr
- Location
- Worldwide
- Engagement
- Contract
- Posted
- Oct 4
Opens Google’s listing on MyRemoteJobs — we don’t collect applications ourselves.
Will your CV get past the filter?
Most applications are rejected by software before a person ever reads them. Scan your CV using our ATS Checker tool, see exactly what's blocking it, then fix it in the CV Builder and apply knowing it'll get through.
Scan my CVGet roles like this by email
Tell us what you do and we'll email you when a similar remote job is posted. No account needed.
See our jobs more often in your search
Add MyRemoteJobs as a Preferred Source on Google. Free, takes seconds, and you stay on this page.
Application Tip
Lead with a concrete example of ambiguous AI output you've evaluated before—what made the judgment hard, how you structured your thinking, and what you concluded. Hiring managers here are looking for evidence that you can articulate *why* one response is better than another with specificity, not just that you can spot obvious errors. If you lack formal AI evaluation experience, bridge it with work in policy review, content moderation, or academic analysis where you had to defend nuanced judgments in writing.
Sourced from MyRemoteJobs · verified against the original posting
Similar open positions
Language Specialist – AI Language Evaluation
Meridial
VerifiedEveryday AI Agent Tester
micro1
VerifiedAI Agent Evaluator
micro1
Verified