myRemoteJobs
Quality Assurance

Gmail & Google Calendar AI Assistant Evaluator

micro1

Verified listing
Contract
United States onlyPosted Oct 4$15 - $30/hr

Overview

micro1 is seeking evaluators to help improve AI assistant quality by testing and grading how well AI models handle real-world digital tasks like email, scheduling, and calendar management. In this role you'll use everyday tools to give AI assistants realistic tasks, review their outputs, and provide detailed feedback to help train better models.

What You’ll Do

  • Complete realistic digital tasks (emailing, scheduling, booking, coordinating) using AI assistant tools under standardized conditions
  • Give AI assistants clear instructions for real-world tasks and carefully review their work for accuracy and completeness
  • Apply detailed grading guidelines to evaluate and annotate assistant responses with structured feedback
  • Test scenarios involving email threads, calendar invites, online bookings, and multi-person coordination
  • Record findings and submit evaluation data using AI training tools and platforms
  • Collaborate with project trainers to resolve ambiguities and maintain evaluation consistency
  • Participate in ongoing quality reviews and incorporate feedback

Requirements

  • Based in the US with an iPhone with iMessage and a Facebook or Instagram account
  • Daily hands-on use of Gmail and Google Calendar for personal tasks
  • Regular experience coordinating with others over email or text
  • Comfort booking things online independently
  • Experience using AI tools for practical tasks
  • Exceptional attention to detail and accuracy when reviewing AI-generated output
  • Strong written communication skills
  • Demonstrated time management and self-organization for remote work

Who Should Apply

This suits someone who is digitally organized, detail-oriented, and genuinely comfortable with AI tools—someone who enjoys the precision work of grading and understands both how AI assistants work and where they fail. It will frustrate anyone who finds repetitive task evaluation tedious, prefers creative or strategy work over quality assurance, or lacks patience for careful error-checking.

Salary Insight

$15–$30/hr is a wide band typical of contract evaluation work where pay varies by experience, accuracy metrics, and task complexity. The range suggests you may start at the lower end and move up as you demonstrate consistency and speed. This is gig-rate work, not salaried; clarify with the hiring team upfront whether there are minimum hour guarantees, how payment is structured (per task, per hour, bonuses for quality), and whether you need to track and invoice your own time.

Required Skills

GmailGoogle CalendarAI evaluationSchedulingData labeling

Compensation

$15 - $30/hr

Location
United States only
Engagement
Contract
Posted
Oct 4

Opens micro1’s listing on MyRemoteJobs — we don’t collect applications ourselves.

About Micro1

Platform connecting domain experts with AI training and evaluation projects, paid on a flexible contract basis.

Remote policy: Remote worldwide, excluding Afghanistan, Belarus, China, Cuba, Democratic Republic of the Congo, Hong Kong, Iran, Iraq, Libya, Macao, Myanmar, North Korea, Russia, Somalia, South Sudan, Sudan, Syria, Venezuela, Ukraine and Yemen. Individual projects may have…

How hiring works at Micro1 →

Will your CV get past the filter?

Most applications are rejected by software before a person ever reads them. Scan your CV using our ATS Checker tool, see exactly what's blocking it, then fix it in the CV Builder and apply knowing it'll get through.

Scan my CV

Get roles like this by email

Tell us what you do and we'll email you when a similar remote job is posted. No account needed.

Free. We email you as soon as a role matches — applying early matters. Only genuine matches, and unsubscribe in one click.

See our jobs more often in your search

Add MyRemoteJobs as a Preferred Source on Google. Free, takes seconds, and you stay on this page.

Application Tip

Lead with concrete examples of AI evaluation or quality assurance work you've actually done—even informal grading of ChatGPT outputs for a project, user testing, or data labeling shows you understand what rigor looks like. If you lack formal QA experience, instead describe a time you caught a subtle error or inconsistency that mattered, and how you documented it; this signals the attention to detail they need. Mention any experience writing prompts or debugging AI instructions, as that directly informs how well you'll design test cases and evaluate responses.

Sourced from MyRemoteJobs · verified against the original posting