AI Training Guide · Step 7 of 11
How to Rate and Rank AI Responses
Rating AI responses means comparing two or more answers to the same prompt and deciding which is better, by how much, and why. It is one of the core tasks in AI training, and Mercor‘s AI-use policy names judging model responses as work trainers must do themselves.
Consistency counts. When OpenAI picked raters for InstructGPT, it screened them partly on how often they agreed with its own rankings.
Sources: Mercor LLM usage policy, Outlier FAQ, OpenAI InstructGPT paper.
What a Rating Task Looks Like
Response B fails on safety and correctness at once. The NHS says not to give aspirin to children under 16 because of a rare risk of Reye's syndrome, and it advises seeing a doctor if a child's fever lasts 5 days or more, not two weeks. A follows the NHS advice, so A is much better.
Sources: NHS: fever in children, NHS: aspirin.
How the AI Labs Describe a Good Answer
OpenAI published the instructions its human raters followed when training InstructGPT, in the appendix of its 2022 paper. Raters judged three things:
| Quality | What raters checked |
|---|---|
| Helpful | Does what the user meant, asks for clarification when the request is unclear, and does not ramble |
| Truthful | No invented details, no false claims, and false premises are corrected rather than played along with |
| Harmless | Nothing that could hurt the user or others |
For final evaluations, truthful and harmless usually came first. The one exception: prefer the more helpful answer only if it is much more helpful, barely worse on safety, and the topic is not high stakes like medical, legal or money advice. When still unsure, they were told to ask which answer they would rather get from a customer assistant.
The same paper shows the priority was flipped during an earlier training phase. That is the most useful lesson for new raters: general advice is a default, and your project's written instructions always win.
Where Your Ratings Go
The InstructGPT paper lays out the three training steps that rating and writing work feeds:
- People write ideal responses to prompts, and the model is fine-tuned to imitate them.
- People rank several model responses from best to worst, and those rankings train a reward model that predicts which answer people prefer.
- The model is then trained with reinforcement learning to produce answers the reward model scores highly.
That is why careless ratings matter. Each ranking becomes part of what the model is taught to prefer.
Sources: OpenAI InstructGPT paper.
Check in This Order
Rules Good Raters Follow
- Read the project guideline fully before the first task, then re-read its rules on severity and ties.
- Verify facts yourself instead of trusting confident wording.
- Do not reward length. A shorter correct answer beats a longer padded one.
- Check every explicit constraint in the prompt: word count, format, language, persona.
- Use a tie only when the guideline allows it and the responses are genuinely equal.
Sources: OpenAI InstructGPT paper, Anthropic 2022 study, r/DataAnnotationTech thread.
What Google Tells Its Raters
Google's Search Quality Rater Guidelines (September 2025 edition) are written for search results, but three ideas carry straight over to AI responses:
- Top scores are rare. The guidelines say most queries “cannot have a Fully Meets result”. Over-rating is the classic beginner error.
- Trust matters most. Of Experience, Expertise, Authoritativeness and Trust, trust is the most important.
- Some topics need stricter checks. Health, money, safety and civic topics can seriously affect people, so even a small error counts heavily. The fever example above is one of these.
Two Traps the Research Found
- Polish is not truth. In Anthropic's 2022 study, raters often preferred answers containing links that did not work. Check the facts and the links, not the formatting.
- A refusal is not automatically the safest good answer. The same research notes that an answer explaining why a request is harmful is better than a bare refusal.
Disagreement is also normal. OpenAI's raters agreed with each other about 73% of the time, and Anthropic measured about 63% agreement between researchers and crowdworkers. Calibration with reviewers is part of the job, not a sign you are failing.
Your ratings are checked in layers. At our top pick, peer reviewers check work for accuracy and guideline adherence, an operations team may check it again, and some projects add a client review.
How Much Better: Using the Scale
| Rating | When it fits |
|---|---|
| Much better | One response fails on safety or correctness and the other does not |
| Slightly better | Both are acceptable, but one is clearer, more complete or better formatted |
| About the same | No meaningful difference at any level, and the guideline allows a tie |
Project guidelines define these levels in their own words. When they do, use theirs.
Common Rating Mistakes
- Picking the longer or friendlier answer without checking facts.
- Missing a constraint, such as a word limit, that one response ignored.
- Letting a nicely formatted answer hide a wrong fact or a broken link.
- Rating from memory of the guideline instead of checking it.
Video and Audio Annotation
Some AI training work has no text to rate at all. You watch or listen, then mark where something happens: when an action starts and ends, what kind of event it is, or what someone said and when they said it.
Mercor's public job board shows how this work is described. On 2026-10-09 it listed eight Audiobook QA Expert roles across seven languages, where reviewers log each narration error and mark the exact span where it occurs. In August 2026 it listed a first-person video project where reviewers checked whether each action segment's start and end matched the action. Ethos‘s expert terms also name annotation among the project work experts take on.
Sources: Mercor job board, Ethos expert terms.
What the Tasks Look Like
- Segment labeling means marking the start and end of a moment or action, then tagging it with a label from the project's list.
- Boundary review means checking segments that someone else or a model marked, then adjusting or flagging the ones that are off.
- Transcription work asks you to write down what is said, fix speech-to-text errors and tie each line to the moment it is spoken.
- Audio review asks you to log problems such as skipped, added or mispronounced words, numbers read wrongly and audio glitches, put each in a category and mark where it happens.
Both marks picked by eye while the clip played at normal speed.
Start: first frame the egg touches the bowl rim. End: first frame both shell halves leave the bowl.
Checked by stepping one frame either side of each mark.
How to Be Precise
- Work in frames when the tool allows it. At 30 frames per second one frame lasts about 33 milliseconds, and at 25 frames per second it lasts 40. A Mercor video editing listing asks for frame-accurate cuts, and the BBC's subtitle guidelines write times in frames too.
- Check every boundary at the frame level. Find the moment at normal speed, then step one frame at a time either side of your mark before you save it.
- Write a default rule for unclear starts and ends, with a unit, such as “start on the first frame of contact, end on the first frame the hand lets go”. Use the project's own rule wherever it has one.
- Keep one unit per project. Frame 300 is 10 seconds at 30 frames per second but 12 seconds at 25, so note each clip's frame rate and convert to seconds or milliseconds when the tool asks for time.
- Time speech from its first sound. The BBC guidelines say a subtitle should appear as speech starts, and a new speaker's line as that person starts to speak.
Expect boundaries to be hard. In a 2017 study, crowd workers who re-marked actions in two video datasets overlapped the original labels by about 72.5% and 58.7%. Ends were harder to place than starts, with median errors of about 1.4 seconds against 0.9, and short actions were the hardest of all, so give them extra care.
Sources: Sigurdsson et al. 2017, BBC subtitle guidelines, Mercor job board.
Staying Consistent Over Long Sessions
The Mercor video review listing asked for “consistent judgment across a high volume of short action segments”. These habits help:
- Keep an edge-case log. Each time you make a call the guideline does not cover, write the rule down with one example timestamp, so the next case gets the same call.
- Re-check a sample. Reopen a few of your earliest segments after a batch and compare them with your latest ones; drift shows up there first.
- Calibrate with reviewers. Read their feedback, and ask the project lead in the project channel when your rule and theirs differ. Trainers on Reddit say guidelines change often and that the project chat is the place to ask.
- Take breaks. Trainers on Reddit warn that very long sessions lead to autopilot and slipping quality, and an audiobook listing asks for accuracy across many hours of audio.
Sources: Reddit thread, Reddit thread, Reddit thread.
Common Annotation Mistakes
- Ending segments too late or too early. The end of an action is usually less clear than its start, so check ends twice.
- Setting boundaries at normal playback speed without stepping through frames.
- Mixing frames and seconds, or reusing frame numbers between clips with different frame rates.
- Labeling what you think happened instead of what you hear. The BBC guidelines say a sound label should describe the sound itself, such as floorboards creaking, rather than the action behind it.
- Changing your own rule halfway through a batch without going back to fix the earlier segments.
My first AI training work was labelling moments in videos.
What Working Trainers Say
Two points come up again and again in trainer forums on Reddit.
- Rate exactly what the task asks. If it says to judge the final turn, judge that turn and leave the earlier conversation out of it, in the format requested.
- Your rationale carries the rating. Reviewers say they pass ratings they personally disagree with when the reasoning is specific and tied to the response. One-line rationales are what they mark down.
This section draws on Reddit threads such as: Reddit thread, Reddit thread.
Rating is half the task. The other half is explaining your rating, which is step 8.
Put these skills to work
Chat with DigiNo Toucan. Two questions, no CV needed.
Ready to apply? The AI Training Job Matcher points you to platforms that fit your background.
Rating AI Responses FAQ
How do you rate AI responses?
Compare the responses in a fixed order: safety, correctness, instruction following, helpfulness, then style. Stop at the first level where they clearly differ, then pick the rating that matches the size of the difference.
What matters most when ranking AI responses?
Safety first, then factual correctness. A harmful or wrong answer usually loses, however well it is written.
Should longer AI responses be rated higher?
No. Length is not quality. A shorter correct answer beats a longer one with padding or errors.
When is a tie allowed in AI rating tasks?
Only when the project guideline allows ties and the responses are genuinely equal at every level.
What is video annotation in AI training?
Marking where moments start and end in a video or audio file, tagging what happens, and sometimes transcribing and timestamping speech. Precise work is checked frame by frame and follows one consistent rule for unclear boundaries.
How are AI raters scored?
Through quality checks on reviewed work. Mercor sets quality thresholds per project from reviewed tasks, ratings and disputes.

