AI Training Guide · Step 8 of 11
How to Write Justifications That Pass Review
A justification is the short note that explains your rating. Reviewers read it to decide whether your rating is right, so a correct rating with a weak justification can still be marked down.
Trainers on Reddit say one-line or canned rationales are what reviewers mark bad. The fix is a simple formula, close to the dimension, evidence and consequence pattern in published rater guides: name the winner, point to the exact problem, name the error type, and say why it decides the rating.
Sources: r/DataAnnotationTech thread, Label Your Data rater guide, NHS: aspirin.
The Justification Formula
The weak version says nothing a reviewer can check. The strong version can be verified line by line.
Why “How Much Better” Matters
When Anthropic built its 2022 training data, it kept only the comparisons where raters expressed more than the weakest preference, according to its published paper. A justification that says which answer wins but not by how much gives the team less to use. State the size of the gap every time.
Rules for Every Justification
- One claim per sentence.
- Point to the line, do not describe the whole response.
- Use the guideline's words for error types and severity.
- Cover the deciding issue first; minor points come after, if at all.
- Keep it short. Padding hides the reason.
Our top pick gives its trainers three habits that fit here:
- If you are unsure, ask after your first task instead of finishing many the wrong way.
- Keep notes on how you decided edge cases, so your justifications stay consistent.
- Check for guideline updates, since guidelines change.
Phrases to Avoid
| Vague phrase | Write instead |
|---|---|
| “A is more helpful” | Say what A does that B does not |
| “B has some issues” | Name the issue and quote it |
| “Both are good but A is better” | State the specific difference that decides it |
| “A is more accurate” | Point to the inaccurate claim in B |
What Working Trainers Say
Trainer forums on Reddit back this up. The written explanation is what the model learns from, and one-line or canned rationales are what reviewers mark bad.
- Follow the length the project sets. A range written with a plus sign allows more; a plain range means stay inside it.
- Skip the essay. Reviewers say they cut over-long rationales down, so split a long one into short paragraphs.
- Tie every point to the responses in front of you. Stock phrases you could paste into any task are the first thing reviewers flag.
This section draws on Reddit threads such as: Reddit thread, Reddit thread.
Write It Yourself
Never draft a justification with an AI tool. Mercor‘s AI-use policy bans using language models to write justifications or to judge model responses, and Mercor does not pay for time spent breaking it. Other platforms have similar rules. News reports in 2026 describe contractors removed for it, with reviewers watching for repetitive wording and unusually fast work.
Sources: Mercor LLM usage policy, Mercor time tracking and pay policy, Outlier community guidelines, Handshake AI support, Tech Times report, 2026.
A Quick Self-Check
- ✓It names the winner and the size of the gap
- ✓It points to an exact line or claim
- ✓It uses the guideline's terms for the error
- ✓It explains why that issue decides the rating
- ✓A stranger could check every sentence
Find work that uses this skill
Chat with DigiNo Toucan. Two questions, no CV needed.
Want to apply these skills? Start with the AI Training Job Matcher.
AI Training Justification FAQ
What is a justification in AI training?
A short written explanation of why you rated responses the way you did. Reviewers use it to check that your rating is right.
How long should an AI rating justification be?
Follow the range the project sets. Without one, keep it short and specific: trainers say reviewers cut over-long rationales down, so name the winner, point to the exact problem and say why it decides the rating.
Why was my AI training task marked down?
Often because the justification was vague. Point to the exact line, name the error type in the guideline's terms and explain why it matters.
Should I mention every difference between responses?
No. Lead with the deciding issue. Minor differences only matter if nothing bigger separates the responses.
What makes a good AI training justification?
It is specific, checkable and uses the project's own terms, so a reviewer can verify every sentence.
