All terms

Inter-Annotator Agreement

A measure of how consistently different human reviewers label or score the same piece of data.

Platforms track inter-annotator agreement to catch unreliable grading - if the same prompt gets wildly different scores from different reviewers, either the rubric is unclear or someone isn't applying it consistently. High agreement is often a factor platforms use when deciding who gets more (or higher-paying) work.

Example

Three annotators independently label the same 100 responses; the platform checks how often they agree, and follows up with anyone whose scores are consistent outliers.

Common questions

Occasional disagreement is normal and expected - platforms usually only flag it as a problem when it's a consistent, systematic pattern rather than a one-off.

Related terms