All terms
Preference Ranking
Ordering two or more AI-generated responses from best to worst, the raw data that RLHF and reward models are built from.
This is usually the actual task behind a listing: you're shown a prompt and 2-4 different model responses to it, and asked to rank them (or pick the best one) based on criteria like accuracy, helpfulness, tone, or safety. It sounds simple but is often the highest-volume task type on evaluation platforms.
Example
Given the same customer-support question answered by two different model versions, you pick which answer you'd rather receive and note why in one sentence.