All terms

Preference Ranking

Ordering two or more AI-generated responses from best to worst, the raw data that RLHF and reward models are built from.

This is usually the actual task behind a listing: you're shown a prompt and 2-4 different model responses to it, and asked to rank them (or pick the best one) based on criteria like accuracy, helpfulness, tone, or safety. It sounds simple but is often the highest-volume task type on evaluation platforms.

Example

Given the same customer-support question answered by two different model versions, you pick which answer you'd rather receive and note why in one sentence.

Common questions

Most platforms still want a forced choice plus a confidence note - ties are allowed but should be the exception, not the default.

Related terms