Work Expert (WE) is an independent publishing and referral website. We are not a recruiter, hiring manager, agent or employer, and we are not affiliated with or endorsed by Mercor. Applying takes you to the platform's own website, where we may be recorded as the referring source. We may receive a referral fee at no additional cost to you. Read our full affiliate disclosure →
About the Role
Fluent Language Skills Required: English & Odia. Native fluency in English and Odia is required for this position. At Mercor, we believe the safest AI is the one that’s already been attacked - by us. We are assembling a red team for this project - human data experts who probe AI models with adversarial inputs, surface vulnerabilities, and generate the red team data that makes AI safer for our customers.
What You'll Do
Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation
Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks
Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent
Document reproducibly: produce reports, datasets, and attack cases customers can act on
Independent contractor engagement, fully remote. Mercor pays weekly via Stripe or Wise.
Common Questions
Mercor has a reliability rating of very reliable based on the listings we track, with an onboarding time of 1-3 weeks (interview + assessment + trial). See our full Mercor review for the details behind that rating.
Applying takes you to Mercor's own website, where the application is completed. We are not a recruiter or employer and do not process applications directly - we may be recorded as the referring source.
$20-$22/h - Independent contractor engagement, fully remote. Mercor pays weekly via Stripe or Wise.
Red team conversational AI models and agents: jailbreaks, prompt injections, misuse cases, bias exploitation, multi-turn manipulation Generate high-quality human data: annotate failures, classify vulnerabilities, and flag systemic risks Apply structure: follow taxonomies, benchmarks, and playbooks to keep testing consistent