Back to Blog

AI People Search Benchmark 2026: Wrangle vs. Exa, Juicebox & More

Aug 20, 20264 min readBy: Wrangle
AI People Search Benchmark 2026: Wrangle vs. Exa, Juicebox & More

Recruiting software vendors describe their products as "AI-powered," "intelligent," and "precision-driven" without publishing results on a shared query set that anyone can reproduce or contest. SCOUT (Sourcing and Candidate Understanding Test) is how Wrangle publishes its results: by running against established benchmarks with defined scoring protocols, shared query sets, and maintainer-published leaderboards, and reporting the numbers.

Two benchmarks are used here. The Exa People Search Benchmark tests hit rate and precision across 1,400 natural-language sourcing queries. PeopleSearchBench (maintained by LessieAI) scores on a composite of relevance precision, effective coverage, and information utility across a recruiting-specific query set. Where a competitor's result appears on the maintainer's published leaderboard, it is labeled as such. Where a competitor self-reported a number that is not on the maintainer's leaderboard, it is labeled vendor-reported. Numbers labeled vendor-reported should be treated as directional until independently confirmed.


Figure 01: Exa People Search Benchmark

The Exa People Search Benchmark runs 1,400 natural-language sourcing queries against a fixed candidate pool. Hit@1 measures whether the correct candidate appears as the first result. Hit@10 measures whether the correct candidate appears within the first ten results. Macro precision measures the share of returned candidates that are relevant across all 1,400 queries. Task completion measures the share of queries where the system returned a result. Qualified results measures the average number of relevant candidates returned per query slot (out of 15).

MetricWrangle (Normal)Wrangle (Turbo)DeltaExa (published)
Hit@194.07% (1,317/1,400)93.71%+0.36 pp72.00%
Hit@1098.07% (1,373/1,400)97.64%+0.43 ppNot published
Macro precision89.81%89.29%+0.53 ppNot published
Task completion100%100%Not published
Qualified results14.80 of 1514.80 of 15Not published

Wrangle Normal is the current evaluation run. Wrangle Turbo is the prior run for comparison. Exa's published Hit@1 of 72.00% comes from the maintainer-published leaderboard. Exa has not published Hit@10, macro precision, task completion, or qualified-results figures for its own system.

Competitor status on this benchmark:

SystemExa leaderboardPublished Hit@1Published macro precision
WrangleYes94.07%89.81%
ExaYes72.00%Not published
JuiceboxNoNot published79% (vendor-reported)
Metaview SourcingNoNot published93.50% (vendor-reported)
PinNoNot publishedNot published
NoonNoNot publishedNot published

Juicebox reports a 79% precision figure from its own run of Exa's benchmark; this number does not appear on Exa's maintainer-published leaderboard. Metaview reports 93.50% precision from its own Exa-derived run and is not listed on the Exa leaderboard; it has not published corresponding Hit@1 or Hit@10 scores.


Figure 02: PeopleSearchBench (Recruiting subset)

PeopleSearchBench is maintained by LessieAI. Overall score is a composite of relevance precision, effective coverage, and information utility. The results below are from Wrangle's Turbo evaluation run.

SystemOverallRelevance precisionEffective coverageInformation utility
Wrangle (Turbo)93.4294.1498.4487.69
Juicebox65.7366.1075.3055.80
PinNo published result
NoonNo published result
Metaview SourcingNo published result

Wrangle and Juicebox rows are from the maintainer-published leaderboard. Pin, Noon, and Metaview do not appear on the current leaderboard and have not released comparable results on this benchmark.


What SCOUT does not measure

Candidate search quality is one dimension of a recruiting platform. SCOUT does not measure outreach deliverability or response rates, scheduling and interview coordination, ATS and CRM automation, diversity sourcing compliance, or time-to-hire outcomes. Those capabilities matter and vary across the platforms listed here. This benchmark measures one thing: when a system is asked to find a candidate, how often does it find the right one.

Ready to get started?

Create an account and start running searches with your free trial, or book a demo for your team.

Level up talent operations

Find candidates and keep every search moving.

Sourcing

Start with a free trial

Get your team running in Wrangle in as little as five minutes.

Sign up