AI People Search Benchmark 2026: Wrangle vs. Exa, Juicebox & More

Recruiting software vendors describe their products as "AI-powered," "intelligent," and "precision-driven" without publishing results on a shared query set that anyone can reproduce or contest. SCOUT (Sourcing and Candidate Understanding Test) is how Wrangle publishes its results: by running against established benchmarks with defined scoring protocols, shared query sets, and maintainer-published leaderboards, and reporting the numbers.
Two benchmarks are used here. The Exa People Search Benchmark tests hit rate and precision across 1,400 natural-language sourcing queries. PeopleSearchBench (maintained by LessieAI) scores on a composite of relevance precision, effective coverage, and information utility across a recruiting-specific query set. Where a competitor's result appears on the maintainer's published leaderboard, it is labeled as such. Where a competitor self-reported a number that is not on the maintainer's leaderboard, it is labeled vendor-reported. Numbers labeled vendor-reported should be treated as directional until independently confirmed.
Figure 01: Exa People Search Benchmark
The Exa People Search Benchmark runs 1,400 natural-language sourcing queries against a fixed candidate pool. Hit@1 measures whether the correct candidate appears as the first result. Hit@10 measures whether the correct candidate appears within the first ten results. Macro precision measures the share of returned candidates that are relevant across all 1,400 queries. Task completion measures the share of queries where the system returned a result. Qualified results measures the average number of relevant candidates returned per query slot (out of 15).
| Metric | Wrangle (Normal) | Wrangle (Turbo) | Delta | Exa (published) |
|---|---|---|---|---|
| Hit@1 | 94.07% (1,317/1,400) | 93.71% | +0.36 pp | 72.00% |
| Hit@10 | 98.07% (1,373/1,400) | 97.64% | +0.43 pp | Not published |
| Macro precision | 89.81% | 89.29% | +0.53 pp | Not published |
| Task completion | 100% | 100% | Not published | |
| Qualified results | 14.80 of 15 | 14.80 of 15 | Not published |
Wrangle Normal is the current evaluation run. Wrangle Turbo is the prior run for comparison. Exa's published Hit@1 of 72.00% comes from the maintainer-published leaderboard. Exa has not published Hit@10, macro precision, task completion, or qualified-results figures for its own system.
Competitor status on this benchmark:
| System | Exa leaderboard | Published Hit@1 | Published macro precision |
|---|---|---|---|
| Wrangle | Yes | 94.07% | 89.81% |
| Exa | Yes | 72.00% | Not published |
| Juicebox | No | Not published | 79% (vendor-reported) |
| Metaview Sourcing | No | Not published | 93.50% (vendor-reported) |
| Pin | No | Not published | Not published |
| Noon | No | Not published | Not published |
Juicebox reports a 79% precision figure from its own run of Exa's benchmark; this number does not appear on Exa's maintainer-published leaderboard. Metaview reports 93.50% precision from its own Exa-derived run and is not listed on the Exa leaderboard; it has not published corresponding Hit@1 or Hit@10 scores.
Figure 02: PeopleSearchBench (Recruiting subset)
PeopleSearchBench is maintained by LessieAI. Overall score is a composite of relevance precision, effective coverage, and information utility. The results below are from Wrangle's Turbo evaluation run.
| System | Overall | Relevance precision | Effective coverage | Information utility |
|---|---|---|---|---|
| Wrangle (Turbo) | 93.42 | 94.14 | 98.44 | 87.69 |
| Juicebox | 65.73 | 66.10 | 75.30 | 55.80 |
| Pin | No published result | |||
| Noon | No published result | |||
| Metaview Sourcing | No published result |
Wrangle and Juicebox rows are from the maintainer-published leaderboard. Pin, Noon, and Metaview do not appear on the current leaderboard and have not released comparable results on this benchmark.
What SCOUT does not measure
Candidate search quality is one dimension of a recruiting platform. SCOUT does not measure outreach deliverability or response rates, scheduling and interview coordination, ATS and CRM automation, diversity sourcing compliance, or time-to-hire outcomes. Those capabilities matter and vary across the platforms listed here. This benchmark measures one thing: when a system is asked to find a candidate, how often does it find the right one.


