Perplexity's WANDR: Benchmarking Research Agents That Actually Research

The search company's new benchmark exposes how badly existing agents handle real research tasks

Published: 23 July 2026 Category: AI Benchmarks / Research Agents Sources: MarkTechPost


The Benchmark

Perplexity AI released WANDR this week: an open benchmark for evaluating research agents on tasks that require genuine investigation. Not question-answering from a provided text. Not summarising a single source. Real research: formulating queries, evaluating sources, synthesising conflicting information, updating conclusions as new evidence emerges.

The results are not flattering for existing models.

The Results

Current frontier models — GPT-5.6, Claude Fable 5, Gemini 3.5 — score between 35-45% on WANDR tasks. That is better than random chance, but it is failing grade territory. The failures are revealing. Models struggle with source evaluation — distinguishing credible from dubious sources. They struggle with synthesis — combining information from multiple sources without contradiction. They struggle most with what WANDR calls "depth awareness": knowing when they have searched enough and when they need to keep going.

The benchmark is deliberately adversarial. Tasks include questions where the answer changed recently (requiring up-to-date information), questions where sources conflict (requiring judgment), and questions where the initial framing is wrong (requiring correction). These are the conditions of real research, and they are conditions that current AI agents handle poorly.

The Analysis

Perplexity has a commercial interest in research agents — the company's core product is an AI search engine. So WANDR is partly competitive positioning. But the benchmark is well-designed and the results are independently verifiable. The research agent space is genuinely underdeveloped relative to the hype.

The problem is structural. Research requires something that current architectures are not good at: iterative refinement. A human researcher does not formulate one query, read one page, and produce an answer. They search, read, search again with new terms, discard initial hypotheses, follow tangents, return to the main thread. Current AI agents can simulate some of this through tool use, but the simulation is shallow. They do not genuinely rethink. They do not experience the uncertainty that drives real research.

WANDR's scoring reflects this. The models that do best are not the largest or most expensive. They are the ones with the most sophisticated tool use — the ability to call search APIs, evaluate results, and reformulate queries based on what they find. Scale helps, but architecture and training matter more.

The Verdict

WANDR is a valuable contribution to AI evaluation. It exposes a real gap between what research agents promise and what they deliver. It gives developers a concrete target to improve against. And it suggests that the research agent space is still early — the winner may not be the current frontier model, but a future model trained specifically for iterative investigation.

Perplexity benefits from WANDR's existence regardless of how models score. If research agents improve, Perplexity's product improves. If they do not, Perplexity can differentiate on being the company that takes research seriously. It is a clever strategic move wrapped in genuine research contribution.

For users, WANDR is a reminder that AI research agents are promising but premature. They can help with research, but they cannot replace it. Not yet. The gap is measurable now, which means it can be closed. But closing it will require more than bigger models. It will require better tools, better training, and a better understanding of what research actually is.


Read next: DeepMind's Video World Models: Seeing Is Understanding