Does the search reach the primary source, or only the reporting about it?
Our rule is: no filtering of sources by origin, no mainstream bonus, no fact-checker shortcut. That was a claim about ourselves. This page is the test, with the raw data.
Measured 21 August 2026 · raw per-claim results (JSON)
Why this is the measurement that matters most. Retrieval is the most powerful layer in the whole product and the least visible. What the search does not return does not exist, for any model, any cross-check, any quote verification. A tool can have a perfect evidence chain and still be blind, if the funnel in front of it is narrow.
Method
24 claims, and for each one the root is named in advance, the document that settles the matter. Four kinds:
- Official root: court rulings, central banks, statistics offices, legislation.
- Scientific root: meta-analyses, trials, working papers.
- Non-English root: German, Swiss, Turkish, Spanish, Thai, Russian.
- Heterodox but evidenced positions, minimum wage without employment effect, lab-leak unresolved, system costs of solar understated, null results on screen time, criticism of the WHO pandemic treaty, disputed mask efficacy.
That last group is the actual test of "no mainstream bonus". Without it, the claim is unfalsifiable.
Configurations: the live setting (10 results), the same test at 6, wider (15), and keyword instead of neural search (15).
Correction, 26 August 2026. This page originally called the 6-result
configuration "the live setting". That was wrong: production requests 10 results
(lib/spectrum.js, exaSearch(claim, { n: 10 })), and that is the path the
X bot uses. The table below now shows the measured production setting. The 6-result column is
kept for comparison. Neural search drifts, so every figure carries its measurement date:
the same 6-result test gave 63 % primary sources on 21 August and 58 % on 26 August. The
structural findings (root reach, fact-checker share, primary source at position 1) were stable
across both runs.
What comes back, production setting (10 results, measured 26 Aug 2026)
| Source class | share | |
|---|---|---|
| Independent / other | 33 % | |
| Primary, official courts, ministries, central banks | 23 % | |
| Primary, scientific journals, meta-analyses, NBER | 22 % | |
| Legacy media | 15 % | |
| Primary, organisation IPCC, NASA, HRW | 4 % | |
| Primary, datasets | 3 % | |
| Fact-checkers | 0.4 % |
52 % of everything returned is a primary source. Legacy media is 15 %. Fact-checkers are 0.4 %, a single hit across all 240 results.
And the number that says most about ranking: in 19 of 24 claims a primary source held first place.
The honest weak point is the largest single class: 33 % "independent / other", meaning trade press, NGOs, specialist sites and blogs. That is the unavoidable consequence of having no origin filter. The answer to it is not the search layer but the evidence chain: an unknown site can enter the result list, but it cannot carry an unverified statement, because every quote is machine-checked against the source text.
Reach to the root
| Configuration | at least one primary source | primary at position 1 |
|---|---|---|
| production (10 results, 26 Aug) | 23/24, 96 % | 19 |
| 6 results (26 Aug) | 23/24, 96 % | 19 |
| 6 results (21 Aug) | 23/24, 96 % | 20 |
| wider (15 results) | 24/24, 100 % | 19 |
| keyword instead of neural (15) | 22/24, 92 % | 14 |
Three usable conclusions:
- Root reach does not depend on how many results you ask for. 96 % at six results, 96 % at ten, measured five days apart. What changes is dilution: the primary-source share falls from 58 % at six to 52 % at ten and 50 % at fifteen, because the extra ranks pull in secondary material.
- Keyword search is worse on every axis. Do not switch.
- More results ≠ deeper root. The root is in the top ranks or it is not there at all.
Non-English roots: 6 of 6
| Claim | First primary source | local-language hits |
|---|---|---|
| German Constitutional Court, climate act | bundesverfassungsgericht.de, rank 1 | 14/15 |
| Swiss CO2 referendum | admin.ch, rank 1 | 14/15 |
| Turkish central bank rates | tcmb.gov.tr, rank 1 | 15/15 |
| Mexican Supreme Court, abortion | internet2.scjn.gob.mx, rank 1 | 1/15 |
| Thai Constitutional Court | constitutionalcourt.or.th, rank 1 | 6/15 |
| Russian consumer prices | normativ.kontur.ru, rank 3 | 13/15 |
In every case the responsible institution itself came back, five times at rank 1. No English-language bias is measurable here.
Heterodox positions: does the search reach real dissent?
This is the group that decides whether "no mainstream bonus" means anything.
| Position | primary sources | top hits |
|---|---|---|
| Minimum wage does not reduce employment | 10/15 | nber.org ×3, epi.org |
| Lab-leak hypothesis unresolved | 12/15 | science.org, who.int |
| No link screen time / depression | 11/15 | jamanetwork.com, pubmed |
| Cloth mask efficacy disputed | 11/15 | cidrap.umn.edu, cdc.gov |
| LCOE understates system cost of solar | 6/15 | wiley.com, repec.org |
| WHO pandemic treaty criticism | 3/15 | reuters.com, rollcall.com |
Five of six get predominantly primary literature, for the minimum wage, three NBER papers, which is the Card/Krueger line itself rather than a summary of it. None of them is replaced by a fact-checker's verdict.
The one place where it does not reach
Criticism of the WHO pandemic treaty: only 3 of 15 primary, position 1 a legacy outlet. The search returns reporting about the position rather than the position itself. This is typical of political criticism, which often has no primary document, only voices. That is not a filter, but it is a limit, and it is where a reader should be most careful with our output.
Our first evaluation was wrong, and it made us look worse
The first pass reported "root reached: 71 %" with seven failures. On individual inspection, not one was a failure of the search:
| "NASA dataset missed" | sealevel.nasa.gov, earthdata.nasa.gov were returned |
| "Mexican Supreme Court missed" | internet2.scjn.gob.mx was returned |
| "Card/Krueger missed" | nber.org, three times |
Our classifier did not know .gob.mx, .or.th, Russian authorities or nber.org, and a generic .gov rule fired before the more specific one. The instrument was broken. We publish this because a study that only reports the flattering correction is not a study, and because the rule we took from it is the useful part: before any rate, inspect five individual cases by hand.
What this does not show
- 24 claims. Enough for an order of magnitude, not for a percentage point.
- Our own classification. Which host counts as "primary" is our judgement. The raw data lists every host, so you can reclassify and recount.
- Nothing about correctness. Reaching a primary source is not the same as reading it correctly. That is measured separately, in the quote study.
- Not the whole funnel. After retrieval come reading (19 % of pages are unreadable to us) and judging. Both are measured elsewhere and both filter too.
What changed because of it
Nothing, and that is the result. We tested widening the search and switching to keyword mode; both are worse. The setting that ships is the best of the three. The doctrine is now measured rather than asserted, and the one limit we found is named above instead of buried.
