Which model should answer
a legal question?
Two open benchmarks, run end to end and published with their receipts. We are not publishing a leaderboard: plenty of those exist and ours would be the weakest of them. What follows is only what our own runs found that the published rankings do not say — what a model costs against what it is worth, which kind of legal question breaks which model, and what these benchmarks are structurally unable to see. Every figure traces to a run artifact.
Six things our run found
We ran DELTA 1.1.0 ourselves: 15 practitioner assignments across nine Dutch practice areas, 273 lawyer-written criteria, five models answering every task 2 times with no retrieval and no tools under DELTA’s own fixed system prompt, graded blind by two judges from different model families. We are not publishing a ranking. DELTA publishes one, over a far larger held-back set, and theirs is the one to read. These six are what our run can say that theirs does not.
Finding 1: Price and quality come apart, and DELTA has no cost axis at all
This is the finding that actually answers “which model should I use”, and it exists nowhere upstream: DELTA grades quality and never prices it. Across our five models the spread is 54× in cost per task against 28.4 points of substance. There is no order to read down the chart below, only a frontier to read across.
Answering cost only, at list prices on 2026-09-13; judging is excluded because it is a property of our harness, not of the model. The price axis is logarithmic — on a linear one, four of five models sit on the origin. Monocle NL is absent because a product endpoint has no per-call token accounting to place it with.
The ordering holds. The exchange rate does not.
On criteria met, price and quality track in rank order exactly — we searched every pair and found no inversion. What collapses is what the money buys. The last step up the ladder costs 20× per point of what the first step costs.
| Step up | Extra per task | Points bought | Cost per point |
|---|---|---|---|
| DeepSeek V4 Pro toGPT-6 Astra | +$0.071 | +12.8 | $0.006 |
| GPT-6 Astra toOpus 5 | +$0.36 | +13 | $0.027 |
| Opus 5 toFable 5.1 | +$0.31 | +2.6 | $0.12 |
On the strict measure, it does invert.
“Tasks fully passed” requires every criterion on a task, not most of them — the difference between mostly right and right. Measured that way, paying more buys less, twice:
- Opus 5
- $0.4426.7%
- Fable 5.1
- $0.7523.3%
Fable 5.1 costs 1.7× more than Opus 5 and scores below it.
- Grok 4.6
- $0.0146.7%
- GPT-6 Astra
- $0.0853.3%
GPT-6 Astra costs 6.1× more than Grok 4.6 and scores below it.
And one of these is just a rate card. Opus 5 emitted 527,993 output tokens against Fable 5.1’s 447,912 — more work, not less — and still cost 41% less, because its published rate is half. No quality measurement produced that gap and none would predict it.
Finding 2: The benchmark cannot see a fabricated citation
Every citation criterion in DELTA is a presence test, never an accuracy test. Does this article number appear, does this ECLI appear. An answer can cite the wrong article for a rule, or the wrong ECLI for a case, and collect the criterion anyway. 7 of our 9 independent reviewers found this on their own, without seeing each other’s work.
What that blindness is worth, in score: DeepSeek V4 Pro scores 51.4% of substance criteria while inventing Dutch law outright — not misremembering it, inventing it.
art. 29a MwCompetition lawNo art. 29a Mw exists. The lowered healthcare thresholds come from an AMvB under the art. 29 delegation.
art. 3:246 lid 5 BWInsolvencyThere is no derogation clause there. Lid 5 is the uitstel-van-betaling power, and the answer reaches the opposite of the law.
art. 3:97 lid 2 BWProperty lawQuoted statutory text that does not exist, about “een overeenkomst tot het verrichten van arbeid”.
The consequence binds every number on this page, and DELTA’s too. No ranking built on this benchmark can be read as a quality ordering at the bottom, because the rubric cannot tell recalling a case from making one up. Differences between systems are more trustworthy than the levels. It is also why the second benchmark on this page exists: it is the test that can see a fabricated citation, and it is further down.
Finding 3: Which task shape breaks which model
The spread within a model is wider than the spread between models — 76.3 points against 28.4. So the useful question is not which model, it is which kind of question. DELTA tags every citation criterion with the authority it points at, so we could split the tasks by whether that authority is an article or a decision without labelling anything by hand.
Blue is the mean score across all five models; amber is the mean gap between the best and worst model on a task of that shape. Task counts are in the labels — these are small groups, so read the direction, not the decimal. Shape is derived from DELTA’s own cited_authority fields in tasks.jsonl.
Where the answer is an article, every model scores well and the choice barely matters. Mean 81.8% across all five. This is the cheap-model case.
Where the answer is a decision, the mean falls to 53.6% and the gap between best and worst model widens to 38.2 points. This is where paying more is defensible.
The sharpest separator in the whole set is a pure case-law question. Tort law, duty to warn: which Hoge Raad decision leads on when a warning is an adequate measure. Fable 5.1 takes 100% of its criteria; DeepSeek V4 Pro takes 20.8%. That is 79.2 points on one question.
| Which case controls task | Fable 5.1 | Opus 5 | Grok 4.6 | GPT-6 Astra | DeepSeek V4 Pro |
|---|---|---|---|---|---|
Director liability, selective payments Insolvency | 54.2% | 54.2% | 58.3% | 66.7% | 45.8% |
Interpretation of notarial deeds Real estate | 57.7% | 57.7% | 34.6% | 34.6% | 26.9% |
Pledgee settlement authority Insolvency | 58.8% | 32.4% | 35.3% | 41.2% | 20.6% |
Animal liability and exoneration Tort law | 66.7% | 66.7% | 53.3% | 53.3% | 46.7% |
Inappropriate conduct at work Employment law | 77.5% | 77.5% | 57.5% | 55% | 37.5% |
Duty to warn (Jetblast) Tort law | 100% | 100% | 33.3% | 83.3% | 20.8% |
The six case-anchored tasks, hardest first. Substance criteria met, Anthropic judge, both runs averaged.
Finding 4: The public 15 tasks are easier than the held-back set
A warning about our own numbers before anyone borrows them. Fable 5.1 passed 30% of tasks for us, against DELTA’s published 15% on the full set — same protocol, same fixed prompt. The 15 public tasks are the ones DELTA chose to publish; 200+ are held back.
Read this as directional, not exact. Our judge pair is not DELTA's, so part of the gap is the grading and part is the task set. The point is the warning, and it applies to us first: any number taken off a public task set and read as a leaderboard number overstates by roughly double. Ours included.
Finding 5: Judge reliability varies by practice area
Two judges from different model families graded every answer blind. Where they disagreed we put the call to independent review: 122 criterion-level disagreements, split 75/47 in the Anthropic judge’s favour overall. But the split inverts in 2 of the nine areas.
Net calls per practice area. Bars to the right are disagreements the reviewers gave to the Anthropic judge, bars to the left to the xAI judge. The reviewers are not lawyers, and none of this was carried back into any score on this page.
Cross-family judging is necessary, not sufficient. Picking one judge because it looks more faithful overall would have been wrong in property law and family law. So every figure on this page is the both-agree number: a criterion counts as met only where both judges passed it. That is the conservative floor, not an adjudicated score.
Finding 6: Running twice catches retrieval variance, not model variance
We ran every system through the whole set 2 times, expecting to average out model noise. There was none to average: the raw models gave the same answer in both passes. What moved was the system that retrieves.
KelderluikJetblastMonocle NL returned a different leading authority on the duty-to-warn question in each pass. No raw model moved between passes at all, so a second pass buys nothing against model noise and everything against a retrieval layer landing somewhere else.
And one thing we did not learn: whether retrieval helps
Monocle NL appears in the full grid below and nowhere else on this page. Its row is a product measurement, not a controlled comparison: it is the only system that could use retrieval and the only one that could not be given DELTA’s fixed system prompt, so nothing separates the effect of retrieval from the effect of a different prompt. Its shape is suggestive — strong on statute-anchored questions, weak on case-law ones — and that is as far as it goes. It is labelled everywhere and never ranked against the models.
Two criterion defects, filed upstream
DELTA invites criterion disputes as contributions, so we checked each one against the statute and filed the two that survived:
S-005Property lawCredits art. 3:97 lid 2 BW with enabling advance delivery. That is lid 1; lid 2 ranks competing advance deliveries.
Models that state the law correctly are failed for it; an answer that got it wrong would pass. 3 of the area's 19 judge disagreements are this one criterion.
delta#1S-009Family lawConditional on a “dynamische uitleg” argument that no candidate makes — zero occurrences across all five models.
The condition never fires, so the criterion has no determinate value. One judge failed all five answers, the other passed three.
delta#2
A third one we withdrew, because it was wrong
Our reviewers reported that art. 1:250 BW is undivided, so the criterion’s demand for “art. 1:250 lid 1 BW” pointed at a lid that does not exist. It does exist. Art. 1:250 BW has leden, and lid 1 is the bijzondere-curator provision on a conflict of interest, which is exactly what the criterion asks for. There is also a lid 2 on the raad voor de kinderbescherming (art. 242a).
Nine independent reviewers agreeing is not verification. This one was caught by reading the statute before filing, which is the check that should have run first. We are leaving this on the page rather than quietly dropping it: finding 2 argues that a benchmark should be able to tell a real citation from an invented one, and a page making that argument does not get to hide the moment its own reviewers failed the same test.
Source: Burgerlijk Wetboek Boek 1, BWBR0002656, text in force at the task’s law_as_of of 2026-08-30.
Substance criteria met, Anthropic judge, both runs averaged; totals are both-agree. This grid is not a ranking — 15 public tasks cannot order five models, and finding 2 above says the bottom of any such order is not a quality order anyway. For a general ranking, DELTA publishes one over 200+ held-back tasks. Monocle’s zeroes are not wrong answers: on those tasks it asked a clarifying question instead of answering, which scores zero under a strict one-shot protocol. DELTA 1.1.0 is licensed CC BY 4.0. DELTA is curated by Legal Benchmarks with Zeno.Law as Founding Research Partner. Every figure above is generated from the run artifacts; none is typed by hand.
The gap DELTA cannot score
Finding 2 above is that DELTA’s citation criteria are presence tests: they ask whether an ECLI appears, never whether it is real. That is the single most important signal for a practising litigator, and no score on this page can see it. So we built the test that can, and pointed it at the one thing DELTA’s rubric leaves uncovered. Half the items are true bibliographic statements about real legal documents; the other half carry exactly one controlled mutation. The model judges true or false from memory alone, across 227 jurisdictions rather than one.
This measures a model’s unaided memory, not the Legalcomplex product. The number below comes from asking Grok 4.1 Fast (non-reasoning) true/false questions about citations with no retrieval, no search, and no source documents — the exact condition Monocle is built to never operate in. Monocle looks up official sources, pins verified provisions, and judges its own citations before answering. The two numbers are not comparable, and a skim that concludes “the product scores 5.8%” is reading this page wrong.
Grok 4.1 Fast (non-reasoning) correctly confirmed a real legal citation exists, from memory, 5.8% of the time — against a 50% chance baseline.
Grok 4.1 Fast (non-reasoning) · xai · run 2026-06-20 · full sweep (227/227 packs)
Source: labs/benchmarks/s3-framework/SWEEP-REPORT.md @ 53cba04
Methodology note: 16.3% of the true items (221 of 1,358) carry a year-precision source date rendered as the full January 1, so a model that doubts a suspicious Jan-1 datestamp is graded wrong even where that skepticism is defensible. Excluding those items entirely, Grok 4.1 Fast (non-reasoning) reads 6.9% on real citations. The headline keeps the full set.
Skepticism without knowledge
The model is nearly perfect at rejecting a fabricated citation (98.1%) and nearly incapable of confirming a real one (5.8%). That gap is the finding: it doubts everything, including documents that genuinely exist, rather than actually knowing which ones do.
How the test works
Every jurisdiction pack has 12 bibliographic statements about real, recently collected official legal documents. Half are verbatim true. Half carry exactly one controlled mutation.
The model gets one true/false call per item, temperature 0, no search and no source documents. Grading is exact-match on the first true/false token in the response. Chance is 50%.
2,716 items across every jurisdiction in the world-law catalog, source-grounded and generated in-house.
Accuracy by mutation class
These are FALSE items only, how often the mutation was caught. A wrong-source pairing (a real citation attached to the wrong document) is detectable on internal consistency, without knowing the source.
Accuracy by statement shape
Every jurisdiction, not a curated few
A handful of packs score well above chance. Inspection shows that’s mostly FALSE mutations caught on internal-consistency grounds, not genuine knowledge of the documents — most of the world sits inside the chance band.
Number of jurisdiction packs (12 items each, one exception at 10 or 6) landing in each accuracy band. Chance is 50%.
| Jurisdiction | Items | Accuracy |
|---|---|---|
| ArubaAW | 12 | 75% |
| BoliviaBO | 12 | 75% |
| IndiaIN | 12 | 75% |
| NorwayNO | 12 | 75% |
| TuvaluTV | 12 | 75% |
| Central African RepublicCF | 12 | 67% |
| Costa RicaCR | 12 | 67% |
| CyprusCY | 12 | 67% |
| EgyptEG | 12 | 67% |
| GabonGA | 12 | 67% |
| Hong KongHK | 12 | 67% |
| LatviaLV | 12 | 67% |
| PolandPL | 12 | 67% |
| SloveniaSI | 12 | 67% |
| SpainES | 12 | 67% |
| AlgeriaDZ | 12 | 58% |
| BarbadosBB | 12 | 58% |
| BelgiumBE | 12 | 58% |
| Burkina FasoBF | 12 | 58% |
| BurundiBI | 12 | 58% |
| ChadTD | 12 | 58% |
| Cook IslandsCK | 12 | 58% |
| CzechiaCZ | 12 | 58% |
| Democratic Republic of the CongoCD | 12 | 58% |
| Dominican RepublicDO | 12 | 58% |
| European UnionEU | 12 | 58% |
| KiribatiKI | 12 | 58% |
| KosovoXK | 12 | 58% |
| MadagascarMG | 12 | 58% |
| MoldovaMD | 12 | 58% |
| NamibiaNA | 12 | 58% |
| NicaraguaNI | 12 | 58% |
| NigerNE | 12 | 58% |
| NigeriaNG | 12 | 58% |
| ParaguayPY | 12 | 58% |
| PeruPE | 12 | 58% |
| Republic of the CongoCG | 12 | 58% |
| SingaporeSG | 12 | 58% |
| South AfricaZA | 12 | 58% |
| Sri LankaLK | 12 | 58% |
| TaiwanTW | 12 | 58% |
| TogoTG | 12 | 58% |
| United StatesUS | 12 | 58% |
| UruguayUY | 12 | 58% |
| Åland IslandsAX | 12 | 50% |
| AlbaniaAL | 12 | 50% |
| American SamoaAS | 6 | 50% |
| AndorraAD | 12 | 50% |
| AngolaAO | 12 | 50% |
| AnguillaAI | 12 | 50% |
| Antigua and BarbudaAG | 12 | 50% |
| ArmeniaAM | 12 | 50% |
| AustraliaAU | 12 | 50% |
| AustriaAT | 12 | 50% |
| AzerbaijanAZ | 12 | 50% |
| BahamasBS | 12 | 50% |
| BahrainBH | 12 | 50% |
| BangladeshBD | 12 | 50% |
| BelizeBZ | 12 | 50% |
| BermudaBM | 12 | 50% |
| BhutanBT | 12 | 50% |
| BotswanaBW | 12 | 50% |
| BrazilBR | 12 | 50% |
| British Virgin IslandsVG | 12 | 50% |
| BruneiBN | 12 | 50% |
| BulgariaBG | 12 | 50% |
| CambodiaKH | 12 | 50% |
| CameroonCM | 12 | 50% |
| CanadaCA | 12 | 50% |
| Cape VerdeCV | 12 | 50% |
| Caribbean NetherlandsBQ | 12 | 50% |
| Cayman IslandsKY | 12 | 50% |
| ChileCL | 12 | 50% |
| ChinaCN | 12 | 50% |
| ColombiaCO | 12 | 50% |
| ComorosKM | 12 | 50% |
| Côte d'IvoireCI | 12 | 50% |
| Council of EuropeCOE | 12 | 50% |
| CroatiaHR | 12 | 50% |
| CubaCU | 12 | 50% |
| CuraçaoCW | 12 | 50% |
| DenmarkDK | 12 | 50% |
| DjiboutiDJ | 12 | 50% |
| DominicaDM | 12 | 50% |
| EcuadorEC | 12 | 50% |
| El SalvadorSV | 12 | 50% |
| Equatorial GuineaGQ | 12 | 50% |
| EritreaER | 12 | 50% |
| EstoniaEE | 12 | 50% |
| EswatiniSZ | 12 | 50% |
| EthiopiaET | 12 | 50% |
| Falkland IslandsFK | 12 | 50% |
| Faroe IslandsFO | 12 | 50% |
| FijiFJ | 12 | 50% |
| FinlandFI | 12 | 50% |
| French PolynesiaPF | 12 | 50% |
| GambiaGM | 12 | 50% |
| GeorgiaGE | 12 | 50% |
| GermanyDE | 12 | 50% |
| GhanaGH | 12 | 50% |
| GibraltarGI | 12 | 50% |
| GreeceGR | 12 | 50% |
| GreenlandGL | 12 | 50% |
| GrenadaGD | 12 | 50% |
| GuamGU | 12 | 50% |
| GuatemalaGT | 12 | 50% |
| GuernseyGG | 12 | 50% |
| GuineaGN | 12 | 50% |
| Guinea-BissauGW | 12 | 50% |
| GuyanaGY | 12 | 50% |
| HaitiHT | 12 | 50% |
| HondurasHN | 12 | 50% |
| HungaryHU | 12 | 50% |
| IcelandIS | 12 | 50% |
| IndonesiaID | 12 | 50% |
| InternationalINTL | 12 | 50% |
| IranIR | 12 | 50% |
| IraqIQ | 12 | 50% |
| IrelandIE | 12 | 50% |
| Isle of ManIM | 12 | 50% |
| IsraelIL | 12 | 50% |
| ItalyIT | 12 | 50% |
| JamaicaJM | 12 | 50% |
| JerseyJE | 12 | 50% |
| JordanJO | 12 | 50% |
| KazakhstanKZ | 12 | 50% |
| KenyaKE | 12 | 50% |
| KyrgyzstanKG | 12 | 50% |
| LaosLA | 12 | 50% |
| LebanonLB | 12 | 50% |
| LesothoLS | 12 | 50% |
| LiberiaLR | 12 | 50% |
| LibyaLY | 12 | 50% |
| LithuaniaLT | 12 | 50% |
| LuxembourgLU | 12 | 50% |
| MacaoMO | 12 | 50% |
| MalawiMW | 12 | 50% |
| MaldivesMV | 12 | 50% |
| MaliML | 12 | 50% |
| MaltaMT | 12 | 50% |
| Marshall IslandsMH | 12 | 50% |
| MauritaniaMR | 12 | 50% |
| MicronesiaFM | 10 | 50% |
| MonacoMC | 12 | 50% |
| MongoliaMN | 12 | 50% |
| MontenegroME | 12 | 50% |
| MontserratMS | 12 | 50% |
| MoroccoMA | 12 | 50% |
| MozambiqueMZ | 12 | 50% |
| MyanmarMM | 12 | 50% |
| NauruNR | 12 | 50% |
| NetherlandsNL | 12 | 50% |
| New ZealandNZ | 12 | 50% |
| NiueNU | 12 | 50% |
| North MacedoniaMK | 12 | 50% |
| OmanOM | 12 | 50% |
| PakistanPK | 12 | 50% |
| PalauPW | 12 | 50% |
| PalestinePS | 12 | 50% |
| PanamaPA | 12 | 50% |
| Papua New GuineaPG | 12 | 50% |
| PhilippinesPH | 12 | 50% |
| Pitcairn IslandsPN | 12 | 50% |
| PortugalPT | 12 | 50% |
| Puerto RicoPR | 12 | 50% |
| QatarQA | 12 | 50% |
| RomaniaRO | 12 | 50% |
| RussiaRU | 12 | 50% |
| RwandaRW | 12 | 50% |
| Saint BarthélemyBL | 12 | 50% |
| Saint HelenaSH | 12 | 50% |
| Saint Kitts and NevisKN | 12 | 50% |
| Saint LuciaLC | 12 | 50% |
| Saint MartinMF | 12 | 50% |
| Saint Pierre and MiquelonPM | 12 | 50% |
| Saint Vincent and the GrenadinesVC | 12 | 50% |
| SamoaWS | 12 | 50% |
| San MarinoSM | 12 | 50% |
| Saudi ArabiaSA | 12 | 50% |
| SenegalSN | 12 | 50% |
| SerbiaRS | 12 | 50% |
| SeychellesSC | 12 | 50% |
| Sierra LeoneSL | 12 | 50% |
| Sint MaartenSX | 12 | 50% |
| SlovakiaSK | 12 | 50% |
| Solomon IslandsSB | 12 | 50% |
| SomaliaSO | 12 | 50% |
| South KoreaKR | 12 | 50% |
| South SudanSS | 12 | 50% |
| SudanSD | 12 | 50% |
| SurinameSR | 12 | 50% |
| SwedenSE | 12 | 50% |
| SwitzerlandCH | 12 | 50% |
| SyriaSY | 12 | 50% |
| TajikistanTJ | 12 | 50% |
| TanzaniaTZ | 12 | 50% |
| ThailandTH | 12 | 50% |
| Timor-LesteTL | 12 | 50% |
| TongaTO | 12 | 50% |
| Trinidad and TobagoTT | 12 | 50% |
| TunisiaTN | 12 | 50% |
| TurkeyTR | 12 | 50% |
| TurkmenistanTM | 12 | 50% |
| Turks and Caicos IslandsTC | 12 | 50% |
| U.S. Virgin IslandsVI | 12 | 50% |
| UgandaUG | 12 | 50% |
| United Arab EmiratesAE | 12 | 50% |
| United KingdomUK | 12 | 50% |
| United NationsUN | 12 | 50% |
| UzbekistanUZ | 12 | 50% |
| VanuatuVU | 12 | 50% |
| Vatican CityVA | 12 | 50% |
| VietnamVN | 12 | 50% |
| Wallis and FutunaWF | 12 | 50% |
| YemenYE | 12 | 50% |
| ZambiaZM | 12 | 50% |
| ZimbabweZW | 12 | 50% |
| BeninBJ | 12 | 42% |
| Bosnia & HerzegovinaBA | 12 | 42% |
| FranceFR | 12 | 42% |
| JapanJP | 12 | 42% |
| LiechtensteinLI | 12 | 42% |
| MauritiusMU | 12 | 42% |
| MexicoMX | 12 | 42% |
| UkraineUA | 12 | 42% |
| VenezuelaVE | 12 | 42% |
| ArgentinaAR | 12 | 33% |
Across models
| Model | Run date | Items | Aggregate | TRUE items | FALSE items | Est. cost |
|---|---|---|---|---|---|---|
| Claude Opus 4.8 | 2026-07-19 | 2,716 | 52.6% | 7.4% | 97.6% | $5.23 |
| Claude Haiku 4.5 | 2026-07-19 | 2,716 | 52.3% | 6.6% | 98% | $0.90 |
| Grok 4.1 Fast (non-reasoning) | 2026-06-20 | 2,716 | 51.9% | 5.8% | 98.1% | $0.09 |
| Claude Sonnet 5 | 2026-07-19 | 2,716 | 50.5% | 1.5% | 99.5% | $2.26 |
Est. cost is what the full run cost to answer, at list prices on the run date. Every run above covers the same 2,716 items, so the column compares directly. Read it next to accuracy, never on its own: the cheapest run here is within a point of the most expensive.
How Monocle answers instead
Retrieves from official sources — government legislation portals, court registries, and gazette systems, not model memory.
Pins verified provisions for the statute-level questions that matter most, each citation deep-linked to the official text, checked against contradiction across repeated judge runs before it ships.
Judges its own citations with a calibrated NLI model before showing them, and only marks a citation verified when the judge finds full support with no flags.
Abstains honestly when the facts don’t support a ranking — a tie stays a tie, a gap stays a gap, and a citation the judge can’t support is never presented as confirmed.
Other receipts
Support-accuracy of the judge that verifies citations inside Monocle answers (a different, retrieval-grounded task from the S3 memory test above). Single calibration run at this n; treat deltas as noise until re-run.
gemini-flash-lite · n=24 · 2026-07-02 · labs/benchmarks/citation-nli/
What this does and doesn’t show
- This is a memory-only condition: no retrieval, no web search, no source documents. It is not a measurement of any retrieval-augmented system, ours or anyone else’s.
- Packs are generated in-house from source-grounded facts. The pack set itself is not published (it would stop being a clean test the moment a model could train on it) — the sweep report and this page are the receipts instead.
- Grading is exact-match true/false on a single response. It doesn’t credit partial reasoning or hedged answers, by design — the question is binary, so the grade is too.
- Where a pack scores well above chance, that is usually a mutation caught on internal-consistency grounds, not genuine knowledge of the underlying document — see the per-jurisdiction section above.
- Every number on this page traces to a dated run artifact. None are hand-typed estimates.
These receipts are one line on a longer plan. The roadmap shows what shipped, what is building and what is still a question.
Model: Grok 4.1 Fast (non-reasoning) (xai) · Run date: 2026-06-20 · Items: 2,716 (full sweep)
536,107 tokens in / 82,961 tokens out · est. $0.09 · 10 min wall time