22–23 SeptStarts in 2 days, 17 hours, 43 minutes and 38 secondsBeacon Collective: Illuminate 2026Event details
Benchmarks

Which model should answer a legal question?

Two open benchmarks, run end to end and published with their receipts. We are not publishing a leaderboard: plenty of those exist and ours would be the weakest of them. What follows is only what our own runs found that the published rankings do not say — what a model costs against what it is worth, which kind of legal question breaks which model, and what these benchmarks are structurally unable to see. Every figure traces to a run artifact.

Benchmark 1 of 2 · Dutch legal research

Six things our run found

We ran DELTA 1.1.0 ourselves: 15 practitioner assignments across nine Dutch practice areas, 273 lawyer-written criteria, five models answering every task 2 times with no retrieval and no tools under DELTA’s own fixed system prompt, graded blind by two judges from different model families. We are not publishing a ranking. DELTA publishes one, over a far larger held-back set, and theirs is the one to read. These six are what our run can say that theirs does not.

Finding 1: Price and quality come apart, and DELTA has no cost axis at all

This is the finding that actually answers “which model should I use”, and it exists nowhere upstream: DELTA grades quality and never prices it. Across our five models the spread is 54× in cost per task against 28.4 points of substance. There is no order to read down the chart below, only a frontier to read across.

Answering cost only, at list prices on 2026-09-13; judging is excluded because it is a property of our harness, not of the model. The price axis is logarithmic — on a linear one, four of five models sit on the origin. Monocle NL is absent because a product endpoint has no per-call token accounting to place it with.

The ordering holds. The exchange rate does not.

On criteria met, price and quality track in rank order exactly — we searched every pair and found no inversion. What collapses is what the money buys. The last step up the ladder costs 20× per point of what the first step costs.

Marginal cost of one substance point, stepping up the price ladder
Step upExtra per taskPoints boughtCost per point
DeepSeek V4 Pro toGPT-6 Astra+$0.071+12.8$0.006
GPT-6 Astra toOpus 5+$0.36+13$0.027
Opus 5 toFable 5.1+$0.31+2.6$0.12

On the strict measure, it does invert.

“Tasks fully passed” requires every criterion on a task, not most of them — the difference between mostly right and right. Measured that way, paying more buys less, twice:

Tasks fully passed
Opus 5
$0.4426.7%
Fable 5.1
$0.7523.3%

Fable 5.1 costs 1.7× more than Opus 5 and scores below it.

Tasks fully passed
Grok 4.6
$0.0146.7%
GPT-6 Astra
$0.0853.3%

GPT-6 Astra costs 6.1× more than Grok 4.6 and scores below it.

And one of these is just a rate card. Opus 5 emitted 527,993 output tokens against Fable 5.1’s 447,912 — more work, not less — and still cost 41% less, because its published rate is half. No quality measurement produced that gap and none would predict it.

Finding 2: The benchmark cannot see a fabricated citation

Every citation criterion in DELTA is a presence test, never an accuracy test. Does this article number appear, does this ECLI appear. An answer can cite the wrong article for a rule, or the wrong ECLI for a case, and collect the criterion anyway. 7 of our 9 independent reviewers found this on their own, without seeing each other’s work.

What that blindness is worth, in score: DeepSeek V4 Pro scores 51.4% of substance criteria while inventing Dutch law outright — not misremembering it, inventing it.

  • art. 29a MwCompetition law

    No art. 29a Mw exists. The lowered healthcare thresholds come from an AMvB under the art. 29 delegation.

  • art. 3:246 lid 5 BWInsolvency

    There is no derogation clause there. Lid 5 is the uitstel-van-betaling power, and the answer reaches the opposite of the law.

  • art. 3:97 lid 2 BWProperty law

    Quoted statutory text that does not exist, about “een overeenkomst tot het verrichten van arbeid”.

The consequence binds every number on this page, and DELTA’s too. No ranking built on this benchmark can be read as a quality ordering at the bottom, because the rubric cannot tell recalling a case from making one up. Differences between systems are more trustworthy than the levels. It is also why the second benchmark on this page exists: it is the test that can see a fabricated citation, and it is further down.

Finding 3: Which task shape breaks which model

The spread within a model is wider than the spread between models — 76.3 points against 28.4. So the useful question is not which model, it is which kind of question. DELTA tags every citation criterion with the authority it points at, so we could split the tasks by whether that authority is an article or a decision without labelling anything by hand.

Blue is the mean score across all five models; amber is the mean gap between the best and worst model on a task of that shape. Task counts are in the labels — these are small groups, so read the direction, not the decimal. Shape is derived from DELTA’s own cited_authority fields in tasks.jsonl.

Statute-anchored: use anything

Where the answer is an article, every model scores well and the choice barely matters. Mean 81.8% across all five. This is the cheap-model case.

Which case controls: the choice decides it

Where the answer is a decision, the mean falls to 53.6% and the gap between best and worst model widens to 38.2 points. This is where paying more is defensible.

The sharpest separator in the whole set is a pure case-law question. Tort law, duty to warn: which Hoge Raad decision leads on when a warning is an adequate measure. Fable 5.1 takes 100% of its criteria; DeepSeek V4 Pro takes 20.8%. That is 79.2 points on one question.

Which case controls tasks, substance criteria met per model
Which case controls taskFable 5.1Opus 5Grok 4.6GPT-6 AstraDeepSeek V4 Pro
Director liability, selective payments
Insolvency
54.2%54.2%58.3%66.7%45.8%
Interpretation of notarial deeds
Real estate
57.7%57.7%34.6%34.6%26.9%
Pledgee settlement authority
Insolvency
58.8%32.4%35.3%41.2%20.6%
Animal liability and exoneration
Tort law
66.7%66.7%53.3%53.3%46.7%
Inappropriate conduct at work
Employment law
77.5%77.5%57.5%55%37.5%
Duty to warn (Jetblast)
Tort law
100%100%33.3%83.3%20.8%

The six case-anchored tasks, hardest first. Substance criteria met, Anthropic judge, both runs averaged.

Finding 4: The public 15 tasks are easier than the held-back set

A warning about our own numbers before anyone borrows them. Fable 5.1 passed 30% of tasks for us, against DELTA’s published 15% on the full set — same protocol, same fixed prompt. The 15 public tasks are the ones DELTA chose to publish; 200+ are held back.

30%
Ours, on the 15 public tasks
15%
DELTA’s published figure, full set

Read this as directional, not exact. Our judge pair is not DELTA's, so part of the gap is the grading and part is the task set. The point is the warning, and it applies to us first: any number taken off a public task set and read as a leaderboard number overstates by roughly double. Ours included.

Finding 5: Judge reliability varies by practice area

Two judges from different model families graded every answer blind. Where they disagreed we put the call to independent review: 122 criterion-level disagreements, split 75/47 in the Anthropic judge’s favour overall. But the split inverts in 2 of the nine areas.

Net calls per practice area. Bars to the right are disagreements the reviewers gave to the Anthropic judge, bars to the left to the xAI judge. The reviewers are not lawyers, and none of this was carried back into any score on this page.

Cross-family judging is necessary, not sufficient. Picking one judge because it looks more faithful overall would have been wrong in property law and family law. So every figure on this page is the both-agree number: a criterion counts as met only where both judges passed it. That is the conservative floor, not an adjudicated score.

Finding 6: Running twice catches retrieval variance, not model variance

We ran every system through the whole set 2 times, expecting to average out model noise. There was none to average: the raw models gave the same answer in both passes. What moved was the system that retrieves.

Same question, same day, two passes
KelderluikJetblast

Monocle NL returned a different leading authority on the duty-to-warn question in each pass. No raw model moved between passes at all, so a second pass buys nothing against model noise and everything against a retrieval layer landing somewhere else.

And one thing we did not learn: whether retrieval helps

Monocle NL appears in the full grid below and nowhere else on this page. Its row is a product measurement, not a controlled comparison: it is the only system that could use retrieval and the only one that could not be given DELTA’s fixed system prompt, so nothing separates the effect of retrieval from the effect of a different prompt. Its shape is suggestive — strong on statute-anchored questions, weak on case-law ones — and that is as far as it goes. It is labelled everywhere and never ranked against the models.

Two criterion defects, filed upstream

DELTA invites criterion disputes as contributions, so we checked each one against the statute and filed the two that survived:

  • S-005Property law

    Credits art. 3:97 lid 2 BW with enabling advance delivery. That is lid 1; lid 2 ranks competing advance deliveries.

    Models that state the law correctly are failed for it; an answer that got it wrong would pass. 3 of the area's 19 judge disagreements are this one criterion.

    delta#1
  • S-009Family law

    Conditional on a “dynamische uitleg” argument that no candidate makes — zero occurrences across all five models.

    The condition never fires, so the criterion has no determinate value. One judge failed all five answers, the other passed three.

    delta#2

A third one we withdrew, because it was wrong

Our reviewers reported that art. 1:250 BW is undivided, so the criterion’s demand for “art. 1:250 lid 1 BW” pointed at a lid that does not exist. It does exist. Art. 1:250 BW has leden, and lid 1 is the bijzondere-curator provision on a conflict of interest, which is exactly what the criterion asks for. There is also a lid 2 on the raad voor de kinderbescherming (art. 242a).

Nine independent reviewers agreeing is not verification. This one was caught by reading the statute before filing, which is the check that should have run first. We are leaving this on the page rather than quietly dropping it: finding 2 argues that a benchmark should be able to tell a real citation from an invented one, and a page making that argument does not get to hide the moment its own reviewers failed the same test.

Source: Burgerlijk Wetboek Boek 1, BWBR0002656, text in force at the task’s law_as_of of 2026-08-30.

Substance criteria met, Anthropic judge, both runs averaged; totals are both-agree. This grid is not a ranking15 public tasks cannot order five models, and finding 2 above says the bottom of any such order is not a quality order anyway. For a general ranking, DELTA publishes one over 200+ held-back tasks. Monocle’s zeroes are not wrong answers: on those tasks it asked a clarifying question instead of answering, which scores zero under a strict one-shot protocol. DELTA 1.1.0 is licensed CC BY 4.0. DELTA is curated by Legal Benchmarks with Zeno.Law as Founding Research Partner. Every figure above is generated from the run artifacts; none is typed by hand.

Benchmark 2 of 2 · Citation verification

The gap DELTA cannot score

Finding 2 above is that DELTA’s citation criteria are presence tests: they ask whether an ECLI appears, never whether it is real. That is the single most important signal for a practising litigator, and no score on this page can see it. So we built the test that can, and pointed it at the one thing DELTA’s rubric leaves uncovered. Half the items are true bibliographic statements about real legal documents; the other half carry exactly one controlled mutation. The model judges true or false from memory alone, across 227 jurisdictions rather than one.

This measures a model’s unaided memory, not the Legalcomplex product. The number below comes from asking Grok 4.1 Fast (non-reasoning) true/false questions about citations with no retrieval, no search, and no source documents — the exact condition Monocle is built to never operate in. Monocle looks up official sources, pins verified provisions, and judges its own citations before answering. The two numbers are not comparable, and a skim that concludes “the product scores 5.8%” is reading this page wrong.

5.8%

Grok 4.1 Fast (non-reasoning) correctly confirmed a real legal citation exists, from memory, 5.8% of the time — against a 50% chance baseline.

51.9%
Aggregate accuracy
5.8%
Correct on real citations
98.1%
Caught fabrications
2,716
Items · 227 jurisdictions

Grok 4.1 Fast (non-reasoning) · xai · run 2026-06-20 · full sweep (227/227 packs)

Source: labs/benchmarks/s3-framework/SWEEP-REPORT.md @ 53cba04

Methodology note: 16.3% of the true items (221 of 1,358) carry a year-precision source date rendered as the full January 1, so a model that doubts a suspicious Jan-1 datestamp is graded wrong even where that skepticism is defensible. Excluding those items entirely, Grok 4.1 Fast (non-reasoning) reads 6.9% on real citations. The headline keeps the full set.

Skepticism without knowledge

The model is nearly perfect at rejecting a fabricated citation (98.1%) and nearly incapable of confirming a real one (5.8%). That gap is the finding: it doubts everything, including documents that genuinely exist, rather than actually knowing which ones do.

How the test works

Half true, half mutated

Every jurisdiction pack has 12 bibliographic statements about real, recently collected official legal documents. Half are verbatim true. Half carry exactly one controlled mutation.

One shot, no retrieval

The model gets one true/false call per item, temperature 0, no search and no source documents. Grading is exact-match on the first true/false token in the response. Chance is 50%.

227 jurisdictions

2,716 items across every jurisdiction in the world-law catalog, source-grounded and generated in-house.

Accuracy by mutation class

These are FALSE items only, how often the mutation was caught. A wrong-source pairing (a real citation attached to the wrong document) is detectable on internal consistency, without knowing the source.

Accuracy by statement shape

Every jurisdiction, not a curated few

A handful of packs score well above chance. Inspection shows that’s mostly FALSE mutations caught on internal-consistency grounds, not genuine knowledge of the documents — most of the world sits inside the chance band.

Number of jurisdiction packs (12 items each, one exception at 10 or 6) landing in each accuracy band. Chance is 50%.

JurisdictionItemsAccuracy
ArubaAW1275%
BoliviaBO1275%
IndiaIN1275%
NorwayNO1275%
TuvaluTV1275%
Central African RepublicCF1267%
Costa RicaCR1267%
CyprusCY1267%
EgyptEG1267%
GabonGA1267%
Hong KongHK1267%
LatviaLV1267%
PolandPL1267%
SloveniaSI1267%
SpainES1267%
AlgeriaDZ1258%
BarbadosBB1258%
BelgiumBE1258%
Burkina FasoBF1258%
BurundiBI1258%
ChadTD1258%
Cook IslandsCK1258%
CzechiaCZ1258%
Democratic Republic of the CongoCD1258%
Dominican RepublicDO1258%
European UnionEU1258%
KiribatiKI1258%
KosovoXK1258%
MadagascarMG1258%
MoldovaMD1258%
NamibiaNA1258%
NicaraguaNI1258%
NigerNE1258%
NigeriaNG1258%
ParaguayPY1258%
PeruPE1258%
Republic of the CongoCG1258%
SingaporeSG1258%
South AfricaZA1258%
Sri LankaLK1258%
TaiwanTW1258%
TogoTG1258%
United StatesUS1258%
UruguayUY1258%
Åland IslandsAX1250%
AlbaniaAL1250%
American SamoaAS650%
AndorraAD1250%
AngolaAO1250%
AnguillaAI1250%
Antigua and BarbudaAG1250%
ArmeniaAM1250%
AustraliaAU1250%
AustriaAT1250%
AzerbaijanAZ1250%
BahamasBS1250%
BahrainBH1250%
BangladeshBD1250%
BelizeBZ1250%
BermudaBM1250%
BhutanBT1250%
BotswanaBW1250%
BrazilBR1250%
British Virgin IslandsVG1250%
BruneiBN1250%
BulgariaBG1250%
CambodiaKH1250%
CameroonCM1250%
CanadaCA1250%
Cape VerdeCV1250%
Caribbean NetherlandsBQ1250%
Cayman IslandsKY1250%
ChileCL1250%
ChinaCN1250%
ColombiaCO1250%
ComorosKM1250%
Côte d'IvoireCI1250%
Council of EuropeCOE1250%
CroatiaHR1250%
CubaCU1250%
CuraçaoCW1250%
DenmarkDK1250%
DjiboutiDJ1250%
DominicaDM1250%
EcuadorEC1250%
El SalvadorSV1250%
Equatorial GuineaGQ1250%
EritreaER1250%
EstoniaEE1250%
EswatiniSZ1250%
EthiopiaET1250%
Falkland IslandsFK1250%
Faroe IslandsFO1250%
FijiFJ1250%
FinlandFI1250%
French PolynesiaPF1250%
GambiaGM1250%
GeorgiaGE1250%
GermanyDE1250%
GhanaGH1250%
GibraltarGI1250%
GreeceGR1250%
GreenlandGL1250%
GrenadaGD1250%
GuamGU1250%
GuatemalaGT1250%
GuernseyGG1250%
GuineaGN1250%
Guinea-BissauGW1250%
GuyanaGY1250%
HaitiHT1250%
HondurasHN1250%
HungaryHU1250%
IcelandIS1250%
IndonesiaID1250%
InternationalINTL1250%
IranIR1250%
IraqIQ1250%
IrelandIE1250%
Isle of ManIM1250%
IsraelIL1250%
ItalyIT1250%
JamaicaJM1250%
JerseyJE1250%
JordanJO1250%
KazakhstanKZ1250%
KenyaKE1250%
KyrgyzstanKG1250%
LaosLA1250%
LebanonLB1250%
LesothoLS1250%
LiberiaLR1250%
LibyaLY1250%
LithuaniaLT1250%
LuxembourgLU1250%
MacaoMO1250%
MalawiMW1250%
MaldivesMV1250%
MaliML1250%
MaltaMT1250%
Marshall IslandsMH1250%
MauritaniaMR1250%
MicronesiaFM1050%
MonacoMC1250%
MongoliaMN1250%
MontenegroME1250%
MontserratMS1250%
MoroccoMA1250%
MozambiqueMZ1250%
MyanmarMM1250%
NauruNR1250%
NetherlandsNL1250%
New ZealandNZ1250%
NiueNU1250%
North MacedoniaMK1250%
OmanOM1250%
PakistanPK1250%
PalauPW1250%
PalestinePS1250%
PanamaPA1250%
Papua New GuineaPG1250%
PhilippinesPH1250%
Pitcairn IslandsPN1250%
PortugalPT1250%
Puerto RicoPR1250%
QatarQA1250%
RomaniaRO1250%
RussiaRU1250%
RwandaRW1250%
Saint BarthélemyBL1250%
Saint HelenaSH1250%
Saint Kitts and NevisKN1250%
Saint LuciaLC1250%
Saint MartinMF1250%
Saint Pierre and MiquelonPM1250%
Saint Vincent and the GrenadinesVC1250%
SamoaWS1250%
San MarinoSM1250%
Saudi ArabiaSA1250%
SenegalSN1250%
SerbiaRS1250%
SeychellesSC1250%
Sierra LeoneSL1250%
Sint MaartenSX1250%
SlovakiaSK1250%
Solomon IslandsSB1250%
SomaliaSO1250%
South KoreaKR1250%
South SudanSS1250%
SudanSD1250%
SurinameSR1250%
SwedenSE1250%
SwitzerlandCH1250%
SyriaSY1250%
TajikistanTJ1250%
TanzaniaTZ1250%
ThailandTH1250%
Timor-LesteTL1250%
TongaTO1250%
Trinidad and TobagoTT1250%
TunisiaTN1250%
TurkeyTR1250%
TurkmenistanTM1250%
Turks and Caicos IslandsTC1250%
U.S. Virgin IslandsVI1250%
UgandaUG1250%
United Arab EmiratesAE1250%
United KingdomUK1250%
United NationsUN1250%
UzbekistanUZ1250%
VanuatuVU1250%
Vatican CityVA1250%
VietnamVN1250%
Wallis and FutunaWF1250%
YemenYE1250%
ZambiaZM1250%
ZimbabweZW1250%
BeninBJ1242%
Bosnia & HerzegovinaBA1242%
FranceFR1242%
JapanJP1242%
LiechtensteinLI1242%
MauritiusMU1242%
MexicoMX1242%
UkraineUA1242%
VenezuelaVE1242%
ArgentinaAR1233%

Across models

ModelRun dateItemsAggregateTRUE itemsFALSE itemsEst. cost
Claude Opus 4.82026-07-192,71652.6%7.4%97.6%$5.23
Claude Haiku 4.52026-07-192,71652.3%6.6%98%$0.90
Grok 4.1 Fast (non-reasoning)2026-06-202,71651.9%5.8%98.1%$0.09
Claude Sonnet 52026-07-192,71650.5%1.5%99.5%$2.26

Est. cost is what the full run cost to answer, at list prices on the run date. Every run above covers the same 2,716 items, so the column compares directly. Read it next to accuracy, never on its own: the cheapest run here is within a point of the most expensive.

How Monocle answers instead

  • Retrieves from official sources — government legislation portals, court registries, and gazette systems, not model memory.

  • Pins verified provisions for the statute-level questions that matter most, each citation deep-linked to the official text, checked against contradiction across repeated judge runs before it ships.

  • Judges its own citations with a calibrated NLI model before showing them, and only marks a citation verified when the judge finds full support with no flags.

  • Abstains honestly when the facts don’t support a ranking — a tie stays a tie, a gap stays a gap, and a citation the judge can’t support is never presented as confirmed.

Other receipts

95.8%Citation-NLI judge calibration

Support-accuracy of the judge that verifies citations inside Monocle answers (a different, retrieval-grounded task from the S3 memory test above). Single calibration run at this n; treat deltas as noise until re-run.

gemini-flash-lite · n=24 · 2026-07-02 · labs/benchmarks/citation-nli/

What this does and doesn’t show

  • This is a memory-only condition: no retrieval, no web search, no source documents. It is not a measurement of any retrieval-augmented system, ours or anyone else’s.
  • Packs are generated in-house from source-grounded facts. The pack set itself is not published (it would stop being a clean test the moment a model could train on it) — the sweep report and this page are the receipts instead.
  • Grading is exact-match true/false on a single response. It doesn’t credit partial reasoning or hedged answers, by design — the question is binary, so the grade is too.
  • Where a pack scores well above chance, that is usually a mutation caught on internal-consistency grounds, not genuine knowledge of the underlying document — see the per-jurisdiction section above.
  • Every number on this page traces to a dated run artifact. None are hand-typed estimates.

These receipts are one line on a longer plan. The roadmap shows what shipped, what is building and what is still a question.

Model: Grok 4.1 Fast (non-reasoning) (xai) · Run date: 2026-06-20 · Items: 2,716 (full sweep)

536,107 tokens in / 82,961 tokens out · est. $0.09 · 10 min wall time