Evidence ledger / 14 Aug 2026
Document extraction model benchmark
Four extraction cases. Exact provider routes. Every score is tied to runtime audit evidence, so a fallback or a response with no verified LLM call stays unscored.
B1
CRM CSV
5 rows
B2
Financial CSV
332 rows
B3
Russian invoice
OCR + Cyrillic
B4
Investor table
32 rows
Latest cohort
OpenRouter agent-session models
All 28 requests completed: 21 cases passed strict runtime identity and seven remain unscored. Asterisks average accepted cases only.
| Model / route | B1 | B2 | B3 | B4 | Valid | Quality | Valid time | Est. cost | Audit |
|---|---|---|---|---|---|---|---|---|---|
DeepSeek V4 Flash deepseek/deepseek-v4-flash-0731 | 100% | Unscored | Unscored | 60% | 2 / 4 | 80.0%* | 198.57s | $0.003084 | partial |
Gemini 3.7 Flash google/gemini-3.7-flash | 100% | 90% | Unscored | 80% | 3 / 4 | 90.0%* | 231.77s | $0.148371 | partial |
GLM 5.2 z-ai/glm-5.2 | 100% | 90% | Unscored | 80% | 3 / 4 | 90.0%* | 187.35s | $0.133615 | partial |
GPT-5.6 Terra openai/gpt-5.6-terra | 100% | 90% | Unscored | 80% | 3 / 4 | 90.0%* | 79.46s | $0.220568 | partial |
Gemini 3.6 Flash google/gemini-3.6-flash | 100% | 90% | Unscored | 80% | 3 / 4 | 90.0%* | 410.10s | $0.312994 | partial |
Muse Spark 1.2 meta/muse-spark-1.2 | 100% | 90% | 0% | 80% | 4 / 4 | 67.5% | 94.67s | $0.246303 | audited |
Claude Sonnet 5 anthropic/claude-sonnet-5 | 100% | 90% | Unscored | 60% | 3 / 4 | 83.3%* | 368.18s | $0.769854 | partial |
Strict partial run / 13 Aug 2026
Muse, Gemma, and Qwen
Asterisks are averages across accepted cases only. Missing cases were rejected by the runtime evidence gate and are not counted as zero.
| Model / route | B1 | B2 | B3 | B4 | Valid | Quality | Valid time | Est. cost | Audit |
|---|---|---|---|---|---|---|---|---|---|
Muse Glimmer 30B meta/muse-glimmer-30b | 100% | 90% | Unscored | 80% | 3 / 4 | 90.0%* | 239.63s | $0.078254 | partial |
Gemma 4 31B google/gemma-4-31b-it | 100% | 90% | Unscored | 100% | 3 / 4 | 96.7%* | 731.96s | $0.028575 | partial |
Qwen3.6 27B qwen/qwen3.6-27b | 100% | 90% | Unscored | Unscored | 2 / 4 | 95.0%* | 630.88s | $0.029375 | partial |
Scoring contract
Measured, not inferred
- 01 / identityThe requested OpenRouter route must appear in runtime audit evidence.
- 02 / executionAt least one audited LLM call must have handled the extraction.
- 03 / qualityGround-truth values, rows, and columns are checked per case.
- 04 / failuresFallbacks and failed cases remain visible but never receive a score.
Benchmark your own documents
Use the same evidence contract on a representative sample from your workflow.