Datatera Logo
DATATERA.ai
Evidence ledger / 14 Aug 2026

Document extraction model benchmark

Four extraction cases. Exact provider routes. Every score is tied to runtime audit evidence, so a fallback or a response with no verified LLM call stays unscored.

B1
CRM CSV
5 rows
B2
Financial CSV
332 rows
B3
Russian invoice
OCR + Cyrillic
B4
Investor table
32 rows

Latest cohort

OpenRouter agent-session models

All 28 requests completed: 21 cases passed strict runtime identity and seven remain unscored. Asterisks average accepted cases only.

Model / routeB1B2B3B4ValidQualityValid timeEst. costAudit
DeepSeek V4 Flash
deepseek/deepseek-v4-flash-0731
100%UnscoredUnscored60%2 / 480.0%*198.57s$0.003084partial
Gemini 3.7 Flash
google/gemini-3.7-flash
100%90%Unscored80%3 / 490.0%*231.77s$0.148371partial
GLM 5.2
z-ai/glm-5.2
100%90%Unscored80%3 / 490.0%*187.35s$0.133615partial
GPT-5.6 Terra
openai/gpt-5.6-terra
100%90%Unscored80%3 / 490.0%*79.46s$0.220568partial
Gemini 3.6 Flash
google/gemini-3.6-flash
100%90%Unscored80%3 / 490.0%*410.10s$0.312994partial
Muse Spark 1.2
meta/muse-spark-1.2
100%90%0%80%4 / 467.5%94.67s$0.246303audited
Claude Sonnet 5
anthropic/claude-sonnet-5
100%90%Unscored60%3 / 483.3%*368.18s$0.769854partial

Strict partial run / 13 Aug 2026

Muse, Gemma, and Qwen

Asterisks are averages across accepted cases only. Missing cases were rejected by the runtime evidence gate and are not counted as zero.

Model / routeB1B2B3B4ValidQualityValid timeEst. costAudit
Muse Glimmer 30B
meta/muse-glimmer-30b
100%90%Unscored80%3 / 490.0%*239.63s$0.078254partial
Gemma 4 31B
google/gemma-4-31b-it
100%90%Unscored100%3 / 496.7%*731.96s$0.028575partial
Qwen3.6 27B
qwen/qwen3.6-27b
100%90%UnscoredUnscored2 / 495.0%*630.88s$0.029375partial

Scoring contract

Measured, not inferred

  1. 01 / identityThe requested OpenRouter route must appear in runtime audit evidence.
  2. 02 / executionAt least one audited LLM call must have handled the extraction.
  3. 03 / qualityGround-truth values, rows, and columns are checked per case.
  4. 04 / failuresFallbacks and failed cases remain visible but never receive a score.

Benchmark your own documents

Use the same evidence contract on a representative sample from your workflow.

Book a benchmark review