IDP Leaderboard benchmark shows Claude Sonnet 4.6 matches Opus 4.6 for document AI tasks

The IDP Leaderboard, an open benchmark for document AI, has published results comparing Claude models on document processing tasks. The benchmark tested 16 models across multiple categories using over 9,000 real documents.
Benchmark Results
The Claude model scores from the IDP Leaderboard:
- Claude Sonnet 4.6: 80.8 overall
- Claude Opus 4.6: 80.3 overall
- Claude Haiku 4.5: 69.6 overall
Sonnet and Opus performed essentially equivalently on extraction tasks including text, tables, formulas, and layout analysis. The radar charts for both models look identical according to the benchmark results.
Cost Comparison
The source notes significant cost differences:
- Sonnet costs $24 per 1,000 pages
- Opus costs $40 per 1,000 pages
For document processing workloads, the benchmark suggests there's no reason to use Opus given the equivalent performance at lower cost.
Important Caveat
One notable finding: Claude models had stricter content moderation that affected performance on certain document types. Old newspaper scans, textbook pages, and historical documents sometimes triggered content filters. This issue only appeared in the OlmOCR and OmniDoc benchmarks.
All predictions from the benchmark are visible in the Results Explorer at idp-leaderboard.org, where you can see exactly what each Claude model output on every document.
📖 Read the full source: r/ClaudeAI
👀 See Also

Netlify CTO Dana Lawson: Writing Code Is No Longer the Job
Netlify CTO Dana Lawson argues that developer work shifts from writing code to orchestrating AI agents. Engineers become experience designers, curating agent outputs and managing system boundaries.

Anthropic Raises Claude Limits and Adds SpaceX Compute Capacity
Anthropic has increased Claude usage limits and secured a compute deal with SpaceX. The Reddit discussion weighs whether this is just infra scaling or a strategic move toward making Claude a better platform for agentic work.

Claude's policy filter blocks bioinformatics work with pathogen names
A computational virology researcher reports Claude's usage policy filter flags legitimate bioinformatics scripts when pathogens are named, requiring workarounds like describing tasks without organism names or downgrading to Sonnet 4. The issue affects Claude Code, claude.ai, and both Opus 4.6 and Sonnet 4.6 models.

AI-Generated Flower Seeds Scam Floods eBay, Amazon, Etsy
Scammers use AI images to sell seeds for plants like 'teddy bear sunflowers' that don't exist. eBay, Amazon, and Etsy struggle to stop the flood.