Subquadratic Debuts 12M Token Context Window for AI Models

Subquadratic has announced a 12-million-token context window, claiming a breakthrough in subquadratic attention mechanisms. This compares to typical 128K-1M token windows in current models. The technique allows models to handle vastly larger contexts without quadratic scaling of compute or memory.
Key Details
- Context window: 12 million tokens (12x larger than GPT-4's 128K tokens)
- Based on subquadratic attention, likely using linear or near-linear complexity in sequence length
- Enables processing entire large codebases, long documents, or multi-hour video transcripts in a single forward pass
- Potential applications: code review of entire repos, long-document analysis, multi-turn dialog with full history
- Compatible with existing transformer-based LLMs via drop-in attention replacement
The approach reduces O(n²) attention to near-O(n) using techniques like state-space models or low-rank factorizations. No specific benchmark numbers are provided in the source, but the claim is that this makes 12M-token windows practical on a single GPU.
Who It's For
AI engineers working on code analysis, document processing, or any task requiring long-context understanding without expensive chunking or retrieval.
📖 Read the full source: HN AI Agents
👀 See Also
OpenClaw v2026.9.8 Ships GPT-6.1 Sol, Fixes Windows Startup and Web UI Tabs
OpenClaw v2026.9.8 lands 43 pull requests from 8 contributors, adding GPT-6.1 Sol via the OpenAI provider plus fixes for Windows file locks, plugin settings, and Web UI tab sessions.

Local vs Cloud Models: Qwen-3.6-27B, Gemma-4-31B, Claude Haiku, Codex-Spark on Hard Code Gen
A user tested Qwen-3.6-27B (q4_k_m) locally on an RTX 5080 against API-based Gemma-4-31B, Claude Haiku 4.5, and Codex-Spark on a complex code task. Only Codex-Spark produced complete code (but with import errors); all others failed partially. Cost: Gemma used $0.112 for 803k input tokens.

Claude Code Postmortem: Three Bugs Caused Quality Degradation, Now Fixed
Anthropic traced recent Claude Code quality complaints to three separate changes: default reasoning effort was lowered, a caching bug dropped session memory, and a verbosity prompt hurt coding quality. All fixed as of April 20 (v2.1.116).

Chrome's Gemini Nano AI Model Consumes 4GB of Disk Space
Google Chrome automatically downloads a 4GB weights.bin file for the Gemini Nano on-device AI model, which may bloat storage without clear user notification. Disabling the On-Device AI toggle in settings removes the file and prevents re-download.