Strands Decider 2B: A 2B-Parameter Decision Model for Agentic Workflows
Strands Labs (AWS) released Strands Decider 2B, a 2-billion-parameter open-source decision model optimized for fast local experimentation and agentic workflows. Unlike general-purpose LLMs, decision models are designed to pick between a fixed set of options (e.g., "Is 'turn on the lights' about the coffee machine? Yes or no.") or assign simple numerical scores (e.g., sentiment between 0 and 1).
What it is and why it matters
Decision models trade flexibility for speed and reliability. They always return an answer from the provided options, run with very low latency, and emit a calibration score ("how sure can I be this is correct?") that frontier LLM APIs don't expose. The tradeoff: they're worse at complex reasoning and can't generate text, so they're unsuited for coding, chatbots, or summarization.
Strands positions this as a new class of "system one" models, following TypeSafe AI's launch of Jev earlier this month. The intent is to drive agentic decision steps in the Strands Harness SDK.
Architecture
The model takes a pre-trained Qwen3.5-2B torso, strips the LM head (removing text generation), and replaces it with a pointer head that scores each offered option. The head compares the hidden state at each option position against the hidden state at the <answer> position. The head is small — just over 1M parameters — and the torso is fine-tuned with a rank-16 LoRA adapter.
This is v19 of the architecture. The first iteration used a slot head, which performed significantly worse; all iteration notes are in the repo.
Benchmarks
Strands measured accuracy on JevBench's public set and calibration via Brier score on the same set:
- Accuracy: 3rd of 33 in the 2B class
- Accuracy excluding just-over-2B models: 1st of 30
- Latency (median): ~115ms on an Nvidia RTX 3090, ~153ms for small tasks on an M3 MacBook
- Latency scales roughly linearly with task size
Availability
- Code on GitHub:
strands-decider-2b - Weights on Hugging Face
- Includes all training data and scripts
- Runs on local CPU or GPU — no API calls required
Who it's for
Developers building agentic workflows who need fast, calibrated yes/no or classification decisions at the edge, and researchers who want a small, hackable base for decision-model experiments.
📖 Read the full source: HN AI Agents
👀 See Also

yburn: Tool to audit and replace unnecessary AI agent cron jobs
yburn is a Python tool that audits AI agent cron jobs and replaces those that don't need LLMs with standalone Python scripts. The creator found 58% of 98 cron jobs were purely mechanical tasks like system health checks and git backups.

Claude Code Logs Every Session to Disk — Here's How to Index and Recall Them
Claude Code writes every session turn to ~/.claude/projects/ as JSONL. One user indexed 1026 sessions (57MB, 76K turns) into SQLite+FTS5 with an MCP server for search and thread recall across sessions.

Legal MCP Server for Claude Provides Access to 4M+ US Court Opinions
A free, open-source MCP server built with Claude Code gives Claude AI access to 4M+ real US court opinions, providing 18 tools for case law search, citation tracing, Bluebook parsing, Clio practice management, and PACER federal filings without hallucinations.

LobsterBoard adds theme system and template gallery
LobsterBoard now includes a theme system with five visual options and a template gallery that allows users to export and import dashboard layouts with automatic sensitive data stripping.