Lathoa: A Kids' Math App Where the AI Is Meant to Be Wrong
Lathoa is a math practice app for kids roughly 10 to 14 years old. A robot character named Errol walks through a math problem step by step, and one of those steps is deliberately wrong. The kid's job is to spot the bad step and explain what's wrong with it. Sometimes nothing is wrong — so just hitting "there's a mistake" every time won't work. There's a playable case on the homepage with no signup.
How scoring works
- If the user finds a real error, she has to enter an explanation to gain extra XP.
- Speed also factors into the score — faster answers earn more points.
- There is no direct interaction or chatting with an LLM; the kid never prompts a model.
The hard part: making the model wrong on purpose
The author calls out the counterintuitive finding: getting an LLM to be wrong on purpose is difficult. Roughly half the time the model gives the correct answer and labels it wrong, or produces a "mistake" that's actually correct. That forced a verification pipeline before any case reaches a kid.
Evaluation pipeline
- A plain arithmetic check redoes the math exactly where possible.
- A second model solves the same problem without seeing Errol's work. If the two models disagree, the case is thrown away.
- Lathoa's harness is described as stable, with many evaluation steps to catch inconsistencies and prompt injections.
- Known weak spot: the second model can make the same mistake as the first. The arithmetic check exists to catch that.
- That arithmetic check currently only works on English cases. German and Greek use a comma for decimals, and the parsing isn't right yet.
The open pedagogical question
The author's real ask is whether finding someone else's mistake teaches something that solving the problem yourself doesn't. He's not sure, and wants to hear from teachers. If you work with this age group, that's the thread to weigh in on.
Worth noting for anyone building similar tools: the two-model + symbolic-check pattern here is a reasonable template for any task where you need a verifiably specific output, not just a plausible one. The failure mode — a verifier model sharing the same blind spot as the generator — is the classic limitation you have to design around, and this project names it openly.
📖 Read the full source: HN LLM Tools
👀 See Also

Using a Local LLM as a Claude Code Subagent to Reduce Context Usage
A developer shares a method to use Claude Code to delegate tasks to a local LLM via LM Studio's API, keeping file content out of Claude's context. The approach uses a ~120-line Python script with tool-calling to read files locally and return summaries.

Claude Code v2.1.144: Background Sessions, /model Scoping, and 15s Startup Timeout
Claude Code v2.1.144 adds /resume for background sessions, scopes /model to current session only, and fixes a 75s startup hang when api.anthropic.com is unreachable with a 15s timeout.

Open-source solo RPG engine uses three Claude instances for parsing, narration, and direction
EdgeTales is an open-source text-based solo RPG engine where dice mechanics determine outcomes and Claude AI generates atmospheric prose. The system uses three Claude instances in a pipeline: Brain (Haiku) for parsing input to JSON, Narrator (Sonnet) for writing prose, and Director (Haiku) for async scene analysis.

ClearSpec: A Spec Generator to Reduce Hallucination in Claude Code
ClearSpec is a tool that generates structured specifications from plain English descriptions, connecting to GitHub repos to reference real file paths and dependencies, then uses those specs as prompts for Claude Code to provide better context.