Ontario Audit: 60% of AI Scribe Systems Mix Up Drugs, 85% Miss Mental Health Details

The Office of the Auditor General of Ontario audited 20 approved AI Scribe systems used by physicians and nurse practitioners, simulating doctor-patient recordings to evaluate accuracy. The results are stark:
- 12 of 20 systems inserted incorrect drug information into patient notes.
- 9 of 20 fabricated information — e.g., claiming “no masses found” or “patient anxious” — that was never discussed.
- 17 of 20 missed key mental health details from the recording.
- 6 of 20 fully or partially omitted mental health issues.
The audit also slammed the evaluation scoring methodology. Accuracy of medical notes accounted for just 4% of the total score, while having a domestic presence in Ontario contributed 30%. Bias controls, threat/risk/privacy assessments, and SOC 2 Type 2 compliance each counted only 2–4%. As the report states, such weightings “could result in the selection of vendors whose AI tools may produce inaccurate or biased medical records.”
While OntarioMD has recommended manual review of AI notes, the audit noted no mandatory attestation feature in any approved system. Ontario’s Ministry of Health said over 5,000 physicians use these tools with no reported patient harm.
📖 Read the full source: HN AI Agents
👀 See Also

Yann LeCun's AMI raises $1B for AI world models, challenges LLM approach
Yann LeCun's startup AMI raised over $1 billion to develop AI world models that understand the physical world, arguing LLMs alone won't achieve human-level intelligence. The company will build systems with persistent memory, reasoning, and planning capabilities for manufacturing, biomedical, and robotics applications.

Claude vs GPT-4o: Same Double Pendulum Prompt, Different Coordinate Conventions
Claude and GPT-4o produce visually different double pendulum simulations because they interpret theta from opposite verticals — top vs bottom — while using the same renderer. The math is correct in both cases, but the mismatch reveals a subtle ambiguity in prompt interpretation.

Claude System Prompt Compliance Degrades in Long Conversations
Claude-based agents show degraded system prompt compliance after 40-50 messages, with formatting rules being ignored and constraints forgotten. The issue stems from system prompts competing with conversation history for attention weight in the context window.

Anthropic Launches 10 Finance AI Agents for Pitchbooks, KYC, Month-End Close
Anthropic released 10 ready-to-run AI agents for financial services and insurance, covering pitchbook creation, KYC screening, and month-end close, delivered via Claude Cowork, Claude Code, and Managed Agents.