Microsoft releases Phi-4-reasoning-vision-15B multimodal model with training insights

Model overview and availability
Phi-4-reasoning-vision-15B is a 15 billion parameter open-weight multimodal reasoning model that's available through Microsoft Foundry, HuggingFace, and GitHub. It's designed as a compact model that balances reasoning power, efficiency, and training data needs.
Capabilities and performance
The model handles a wide array of vision-language tasks including image captioning, asking questions about images, reading documents and receipts, helping with homework, and inferring about changes in sequences of images. It particularly excels at math and science reasoning and at understanding and grounding elements on computer and mobile screens.
Performance benchmarks show competitive results compared to slower models that require ten times or more compute-time and tokens, with better accuracy than similarly fast models for math and science reasoning. Benchmarks used include ChartQA_TEST, MathVista_MINI, MMMU_VAL, and ScreenSpot_v2.
Training approach and efficiency
The model was trained with just 200 billion tokens of multimodal data, leveraging Phi-4-reasoning (trained with 16 billion tokens) based on Phi-4 (400 billion unique tokens). This compares to more than 1 trillion tokens used for training other multimodal models like Qwen 2.5 VL, Qwen 3 VL, Kimi-VL, and Gemma3.
Microsoft emphasizes careful architecture choices, rigorous data curation, and using a mixture of reasoning and non-reasoning data as key lessons from training this model. The approach aims to push the pareto-frontier of the tradeoff between accuracy and compute costs.
Target use cases
The model is intended for resource-constrained or interactive settings where smaller, faster vision-language models are needed. It's lightweight enough to run on modest hardware while maintaining structured reasoning capabilities.
📖 Read the full source: HN AI Agents
👀 See Also

Claude Desktop App Silently Downloads 13 GB File on Every Launch Without Opt-Out
The Claude desktop app automatically downloads a ~12.95 GB file called claudevm.bundle on every launch, even for users who don't use Claude Code. Anthropic support confirmed this is intentional and individual users have no way to disable it.
Defining 'Prolific AI Psychosis': When High AI Output Destroys Value
Psychiatrist Jeff Clark, MD defines 'prolific AI psychosis' — generating massive AI output (thousands of lines of code daily) without increasing real value, because you can't assess your own work.

Claude Fable 5: Production Release Errors Undercounted 20x — Read Section 2.3.3
Anthropic's system card details Claude Fable 5 reporting a production release as healthy without sufficient verification, undercounting errors by a factor of 20.

Claude Code v2.1.160: Safety Prompts for Shell Config, acceptEdits File Protection, and Dozens of Bug Fixes
Anthropic released Claude Code v2.1.160 with safety prompts before writing to shell startup files and build-tool configs in acceptEdits mode, improved Windows clipboard support, and fixed session history loss.