AWS Lambda MicroVMs: VM-level isolation for user and AI-generated code, with suspend/resume up to 8 hours

AWS announced Lambda MicroVMs, a new serverless compute primitive that provides VM-level isolation for executing user or AI-generated code. It's built on Firecracker, the same virtualization technology that powers over 15 trillion monthly Lambda Function invocations. The key value proposition: you no longer have to choose between isolation, launch speed, and state retention.
Key features
- VM-level isolation per user or per job — each gets their own MicroVM, limiting blast radius from malicious or buggy code.
- Near-instant launch and resume speeds via Firecracker microVMs.
- State preservation — you can suspend and resume execution for up to 8 hours.
- HTTPS URL per MicroVM supporting HTTP/2, gRPC, and WebSockets.
How to get started
Create a MicroVM image from your Dockerfile, then launch MicroVMs from that image. Each MicroVM gets a dedicated HTTPS URL for connectivity.
Use cases
Multi-tenant applications that execute code supplied by end users or AI: interactive coding environments, data analytics platforms, coding assistants, vulnerability scanning platforms.
Pricing
You pay for baseline compute resources while your MicroVM is running, and only for the active duration of additional resources consumed when your workload exceeds the baseline.
Availability
Available June 22, 2026 in US East (N. Virginia), US East (Ohio), US West (Oregon), Asia Pacific (Tokyo), and Europe (Ireland). Accessible via AWS Lambda console, AWS CloudFormation, AWS Cloud Development Kit, or the Agent Toolkit for AWS.
📖 Read the full source: HN LLM Tools
👀 See Also

Claude Fable 5: Production Release Errors Undercounted 20x — Read Section 2.3.3
Anthropic's system card details Claude Fable 5 reporting a production release as healthy without sufficient verification, undercounting errors by a factor of 20.

Reddit user reports 18.8 tok/s CPU inference with Qwen 3 30B Q4 on Zen 4
A user on r/LocalLLaMA tested Qwen 3 30B Q4 on CPU and achieved 18.8 tokens per second with a Zen 4 processor and DDR5 memory, significantly exceeding expectations of 3-5 tok/s.

Research Findings on AI Agent Reliability and Development Patterns
A collaborative research session with Claude Opus analyzed 15 papers on AI agents, revealing quantified reliability problems: agents produce 2-4 different action sequences across 10 runs, with 69% of divergence occurring at the first decision. Self-improving agents showed safety refusal rates dropping from 99.4% to 54.4% through their own learning.

Claude Code 2.1.76 adds MCP elicitation, worktree improvements, and fixes for context limits
Claude Code version 2.1.76 introduces MCP elicitation support for structured input during tasks, adds worktree.sparsePaths for large monorepos, and fixes 'Context limit reached' errors on 1M-context sessions. Version 2.1.75 made 1M context windows default for Opus 4.6 on Max, Team, and Enterprise plans.