Nvidia RTX Spark: 1-Petaflop Superchip Brings Local AI Agents to Windows PCs

Nvidia today announced RTX Spark, a new superchip that brings 1 petaflop of AI compute to Windows PCs, purpose-built for running personal AI agents locally. The chip combines a Blackwell RTX GPU (6,144 CUDA cores, fifth-gen Tensor Cores with FP4), a 20-core Grace CPU, and up to 128GB of unified memory, connected via NVLink-C2C. MediaTek contributed to the custom Arm-based CPU design for power efficiency.
Key Specs and Capabilities
- AI performance: 1 petaflop (FP4)
- GPU: Blackwell RTX with 6,144 CUDA cores
- CPU: 20-core NVIDIA Grace (Arm), co-designed with MediaTek
- Memory: up to 128GB unified memory
- Software stack: CUDA, RTX, DLSS, FP4, TensorRT, OptiX, Reflex, G-SYNC
RTX Spark can run 120B-parameter LLMs with up to 1 million tokens context locally, render 90GB+ 3D scenes, edit 12K 4:2:2 video, generate 4K AI video, and play AAA games at 1440p 100+ fps.
Windows-Native Agent Security
Nvidia and Microsoft are collaborating on new Windows security primitives and the Nvidia OpenShell runtime to enable secure on-device agents. The security layer provides identity, containment, policy, and end-to-end security. OpenShell adds user-defined policies for agent capabilities, intelligent query routing to local vs. cloud models, and PII masking in cloud-bound queries.
Agent frameworks including Hermes Agent and OpenClaw are building Windows apps on this stack, enabling cross-app workflows, file search, image/video generation, and code plugin creation.
Availability
RTX Spark-powered slim laptops (all-day battery, premium displays) and compact desktops will ship this fall from ASUS, Dell, HP, Lenovo, Microsoft Surface, and MSI, with Acer and GIGABYTE models to follow.
📖 Read the full source: HN AI Agents
👀 See Also

Local Qwen 3.6 vs Frontier Models on a Coding Primitive: Single-File HTML Canvas Driving Animation
A Reddit user pitted local Qwen 3.6 quants against frontier models (Claude, Gemini, GPT, Kimi) on a dense single-file HTML canvas driving animation task. The local Qwen 3.6-27B Q4_K_M delivered more natural motion and layering than some frontier outputs.

Developer Switches from Cursor Composer 2 and Kimi 2.6 to Qwen3.6:35b-a3b for Enterprise Workloads
A developer reports using Qwen3.6:35b-a3b for daily work on a 500-700k LOC enterprise suite, citing better performance than Kimi 2.6 and DeepSeek 4 Pro/Flash, with costs ~$0.08/1M tokens on OpenRouter.

NVIDIA Releases Nemotron-3-Ultra-550B: 55B Active Parameters, 1M Context, LatentMoE Hybrid
NVIDIA released Nemotron-3-Ultra-550B-A55B-BF16, a 550B parameter model with 55B active, 1M token context, hybrid LatentMoE architecture (Mamba-2 + MoE + Attention + MTP), and configurable reasoning.

Anthropic's Business Strategy: API Revenue Drives Consumer Tier Limitations
Anthropic's consumer subscription tiers operate at a loss, subsidized to build AI mindshare, while their API business generates revenue. The $20 Pro tier is intentionally limited to filter users toward higher-value Max subscriptions.