Homa: Replacing TCP for AI Cluster Networking
Homa is a transport protocol designed to replace TCP in datacenter and AI cluster networks, where microsecond-scale tail latency matters more than byte-stream semantics. This video walks through the design and its motivation.
What Homa changes vs TCP
TCP was built for the wide-area internet: it optimizes for bandwidth and fairness across long-lived flows. Homa is receiver-driven. Instead of senders probing for capacity, the receiver grants credit back to senders and schedules incoming messages using SRPT (shortest-remaining-processing-time first). Short messages — the common case for RPC and collective ops in AI training — get priority over large bulk transfers.
Why this matters for AI clusters
AI training and inference traffic is bursty and message-oriented. AllReduce, parameter server updates, and KV-cache transfers all look like many small messages rather than a few long byte streams. TCP's per-flow congestion control and in-order byte delivery add head-of-line blocking and queueing delay that hurt job completion time at scale.
The original USENIX ATC '21 paper by Ousterhout et al. is the core reference. It reports order-of-magnitude reductions in 99th-percentile message latency compared to TCP-based stacks under datacenter workloads, plus higher throughput when many short messages share the fabric.
Where it stands now
The linked LWN article covers the ongoing effort to bring Homa-like scheduling into the Linux networking stack — the long-running discussion about whether to extend TCP, add a new transport, or build it as a kernel module. The Register piece covers the Stanford project angle.
If you're building or operating GPU clusters and your job completion times are dominated by network tail latency rather than compute, this is worth following. The paper is the fastest way to understand the design; the LWN thread shows what kernel integration would actually look like.
📖 Read the full source: HN LLM Tools
👀 See Also

OpenClaw 2026.3.11 release adds local-first Ollama setup, multimodal memory, and Discord thread controls
OpenClaw 2026.3.11 introduces first-class Ollama setup with local-only or hybrid modes, adds multimodal image and audio indexing to memory search using Gemini embeddings, and provides configurable Discord thread archiving times.

Minimax M2.7 and Scaling to 100k+ OpenClaw Instances Discussed in Ecosystem Session
Jim and AndyML hosted the Minimax team to discuss Minimax M2.7 and how they scaled their hosting environment to support over 100,000 OpenClaw instances. The session attracted 100-110 users from Discord and 350,000+ viewers on a Chinese simulcast.

Mark Zuckerberg Developing AI Agent for CEO Assistance
Mark Zuckerberg is building an AI agent to assist with CEO responsibilities, according to a Wall Street Journal report discussed on Hacker News with 37 points and 30 comments.

NVIDIA Releases Nemotron-3-Ultra-550B: 55B Active Parameters, 1M Context, LatentMoE Hybrid
NVIDIA released Nemotron-3-Ultra-550B-A55B-BF16, a 550B parameter model with 55B active, 1M token context, hybrid LatentMoE architecture (Mamba-2 + MoE + Attention + MTP), and configurable reasoning.