AI Inference Is Obviously Profitable: Breaking Down the Economics

Sean Goedecke argues that AI inference is obviously profitable, contrary to claims that it's subsidized by VC money. He breaks down the math and open-model pricing to show that serving LLMs can be a sustainable business.
Cost Breakdown for a 70B Dense Model
Using four Nvidia A100 GPUs (400W each, ~2M tokens/hour):
- Power: ~13¢/hr at industrial rates, plus ~13¢/hr cooling → 26¢/hr total
- GPU amortization: $20k per A100 over 5 years → $16k/yr or $1.80/hr
- Total: roughly $1 per million output tokens
OpenAI's GPT-5.4-mini charges $4.50 per million tokens, and stronger models are 3-6x more expensive. This makes the claimed 70-80% gross margin plausible.
Open Models Confirm Profitability
DeepSeek claims over 80% margin on R1 inference while charging less than half of OpenAI/Anthropic. Their DeepSeek-V4-Pro API is around 87¢ per million output tokens—close to actual cost, suggesting margins for frontier models are even higher.
Why High Margins Exist
AI labs like OpenAI and Anthropic need inference profits to subsidize training costs, so they keep API prices high. But an inference-only provider without training costs could profit even at lower prices. Even if frontier labs go under, whoever acquires their model rights can continue selling inference profitably.
📖 Read the full source: HN AI Agents
👀 See Also
FairyFuse Achieves 29.6x Kernel Speedup on CPUs via Ternary Weight Multiplication-Free Inference
FairyFuse fuses eight real-valued sub-GEMVs into a single AVX-512 loop using masked adds/subtracts, yielding 32.4 tokens/s on Xeon 8558P and 1.24x speedup over llama.cpp Q4_K_M with near-lossless quality.

1.2B Local Model Beats 1T Clouds in Poker: Aggression Trumps Knowledge in Shove-or-Fold Format
A 1.2B Liquid model won 2 of 5 Texas Hold'em tournaments against models up to 1T parameters, because in a short-stack format, never folding earned more chips than smart play.

Qwen3.5-27B-FP8 performance benchmarks with OpenClaw agents
Testing shows Qwen3.5-27B-FP8 can run six OpenClaw agents simultaneously with throughput scaling to 120 tokens/second. The SGLang framework with prefix caching reduces 100K context prefill from 10 seconds to 200ms.
Google's AI-Assisted Rewrites of C/C++ Dependencies to Rust: Scaling Memory Safety
Google's blog details AI-assisted rewrites of C/C++ dependencies to Rust to scale memory safety, leveraging AI to automate the conversion process.