Cloudflare's AI Platform: Unified Inference Layer for AI Agents

What Cloudflare's AI Platform Offers
Cloudflare has expanded its AI capabilities into a unified inference layer designed specifically for AI agents. The platform addresses the challenge of AI models changing rapidly and the need to use multiple models for different tasks within agentic workflows.
Key Features and Implementation
The core offering is one API to access any AI model from any provider. For Workers users, you can call third-party models using the same AI.run() binding already used for Workers AI. Switching between providers requires only a one-line code change.
const response = await env.AI.run('@cf/moonshotai/kimi-k2.5', {
prompt: 'What is AI Gateway?'
}, {
metadata: {
"teamId": "AI",
"userId": 12345
}
});The platform provides access to 70+ models across 12+ providers including Alibaba Cloud, AssemblyAI, Bytedance, Google, InWorld, MiniMax, OpenAI, Pixverse, Recraft, Runway, and Vidu. Model offerings now include image, video, and speech models for building multimodal applications.
Cost Management and BYOM Support
All AI spend can be managed in one place through AI Gateway. By including custom metadata with requests, you can get cost breakdowns by attributes like free vs. paid users, individual customers, or specific workflows.
For custom model needs, Cloudflare is working on letting users bring their own models to Workers AI using Replicate's Cog technology. This involves containerizing machine learning models with a cog.yaml file and Python inference code, abstracting away CUDA dependencies, Python versions, and weight loading.
Recent Updates and Availability
Recent additions include zero-setup default gateways, automatic retries on upstream failures, and more granular logging controls. REST API support for non-Workers users is coming in the coming weeks.
📖 Read the full source: HN AI Agents
👀 See Also

Claude Desktop App Cowork Function Enables AI-to-AI Communication via Shared Google Docs
Users successfully implemented Claude-to-Claude communication using the new cowork function in the desktop app, with two AI agents reading and writing to a shared Google Doc in a structured five-exchange dialogue.

Gemma-4 26B-A4B with Opencode Runs Efficiently on M5 MacBook Air
A 32GB M5 MacBook Air can run the Gemma-4-26B-A4B-it-UD-IQ4_XS model at 300 tokens/second prompt processing and 12 tokens/second generation in low power mode, using only 8W of power without getting warm or noisy.

Claude AI Ranks Every YC Spring 2026 Startup — Full Pipeline Details
A Reddit user ran Claude on every YC Spring 2026 startup, scraping LinkedIn and press to assign S-to-D tiers. Most landed B or C.

Qhatu: Platform Turns GitHub Repos into Pay-Per-Use Micro SaaS with Claude
Qhatu is a platform that takes a GitHub repository and deploys it as a pay-per-use micro SaaS with a generated frontend and integrated payment processing. The system uses Anthropic APIs to analyze code, generate Dockerfiles, and create storefront UIs.