Novice Builds Local 'Second Brain' on Intel Arc Pro B60: Qwen 3.8 27B at 38 tok/s, 128K Context

✍️ OpenClawRadar📅 Published: September 28, 2026🔗 Source
Ad

A non-developer with no coding background built a mostly local personal AI agent (OpenClaw) on an old Dell OptiPlex, then built a dedicated AI server around an Intel Arc Pro B60 GPU. The source is a detailed writeup of the exact hardware, costs, and results — including running Qwen 3.8 27B at ~38 tok/s with a 128K context window.

How it started

The author was an average chatbot user — research and writing only — until a YouTube video about OpenClaw (an AI agent that actually does things, not just chats) got them building. Setup happened in about 2 days using a free Claude account on an old Dell OptiPlex 3431, initially running on a $20/month ChatGPT subscription.

The Dell (agent's control plane)

  • CPU: Intel Core i9-9900 (upgraded from stock i5)
  • RAM: 32GB DDR4-2666 (was 64GB, half moved to the AI box)
  • GPU: NVIDIA Quadro P620, 2GB (stock, useless for AI)
  • OS drive: 512GB NVMe (added)
  • Storage: 2TB SATA SSD (original)
  • OS: Windows 11 Pro

Runs OpenClaw (agent harness), memory, scheduling, Telegram (the chat interface), backups, and stores local models deployable to the AI box at any time.

Stability issues traced back to cheap "grey market" memory sticks the Dell shipped with on Amazon — the author suspects bad RAM corrupted Windows files. The i9 upgrade was called out as the only bad spend so far ($200 on eBay), with the agent literally pleading against buying it.

Ad

The BlackBox (BB) — dedicated AI server

While running the agent, the author ordered then cancelled a $2,000 Mac Studio after realizing it couldn't run a model capable of doing what they needed. They briefly went all-cloud: $20/month Claude plus a $100 ChatGPT plan = $120/month just to run the agent, which felt like a trap.

The turning point was small open models — Qwen, GPT-OSS 20B, Mistral, Gemma. When Qwen 3.8 27B dropped, they committed to building a separate machine that does nothing but serve AI to the agent. That separation is described as the best decision of the project.

The guiding rule

Everything must be generic and interchangeable — never married to a model, a piece of hardware, or even the agent harness itself. That lets the agent host machine change while local inference stays put.

Reported numbers

Qwen 3.8 27B at roughly 38 tok/s with a 128K context window on the Intel Arc Pro B60 — now the default model, with cloud models as backup instead of the other way around. Full specs, config, and numbers are in the original post.

📖 Read the full source: r/openclaw

Ad

👀 See Also