Agent-Desktop: Structured Desktop Automation via OS Accessibility Trees

Agent-desktop is a native desktop automation CLI built with Rust, designed for AI agents that need to control desktop applications programmatically. Instead of the common screenshot-based approach (take screenshot, predict pixel coordinates, click, repeat), it interacts through operating system accessibility trees — the same structured data screen readers use. This means the model sees element roles, names, hierarchy, and state directly, making interactions faster, cheaper, and more robust to UI shifts.
Key Features
- Single Rust binary (~15 MB), no runtime dependencies
- 53 commands covering observation, interaction, keyboard, mouse, notifications, clipboard, and window management
- JSON output — machine-readable with error codes and recovery hints
- Accessibility-first activation chain: uses pure accessibility API strategies before falling back to mouse events
- Deterministic element references (e.g.,
@e1,@e2) with optimistic re-identification across UI shifts - Progressive skeleton traversal: shallow tree first (depth ~3), annotated with
children_count, then drill-down into specific regions - Support for windows, menus, sheets, popovers, alerts, and notifications
- Special handling for Chromium/Electron accessibility trees to reduce noise
- C ABI via cdylib — can be loaded directly from Python, Swift, Go, Node, Ruby, or C without shelling out per command
Typical Workflow
For dense apps like Slack or VS Code, use progressive skeleton traversal to minimize token usage:
# 1. Shallow overview — depth-3 map, truncated containers show children_count
agent-desktop snapshot --skeleton --app Slack -i --compact
2. Drill into a region of interest (named containers get refs)
agent-desktop snapshot --root @e3 -i --compact
3. Act on an element found in the drill-down
agent-desktop click @e12
4. Re-drill the same region to verify state change
agent-desktop snapshot --root @e3 -i --compact
For simpler apps, a full snapshot works fine: agent-desktop snapshot --app Finder -i.
Installation
npm install -g agent-desktop
# Or use npx: npx agent-desktop snapshot --app Finder -i
# From source: cargo build --release
Performance Stats
In practice, the progressive skeleton approach reduced token usage by 78% to 96% compared to full-tree dumps in Electron apps like Slack, VS Code, and Notion. For example, Slack's full accessibility tree can exceed 50,000 tokens — impractical for most LLM contexts.
Who It's For
Developers building desktop agents, internal automation tools, or research prototypes who want to avoid the cost and fragility of screenshot-based control loops.
📖 Read the full source: HN AI Agents
👀 See Also

RouteLLM Setup for Cost-Effective AI Task Routing
A Reddit user shares a Docker Compose configuration that combines Ollama's local Qwen3.5:4b model with GitHub Copilot via OpenWire, using RouteLLM to route complex tasks to GPT-4o while handling simpler tasks locally.

YouTube Transcript MCP Improves Claude Research Workflow
A YouTube transcript MCP allows Claude to pull full transcripts with timestamps from YouTube links, eliminating manual tab switching and copy-pasting. The user reports significantly better answers when Claude has actual transcripts versus user summaries.

Benchmark Results: Claude Agent Swarm with Memory System Shows 30-43% Token Cost Savings
A developer tested a 6-agent Claude swarm on a 40-point coding task with and without a custom memory system called Stompy. Results show Sonnet 4.6 with memory achieved perfect scores at $3.98 vs $7.04 without, while Haiku 4.5 failed completely without memory but scored 39/40 with it.

Three Repositories for RAG and AI Agent Development
A Reddit post highlights three repositories for developers building with RAG and AI agents: memvid for agent memory, llama_index for RAG pipelines, and Continue for coding assistants. The author notes that pure RAG works best for knowledge retrieval, while memory systems are better for agents, with hybrid approaches being common in real tools.