Using Claude Code to Automate AI Research Experiments for 12 Hours

Automated AI Research with Claude Code
A developer documented using Claude Code to automate AI research experiments for 12 hours straight. The project focused on CLaaS, a real-time continual learning framework that moves context into weights using self-distillation.
Experimental Setup
The goal was to tune self-distillation training runs to maximize a model's compliance to different preference verifiers, such as concise responses and no emojis. Experiments ran locally on an RTX 5090 overnight.
System Architecture
The repository was built to be highly configurable:
- Every tunable parameter exposed via CLI using Hydra config management
- HTML dashboards for every training step and evaluation run
- Metrics, inputs, and outputs made observable through dashboards
- Claude Code could query dashboards via curl requests to check progress
Experiment Management
The workflow was controlled by a local EXPERIMENTS.md file with specific rules:
- Each experiment could change at most one variable or make one code change
- Between experiments, the model had to either accept or revert the previous change based on results
- Any new code changes had to be exposed via config for later tuning
- The model recorded all progress, hypotheses, and outcomes in the file as a running journal
- Used a "Ralph Wiggum loop" with the goal of maximizing preference compliance
Results
Over 12 hours, the system ran 9 experiments:
- Found and fixed a model collapse bug on the first run
- Tuned gradient steps per batch to 4
- Tuned learning rate to 3e-5
- Compliance improved from 0.000 to 1.000
- Token usage was surprisingly low because most time was spent waiting for training runs between experiments
The same task was also run with Codex for 2 hours using a plain prompt, and it independently converged on the same hyperparameters.
Project repository: https://github.com/kfallah/CLaaS
📖 Read the full source: r/ClaudeAI
👀 See Also

Field Report: AI Research Partner Fails Peer Review, Prompting Methodology Codification
A geologist/geophysicist using Claude Opus for complex multi-file projects discovered the AI produced a flawed critical analysis of an offshore wind study, with four of six points failing verification despite real citations. The user rebuilt the evidence and codified a methodology for future evaluations.

Multi-pane Claude Code setup with role separation and execution hooks
A developer shares a setup using four iTerm2 panes with separate Claude Code instances for implementation, auditing, planning, and prompt refinement, plus pre- and post-tool use hooks for safety and a session log for context retention.

Self-hosted OpenClaw AI agent creates passive accountability system for developers
A developer running OpenClaw on a Mac mini 24/7 reports the AI agent's persistent memory of tasks and projects creates an effective accountability system, helping complete projects that previously stalled.

Migrating from OpenClaw to Cowork + Claude Code: A Developer's Experience
A developer migrated from OpenClaw to Anthropic's Cowork with Claude Code sessions, citing better cron jobs, dispatch routing, and persistent memory. The setup uses a three-layer context design with Cowork handling orchestration and Claude Code executing code in repositories.