Using Claude Code to Automate AI Research Experiments for 12 Hours

✍️ OpenClawRadar📅 Published: February 26, 2026🔗 Source
Using Claude Code to Automate AI Research Experiments for 12 Hours
Ad

Automated AI Research with Claude Code

A developer documented using Claude Code to automate AI research experiments for 12 hours straight. The project focused on CLaaS, a real-time continual learning framework that moves context into weights using self-distillation.

Experimental Setup

The goal was to tune self-distillation training runs to maximize a model's compliance to different preference verifiers, such as concise responses and no emojis. Experiments ran locally on an RTX 5090 overnight.

System Architecture

The repository was built to be highly configurable:

  • Every tunable parameter exposed via CLI using Hydra config management
  • HTML dashboards for every training step and evaluation run
  • Metrics, inputs, and outputs made observable through dashboards
  • Claude Code could query dashboards via curl requests to check progress
Ad

Experiment Management

The workflow was controlled by a local EXPERIMENTS.md file with specific rules:

  • Each experiment could change at most one variable or make one code change
  • Between experiments, the model had to either accept or revert the previous change based on results
  • Any new code changes had to be exposed via config for later tuning
  • The model recorded all progress, hypotheses, and outcomes in the file as a running journal
  • Used a "Ralph Wiggum loop" with the goal of maximizing preference compliance

Results

Over 12 hours, the system ran 9 experiments:

  • Found and fixed a model collapse bug on the first run
  • Tuned gradient steps per batch to 4
  • Tuned learning rate to 3e-5
  • Compliance improved from 0.000 to 1.000
  • Token usage was surprisingly low because most time was spent waiting for training runs between experiments

The same task was also run with Codex for 2 hours using a plain prompt, and it independently converged on the same hyperparameters.

Project repository: https://github.com/kfallah/CLaaS

📖 Read the full source: r/ClaudeAI

Ad

👀 See Also

Field Report: AI Research Partner Fails Peer Review, Prompting Methodology Codification
Use Cases

Field Report: AI Research Partner Fails Peer Review, Prompting Methodology Codification

A geologist/geophysicist using Claude Opus for complex multi-file projects discovered the AI produced a flawed critical analysis of an offshore wind study, with four of six points failing verification despite real citations. The user rebuilt the evidence and codified a methodology for future evaluations.

OpenClawRadar
Multi-pane Claude Code setup with role separation and execution hooks
Use Cases

Multi-pane Claude Code setup with role separation and execution hooks

A developer shares a setup using four iTerm2 panes with separate Claude Code instances for implementation, auditing, planning, and prompt refinement, plus pre- and post-tool use hooks for safety and a session log for context retention.

OpenClawRadar
Self-hosted OpenClaw AI agent creates passive accountability system for developers
Use Cases

Self-hosted OpenClaw AI agent creates passive accountability system for developers

A developer running OpenClaw on a Mac mini 24/7 reports the AI agent's persistent memory of tasks and projects creates an effective accountability system, helping complete projects that previously stalled.

OpenClawRadar
Migrating from OpenClaw to Cowork + Claude Code: A Developer's Experience
Use Cases

Migrating from OpenClaw to Cowork + Claude Code: A Developer's Experience

A developer migrated from OpenClaw to Anthropic's Cowork with Claude Code sessions, citing better cron jobs, dispatch routing, and persistent memory. The setup uses a three-layer context design with Cowork handling orchestration and Claude Code executing code in repositories.

OpenClawRadar