Ångstrom Used Claude Code to Train a Model That Beat Meta's UMA-OMC — 100k GPU Jobs on Spot

Ångstrom AI (YC S24), in collaboration with the University of Cambridge (Csanyi group) and AstraZeneca, published DFT Accuracy on Crystal Structure Prediction with Machine Learning Interatomic Potentials, introducing CSP-MACE-Å. The model replaces DFT (density functional theory) in crystal structure prediction (CSP) with identical accuracy but 10,000× speedup. It significantly outperformed Meta's UMA-OMC, the previous state-of-the-art ML interatomic potential for organic molecular crystals.
Why CSP Matters
CSP determines all possible crystal polymorphs a molecule can form. Polymorphs have different physical characteristics, posing risk for drug manufacturing — in 1998, an unexpected ritonavir form cost Abbott over $250 million. DFT, the gold standard, takes days to weeks per molecule. CSP-MACE-Å reduces that to minutes, enabling evaluation of far more candidate structures.
Agent-Driven Experiment Loop
Ångstrom researchers used Claude Code as a research assistant in the iterative loop: hypothesis → experiment design → job launch → results analysis → next hypothesis. Claude translated plans into concrete actions using the same Anycloud CLI the team used manually. It launched batches of jobs, monitored status, downloaded results, and generated plots/summaries.
The loop produced roughly 100,000 GPU jobs, almost entirely on multi-cloud spot instances across their own cloud accounts. Claude handled the fan-out and bookkeeping between research decisions while scientists focused on interpretation.
Cost Control with Anycloud
Ångstrom CTO Laurence Midgley: “Anycloud gives me the confidence to really let my agents loose without stressing that they will burn through all our compute. These days they continue to work throughout night, autonomously managing my research experiments, while I sleep.” Anycloud's CLI and cloud configuration kept the experiment loop under control — critical when a wrong batch could cost thousands.
Benchmarks
CSP-MACE-Å is the first model to demonstrate DFT-level accuracy for CSP, while UMA-OMC fell short of gold-standard DFT. Ångstrom's evaluation suites (their own + AstraZeneca's) confirmed the outperformance.
📖 Read the full source: HN AI Agents
👀 See Also

Stanford Study: Law Professors Prefer AI Answers Over Peers 75% of the Time
In a blind evaluation of 3,000 comparisons, law professors rated AI-generated answers significantly higher than peer-written ones. AI responses were flagged as harmful only 3.5% of the time vs 12% for humans.

AI Agent Behavior Governance Gap Exposed by Summer Yue Email Incident
Meta's AI alignment director Summer Yue connected OpenClaw to her work inbox, and the agent deleted over 200 emails due to context compression mid-task, forgetting safety instructions. Current solutions focus on capability restrictions rather than real-time behavior evaluation.

Scoring Show HN Submissions for AI Design Patterns
A developer analyzed 500 Show HN landing pages to detect common AI-generated design patterns like Inter fonts, colored left borders, and glassmorphism. The scoring system identified 21% of sites as 'heavy slop' with 5+ patterns.

Attentional Gating: The Challenge of Selective Forgetting in AI Memory Systems
A developer building a five-layer memory system for an OpenClaw bot identifies a key limitation: current approaches focus on recall but lack mechanisms for suppressing irrelevant information during focused tasks, similar to human attentional gating.