Newsletter post covers training AI systems to autonomously replicate scientific research findings.
Faraday beats Opus 4.8 and GPT-5.5 on scientific experiment replication
A 27B model that outperforms larger frontier models on scientific replication shows that task-specialized post-training can close capability gaps against much larger systems. The authors argue the skills Faraday uses to fill in experimental details are the same skills that could allow a model to autonomously advance research by designing its own experiments.
The full picture
Inherent, an AI startup, has built Faraday, a 27B supervisory model post-trained on Qwen-3.6-27B that directs a coding agent to replicate scientific experiments. Faraday outperforms Opus 4.8 and GPT-5.5 on 73% of in-distribution ML tasks and 60% of held-out AI-for-science tasks. The system uses OpenAI Codex as a coding tool and acts as a supervisory layer over larger frontier models. To train and evaluate Faraday, the team built Replica, a dataset of 310 replication tasks drawn from 100 ML and AI-for-science papers published between 1990 and 2026. Faraday was trained using a modified GRPO algorithm with rubric-based reward signals generated by Claude Opus 4.7.
How it developed
Import AI newsletter reports Inherent's Faraday model and its performance on the Replica benchmark.
Sources
Want this in your inbox?
I send a short email each morning with the stories that moved. If you would rather just read here, that works too.
Subscribe free