The Information Machine
Concluded·following since 10 Aug 2026·Day 4·5 sources·updated 14 Aug 2026

Intology's Locus agent on PostTrainBench+

The gist

Intology's Locus agent scored 51.6% on PostTrainBench+, edging past the 51.1% human baseline using over 4,000 H100-hours.

AI Chat Daily describes this as the first time an automated agent has exceeded the human bar on a benchmark measuring AI ability to improve open-weight model performance. The compute requirement of over 4,000 H100-hours means the result is a proof-of-concept demonstration rather than a practical automated research workflow.

The full picture

Intology's Locus agent, running on Opus 5, scored 51.6% on PostTrainBench+, surpassing the 51.1% human baseline at the cost of over 4,000 H100 GPU-hours. On the standard PostTrainBench, Locus scored 44.7%, outperforming Opus 5 without the Locus harness (34.1%) and Fable 5 (41.8%). On PostTrainBench+, Locus also beat Opus 4.8 (44.3%) and GLM 5.2 (42.7%). The results were externally verified by the PostTrainBench authors and underwent contamination and cheating checks. Benchmark scores on this task have risen from 9.9% for Claude Sonnet 4.5 in September 2025 to 23.2% for Opus 4.6 in March 2026 to 44.7%. In a production deployment, Locus built a language model for no-code startup Bubble, achieving 2.8x lower error, 5.4x lower latency, and 105x lower cost. The Import AI author predicted the human baseline on PostTrainBench v1.1 will be exceeded before the end of 2026.

How it developed
14 August 2026

Intology published results on August 13 showing its Locus agent, running on Opus 5, scored 51.6% on PostTrainBench+, clearing the 51.1% human baseline at a cost of more than 4,000 H100-hours; on the standard benchmark, Locus scored 44.7% against 34.1% for Opus 5 without the harness.

A second source corroborated the result, added competitor scores (Opus 4.8 at 44.3% and GLM 5.2 at 42.7% on PostTrainBench+), and characterized the compute cost as proof-of-concept territory rather than a practical research workflow.

13 August 2026

ScientistOne paper proposes chain-of-evidence framework for human-level autonomous scientific research

10 August 2026

Import AI reported Locus scored 44.7% on PostTrainBench and 51.6% on PostTrainBench+, surpassing the 51.1% human baseline after over 4,000 H100 GPU-hours.

6 August 2026

Early case-study evidence on AI agents conducting open-ended research described as preliminary

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free