The Information Machine
The edition

Thursday 27 August 2026

What moved

01
Day 2

Qwen3.8-Flash and Qwen4 architecture preview

  • Alibaba's Qwen team released Qwen3.8-Flash and Qwen3.8-Flash-Next on August 26 as open-weight previews of the Qwen4 architecture, with an FP8 variant also open-sourced at launch.
  • The 125B-parameter models activate only 6B per token, add Qwen Sparse Attention, gated residual connections, and a 51B N-gram embedding table that offloads to host RAM; Qwen reports training cost one-ninth that of Qwen3.7-Plus.
  • Same-day support arrived from SGLang, vLLM, UnslothAI, and TokenSpeed; the QwenCloud API starts at $0.16 per million input tokens, and the release topped Hacker News with 272 points.
The gist

The release combines a novel architecture preview of Qwen4 with competitive cost efficiency claims and broad same-day inference framework support, giving developers immediate access to weights and multiple deployment paths. The N-gram embedding design, if it delivers on the capacity-without-compute claim, represents a departure from standard MoE scaling approaches.

02
Day 5

Z.ai GLM-5.3 versus frontier coding models

  • Together AI published 904-rollout DeepSWE comparisons August 27 showing Z.ai's GLM-5.3, released August 14, tied Fable 5 on pass@1 and beat it on pass@4 at 5.4x lower cost; KingBench 3 placed it at 91.25%, ahead of Opus 5 and Kimi K3.
  • Bloomberg Intelligence analyst Robert Lea said Z.ai 'remains on a completely unsustainable commercial footing,' with shares down 9% on the launch day, and founder Tang Jie argued post-training is now the highest-value scaling dimension.
The gist

Independent benchmarks from Together AI place GLM-5.3 at cost parity or better with closed frontier models on multi-attempt coding tasks. GLM-5.3-Flash operates at 320B parameters on domestic Chinese chips, showing Z.ai can run large MoE models at scale without NVIDIA hardware.

03
Day 4

OpenAI's Jalapeño inference chip

  • Coverage following the August 25 benchmark release confirmed Jalapeño was built in partnership with Broadcom and fabricated by TSMC.
  • The chip handles inference only, leaving OpenAI still dependent on Nvidia for training workloads; a small batch of Jalapeño-powered systems is planned for 2026, more in 2027, and two additional chip generations are in development.
The gist

Jalapeño gives OpenAI a path to run inference at substantially lower power cost per token than Nvidia hardware, if the benchmarks hold. Its inference-only design leaves OpenAI dependent on Nvidia for model training.

04
Concluded today

Anthropic's Claude Code and Cowork updates

  • Anthropic on August 26 added a built-in browser to Cowork, its agentic task product, that opens automatically in the side panel when a task involves a website, letting Claude navigate pages, fill forms, and complete web-based jobs end-to-end.
  • Cowork is available on all paid plans on mobile and web.
The gist

The browser integration lets Claude complete web-based tasks from start to finish inside Cowork without a separate tool. The unified memory means context built up in chat conversations carries directly into agentic task execution.

3 sources
05
Day 6

NVIDIA Vera Rubin and NVLink Fusion platform

  • NVIDIA on August 27 announced NVHBM as part of NVLink Fusion, a custom memory architecture that embeds the memory controller in the HBM base die rather than on the XPU die, freeing silicon area for compute.
  • NVHBM delivers up to 30% greater bandwidth and 15% lower power than standard HBM4E while freeing up to 25% more XPU die area.
  • Amazon's Annapurna Labs became the first named partner, committing to support NVHBM in its Trainium4 chips.
The gist

Vera Rubin entering production at Azure means NVIDIA's next-generation GPU platform is deployed, not merely announced. NVHBM with Amazon as first partner extends NVLink Fusion's reach to a major cloud provider's custom silicon program.

06
Day 8

AI data center opposition across U.S. states

  • Kimmeridge estimated August 26 that up to 50% of planned US data centers could face delays or cancellations, naming power connections, permits, construction, and local approval as bottlenecks beyond chips and capital.
  • Texas has paused grid-connection approvals pending an audit and Pennsylvania now requires local sign-off before state permits can issue, while Kimmeridge cited opposition at 61% of Americans, up from 49% earlier this year.
  • Policy researcher Dean W.
  • Ball argued public debate remains stuck on whether to regulate AI at all, while specific state laws are already passing without useful public input.
The gist

A concrete estimate that half of planned US data center projects may be delayed or cancelled, combined with active state-level restrictions in New York, Texas, and Pennsylvania, creates real uncertainty for AI infrastructure expansion. The $3.2 trillion already committed by major tech firms as contractual obligations means companies cannot easily reverse course even as the regulatory and political environment tightens.

4 sources
07
Day 6

Gates' AI jobs and labor tax proposal

  • Goldman Sachs found this week that AI is already having a measurable negative effect on employment in developed economies, with call centers, software publishing, consulting, advertising, and entry-level work showing the clearest impact, corroborating a Gates essay published in late August arguing AI will cause permanent, not temporary, job displacement across sales, software engineering, finance, and healthcare.
  • Gates also proposed taxing AI tokens and robots to correct a tax asymmetry that incentivizes replacing human workers with machines, and set a ceiling of 40% for 'Human Reserved' jobs.
The gist

A prominent tech figure is publicly advocating for structural tax reform to slow AI-driven job displacement. Goldman Sachs data shows the employment effects Gates warns about are already measurable in developed economies.

08
Day 2

Google's Gemini 3.5 Transcribe

  • Google DeepMind launched Gemini 3.5 Transcribe on August 26 to replace Chirp 3, cutting transcription time by 70% and producing formatted output across 85+ languages.
  • Ars Technica measured a live-speech error rate of 5.5%, down from 7.32% for Chirp 3; Artificial Analysis found a Word Error Rate of 2.6% for non-streaming.
  • Developers access it through the Live API for real-time streaming capped at 10 minutes, and the Interactions API for pre-recorded audio up to one hour, with timestamps and speaker identification.
The gist

The model's accuracy gains and polished output bring speech recognition closer to specialized transcription tools, while broad API and AI Studio access makes it available to a wide developer base. Consumer rollout through Gboard and other Google surfaces extends its reach to general users.

09
Day 3

Apple M6 and M5 Ultra Macs

  • The Neuron Daily reported August 26 that the Mac mini M6 delivers on-device AI performance up to 4x faster than its predecessor, a concrete figure absent from Apple's initial coverage of the August 25 announcement introducing the M6 as the company's first 2nm chip in its Mac lineup.
The gist

The unified memory architecture, 2nm process, and high memory bandwidth make these machines capable of running large AI models locally without cloud infrastructure. The M5 Ultra's 512GB memory ceiling and multi-system clustering extend that to very large models.

10
Day 2

Perplexity and NVIDIA Portable Computer launch

  • Perplexity launched Portable Computer on August 25, a local AI agent product running entirely on NVIDIA DGX Spark hardware with no cloud dependency.
  • The full pipeline, including orchestrator LLM, subagent LLM, and agent harness, runs on-device using a 27B parameter model, meaning users pay no per-token costs.
  • Perplexity said its harness scores 82.6% on real knowledge work benchmarks, outperforming open-source harnesses Pi and Hermes, and that a post-trained PPLX 27B model reaches 85.4% on the same benchmarks.
The gist

Running a full AI agent pipeline on local hardware removes cloud dependency and eliminates per-token costs. Perplexity said the system scores 82.6% on real knowledge work benchmarks, which it said beats open-source alternatives Pi and Hermes.

11
Day 4

ChatGPT Work's remote-browser authentication

  • OpenAI launched ChatGPT Work on August 26, an autonomous agent that operates inside login-gated sites by pausing so the user signs in through an isolated remote browser; session cookies then persist for repeat tasks and the model never handles credentials, with screenshots disabled during sign-in.
  • Targeted tasks include booking appointments, managing insurance, and submitting business paperwork; usage is metered on the same structure as Codex, and Enterprise and Edu admins receive workspace spend controls in the Admin Console.
The gist

The feature lets an AI agent operate across the authenticated web, not just public sites, enabling multi-step tasks inside logged-in accounts. The credential-isolation architecture keeps passwords with the user while the agent retains task control, though website policies and human review remain external constraints.

12
Day 2

Google Gemini Enterprise for Legal

  • Google announced Gemini Enterprise for Legal on August 26, targeting law firms and in-house legal teams with a product that integrates with nine platforms, including iManage, Microsoft 365, DocuSign, and Harvey, via MCP connectors that inherit existing permissions.
  • The product includes reusable legal skills that learn firm-specific playbooks and citation rules, and can run agents for contract review, regulatory scanning, redaction, and drafting.
  • Weil Gotshal is among early users; Google said Legal is the first of planned industry-specific Gemini Enterprise solutions, with Financial Services also in the initial launch.
The gist

The product connects Google's AI to document management and productivity systems already in use across law firms, reducing the need to move documents into separate tools. Google is extending this vertical-specific approach beyond legal to financial services and additional industries.

13
Day 3

Anthropic's wellbeing grants and external Claude data access

  • Anthropic announced on August 26 a $5 million grant program for independent research into how AI affects user wellbeing, and for the first time gave external researchers access to real, privacy-preserved Claude usage data.
  • The grants include direct funding, model access, and technical support; grantees work independently and publish results as open-source projects any developer can use.
  • Anthropic said "to date, this work has only been possible within AI labs" and that "we can't tell the whole story alone, so we opened up our tools".
The gist

External researchers can now study AI's impacts using real usage data rather than relying on findings AI labs self-report. The open-source publication requirement means the resulting evaluations are available to any developer building AI systems.

What is moving now · Every edition · Every story

The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free