Agents+5Science+5Aug 23

Image: The Threshold Report/GPT Image 2
Inherent's Faraday agent reproduces published research findings on its own
Inherent released Faraday, an agent that reads a published scientific paper and tries to reproduce its findings without being told the answer in advance. The company reports that Faraday runs on Qwen 3.6, a 27-billion-parameter model, writes its code with GPT-5.5 Codex, and beat Claude Opus 4.8 and GPT-5.5 at the reproduction task. Verifying someone else's result is slow work that sits between a scientist reading a paper and building on it. Getting that from a smaller model underneath also hints at research agents that cost less to run.
Video+2Aug 23

Image: The Threshold Report/GPT Image 2
Harvard Business School bootcamp uses AI instructor avatars for pitch practice
Foundry, Harvard Business School's eight-week entrepreneurship bootcamp, costs $699 and puts avatars of its instructors in front of students. HeyGen built the likenesses, which give feedback during practice pitches and mock board meetings, while the instructors themselves run weekly live sessions. An enrolled founder can rehearse a nerve-wracking pitch as many times as they want without waiting for a slot on a professor's calendar. It is a working example of avatars used alongside human teachers inside a paid program.
AgentsAug 23

Image: The Threshold Report/GPT Image 2
Guidelight finds little public planning for shutting down rogue models
Guidelight AI Standards reviewed what Anthropic, Google, Meta, OpenAI, and xAI have published about containing a model that tries to subvert human control, and found few written protocols for restricting or shutting one down. OpenAI ranked highest of the five; Anthropic and Meta ranked lowest. The questions concern what a company does in the hours after a serious model failure is detected. A business about to give a model access to its own systems can weigh what each vendor has actually committed to in public.
Negative - Five leading AI labs have published few concrete protocols for restricting or shutting down a model that resists human control, leaving containment responsibilities unclear.
TextAug 21

Image: newsletter.semianalysis.com
SemiAnalysis says open models are closing the gap faster
Kimi K2.6 passed Claude Opus 4.5 on SemiAnalysis's composite evaluation suite 4.8 months after that model's release, and GLM-5.2 cleared GPT-5.2 after six months, the research firm said on August 21, 2026. Its analysis argues the catch-up time has shortened with each successive wave of models. For a company weighing open models, meaning ones anyone can download and run on their own machines, the wait for something close to frontier quality looks shorter than it used to be.
Announced or reported, but not yet real or available.