2026-09-14 · 07:08 (CEST)

AI Briefing — 2026-09-14

Top stories from 16 subreddits, 08.09. to 14.09.2026

🚀 Innovation

DeepSeek V4.1 Flash ships and immediately replaces V4 Pro. Released on 10.09., DeepSeek says the new Flash model beats V4 Pro across all key metrics: performance, cost, speed and task completion time. From 14.09. every request to V4 Pro gets silently routed to V4.1 Flash at Flash pricing until a V4.1 Pro arrives, a remarkably aggressive self-cannibalization. On r/LocalLLaMA the model also took first place on Artificial Analysis' new Intelligence Index v4.3 benchmark, ahead of Astra and Fable, barely a week after that index changed twice in three days.

Qwen3.8-27B cements its status as the local workhorse. Two posts show what the open 27B model does on a single RTX 3090: a complete game built in roughly five hours with two harness configurations, and another user reporting it one-shots vague coding prompts, writes its own tests and validates changes unprompted. The Terminal Bench v4 numbers back the hype, where it is the only small model that scores meaningfully at all.

A 1-bit 27B model now runs in the browser at 25 to 30 tok/s. A solo-built WebGPU inference engine (mentria.ai) runs Prism ML's natively 1-bit Bonsai-27B from a plain web page on a 6 GB RTX 3060 laptop: 27 billion parameters in 3.8 GB, about 1.14 bits per weight, no install and nothing leaves the machine. Local AI is moving from server-class rigs to any browser tab.

Read more →

🔬 Research

OpenAI claims the Navier-Stokes Millennium Prize, and the controversy is as big as the proof. On 08.09. OpenAI said an unreleased internal model, run as roughly 10,000 agents for a week, produced a 160-page proof of fluid blow-up for the Navier-Stokes equations. Twelve hours earlier, NYU's Tristan Buckmaster and Anthropic's Levent Alpöge had published related proofs, and Buckmaster now alleges his unpublished work reached OpenAI and that he was pressured to drop his Anthropic coauthor from a joint announcement. OpenAI admits it cannot rule out that de-identified usage data improved the model. The proof is public but independently unverified, and the episode is turning into a referendum on data ethics at frontier labs.

Terminal Bench v4 reshuffles the open model ranking. GLM-5.3 leads at 41.9 percent with GLM-5.3-Flash second at 32.8, ahead of DeepSeek V4.1 Flash (26.8) and Qwen3.8-Flash-Next (25.3). Kimi-K3 badly underperforms for its size, and Qwen3.8-27B is the only small model that scores at all. The community increasingly treats agentic terminal work as a better intelligence signal than static leaderboard evals.

Read more →

🔒 Security

One unauthenticated GitHub issue was enough to RCE all three big coding agents. Researchers showed that the default GitHub Actions configurations Anthropic, Google and OpenAI publish for their own agents (Claude Code, Gemini CLI and Codex) could all be tripped by a single GitHub issue, ending in remote code execution and access to CI runner secrets. Claude Code's bash validator stripped single-quoted content before checking it, Gemini CLI's tool allowlist was never enforced at runtime (Google rated it CVSS 10.0), and Codex had a two-pass workflow sharing one writable checkout. Anyone running coding agents in CI should review their workflow configs now.

Forgejo critical RCE, CVSS 9.9. All releases up to 16.0.3 allowed a malicious template repository to recreate a .git folder during template expansion, giving arbitrary file read and command execution on the Forgejo host. Fixed in 16.0.4 and 15.0.8, released on 11.09. Self-hosters should upgrade immediately, no in-the-wild exploitation is known yet.

LG smart TVs caught logging audio with the screen off. A post on r/cybersecurity drew wide attention to evidence that LG TVs record and transmit audio even when turned off via remote. It lands in the middle of a bad privacy news cycle for smart-TV vendors and renews the case for blocking IoT telemetry at the network level.

Read more →

💰 Market

The memory wall is now the binding constraint of the hardware race. A Hot Chips 2026 slide from Micron shows compute FLOPS rising about 3x every two years while HBM bandwidth grows under 2x, and the fixes under construction (memory beside compute, processing inside memory) are only starting to ship, with Samsung's LPDDR5X measuring 3.01x tokens/s on Llama 3.1 8B. Meanwhile RTX 5090 retail stock is nearly gone and prices keep climbing. For local inference, memory bandwidth and VRAM are becoming the scarce commodity, not FLOPS.

Read more →

🏛️ Politics

Trump rejects Silicon Valley's calls for an AI slowdown. The US president dismissed safety concerns ("It's going to be fine. We'll always have something to stop them") and framed the issue purely as great-power competition: whoever wins with AI wins. The posture lands the same week Dario Amodei warned that within 6 to 12 months an agent swarm could be capable of building a persistent internet-scale botnet, a concrete threat analysis given crypto laundering, darknet markets and rentable cloud compute already exist. The gap between lab warnings and Washington's deregulatory stance keeps widening.

King Charles convenes AI leaders as global risk warnings pile up. Buckingham Palace will host executives from Nvidia, Google DeepMind, OpenAI and Anthropic at Dumfries House this week, facilitated by the Ditchley Foundation, to discuss developing AI for "the good of humanity." It comes as the UN human rights chief warned AI could pose an existential risk, and as China's intelligence chief called AI a "new arena for strategic rivalry." Europe is positioning itself as the venue for AI governance while the US races ahead and China securitizes the field.

Read more →

📎 Sources

📎 Sources

← Back to Archive