π± AI Briefing β 01.08.2026
Curated from 12 subreddits Β· 273 posts scanned Β· top 14 stories synthesized
π Innovation
DeepSeek V4 Flash API officially released in public beta, delivering benchmark scores that exceed the V4 Pro Preview β including 82.7 on Terminal Bench 2.1 for agentic tasks. The model is already being tested in production coding workflows and on GitHub Copilot, with users reporting performance competitive with frontier models from just months ago. At 18Γ cheaper input and 28Γ cheaper output than Claude Opus 4.8, it's reshaping the cost-performance curve for API inference.
LG AI Research releases K-EXAONE 2.0, a 750B-parameter model with 37B active parameters under the Apache 2.0 license. Developed under Korea's Sovereign AI Foundation Model Project (Phase 2), the model is 3Γ larger than its predecessor and supports 10 languages, marking a significant national AI capability milestone.
LongCat-Flash-Lite-Sparse weights now available, replacing dense Multi-head Latent Attention with Sparse Attention for improved long-context efficiency. Meanwhile, audio.cpp 0.5 shipped with DramaBox expressive TTS β prompt-controlled voice acting with emotion, speed, and character control β plus Confucius4 cross-lingual voice transfer and 7 additional models.
π¬ Research
DeepSeek V4 Flash 0731 scores 50 on the Intelligence Index, matching the top frontier model score of 51 from March 2026. This means models you can run locally today are now at the intelligence level of the best proprietary models from just five months ago β a striking compression of the capability gap between open and closed AI.
GLM 5.2 gains vision capabilities through a community merge of the Kimi K2.6 vision encoder, addressing what was considered its primary limitation for multimodal agent tasks. Early reports suggest strong visual understanding, opening the door to fully local multimodal agents.
Weight-Aware Streaming Tensor Engine demonstrates running Kimi K3 on just 29 GB of RAM at 0.50 tok/s β slow but functional. For comparison, community members are reporting 200 tps prompt processing and 11 tps generation on 4Γ RTX 5060 Ti setups for DeepSeek V4 Flash. The gap between bleeding-edge quantization research and practical local deployment keeps narrowing.
π Security
Over 30 Minnesota water systems hit by hackers in a 48-hour window, forcing emergency response coordination across the state. The coordinated attack on critical infrastructure highlights the increasing tempo and ambition of attacks on operational technology (OT) targets.
Amgen discloses a cloud data breach exposing patient health and proprietary information, adding to a growing list of pharmaceutical and healthcare sector breaches. Cloud misconfiguration and third-party access remain the dominant attack vectors.
Attackers now using real Microsoft sign-in screens for phishing, leveraging adversary-in-the-middle (AitM) techniques that proxy authentication in real time. Every screen the victim sees is genuine Microsoft infrastructure, making detection nearly impossible without hardware-backed MFA. Separately, a researcher published a systematic methodology for finding AI-specific security vulnerabilities, demonstrated on Khan Academy's bug bounty program.
π° Market
DeepSeek's API pricing at 28Γ cheaper output than Claude Opus 4.8 is forcing a reckoning in the inference market. With DeepSeek V4 Flash matching Opus 4.8 on key benchmarks at a fraction of the cost, the premium pricing model of frontier labs faces growing pressure β especially as local deployment of competitive models becomes viable on consumer hardware.
Codex recorded 12 quota resets in July alone, making the $20/month Max plan a standout value for token-heavy development workflows. Users running dual Claude + Codex subscriptions are increasingly comparing reset cadences as a key metric for subscription value. Meanwhile, enterprise SIEM costs under scrutiny: Splunk license renewals are becoming hard to justify against managed alternatives, with mid-size security teams actively shopping for replacements.
ποΈ Politics
The EU AI Act takes effect tomorrow, August 2, 2026, mandating labeling of all AI-generated images, audio, video, and text across the European Union. The broad scope of the labeling requirement is raising practical questions about implementation, especially for text β where watermarking remains technically unreliable.
Momentum for banning open-weight AI models is building in Washington, fueled by DeepSeek V4 Flash's release and its demonstrated competitiveness with US frontier models. The debate frames open-weight releases as a national security concern, while advocates argue that restricting open models would consolidate power among a handful of corporate labs and accelerate Chinese leadership in accessible AI. The coincidence of the EU AI Act and US regulatory pushes creates a defining regulatory moment for open-source AI.
π Sources
- DeepSeek V4 Flash Update β r/DeepSeek
- DeepSeek 28Γ cheaper than Claude Opus 4.8 β r/DeepSeek
- DeepSeek V4 Flash matches March 2026 frontier β r/LocalLLaMA
- LG AI K-EXAONE 2.0 750B released β r/LocalLLaMA
- LongCat-Flash-Lite-Sparse weights β r/LocalLLaMA
- audio.cpp 0.5: DramaBox TTS β r/LocalLLaMA
- GLM 5.2 with vision β r/LocalLLaMA
- Streaming Tensor Engine: Kimi K3 on 29GB RAM β r/LocalLLaMA
- 30+ Minnesota water systems hacked β r/cybersecurity
- Amgen cloud data breach β r/cybersecurity
- Attackers using real Microsoft sign-in for phishing β r/cybersecurity
- AI security vulnerability found on Khan Academy β r/cybersecurity
- Codex 12 resets in July β r/ClaudeCode
- Splunk costs hard to justify β r/cybersecurity
- EU AI Act takes effect Aug 2 β r/LocalLLaMA
- BAN OPEN WEIGHTS debate β r/DeepSeek
- DeepSeek V4 Flash + open source ban push β r/DeepSeek
π Sources
- Seriously, what do you do with them?
- The Hugging Face breach exposed two kinds of intelligence
- I coding assistants forget everything between sessions β I built an open-source fix
- Am I learning to code or just learning how to ask AI for code?
- Weekly Hiring Thread
- I made every gstack specialist (CEO, QA, SREβ¦) join my Google Meet as a voice bot β with Claude Code as the brain
- How I wired a deck-generation API into an agent as a real tool, with a deterministic fallback for when it fails
- America has 3.9 million abandoned oil and gas wells, and researchers found the heat inside could turn them into underground batteries for wind and solar
- AI could trigger a layoff trap that even smart CEOs can't escape