π± AI Briefing β 29.07.2026
Curated from 185 unread posts across 12 AI & tech subreddits
π Innovation
Kimi K3 Weights Land β 2.8 Trillion Parameters Now Open
Moonshot AI released the full open weights for Kimi K3 on July 27, making it the first open-source model to break the 3-trillion-parameter class. The model features Kimi Delta Attention (KDA) architecture, a 1-million-token context window, native vision support, and a highly sparse Mixture-of-Experts design. Early community tests show it running β slowly but successfully β on configurations ranging from dual RTX 6000 Pro workstations to patched laptop GPUs streaming weights off NVMe SSDs at 0.1 tok/s. Why it matters: Frontier-level performance is now available for self-hosting, closing the gap between open and proprietary models at a scale previously thought impossible outside megaclusters.
South Korea's A.X-K2 Debuts β Sovereign AI Goes Industrial
SK Telecom released A.X-K2, a 688-billion-parameter model built under South Korea's β©530 billion ($360M) sovereign AI project. The model delivers a 32.2% improvement over its predecessor using a novel Sparse Gate Attention (SGA) architecture and already has deployments locked in across manufacturing (KG Steel), defense (quantized models for the military), and biotech (cutting drug discovery timelines at SK Biopharmaceuticals). Why it matters: This is the clearest example yet of a nation-state treating AI infrastructure as strategic industrial policy rather than pure research β and shipping working models into production sectors.
Hermes Agent v0.19.0 "Quicksilver" β The Speed Spine
Nous Research shipped the Quicksilver release on July 20, with ~2,245 commits from over 450 community contributors. First-turn time-to-first-token dropped ~80% across all platforms, reasoning now streams live by default, and the desktop app got a 14Γ speedup on streaming markdown. New features include live subagent transcripts, Bitwarden/1Password integration, and durable delivery that survives gateway crashes. Why it matters: Agent infrastructure is maturing from "does it work" to "how fast and reliable is it" β the same transition that turned web apps from novelties into infrastructure.
π¬ Research
1.56TB MoE Model Runs on a 6GB Laptop GPU
A community member demonstrated running a 1.56TB Mixture-of-Experts checkpoint (93 layers, 896 experts/layer, MXFP4 quantization) on an RTX 4050 laptop with just 16GB RAM. By patching the runtime to stream non-cached dense weights directly off the NVMe alongside active experts, the setup achieved 0.106 tok/s β each token requiring ~33GB of disk I/O. Why it matters: This isn't about practicality; it's a proof point that extreme model compression and streaming architectures are viable. If a 1.56TB model can crawl on a gaming laptop today, tomorrow's hardware-aware inference stacks will make trillion-parameter local models routine.
Kimi K3's Architectural Innovations Set a New Scaling Baseline
Kimi K3 introduces Kimi Delta Attention (KDA), which decouples attention computation from sequence length to enable million-token contexts without quadratic compute blowup, and Attention Residuals (AttnRes), which selectively retrieve representations across depth rather than uniformly accumulating them. Combined with Stable LatentMoE, the architecture achieves near-GPT-5.6 and Fable-5 level benchmark performance with dramatically lower per-token cost. Why it matters: These aren't incremental tweaks β KDA was first published 9 months ago but only now deployed at scale, suggesting the research-to-production pipeline for architectural breakthroughs is accelerating.
GPT-5 Now Matched by Mid-Tier Open Models
A widely-noted observation this week: GPT-5, considered the world's best model just one year ago, now trails models like Qwen3.6 27B β a 27-billion-parameter open-weight model β on multiple benchmarks. Why it matters: The rate at which frontier capability diffuses downward is accelerating. What was a $200/month subscription model 12 months ago now runs on commodity hardware, fundamentally reshaping the economics of AI deployment.
π Security
Critical vBulletin RCE β CVSS 9.8, No Auth Required
CVE-2026-61511, disclosed July 27, allows unauthenticated remote code execution on vBulletin 5.x through 6.2.1 via an eval injection in the template runtime's runMaths() method. A public PoC is already available, and the vulnerability can be triggered through a publicly accessible AJAX endpoint with no credentials or user interaction. Why it matters: vBulletin powers thousands of forums worldwide, many running older versions. Previous vBulletin RCE vulnerabilities have been mass-exploited within days of disclosure β patch immediately if you self-host.
Volvo/Eicher Fleet Platform Hack β Full Vehicle Control Achieved
Security researcher Eaton Z revealed how API vulnerabilities in Volvo/Eicher's "My Eicher" fleet management platform allowed taking control of all registered commercial vehicles in India β including real-time tracking, live gauge clusters, and geofence manipulation. The vulnerability was reported in November 2025 but took 8 months to fully remediate. Why it matters: As vehicles become "datacenters on wheels," the attack surface grows exponentially. This isn't a theoretical threat β it affected real trucks and buses with real drivers, and the disclosure timeline shows how slowly OEMs move on security.
GitHub Posts $100K RCE Bounty; WPForms Flaw Found by 11 Researchers
GitHub issued a $100,000 bug bounty for critical remote code execution vulnerabilities, signaling how seriously platforms are taking supply-chain security. Separately, CVE-2026-4986 β an authentication bypass in WPForms PayPal Commerce webhooks β was independently discovered by at least 11 researchers before being patched, raising questions about whether duplicate report volume should accelerate vendor response timelines. Why it matters: When 11 researchers independently find the same auth bypass, it's likely attackers found it too. The security community is wrestling with how to prioritize fixes when multiple discoveries converge.
π° Market
Nvidia RTX Prices Surge Up to 30% β Third Hike This Year
Nvidia implemented its third price increase of 2026, pushing RTX 50-series cards up by 15β30%. The RTX 5090 now retails at ~$4,329 on Amazon (vs. a $1,999 MSRP), with premium AIB models exceeding $5,000. MSI called 2026 its "most difficult year," citing 20% supply cuts and GDDR7 memory shortages. Why it matters: Consumer GPU pricing is now structurally decoupled from gaming β AI demand from developers, researchers, and local model enthusiasts is the primary price driver, and there's no sign of relief through 2027.
SK Hynix Stock Crashes ~40% β Semiconductor Rout Deepens
SK Hynix shares plummeted 40% in 30 days, with trading halted on the Korean exchange on July 28 as the semiconductor selloff intensified. Samsung fell over 8% in the same session. The rout reflects growing investor anxiety about whether the AI capex boom can sustain current valuations, even as SK Hynix's CEO forecasts "the most severe supply shortage in 2027." Why it matters: This is a classic AI infrastructure paradox β insiders see insatiable demand for the next 18 months, while markets price in a potential bubble. LocalLLaMA community members are cautiously optimistic this could eventually mean cheaper RAM and GPUs.
South Korea Pours β©530 Billion into Sovereign AI Race
The Korean government is funding a multi-year sovereign AI competition where companies face elimination every 6 months. Three consortia remain in Phase 2 β SKT, LG AI Research, and Upstage β with the next cut coming in August 2026. Why it matters: Sovereign AI is becoming as much about economic competitiveness as national security. The model of government-backed, competition-driven development is being watched closely by the EU, Japan, and other regions considering similar programs.
ποΈ Politics
Zuckerberg Enters the AI Manifesto Wars β Full-Throated Pro-Diffusion
Mark Zuckerberg published a WSJ op-ed on July 28 arguing that AI should "primarily be understood as a tool for expanding individual agency β not as a force from which institutions must protect humanity." His position is the most pro-diffusion of four emerging AI policy stances, directly countering the "Pacing the Frontier" letter signed by 1,100+ employees and Dario Amodei's threshold-based restrictions. Why it matters: The AI policy landscape is crystallizing into four distinct camps, and Meta β with its open-weight Llama strategy β is anchoring the "accelerate diffusion, regulate concrete harms" position with real economic and political weight behind it.
AI Governance Tools Emerge as Compliance Becomes Real
Two new platforms surfaced this week: Maetra, which maintains a continuous compliance record across the AI lifecycle so regulatory changes don't invalidate past decisions, and HAIEC, a free tool that scans AI application websites for privacy, data usage, and safety disclosures. Why it matters: As AI regulations move from proposals to enforceable requirements (EU AI Act implementation, potential US executive actions), the compliance tooling ecosystem is forming. Early movers in governance infrastructure may define how the industry operationalizes transparency.
South Korea Deploys Sovereign AI to Defense Sector
SK Telecom signed an MOU with South Korea's Ministry of National Defense to deploy quantized versions of A.X K1 and K2 models for defense applications, including administration and operational support. This is the first publicly acknowledged instance of a nation-state explicitly integrating sovereign foundation models into military workflows. Why it matters: It crosses a symbolic line β sovereign AI isn't just for chatbots and industry anymore; it's entering the defense domain, with all the geopolitical implications that entails.
π Sources
- Kimi K3 β Anyone tried the Q1 yet?
- A.X-K2 Released
- Hermes Agent v0.19.0 Quicksilver Release
- 1.56TB MoE on 6GB RTX 4050 Laptop
- GPT-5 Now Inferior to Qwen3.6 27B
- vBulletin CVE-2026-61511 β Unauthenticated RCE
- Volvo/Eicher Fleet Management Exploit
- GitHub $100K RCE Bounty
- WPForms PayPal Webhook CVE-2026-4986
- Nvidia RTX GPU Price Hike Up to 30%
- SK Hynix Stock Crash
- Zuck: The AI Future Is for Everyone
- AI Governance β Is Your AI App Safe?
- Maetra β AI Governance & Compliance Platform
- SK Telecom Γ Defense MOU β Sovereign AI for Military
- Linus Torvalds: Linux Is Not Anti-AI
π Sources
- Is your AI App Safe ? How do you make AI APP ready for vendor onboarding.
- AI governance and compliance should remain connected as systems and regulations change
- Linus Torvalds Reaffirms That Linux Is Not "Anti-AI" And Not A "Social Warrior" Project
- Release Hermes Agent v0.19.0 (2026.7.20) β The Quicksilver Release
- Exploiting Volvo/Eicherβs fleet management platform to gain control over all users and vehicles
- New vBulletin Vulnerability!
- GitHub issues $100,000 bounty for critical RCE vulnerability
- I was reporter #11 for a WPForms PayPal webhook vulnerability (CVE-2026-4986)
- GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B and most todayβs low-tier models
- SK Hynix stock fell some 40% in the last 30 days, finally cheap RAM and GPUs again?
- Zuck's opinion: The AI Future Is for Everyone
- Nvidia is expected to raise GeForce RTX GPU prices again by up to 30%
- A.X-K2 released
- I tried running a 1.56TB MoE model on a 6GB RTX 4050 Laptop, Hereβs the result
- Anyone tried the Q1 Kimi K3 yet? (555GB)