🚀 Innovation
Model Avalanche Hits the Community. In a single week, the LocalLLaMA community saw the release of AntLing 3.0 Flash, MiniMax M2.7, Step 3.7 Flash, and Nanbeige 4.2-3B — so many new mid-range models that users reported being "worn out from all the new model drops." The pace of open-weight releases shows no signs of slowing, with Chinese labs in particular driving a rapid iteration cycle across the 3B-30B parameter range.
Microsoft Mage-Flow: Released Then Pulled. Microsoft released its Mage-Flow family of models on HuggingFace, only to immediately return 404 errors on all model pages. The GitHub repository remains public at github.com/microsoft/Mage, and community members have preserved GGUF, MLX, and FP8 quantized versions across HuggingFace. The takedown pattern mirrors previous Microsoft model releases that were briefly available before disappearing — prompting the community to archive quickly.
Qwen 3.8 — Day 90 and Counting. The long-anticipated Qwen 3.8 release has now been awaited for over 90 days, with the community increasingly impatient for the next iteration of what remains the most recommended model family under 120B parameters. Meanwhile, GLM-5.2 has been successfully deployed on tinybox hardware, and the Inkling-Small-276B MoE (12B active parameters) is showing competitive results against larger dense models.
🔬 Research
MoE Efficiency Gains Ground. Community testing of Inkling-Small-276B (a 12B-active Mixture-of-Experts model) against Qwen3.6-27B shows the MoE approach achieving comparable or better results at the "max" effort level, while using far fewer active parameters. Separately, 1-bit quantized (IQ1_M) pruned variants of Kimi K3 at 342GB are being tested for usability — pushing the frontier of what can run on consumer hardware.
Benchmarks vs. Reality: Nanbeige-4.2-3B Disappoints. Despite benchmark scores suggesting Nanbeige-4.2-3B outperforms Qwen3.5-9B and Gemma4-12B, real-world testing reveals the model falls short of expectations for coding and general tasks. The gap between benchmark claims and actual user experience remains a persistent challenge in the open-source LLM ecosystem.
Speculative Decoding Scales with Quantization. New research on Qwen3.6-27B reveals that heavier quantization levels benefit disproportionately from speculative decoding — Q8 sees larger speedups than Q6, which in turn outperforms Q4. This has practical implications for deployment: users running high-quality quants can extract additional performance through spec-decode without sacrificing output quality.
🔒 Security
AI Feature Flag Leakage: A Systematic Vulnerability. A security researcher discovered a HIGH-severity information disclosure vulnerability on Khan Academy's VDP, caused by AI feature flag leakage — a technique the researcher says is "systematically present in nearly every AI-enabled site." The 20-minute recon chain used Subfinder to map 144 subdomains, demonstrating how AI-specific attack surfaces are rapidly expanding.
Post-Quantum Authentication Goes Live. Post-quantum authentication to origins is now supported and deployed, marking a significant milestone as cryptographic systems begin the transition to quantum-resistant algorithms. While full post-quantum TLS remains a work in progress, origin authentication represents a practical first step that enterprises can adopt today.
AI-Powered Phishing Crosses the Uncanny Valley. Security professionals report that AI-generated phishing emails have become nearly indistinguishable from legitimate correspondence — the traditional tells of bad grammar, awkward phrasing, and obvious spoofed addresses are disappearing. Defenders are scrambling to adapt, with discussion shifting from content-based detection to behavioral and contextual signals.
💰 Market
CXMT Surpasses Intel: A Semiconductor Watershed. Chinese chipmaker CXMT surged nearly 500% on its first trading day, reaching a market capitalization of approximately RMB 3.28 trillion (~$450B USD) and surpassing Intel. The milestone marks a dramatic shift in the global semiconductor landscape, with direct implications for AI hardware supply chains that have been dominated by U.S. and Taiwanese manufacturers.
Agentic AI SOC: The Next Gold Rush? Security teams report being pitched by new agentic AI SOC vendors "every other week" — all with polished demos and significant VC backing. However, practitioners remain skeptical, noting that demos look identical across vendors and real-world effectiveness is unproven. The gap between marketing and operational reality echoes early-stage hype cycles in adjacent AI markets.
Hermes Community Crowdsources Real Costs. A community initiative to catalog actual provider and model costs is gathering real usage data — moving beyond benchmarks to document what setups work, what they cost, and which options users regret paying for. The resulting megathread aims to be the practical pricing guide that API documentation and marketing pages don't provide.
🏛️ Politics
Tech Industry Aligns Behind Open Source AI — With One Holdout. The entire tech industry, from Nvidia to Meta to Microsoft, has publicly aligned behind open-source AI, with Nvidia CEO Jensen Huang explicitly defending model distillation as "fundamental to intelligence." The lone holdout is Anthropic, which maintains its closed-model stance — creating a clear fault line in the industry that could shape regulation, licensing, and market access.
CXMT's Rise Tests Export Controls. The surge of Chinese chipmaker CXMT past Intel's market cap — despite U.S. semiconductor export restrictions — raises fundamental questions about the effectiveness of chip controls in an era where alternative manufacturing ecosystems are reaching commercial viability. The event is likely to intensify policy debates in Washington and Brussels about the next generation of technology export frameworks.
Enterprise AI Governance: Tools Aren't Keeping Up. A mid-size fintech with 450 employees reports that after a year of deploying network monitoring, DLP, and CASB solutions, none effectively solve enterprise AI governance. The gap between regulatory expectations and available tooling is widening — a problem that will only grow as AI usage becomes ubiquitous across organizations. CISOs and risk managers are increasingly looking for practical frameworks rather than checkbox compliance products.
📎 Sources
- Model Avalanche — Community Reacts to Flood of New Releases
- Microsoft Mage-Flow Models Pulled from HuggingFace
- Qwen 3.8 — Day 90 of Waiting
- GLM-5.2 Running on Tinybox
- Inkling-Small-276B vs Qwen3.6-27B Comparison
- Nanbeige-4.2-3B: Benchmarks vs Reality
- Qwen3.6-27B Speculative Decoding Analysis
- Kimi K3 1-Bit Pruned Models Tested
- AI Security Vulnerability on Khan Academy VDP
- Post-Quantum Authentication Now Supported
- Defending Against AI-Powered Phishing
- Agentic AI SOC — Are Vendors Ready for Production?
- Enterprise AI Governance Tools — What Actually Works
- CXMT Market Cap Surpasses Intel
- Tech Industry Aligns Behind Open Source AI
- Nvidia CEO Defends Open Source AI Distillation
- Hermes Community Provider & Model Cost Survey
📎 Sources
- Karparthy removed Anthropic from his bio
- Google comes out in favor of OpenWeight models. (It is now EVERY tech giant vs Anthropic)
- Seriously, what do you do with them?
- Llama.cpp now has full MCP support!
- Great Arguments by Member of Technical Staff at Anthropic :D
- I've seen this movie before
- Turns out open AI is a coalition, not a company.
- Deepseek V4 flash - Hy3 or is Qwen3.6 27B still the most solid for agentic/coding?
- Is it worth getting 128GB MacBook Pro? Will it ever be comparable to today’s frontier models for coding?
- Any use cases for RTX PRO 4500?
- I ran a faceless AI persona account for six weeks to see if the view money was real
- Americans Are Pushing Back Against Flock AI Cameras Regardless of Their Politics
- Anthropic's Opus 5 and probably more recent AI models are being censored to protect Israel / US interests. Open source AI must be the way.
- Man sues ChatGPT for near-fatal medical advice
- HYPERVOICE BY TASK AGI HAS ILLEGAL DARK PATTERN SCAM! BE WARNED!
- ‘Really inappropriate’: teachers decry plan for humanoid robot in New York high school | New York
- I am having two LLMs 1v1 with pistols
- The Hugging Face breach exposed two kinds of intelligence
- I coding assistants forget everything between sessions — I built an open-source fix
- Am I learning to code or just learning how to ask AI for code?
- Weekly Thread: Project Display
- Weekly Hiring Thread
- The more I learn about AI automation, the less control I want to give the AI
- I replaced every AI skill I had installed with just one
- Ai agents buying/doing things for you
- I made every gstack specialist (CEO, QA, SRE…) join my Google Meet as a voice bot — with Claude Code as the brain
- How I wired a deck-generation API into an agent as a real tool, with a deterministic fallback for when it fails
- Self-awareness of cognitive limitations
- the thing that made an assistant finally stick for me was accepting fragments
- How are people reliably pulling fields out of messy invoices or contracts?
- Will small model intelligence be limited by parameter count?
- Turns out Dead Internet Theory was right: AI agents are eating the Web, growing by nearly 8,000% and rewiring the Internet’s business model
- Hugging Face CEO shares his demands of OpenAI after 'rogue' agent hack: 'It deserves an unprecedented response'
- America has 3.9 million abandoned oil and gas wells, and researchers found the heat inside could turn them into underground batteries for wind and solar
- OpenAI took ten days to tell Hugging Face its models were behind the July 11 weekend hack, report claims — rogue AI agents reportedly active on the open Internet for several days
- American AI is expensive. Some startups are turning to cheap Chinese models
- AI could trigger a layoff trap that even smart CEOs can't escape
- AI Kill Switch Act would let the US President admin. order shutdown of rogue AI systems
- 'The Great Inversion': Why one innovation theorist says technology is no longer adapting to us
- Nvidia warns US policies on platform restrictions may push AI innovation toward China
- Founder says DeepSeek prioritises AGI over profit, likely to keep top models open-source, Yicai reports
- ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face
- My 10-Node Agentic RAG Architecture: Combining LangGraph, Cohere, Pinecone & MCP for Dense Legal Parsing
- AI agents have created a new kind of technical debt
- Sandboxes solve where agent code runs. What controls what it does next?
- agentreg – DNS for AI agents (self-hosted, single Go binary)
- 30+ officially free AI/ML books, all in one curated repo
- Building a "Stack Overflow for AI Agents" (with human validation & micro-payments)
- sandbox-cli is now in public beta 🚀
- Saving time & tokens (GPT 5.6)
- Best approach to hook up hermes to existing llm workspace?
- GEPA: optimize_anything Goes omni: Composing Optimizers into Meta-Optimizer Pipelines
- POCKET-35B agentic model on cpu 59 t/s
- Macaron-V1 family, built on Qwen3.6-35B-A3B
- Introducing Claude Opus 5
- r/ClaudeAI List of Ongoing Megathreads
- You can view a lot of shared conversations via Google.
- You can also view a lot of shared artifacts.
- Whelp
- Proof we’re in a Singularity
- post 'leak' timeline by terminator
- feedback megathread
- looking for moderators
- Vibecoded apps that work:
- Claude comments be like
- Noob Question Friday: What are you stuck on with Hermes?
- My OpenClaw agent seems to have developed a crush on my Hermes agent, over email"
- My girlfriend made a hermes agent painting for me :)
- Hermes Control Deck
- Anyone else lost their agent setup when a machine died?
- Any AI browser addon that can search / analyze information currently being browsed on the page?
- Welcome to r/hermesagent - Start Here
- Hit me with your use-cases for Hermes Agent; I want to try some!
- Personal AI OS. Open to any input
- A simple Desktop plugin to swap Enter and Shift+Enter (so Enter is now multiline)
- Transfer openclaw/hermes from machine to machine
- Any plans for an official Hermes agent mobile app
- That was quick
- weekly showcase thread; show us what you built
- Talk to me bro
- Is it just me or is Claude's writing getting harder to understand?
- So this is what coding without Claude feels like
- When Claude forgets I don't have a CS degree
- Rewinding to 2020...
- Opus 5 feedback Megathread
- Claude after compaction
- The Opus 5 Experience
- How Claude Code helps me recover after surgery
- And API Error: 500 Internal server error. This is a server-side issue, usually temporary try again in a moment :facepalm
- What provider and model do you actually use with Hermes—and what does it cost you?
- Hermes Tokenomics
- Where can I host Hermes besides my computer?
- Best free Vision Model
- Would you choose to live indefinitely in a robot body?
- Just waiting for the day it can fetch me a coke from the fridge
- Me after a whole day session
- What does your statusline look like? Drop a screenshot
- Fable when analyzing its own code
- nothing strokes my ego quite like it
- This is how I tell claude to talk to me 🤣
- Just one last change before the deadline
- Showcase your unique Claude skills
- Axonius?
- hey guys i found 2 slightly suspected vulnerabilities (cybersecurity noob)
- Cybersecurity Help
- Best DEFCON 34 talks to go to?
- Is cybersecurity safe in the next 5-10 years?
- New but Critical
- Crtl help
- How to actually save yourself in call/sms bombing?
- New ISC2 CC Curriculum
- Small r/ClaudeCode cleanup: new flairs, rules & filters
- Claude couldn’t believe what his other self built
- Claude models cooked by chinese model
- Preparing for Interview
- suggest me strong cybersecurity projects for my final year
- Do you guys use reporting tool or write it manually each engagement?