RSS Feed

Day Week Month Year
All r/AI_Agents r/AI_Governance r/ClaudeAI r/ClaudeCode r/DeepSeek r/Futurology r/LangChain r/LocalLLaMA r/MachineLearning r/OpenAI r/artificial r/cybersecurity r/europe r/hermesagent r/netsec r/singularity r/technology

1212 items in r/LocalLLaMA

30.08.2026
r/LocalLLaMA Ran Qwen3.8-Flash-Next (79 GB, 2-bit) at 350K ctx for 3.5 hours on a 128 GB M5 Max โ€” speed vs context depth, 100 turns, one graph ๐Ÿ”—
29.08.2026
r/LocalLLaMA llama.cpp Open PRs list - CPU/RAM/Disk/Hybrid Related - Better for CPU-only & Hybrid inference ๐Ÿ”—
r/LocalLLaMA Tencent compressed Hy4-preview from 1.5TB to about 200GB GGUF and kept about 98% performance. ๐Ÿ”—
r/LocalLLaMA Someone tested various Models on the Political Compass test... ๐Ÿ”—
r/LocalLLaMA Exo labs claiming 4.8 tb/s memory bandwidth through m5u Mac Studio clustering ๐Ÿ”—
r/LocalLLaMA Qwen 3.8 27B at 50 tok/s with 100k Context on a 16GB GPU! (beellama.cpp) ๐Ÿ”—
r/LocalLLaMA How important is it for Chinese LLMs to reach the Opus 4.8 level? ๐Ÿ”—
r/LocalLLaMA If your t/s is low enough, you can see speculative decoding with your own eyes ๐Ÿ”—
r/LocalLLaMA Different Qwen thinking levels ๐Ÿ”—
r/LocalLLaMA I always wonder how much more speed and/or context they'd be getting.. ๐Ÿ”—
r/LocalLLaMA Terminal Bench 4.0 just dropped, GLM-5.3 is at the same level as Fable 5, accounting for margin of error ๐Ÿ”—
r/LocalLLaMA Saved my fiances phone with qwen 3.8 27b ๐Ÿ”—
r/LocalLLaMA Qwen 3.8 Flash Next ngram look up table offloaded to SSD and streamed in SGLang ๐Ÿ”—
28.08.2026
r/LocalLLaMA Is it worth running Qwen 3.8 Flash Next on 4x3090 vs 27B? ๐Ÿ”—
r/LocalLLaMA Today I hit 181 toks/s (aggregate) on Qwen3.8-Flash-Next on 2x DGX Sparks ๐Ÿ”—
r/LocalLLaMA [Release] SOTA GGUFs for Qwen3.8-27B: GSQ-RCO at 2.5 to 3.0 bpw ๐Ÿ”—
r/LocalLLaMA I audited 443 GGUF quants across 25 repos. 64 of them can't be the quant their filename claims. ๐Ÿ”—
r/LocalLLaMA Breeze-TTS-2 initial impressions: genuinely 'frontier' TTS ๐Ÿ”—
r/LocalLLaMA ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI ๐Ÿ”—
r/LocalLLaMA ds4 branch with GLM 5.3 Flash support ๐Ÿ”—
r/LocalLLaMA Qwen3.8-Flash on RTX3090 + 64GB RAM (but you only need 12GB VRAM) ๐Ÿ”—
r/LocalLLaMA zai-org/GLM-5.3 ยท Hugging Face ๐Ÿ”—
r/LocalLLaMA It's unbelievable! I used the mmap function in llama.cpp to fit Qwen3.8-Flash-Next IQ3_XSS into 16G+64G RAM, and the speed still reached 26t/s. ๐Ÿ”—
r/LocalLLaMA Qwen3.8-Flash-Next (UD-IQ4_XS) on 2x RTX 3060 + 7800X3D, from initial 36 tps prefill to 400 tps and other benchmarks (-sm tensor trap) + VRAM/RAM usage ๐Ÿ”—
r/LocalLLaMA Micron: HBM Requires Three Times More Wafer Area Than DDR5 ๐Ÿ”—