RSS Feed

Day Week Month Year
All r/AI_Agents r/AI_Governance r/ClaudeAI r/ClaudeCode r/DeepSeek r/Futurology r/LangChain r/LocalLLaMA r/MachineLearning r/OpenAI r/artificial r/cybersecurity r/europe r/hermesagent r/netsec r/singularity r/technology

1251 items in r/LocalLLaMA

02.08.2026
r/LocalLLaMA All Qwen model oneshots: 1109 outputs to look at and compare! ๐Ÿ”—
r/LocalLLaMA Are you ready for Le Chaton FAT or still wasting money on GPUs? ๐Ÿ”—
r/LocalLLaMA Deepseek v4 flash - 100-150 faster t/s in prefill/pp. ๐Ÿ”—
r/LocalLLaMA Deepseek-V4-Flash-0731 Dwarfstar on Mac ๐Ÿ”—
r/LocalLLaMA Conclusion: r/LocalLLaMA still has brilliant open-weight research, but finding it requires wading through endless benchmark drama, non-local Discussion Points and repetitive hardware flexes. ๐Ÿ”—
r/LocalLLaMA Has Qwen 3.8 has dropped yet? Day 90... ๐Ÿ”—
r/LocalLLaMA llama.cpp just added MTP / DSpark support for DeepSeek V4 Flash ๐Ÿ”—
r/LocalLLaMA Vacuum 16T ๐Ÿ”—
r/LocalLLaMA Xberg v1 is out ๐Ÿ”—
r/LocalLLaMA [Release] WinterMix โ€” Qwen3.5-122B-A10B in native MLX: an 82 GiB build that beats 94โ€“95 GiB quants, plus a 68 GiB build for agent swarms ๐Ÿ”—
r/LocalLLaMA Setting up of a 16xGB10 (DGX Spark) cluster ๐Ÿ”—
r/LocalLLaMA DeepSeek-V4-Flash 284B on 5.3GB of memory ๐Ÿ”—
r/LocalLLaMA PSA for DeepSeek-V4-Flash-0731 users โ€” don't blow out your prompt cache with system role messages mid-conversation ๐Ÿ”—
r/LocalLLaMA I pushed Kimi K3 onto one CPU with 8 GB of RAM ๐Ÿ”—
r/LocalLLaMA DeepSeek-V4-Flash-0731 UD-Q8_K_XL 17.20~ t/s on A6000 + 256GB DDR4 ๐Ÿ”—
r/LocalLLaMA Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TG ๐Ÿ”—
r/LocalLLaMA Expert-only IQ3 requant of DeepSeek-V4-Flash-0731: better KLD than UD-IQ3_S, 1.4x decode on a CPU-spill rig ๐Ÿ”—
r/LocalLLaMA Why are almost all new benchmarks and leaderboards coding focused? ๐Ÿ”—
01.08.2026
r/LocalLLaMA Koboldcpp v1.118 released ๐Ÿ”—
r/LocalLLaMA DeepSeek-V4-Flash-0731 UD-IQ3_S 12.5 tok/s on RTX 3090 +128GB DDR5 ๐Ÿ”—
r/LocalLLaMA Is there a point where models just cannot get any smaller without losing intelligence? ๐Ÿ”—
r/LocalLLaMA Fix for Deep Seek v4 Flash 0731 tool calling has been added to llama cpp ๐Ÿ”—
r/LocalLLaMA I've had ling-3.0-flash and glm-5.2 both in my executor slot for a few weeks. They don't split the way the benchmarks predict ๐Ÿ”—
r/LocalLLaMA Deepseek v4 flash 0731 still not holding up. ๐Ÿ”—
r/LocalLLaMA DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM โ‰ˆ 3.5 tok/s. ๐Ÿ”—