RSS Feed

Day Week Month Year
All r/AI_Agents r/AI_Governance r/ClaudeAI r/ClaudeCode r/DeepSeek r/Futurology r/LangChain r/LocalLLaMA r/MachineLearning r/OpenAI r/artificial r/cybersecurity r/europe r/hermesagent r/netsec r/singularity r/technology

1216 items in r/LocalLLaMA

21.08.2026
r/LocalLLaMA Chinese Models? Someone's going to sleep on couch tonight... ๐Ÿ”—
r/LocalLLaMA Buying a V100/older NVIDIA GPU? Run this to check for older memory issues ๐Ÿ”—
r/LocalLLaMA When 9 of top 12 models are just exactly one. ๐Ÿ”—
r/LocalLLaMA Ornith-1.5-35B-A3B-NInfer - 250 tok/s, 5-8k prefill, 5090 ๐Ÿ”—
r/LocalLLaMA DeepSeek-V4-Flash-Vision-Exp ๐Ÿ”—
r/LocalLLaMA Fastest NVFP4 quant of Qwen3.8 27B out there ๐Ÿ”—
r/LocalLLaMA Qwen3.8-27B at 262K context on a Strix Halo + RTX 3090 Ti: 9.5 -> 153 tok/s, and it beats a dual-3090 vLLM box on HumanEval ๐Ÿ”—
r/LocalLLaMA Ox Alpha stealth model: GLM5 Air, Mimo V3 or ? ๐Ÿ”—
r/LocalLLaMA SenseNova U1.5-Lite full release: expert training, OPD distillation, one model at inference ๐Ÿ”—
r/LocalLLaMA I did it! I'm free! It's been 7 hours since I used claudecode ๐Ÿ”—
r/LocalLLaMA Qwen 3.8 27b - PI AGENT vs OPENCODE ๐Ÿ”—
20.08.2026
r/LocalLLaMA This benchmark is getting out of hand ๐Ÿ”—
r/LocalLLaMA Gonna be huge for US open source ๐Ÿ”—
r/LocalLLaMA If you are wondering why Ornith 1.5 35B A3B with MTP is so slow, this is why ๐Ÿ”—
r/LocalLLaMA NVIDIA dropped an NVIDIA-hosted CUDA MCP for AI-assisted CUDA operations, such as searching official, up-to-date documentation, writing optimized GPU code, and analyzing performance data ๐Ÿ”—
r/LocalLLaMA Any speculation on whether or not Google will announce a new Gemma model at the Gemma SF Celebration tonight? ๐Ÿ”—
r/LocalLLaMA AQuA's "self-improvement" updates research state, not the agent LM. What should a local port freeze? ๐Ÿ”—
r/LocalLLaMA Unsloth Dynamic 3.0 GGUFs ๐Ÿ”—
r/LocalLLaMA Ladies and gentlemen I present to you Qwen3.8 27b 1bit brain damage quant ๐Ÿ”—
r/LocalLLaMA Ling-3.0 released all 6 base checkpoints: 2 sizes ร— 3 stages ๐Ÿ”—
r/LocalLLaMA QwenMix-3.7: Kept seeing posts about Qwen3.8 and 3.6 sharing the same structure.. so I had Qwen3.8 combine them. ๐Ÿ”—
r/LocalLLaMA Getting better at coding doesn't make a model better at everything else ๐Ÿ”—
r/LocalLLaMA [MASSIVE TINY RELEASE] - Supra2-Medium-Base - a tiny 25M parameters model competing heavily with our previous 50M model! ๐Ÿ”—
r/LocalLLaMA The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches ๐Ÿ”—
r/LocalLLaMA Aurora-80K releases! A modern tiny language model. ๐Ÿ”—