RSS Feed

Day Week Month Year
All r/AI_Agents r/AI_Governance r/ClaudeAI r/ClaudeCode r/DeepSeek r/Futurology r/LangChain r/LocalLLaMA r/MachineLearning r/OpenAI r/artificial r/cybersecurity r/europe r/hermesagent r/netsec r/singularity r/technology

1209 items in r/LocalLLaMA

07.09.2026
r/LocalLLaMA when will open source LLM catch up to Astra I wonder? ๐Ÿ”—
r/LocalLLaMA New Benchmark: The Struggle Bench ๐Ÿ”—
06.09.2026
r/LocalLLaMA Trying to create my own server and consuming it for code with my phone remotely (Mac OS) ๐Ÿ”—
r/LocalLLaMA I built an LLM benchmark harness that lets you browse and compare how models answered each question ๐Ÿ”—
r/LocalLLaMA DeepSeek-V4-Flash-Vision Q8 vs Qwen3.8-Flash-Next Q8 ๐Ÿ”—
r/LocalLLaMA Expert expansion with llama.cpp ๐Ÿ”—
r/LocalLLaMA 2x R9700, 64 GB DDR5 is an absolute beast machine with vLLM Radiance / R9V and Qwen 3.8 27b and Flash next ๐Ÿ”—
r/LocalLLaMA Qwen 3.8-27B NVFP4 actually beats Q5_K_M and touches official BF16 levels, but only if you change 2 sampling params. Also almost 3x faster. ๐Ÿ”—
r/LocalLLaMA [Model] Support for Spark2_5ForCausalLM implementation by KnightYao ยท Pull Request #27868 ยท ggml-org/llama.cpp ๐Ÿ”—
r/LocalLLaMA Using GPT Astra to teach Qwen Next how to sculpt in 3D in Blender. ๐Ÿ”—
r/LocalLLaMA 8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics ๐Ÿ”—
r/LocalLLaMA Coding benchmarks that are quickly showcasing deep capability ๐Ÿ”—
r/LocalLLaMA vibeblending locally with Qwen 3.8 27B ๐Ÿ”—
r/LocalLLaMA Validate your local LLM advertised KV cache against real pressure; see exactly how old contexts get evicted from cache ๐Ÿ”—
r/LocalLLaMA Villager Simulation Game POC Created with Qwen3.8-27B-UD-Q3_K_XL.gguf - 16GB VRAM ๐Ÿ”—
r/LocalLLaMA Qwen3.8-Flash-Next-oQ4e-mtp: 45 tok/s on M4 Max, 25 tok/s on M2 Ultra for local inference โ€” llm-bench.io ๐Ÿ”—
r/LocalLLaMA Block KV cache streaming: bound VRAM at long context via a shared CUDA phase arena by giveen ยท Pull Request #357 ยท TheTom/llama-cpp-turboquant ๐Ÿ”—
r/LocalLLaMA Qwen3.8-27B "Unhacked" my PC ๐Ÿ”—
r/LocalLLaMA Qwen 3.8 Flash Next (Max) is impressive just to talk with. ๐Ÿ”—
05.09.2026
r/LocalLLaMA Which agent harness do you use and why? ๐Ÿ”—
r/LocalLLaMA Your opinion on Ling 3.0 tiny on CPU? ๐Ÿ”—
r/LocalLLaMA Otaku โ€” an LLM frontend ๐Ÿ”—
r/LocalLLaMA Qwen3.8 Flash Next - Templates Comparison ๐Ÿ”—
r/LocalLLaMA The gap has closed, open source will win ๐Ÿ”—
r/LocalLLaMA NInfer vs llama.cpp vs vLLM: quality + speed comparison for Qwen3.8-27B NVFP4 on RTX 5090 ๐Ÿ”—