RSS Feed

Day Week Month Year
All r/AI_Agents r/AI_Governance r/ClaudeAI r/ClaudeCode r/DeepSeek r/Futurology r/LangChain r/LocalLLaMA r/MachineLearning r/OpenAI r/artificial r/cybersecurity r/europe r/hermesagent r/netsec r/singularity r/technology

1211 items in r/LocalLLaMA

01.09.2026
r/LocalLLaMA MTP released for Qwen3.8-Flash-Next-GGUF ๐Ÿ”—
r/LocalLLaMA A very confusing report from Puget Systems ๐Ÿ”—
r/LocalLLaMA Mac โ† USB-C cable โ†’ Linux box is becoming a thing. ๐Ÿ”—
r/LocalLLaMA Here is your chance to take over the world: glm 5.3 abliterated ๐Ÿ”—
r/LocalLLaMA I finished upcycling of gemma4-12B ๐Ÿ”—
31.08.2026
r/LocalLLaMA Deepseek v4 Flash Vision is out... ๐Ÿ”—
r/LocalLLaMA Don't sleep on Vision support for coding! ๐Ÿ”—
r/LocalLLaMA The state of open source LLM (08/31/2026) ๐Ÿ”—
r/LocalLLaMA Doesn't this look like NVIDIA is price fixing? ๐Ÿ”—
r/LocalLLaMA AVX2: Speed up large batch size prompt processing of IQ models by bartowski1182 ยท Pull Request #27402 ยท ggml-org/llama.cpp ๐Ÿ”—
r/LocalLLaMA GLM 5.3 and GLM 5.3 Flash ran locally on RTX PRO 6000 WS and built a penthouse using BlenderMCP ๐Ÿ”—
r/LocalLLaMA SlopTV: an infinite livestream of AI slop generated from youtube chat comments, Minimax H3 on 2x5090 ๐Ÿ”—
r/LocalLLaMA What are your hopes for the new Mistral? ๐Ÿ”—
r/LocalLLaMA First time running local models ๐Ÿ”—
r/LocalLLaMA The Chrono Trigger plot challenge - Crono awakens in his modest bedroom of 2095... ๐Ÿ”—
r/LocalLLaMA How bad do you think models like Qwen3.8-27B or GLM-5.3-Flash would be with H-Neurons disabled? ๐Ÿ”—
r/LocalLLaMA vote for the Qwen 3.8 ๐Ÿ”—
r/LocalLLaMA CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were restricted to 1 token by ynankani ยท Pull Request #27621 ยท ggml-org/llama.cpp ๐Ÿ”—
r/LocalLLaMA deepseek-ai/DeepSeek-V4-Flash-Vision-Exp ยท Hugging Face ๐Ÿ”—
r/LocalLLaMA pipecat-ai/phonellm-alpha-1: GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost ๐Ÿ”—
r/LocalLLaMA How I got Qwen 3.8 27b running at ~75t/s decode on 16GB RTX 5080 ๐Ÿ”—
r/LocalLLaMA Could this affect M5 Ultra price/availability? ๐Ÿ”—
r/LocalLLaMA Qwen3.8-Flash-Next-NVFP4 vs Qwen3.8-27B-FP Test Results ๐Ÿ”—
30.08.2026
r/LocalLLaMA I collected every single LLM coding benchmark, and computed their Intelligence Density ๐Ÿ”—
r/LocalLLaMA Qwen 3.8 27B - Fantastic German capabilities ๐Ÿ”—