RSS Feed

Day Week Month Year
All r/AI_Agents r/AI_Governance r/ClaudeAI r/ClaudeCode r/DeepSeek r/Futurology r/LangChain r/LocalLLaMA r/MachineLearning r/OpenAI r/artificial r/cybersecurity r/europe r/hermesagent r/netsec r/singularity r/technology

1209 items in r/LocalLLaMA

09.09.2026
r/LocalLLaMA MiMo-X-Pro-Preview and MiMo-X-Flash-Preview - New Mimo Model found in Mimo Desktop preview announcement πŸ”—
r/LocalLLaMA Is there a dummies guide for setting up qwen 27B with dflash2 and n-gram? πŸ”—
r/LocalLLaMA What OpenBMB 1B version is this? πŸ”—
r/LocalLLaMA US accuses Chinese AI firms of 'malicious' copying of AI technology πŸ”—
r/LocalLLaMA Qwen3.8-Flash-Next on MLX-serve, 1m context is released! πŸ”—
08.09.2026
r/LocalLLaMA A hilarious comment about llama.cpp: β€œIt’s a FB business using the pipeline to make profits” πŸ”—
r/LocalLLaMA Qwen 3.8 27b with PI agent - pushed to its 3D graphic game limits πŸ”—
r/LocalLLaMA Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines. πŸ”—
r/LocalLLaMA Qwen 3.8 27b with PI agent - pushed to its 3D graphic game limits πŸ”—
r/LocalLLaMA Qwen/Qwen-Drive-1.0-4B Β· Hugging Face πŸ”—
r/LocalLLaMA nex-agi/Nex-N2.5-mini - 35b πŸ”—
r/LocalLLaMA inclusionAI/Ling-3.0-flash-VL Β· Hugging Face πŸ”—
r/LocalLLaMA AA updated yet again, here's how the frontier ranks. πŸ”—
r/LocalLLaMA GPU guide (GB per dollar, bandwidth) πŸ”—
r/LocalLLaMA Qwen3-0.6B (400 MB) on a Samsung Note 8 (2017) phone drives a real desktop Chrome πŸ”—
r/LocalLLaMA OpenAI alleged of stealing mathematicians work πŸ”—
r/LocalLLaMA DeepSeek Flash 4.1 is already being tested via API and rolling out. πŸ”—
r/LocalLLaMA Fallout 2 x Fallout: Bakersfield x H3 as Interactive \ Reactive World Model, Let's go! πŸ”—
r/LocalLLaMA Which local model is actually good at knowing when to stop and ask you a question? πŸ”—
r/LocalLLaMA What are some practical tasks I can assign to my local AI models? πŸ”—
r/LocalLLaMA I made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop) πŸ”—
r/LocalLLaMA Voice conversations between Gemma4 12B and E2B on GPU and Jetson Orin πŸ”—
r/LocalLLaMA For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput πŸ”—
r/LocalLLaMA WSJ: Unregulated Open-Weight AI Is an Invitation to Disaster πŸ”—
07.09.2026
r/LocalLLaMA I made Warrior Quest, a local LLM-powered dark-fantasy RPG where the model only plays NPCs and the actual game state stays deterministic πŸ”—