RSS Feed

Day Week Month Year
All r/AI_Agents r/AI_Governance r/ClaudeAI r/ClaudeCode r/DeepSeek r/Futurology r/LangChain r/LocalLLaMA r/MachineLearning r/OpenAI r/artificial r/cybersecurity r/europe r/hermesagent r/netsec r/singularity r/technology

1209 items in r/LocalLLaMA

07.09.2026
r/LocalLLaMA My Qwen3.8-27B task-aware quant reaches 99% of BF16 reasoning performance at 15% of the size. ๐Ÿ”—
r/LocalLLaMA ExLlamaV3 is underrated ๐Ÿ”—
r/LocalLLaMA Friends Don't Let Friends Use Ollama ๐Ÿ”—
r/LocalLLaMA Cybersecurity is local AI model's killer use case ๐Ÿ”—
r/LocalLLaMA DeepSeek-V4-Flash-Vision-Exp is amazing at creating game worlds! ๐Ÿ”—
r/LocalLLaMA I REALLY hope the new gemma 5 family sticks to the "chat model first" philsophy and doesn't fall into the Qwen trap ๐Ÿ”—
r/LocalLLaMA After over a year of my nights and weekends, the Jenny app is done! ๐Ÿ”—
r/LocalLLaMA Are you running Qwen 3.8 27b or Qwen Flash Next? ๐Ÿ”—
r/LocalLLaMA How to squeeze out every last drop of your precious RAM on your Mac - Use iPhone mirroring ๐Ÿ”—
r/LocalLLaMA Why are the SOTA open-weight models scoring (relatively) low scores on AA-Omniscience Index ๐Ÿ”—
r/LocalLLaMA MiniCPM5-2B Release Day ๐Ÿ”—
r/LocalLLaMA 9 easy steps for llama.cpp, a local model, Freecad (and pi coding agent) to generate solid objects that sound mechanically good and can be also be 3D printed/milled ๐Ÿ”—
r/LocalLLaMA little-coder vs just Pi ๐Ÿ”—
r/LocalLLaMA Easy local Copilot with VS Code and Lemonade ๐Ÿ”—
r/LocalLLaMA Qwen Next on 24 + 64 GB VRAM? ๐Ÿ”—
r/LocalLLaMA Best local models for hardware programming? ๐Ÿ”—
r/LocalLLaMA Which models are you running on 32Gb VRAM (16+16) and 128Gb RAM? ๐Ÿ”—
r/LocalLLaMA Me trying to keep up with all the new AI models being released ๐Ÿ”—
r/LocalLLaMA tencent/EVIE-8B and EVIE-4.5B (High-Capacity Visual Document Retrieval) ๐Ÿ”—
r/LocalLLaMA Dual R9700 on Asus X570 VIII Motherboard ๐Ÿ”—
r/LocalLLaMA Qwen 3.8 Next Flash is really really REALLY verbose.. ๐Ÿ”—
r/LocalLLaMA Bifurcation and riser cables suggestions ๐Ÿ”—
r/LocalLLaMA Thinking about grabbing an RTX 2000 Ada 16gb to add to my gaming pc for inference due to Wattage constraints, any advice? ๐Ÿ”—
r/LocalLLaMA Lit Review on Benchmarking LLMs Running in your phone!: MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments ๐Ÿ”—
r/LocalLLaMA Benchmarking calories evaluation with LLMs ๐Ÿ”—