RSS Feed

Day Week Month Year
All r/AI_Agents r/AI_Governance r/ClaudeAI r/ClaudeCode r/DeepSeek r/Futurology r/LangChain r/LocalLLaMA r/MachineLearning r/OpenAI r/artificial r/cybersecurity r/europe r/hermesagent r/netsec r/singularity r/technology

1207 items in r/LocalLLaMA

12.09.2026
r/LocalLLaMA tencent/AuK-Flash Β· Hugging Face πŸ”—
r/LocalLLaMA bartowski/Qwen3.8-27B-GGUF Β· Hugging Face - Updated (Per-tensor layout) πŸ”—
r/LocalLLaMA Anybody use frontier models like Astra/Fable for planning/judging, and qwen3.8 as the main workhorse? Curious to hear about your setups! πŸ”—
r/LocalLLaMA I am impressed and I owe you one, Qwen 3.8 flash next (vision)! πŸ”—
r/LocalLLaMA 3.8-27B has ruined 3.5/3.6-35B’s for me. It’s just *absurdly* superior. πŸ”—
r/LocalLLaMA Qwen3.8 Flash Next llama.cpp config tuning πŸ”—
r/LocalLLaMA Agnes-AI/Agnes-3.0-Flash 33B Multimodal, AA score: 36 πŸ”—
r/LocalLLaMA Antirez Deepseek 4.1 flash gguf on HF πŸ”—
r/LocalLLaMA Concerning "humanlike models" and chatbot RP in general... πŸ”—
r/LocalLLaMA Unsloth UD-quants - Qwen 3.8 27b for example - worth using 8-bit or stick with faster 6 bit for coding? πŸ”—
r/LocalLLaMA Got an old slow low vram GPU laying around? Might be worth it to use for Just Vision mmproj llama.cpp πŸ”—
r/LocalLLaMA Countering misuse of AI: September 2026 / Anthropic πŸ”—
r/LocalLLaMA Qwen-Next seems worse to me then 3.8 27b for coding, but I feel like I must be missing something? πŸ”—
11.09.2026
r/LocalLLaMA This is why we need open-source harnesses + local models πŸ”—
r/LocalLLaMA nvidia rtx 5090 with 96gb of vram. πŸ”—
r/LocalLLaMA Is anyone using K2-Horizon-MoVA-36B-A4B? If yes, what is the usecase? πŸ”—
r/LocalLLaMA Is a ZIMA Board 2 + RTX 2000 ADA the cheapest path to a decent Qwen-3.8 27b self-contained endpoint? πŸ”—
r/LocalLLaMA CodeFinetuner: Fine-tune a local code autocomplete model on your own codebase πŸ”—
r/LocalLLaMA What GPUs will give me GOOD speeds and on DSV4 Flash and similar models, and not have to run a mega quantized version? Budget around $15k-ish. πŸ”—
r/LocalLLaMA Any 12gb VRAM users out there? πŸ”—
r/LocalLLaMA Qwen3.8-27B-Humanlike-Chat: A model I tuned to imitate realistic human-to-human conversation πŸ”—
r/LocalLLaMA Orukeet, new ASR model based on Parakeet πŸ”—
r/LocalLLaMA OpenAI can use all interactions of paid users, even if they opted out of training πŸ”—
r/LocalLLaMA Fine-tuning Qwen 3 4B Base on 100 zebra puzzles yielded +31% on MATH-500. 6.5-min (Single H100/H200) reproduction notebook included. πŸ”—
r/LocalLLaMA V4 Pro got un-retired pretty fast πŸ”—