RSS Feed

Day Week Month Year
All r/AI_Agents r/AI_Governance r/ClaudeAI r/ClaudeCode r/DeepSeek r/Futurology r/LangChain r/LocalLLaMA r/MachineLearning r/OpenAI r/artificial r/cybersecurity r/europe r/hermesagent r/netsec r/singularity r/technology

1223 items in r/LocalLLaMA

16.08.2026
r/LocalLLaMA Qwen 3.8 27b vs 3.6 27b - how good is with a Turtle library. πŸ”—
r/LocalLLaMA Dario Amodei defends his policy proposals, warns open weights won't decentralize power, endorses pre-launch vetting, says real accomplishments will earn trust πŸ”—
r/LocalLLaMA Qwen 3.8 9b? πŸ”—
r/LocalLLaMA Why are RTX 6000 PROs still getting bought at 16000+ USD? And who are buying them? πŸ”—
r/LocalLLaMA Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72 πŸ”—
r/LocalLLaMA Qwen 3.8 distillations πŸ”—
r/LocalLLaMA Based on an accelerating frontier -> local trajectory, expect a ~30b param 'Mythos at home' by as soon as Jan 2027 (rationalisation below) πŸ”—
r/LocalLLaMA Let’s all thank Georgi Gerganov who gave use llama.cpp πŸ”—
r/LocalLLaMA Qwen3.8-27B Hybrid IQ4_XS quantization for 16GB gang πŸ”—
r/LocalLLaMA Newer commits removed the Qwen 35B πŸ”—
r/LocalLLaMA The dream is to reach 200GB VRAM πŸ”—
r/LocalLLaMA Qwen 3.8 27b with DSH(DeepSeek Harness) is Amazing!! Experiences so far and perfomance. πŸ”—
r/LocalLLaMA Genie-style playable world model running 720p at 16 FPS on a single 5090 in 19GB VRAM πŸ”—
r/LocalLLaMA Paper claims RL for reasoning only changes 1-3% of tokens, and they replicate the gains without RL at ~1000x less compute πŸ”—
r/LocalLLaMA Qwen3.8 27B reasoning effort low/medium/xhigh comparison πŸ”—
r/LocalLLaMA If you are at the lowest budget, which you can think of.Which hardware would you recommend to run? qwen 3.8 27b oWith like 50 tokens per second. I currently have a RTX 5070 Ti. πŸ”—
r/LocalLLaMA Qwen3.8-27B abliterated FP8: refusal 64–99% β†’ 0–6%, and MMLU/GSM8K move less than 1.3 points πŸ”—
r/LocalLLaMA How many people have 24gb over gpu here? πŸ”—
r/LocalLLaMA Show-off Saturday: Intel Arc B140 build. πŸ”—
r/LocalLLaMA Qwen3.8-27B vs Qwen3.6-27B writing ray-tracers in BASIC πŸ”—
15.08.2026
r/LocalLLaMA SOTA Apple Silicon Inference (August 15, 2026) πŸ”—
r/LocalLLaMA The perfect way for Google to screw over OAI and Anthropic is by releasing a 120B dense multimodal Gemma model πŸ”—
r/LocalLLaMA club-5060ti refresh: tested RTX 5060 Ti presets, a proper high-context harness, and Qwen3.8 27B πŸ”—
r/LocalLLaMA A day after Qwen3.8 release, OpenAI announces ads in their models πŸ”—
r/LocalLLaMA 880 tok/s on one 5090 Qwen3.8-27B in 4-bit NVFP4, full 262k context πŸ”—