RSS Feed

Day Week Month Year
All r/AI_Agents r/AI_Governance r/ClaudeAI r/ClaudeCode r/DeepSeek r/Futurology r/LangChain r/LocalLLaMA r/MachineLearning r/OpenAI r/artificial r/cybersecurity r/europe r/hermesagent r/netsec r/singularity r/technology

1210 items in r/LocalLLaMA

04.09.2026
r/LocalLLaMA I built a server with 768GB VRAM for frontier, but all new frontier open source models are likely to be two trillion or above now, including next GLM 6, am I cooked? ๐Ÿ”—
r/LocalLLaMA Model: add Tencent Hy 4 (hy_v4) preview architecture support by Little0o0 ยท Pull Request #28127 ยท ggml-org/llama.cpp ๐Ÿ”—
r/LocalLLaMA Qwen 3.8 Flash Next Can Build Funny Games ๐Ÿ”—
r/LocalLLaMA Even Qwen3.8 followed the instruction inside my translation data, and Gemma 4 beat the translation specialists I tested ๐Ÿ”—
r/LocalLLaMA MINISFORUM MS-S1 MAX-P495 ๐Ÿ”—
r/LocalLLaMA K2-Horizon-MoVA-36B-A4B-MLX-4bit: up to 49.1 tok/s for local inference โ€” llm-bench.io ๐Ÿ”—
r/LocalLLaMA NVIDIA's $12,930,300,000.00 acquisition of Hugging Face contains an easter egg. The first 6 numbers of the acquisition price represent the decimal conversion of Unicode character U+1F917. The ๐Ÿค— emoji. ๐Ÿ”—
r/LocalLLaMA HF easter egg ๐Ÿ”—
r/LocalLLaMA We open-sourced Paddock, our Rust/C++ inference engine with its own CUDA kernels (MIT/Apache-2.0) ๐Ÿ”—
r/LocalLLaMA On GPT-6 Astra 98.6% ARC AGI-3: don't fall for the hype ๐Ÿ”—
r/LocalLLaMA Has anyone already tried IFM's new K2-Horizon-MoVA-36B-A4B? ๐Ÿ”—
r/LocalLLaMA Spark-2.5-4B is an interesting model for 8GB Jetson Orin Nano Super SoC. ๐Ÿ”—
03.09.2026
r/LocalLLaMA Can the bubble pop please? ๐Ÿ”—
r/LocalLLaMA The benchmarks the big labs don't want you to see ๐Ÿ”—
r/LocalLLaMA I released sanoTTS: smallest complete TTS stack in 294k params (337 KB) that runs on $3 microcontroller and a 1.46m one that beats models 3x and 10x it's size ๐Ÿ”—
r/LocalLLaMA Increasing active parameters per token in MOE (Qwen 35B A4B+) reduce reasoning token by 8.5% - and you don't need to train or finetune! ๐Ÿ”—
r/LocalLLaMA Qwen3.8-27B at ~130 tok/s with full 262k context, on Kaggle's free TPU. OpenAI-compatible endpoint. ๐Ÿ”—
r/LocalLLaMA Need to decide: DGX spark vs framework desktop vs Mac mini/studio ๐Ÿ”—
r/LocalLLaMA Bernie Sanders proposes to ban AI ๐Ÿ”—
r/LocalLLaMA Feature/adaptive kv stream integration by giveen ยท Pull Request #326 ยท TheTom/llama-cpp-turboquant ๐Ÿ”—
r/LocalLLaMA Megathread for listing latest open source projects, research papers that are helping optimizations, efficiencies and accessibility to Open Source LLM and related hardware, software ? ๐Ÿ”—
r/LocalLLaMA Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 โ†’ 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070 ๐Ÿ”—
r/LocalLLaMA "ModelScope" Is a Hugging Face Alternative now that Nvidias deal is a Go ๐Ÿ”—
r/LocalLLaMA Ling-3.0-flash-Fin weights released ๐Ÿ”—
r/LocalLLaMA Apparently ChatGPT, Claude, and Grok were down ๐Ÿ”—