RSS Feed
All
r/AI_Agents
r/AI_Governance
r/ClaudeAI
r/ClaudeCode
r/DeepSeek
r/Futurology
r/LangChain
r/LocalLLaMA
r/MachineLearning
r/OpenAI
r/artificial
r/cybersecurity
r/europe
r/hermesagent
r/netsec
r/singularity
r/technology
1210 items in r/LocalLLaMA
04.09.2026
r/LocalLLaMA
I built a server with 768GB VRAM for frontier, but all new frontier open source models are likely to be two trillion or above now, including next GLM 6, am I cooked?
๐
r/LocalLLaMA
Model: add Tencent Hy 4 (hy_v4) preview architecture support by Little0o0 ยท Pull Request #28127 ยท ggml-org/llama.cpp
๐
r/LocalLLaMA
Even Qwen3.8 followed the instruction inside my translation data, and Gemma 4 beat the translation specialists I tested
๐
r/LocalLLaMA
K2-Horizon-MoVA-36B-A4B-MLX-4bit: up to 49.1 tok/s for local inference โ llm-bench.io
๐
r/LocalLLaMA
NVIDIA's $12,930,300,000.00 acquisition of Hugging Face contains an easter egg. The first 6 numbers of the acquisition price represent the decimal conversion of Unicode character U+1F917. The ๐ค emoji.
๐
03.09.2026
r/LocalLLaMA
I released sanoTTS: smallest complete TTS stack in 294k params (337 KB) that runs on $3 microcontroller and a 1.46m one that beats models 3x and 10x it's size
๐
r/LocalLLaMA
Increasing active parameters per token in MOE (Qwen 35B A4B+) reduce reasoning token by 8.5% - and you don't need to train or finetune!
๐
r/LocalLLaMA
Qwen3.8-27B at ~130 tok/s with full 262k context, on Kaggle's free TPU. OpenAI-compatible endpoint.
๐
r/LocalLLaMA
Feature/adaptive kv stream integration by giveen ยท Pull Request #326 ยท TheTom/llama-cpp-turboquant
๐
r/LocalLLaMA
Megathread for listing latest open source projects, research papers that are helping optimizations, efficiencies and accessibility to Open Source LLM and related hardware, software ?
๐