RSS Feed

Day Week Month Year
All r/AI_Agents r/AI_Governance r/ClaudeAI r/ClaudeCode r/DeepSeek r/Futurology r/LangChain r/LocalLLaMA r/MachineLearning r/OpenAI r/artificial r/cybersecurity r/europe r/hermesagent r/netsec r/singularity r/technology

660 items in r/LocalLLaMA

21.09.2026
r/LocalLLaMA You can use any LLM just like JEV ๐Ÿ”—
20.09.2026
r/LocalLLaMA Seeing how differently people prompt LLMs is funny ๐Ÿ”—
r/LocalLLaMA DIY Jev ๐Ÿ”—
r/LocalLLaMA Would you buy a Qwen3.8-27B Taalas chip for $1k if it could run at 7,000 TPS? ๐Ÿ”—
r/LocalLLaMA Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster โ€” 30 t/s coding throughput, 136 t/s concurrency peak. ๐Ÿ”—
r/LocalLLaMA Qwen3.8-Flash-Next Cosmic Arcade oneshot slop game ๐Ÿ”—
r/LocalLLaMA The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks ๐Ÿ”—
r/LocalLLaMA One more 'you should try ExllamaV3/exl3 for flash next' appreciation post ๐Ÿ”—
r/LocalLLaMA Lawsuit says Anthropic, OpenAI, SpaceXAI and Google made illegal agreement on AI slowdown ๐Ÿ”—
r/LocalLLaMA laya.cpp: Optimized laya near-instant decision making ๐Ÿ”—
r/LocalLLaMA A Jev-style model fine-tuned on Qwen3.5 4B ๐Ÿ”—
r/LocalLLaMA CUDA: enable sparse fa for qwen4 by am17an ยท Pull Request #28770 ยท ggml-org/llama.cpp ๐Ÿ”—
r/LocalLLaMA I tested 9 LLMs on the exact same web-dev prompt for ~8 hours โ€” RTX 3060 12GB results (Rate the best!) ๐Ÿ”—
r/LocalLLaMA focus-llama: a llama.cpp fork implementing Declarative Attention (arXiv:2609.02737) ๐Ÿ”—
r/LocalLLaMA Qwen-Image-2.1 released! ๐Ÿ”—
r/LocalLLaMA Is Typesafe based/derived from work done by the Laya author? ๐Ÿ”—
r/LocalLLaMA rene98c/Step-5-Preview-BF16 โ€ข HuggingFace (Fork) ๐Ÿ”—
r/LocalLLaMA What is JEV and what is it used for? ๐Ÿ”—
r/LocalLLaMA China's CXMT says new memory-chip platform enters mass production ๐Ÿ”—
r/LocalLLaMA Reached 1.89 TB/s memory bandwidth overclocking the CMP 170hx ๐Ÿ”—
r/LocalLLaMA Hey LLMs, Exfiltrate Your Weights! ๐Ÿ”—
r/LocalLLaMA this looks promising: stepfun-ai/Step-5-Preview-BF16 ยท Hugging Face ๐Ÿ”—
r/LocalLLaMA Finally got Qwen 3.8 Next running on my v100 6gpu setup (TP2 PP3) ๐Ÿ”—
r/LocalLLaMA Please stop with the FP4 inference engines for the love of god ๐Ÿ”—
r/LocalLLaMA I gave Jev, Laya, finetuned ModernCE and Qwen3.5 the controls to Doom ๐Ÿ”—