1h agoHas anyone tried Qwen3.7 flash on openrouter? How does it compare to our Qwen 3.6 27B?1h ago·reddit.com
2h agoThose who use many layers in CPU/RAM and some in GPU - what are your specs and speeds?2h ago·reddit.com
3h agoI keep coming back to Qwen... Over and Over. Is there really nothing better under 120B?3h ago·reddit.com
4h agoThe idea: on a CPU the decode speed depends on the active params per token, not the total. My objective is trying to run a 10B at 100tok/s on a mid level PC (No GPU).4h ago·reddit.com
4h agoUnderstand Kimi K3 from first principles: a recommended order for anyone trying to understand this beast4h ago·reddit.com
4h agoA slide deck you can edit with a local model or in Chrome — the whole deck is a JSON block in one HTML file (~640KB with editor and viewer included)4h ago·reddit.com
5h agomodel: add NextN/MTP speculative decoding support for GLM_DSA (GLM-5.2)- #25980 MERGED!5h ago·reddit.com
8h agoI built a GBNF grammar compiler that makes 8B models reliably call tools - here's how it works (deep dive)8h ago·reddit.com
10h agoBuilt and released BetterGPT-150M – A compact 150M parameter completion model (+ live HF Space demo)10h ago·reddit.com
16h agoI tried running a 1.56TB MoE model on a 6GB RTX 4050 Laptop, Here’s the result16h ago·reddit.com