Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

you'd be surprised how good small models have gotten. Size of the model isnt all that matters.


My experience with qwen-3.6:35B-A3B reinforces this, gonna give this a spin when unsloth has quants available

Gemini flash was just as good as pro for most tasks with good prompts, tools, and context. Gemma 4 was nearly as good as flash and Qwen 3.6 appears to be even better.


> when unsloth has quants available

https://huggingface.co/unsloth/Qwen3.6-27B-GGUF


That was quick (compared to the 1T Kimi-2.6, not surprising)


Haha :) We had some issues with Kimi-2.6 since it was int4 and we were investigating how to handle it :)


Appreciate what y'all do! We were slacking about how many HGX-B300 it would take to run Kimi and it looks like we could actually fit 2-3 Kimis on a single HGX.


Sorry on the delay - oh haha that would be cool :) We did release 2bit dynamic ones, but unsure if they'll be helpful


> Size of the model isnt all that matters.

What matters is the motion in the tokens


Plus you can control thinking time a lot more, so when Anthropic lobotomizes Opus on you...




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: