I'm not paying $20 for ChatGPT or Claude because a free local LLM does(xda-developers.com) |
I'm not paying $20 for ChatGPT or Claude because a free local LLM does(xda-developers.com) |
Im convinced AI is local now. Halogen loads QFN in 30 GB at 4bits. Its token gen is slightly faster than readable. You can follow along and leave it alone based on its thinking traces.
It gets 256k context but using dynamic context pruning, i got it to loop between 64k and 128k to maximize its speed vs prefill.
Its real. If the memory cartel is busted, the moat is busted.
"That stop being a complaint"? That's a strange construction.
I wonder if qwen 4 will make these smaller cards more viable by allowing the ngram storage to be hosted on CPU RAM.
Local model can be a local assistant, take care of your todos, calendar, private stuff in general. Things you shouldn't share with OpenAI or other Big Tech. Big Tech models for anything else.
I'm doing Qwen 3.6 35BA3B-NVFP4 (a larger ~20GB model) in 8GB VRAM and 20 GB RAM (I also get more than 16k context, and more t/s).
MoE models are kind of bonkers.
I cannot wait until we get "AI Appliance"-type things we can just plug in and use (aka a server). I know there are a few examples of this but few are plug-and-play currently. I'm thinking more "Here is your hardware Opus 5.5, it costs $X,XXX, it can do 100tps and X parallel streams with a 1M context window".
Maybe I'll change my tune when that's available due to the frontier being, potentially, still 6mo+ ahead with faster inference and larger context windows but I do pine for local inference for at least the "Execute the plan a bigger/better/smarter model wrote for you and the same model will check your work". Currently I'm trialling DS Flash 4.1 as my "worker" model and seeing positive results with it being wrapped by Opus 5.5 High.
Throughout the entire crypto-craze I'd been squeeking by gaming on an OC'd 1080 Ti, so when prices finally fell I was more than ready. Since I was getting interested in maybe playing around with AI soon, I spent up from my budget of around $550-$600 and am very glad I did since it's now clear I'll be squeeking by on this card for more years than I'd planned, just like the 1080 Ti.
Dude I can run Astra on my 20 dollar subscription. The leverage with that vs whatever Chinese model you can run in one GPU is not even comparable.
I almost always reach for the most performant model for any of my tasks because my return almost always pays off.