Ask HN: If I cancel Codex today whats the next best local inference agent? better place to ask over /r/LocalLLaMA |
Ask HN: If I cancel Codex today whats the next best local inference agent? better place to ask over /r/LocalLLaMA |
I'm using qwen3.6 on a DGX spark, llama-cpp has prompt cache bugs for qwen/gemma models (among more being reported). Using my OpenCode-go sub when I want a bigger / more capable model
But if you're thinking about buying new hardware just for local inference: don't.
You'll spend more upfront for rapidly out of date and depreciating hardware while running inferior models compared to the latest from Anthropic and OpenAI.
Unless your requirement is prompt privacy, or you have idle hardware sitting around, the ROI just isn't there compared to a monthly subscription.