This is a bit weird because it automatically changes the model depending on the amount of VRAM available, and some of the smaller models are more expensive (presumably because they're being provided via OpenRouter by someone with some GPUs in a colo or smaller providers). It also doesn't allow changing the tokens per second (my 5090 can get like 75-130 tokens per second [assuming I can fit the model in RAM]); which, then results in woefully under-estimated limit on how many tokens a day I can consume.
Some improvements that I think would make this more useful:
1. Allow manually setting tokens per second, or as an alternative, let me jack up the number of tokens a day.
2. A sort of backwards flow "if you want to run this, at X tokens per second, with Y context, you'd have to spend Z."
3. Add support for configuring multiple RTX 6000 variants.
When I was making heavy use of DeepSeekV4-pro I was burning somewhere around 1.5 billion tokens a month, and that was just using it in my free time on random projects. It was something like $24 at the time because of the initial discount/promo period. I don't think there's anyway in hell I could ever run that (on current hardware) for less money.
I think the calculator shows from a purely financial standpoint what we all know... that yeah, it's definitely not worth it if money is your only concern. That (cost per token) will eventually change. Models will get better, more efficient, VRAM prices will come down, VRAM capacity will rocket upwards, and the economics of it all will change. It would just be really cool to have the calculator show me exactly how cheap they'd have to get for it to make sense.
I need to finish up some work and make dinner, and if no one else beats me to it (anyone is welcome to) I'll ask Fable or Opus to knock that out.