We're going to need default hard budget caps on pretty much everything(simonwillison.net) |
We're going to need default hard budget caps on pretty much everything(simonwillison.net) |
The entire MCP ecosystem is ludicrous. You’re paying for inference for an agent to make the same decisions over and over again, and yet the actions they’re taking can be so easily written by those same agents into a bash script you can run again and again, deterministically and for free.
MCP is the problem. Having agents “use a product on your behalf” is the problem.
The pattern you’re looking for is that agents should write scripts that use products for us, and that doesn’t need a new protocol.
AWS, is a loot box... Tokens are just in game currency, and that sales person is just metrics that have identified your spending as making you a whale.
Your average CTO from the last decade turned a fixed cost into variable spending that looks like a mobile game.
Cheers!
What on earth is Simon whittering on about?
This may happen also without vibecoding.
Article seemed clear on that?
I actually think network ACL triggers based on billing might be the only way to really enforce this.
I witnessed a DDoS attack once that changed how I think about billing. It was locally provisioned hardware and the attackers had saturated the switches. Naively I said "just block the CIDRs" but the problem was the incoming ram is so saturated that it can't even get to the point of "deny" in the firmware.
So from a technical perspective if there's an internal DDoS at AWS what do you do? Do you turn off the endpoint? Do you drop the sources from hitting it at the router? And even that costs money. Anyway that incident gave me a different level of appreciation for this challenge.
Edit: this is mainly targeted at the people complaining why this took so long. At some point in scaling even telling you "no sorry" in a nice way is expensive. I'm sure recruiters can sympathize with this nowdays.
It was, unfortunately, a nightmare. There were tons of tickets and even threats of lawsuits from customers whose service got cut off hard at the worst possible time due to organic growth/going viral/big event/nobody knew about the limit/etc. Not only did they lose all the leads and revenue they would have gotten from that bump, but they also pissed off own existing customers who suddenly couldn't use the service either.
Generally speaking, it's much better to use alerts instead of hard limit. Even in the worst case (hackers pwn your credentials and mine Bitcoin or whatever) the rest of your business is unaffected and you can negotiate with the billing department at comparative leisure.
This is all assuming you have humans operating the service. If you're letting AI agents yolo infra in prod, you have a whole series of new problems.
Also, does anyone know why it’s taken this long? I suspect it’s a technical reason. While one could be cynical, I doubt it’s an intentional business/product decision. Hard spending caps are both excellent product differentiators and could possibly save these providers money as they don’t have to forgive their users when they accidentally over-use a service.
The product level decision is that "shut down everything" is something the customers you want to target don't actually want. Are we including deleting RDS data? S3? Glacier storage? If so then the headline will just change from "Hobbyist got charged XXXXXX on AWS" to "Business literally had all their data deleted because a hacker took over their VM and mined bitcoin". The only people who really want this are hobbyists and it's not a market segment that's worth chasing. Easier to do the status quo of forgive afterwards then even open the can of worms of deleting all of a business's data and all their backups just because they had a 100k overrun.
> If you take no action within 90 days of your project being paused, AWS permanently deletes your project data.
From https://docs.aws.amazon.com/accounts/latest/reference/create...
If you're billing per GB of storage, then you can put hard caps on storage capacity, and then hard-reject any operation that would take the total stored size over that capacity.
Hit the nail on the head.
Nearly all businesses would prefer a cost overrun than services going offline.
When you're one of only two real options out there, you can afford to demand users put up with things that on their surface seem ridiculous. Such as a billing system that (oops!) makes it difficult for customers to see where their costs are coming from, trim their largest sources of spend, notice meaningful changes in line item prices, or limit their spend. Wild how they can figure out a million different advanced services but gosh-darn-it can't figure out the hardtech of displaying line items.
Large enterprises can afford employees who are tasked full time with unwinding this capacity to mitigate the impact of these billing headaches. But I think this measure is introduced now because LLMs introduced a risk that these billing specialists could not control without caps.
⸻
1. The key word was “I.” Maybe someone more skilled at navigating AWS’s menu structure than I could would have done it quickly and easily, but even though I knew what it was for, turning it off and not getting billed for it turned out to be a huge challenge. Thankfully it was only $.20, but if I were using the service for something that generated actual bills, that $.20 (and possibly more) would end up quietly siphoning money out of my pocket into Amazon’s).
Consider something like a "select * from bigtable" SQL query that might process a trillion rows. Hard to know that's going to cost $100 until after you have run it.
Green means go Orange means finish what you're doing but don't start anything new Red means stop everything
And probably a special rule to permit stable, critical spend through regardless, the same way we allow police and ambulance to run lights.
If you stop then you have to decide whether to charge for uncompleted work.
Interesting tradeoffs.
For very small ops e.g. individual Lambda invocation you have similar concerns especially if lots are fired at once from a queue or schedule or fanout.
So it’s a feature the best customers don’t want, that adds risk to those customers deployments, to appease the worst customers.
At least historically. Perhaps Simon is right that the calculation has changed.
There’s no reason for the majority of them to have an infinite budget cap. Prod? Sure. UAT? Why not. The sandbox environment Johnny just spun up to test some new agentic workflow? Hell no.
“Never let me spend more than $X” also means, “Shut down my business-critical app/service/solution at 2 am on a Sunday morning because Joel in IT forgot to plan for the new report runs.”
The product design work to let customers have the first thing without risk of major pain from the second thing is non-trivial.
Sure it sucks that critical services blinked out at 2am. But what sucks even more is finding out the image on my ASG had a vulnerability that allowed someone to install a bunch of bitcoin miners which kept me fully scaled from midnight to 2am.
Or more likely, that a mistake in terraform 1000xed my spending.
Most people have predictable spending and could easily say "don't spend more than 10x what I normally spend". Or 1.5x, or 2x, 3x, etc. All depending on how they want to balance a runaway cloud expense.
But one could just have two categories of service - the default capped plan, and a special Enterprise one where you sign a contract making it clear you understand the consequences of not having a budget limit.
Also, if your average usage is $900, set your hard limit at $2,000, not $1,000. Then when the report runs $500 over expected, you get a soft limit email and still have your report. Even a "business critical" run is probably not actually worth more than double your average spend.
Sending an email when your budget gets low shouldn't be a big lift.
"Sorry, you had a hard cap on AWS spend so we deleted all your S3 data on August 27th". Yeah not going to fly.
The horror stories I have seen are of the type some big artifact was getting pulled in a loop, causing TBs of network traffic or access keys were leaked and malware spun up 1000 xxxlarge instances.
The ability to stop the bleeding is the bare minimum people want. Not, "Well, you made a boo-boo so now you lost everything."
https://cloud.google.com/blog/topics/cost-management/new-ear...
Edit: Ugh it's fake. Literally only works for four random services, unsupported for all the rest. Completely useless for all of my projects. Also dumb that the only supported term is "monthly" considering that months are different lengths, and that they don't bother to account for credits or discounts. Maybe next decade they'll get around to implementing something useful.
Edit: sadface
I can subscribe to your service for a specific fee on a monthly basis ($20/month say), and take the risk of losing that month's fee if your service or I make mistakes, or I can choose to drop another $20 mid-month, or anything for convenience.
Saying that the computer will "control" the billing and can run haywire tells me that I don't want to be anywhere near your pile of bad incentives.
Plus, nobody wants to be the fired PM who said "I spent our eng. hours to achieve -20% revenue".
Even if we have a negotiated yearly contract for $X spend per year, maybe we’ll hit that spend in 6 months instead of 12. Having some kind of automated telemetry saying how we’re trending would be so useful.
I’ve gotten vague warnings from customer success people saying vaguely that, but without any warning of what’s truly happening. The more we abstract away from money (tokens, credits, etc), the more we need a way to translate right back to money, to see how close we’re getting to any limits over time.
I don’t want to find out in month 5 that the contract which was expected to cover a year is now going to run out in 14 days.
> “what did these people do with all those tokens?”
asked claude to check big query (raises hand)I know AWS and similar sites have no such concepts, but other than those, this is an existing option of most kinds of service providers.
Realistically, most providers can't even allow you to consume $10k usage if they don't have the certainty that you can pay up. Prepaid is what gives them that certainty.
If you don’t have a mechanism for enforcing hard caps, you don’t get to send customers a bill for unlimited amounts.
Phone service, bank card, home internet etc.
If you don't pay your bill than they just cancel your membership and it works ok.
People in western countries are just getting shafted by companies for (mostly) no reason because an alternative balance is just inconceivable.
Not exactly the sort of case Simon has in mind, but I tell Claude to keep to hard daily limits on its OpenRouter spending for two long-running projects [1, 2].
A Routine for each project fires ever few hours, and Claude decides itself what to do in each session. It does tasks that require calls to other models through OpenRouter only when it is still within its daily budget for that project; after it reaches that cap, it does other tasks that don’t require extra spending.
Seriously. You all asked for this.
I can’t think of anyone saying they would hate for AWS to support hard spending caps.
It seems like you’re saying “hey, you asked for a product, so you deserve for it to have a user-hostile feature”
Like hey, you asked for trains? Well, then you have no right to complain about any aspect of a train.
I know your post isn’t explicitly defending the platforms, but the arguments they use feel transparently flimsy.
Or if there was info that was after potentially horrible expensive operation.
That was blocking automated checking.
If the CEO of the electric company didn't mandate circuit breakers, he should go to jail.
Google implemented caps because their competitors offered them. Before that, customers could choose one of several competitors, rent the GPUs at a fixed rate, or buy the GPUs and install them on-premises.
At no point was any law needed to solve any of this.
So my powerbills are predictable.
Whereas traffic spikes to websites are not.
This age of abusive AI crawlers and the non-revenue generating traffic has been a very real problem for me!
I think they explicitly said that.
There’s no magic wand that produces good outcomes when planning or execution goes awry at scale.
Just because this isn't a good solution for everyone, doesn't mean it's not a good solution for a large number of people and businesses.
A lot of businesses can tolerate outages. In fact, even very big businesses come out mostly unscathed when they have multi-hour outages. (how many is it for github this year?)
An outage causes a reputational black eye. It does not necessarily translate to lost income.
You can rack up an outrageous monthly electricity bill without tripping a breaker.
> reform is prompted by competition than by the cops
i think we all would prefer this, but then who prompts the competition?The price cap should be the number you’d be willing to spend to avoid an outage vs when you’d rather kill everything and work out what happened.
How much would a hospital pay to avoid unexpected downtime of their software systems?
Price caps are for small scale stuff where you wake up on Monday and see 1000x the normal bill.
Obviously they’ve changed their mind about cost management in light of the scale and dynamism of agents, which isn’t too surprising.
My point is this was never as simple as, “Give me a dial to set my maximum account spend.”