Most employees can't tell you what database to use, what software programming framework to use, what document management framework to use, but they are expected to know which of the 25 models available to use for a task, budget appropriately, monitor efficacy, update models to the most relevant for a task, continue to manage architecture patterns??? for LLM agents, this list goes on.
This is getting stupid folks.
Shouldn't software engineers have informed opinions on databases, programming languages, frameworks, etc.?
Seriously. Forget coding agents for a moment, and just consider OEMing a model as part of a more constrained machine learning application. By the time the data science team I was on had a solid understanding of GPT-4o's capabilities, strengths and weaknesses, and best practices for using it well, it was already into its deprecation period. Worst, most of our experimental results couldn't be replicated on any of the newer "long" term support models we had available to replace it. The relevant behaviors had all changed enough to force a considerable re-evaluation.
Combing back to coding agents, where they're releasing new models and harness tweaks multiple times per month, and vibes are the only - let's not say sensible, maybe realistic - thing a developer reasonably has to go on.
In practice, I think that kind of effort is mostly a hindrance at the moment, because of how fast things are moving, and because so much about this is subjective and everyone has different preferences.
"But models will get better". Maybe. But there has been a decline in the speed of development
On these things, yes. But in the case of LLMs, this is impossible. It's not possible to understand what the models are doing, and for several different reasons. You can at best evaluate them, like you do with your fellow engineers when hiring. But that's not a guarantee of anything.
And that's why I don't think this is an engineering renaissance. Engineering progress goes with better understanding of our tools, and adoption of more rigorous practices. LLMs go in the opposite direction.
If you enjoy being an expert in which LLM is best, then I'm all for it. However that isn't an interesting problem for me or my boss. I'm happy someone else figured it out so I can work on interesting problems.
The above likely scares all the model providers: they are a commodity with easially substitute competition. There is a minimum quality standard, but once you meet that there is nothing to differentiate you and so price matters and it becomes a race to the bottom.
All while suffering from the increased the mental load/lack of focus from even more workload. "You've got AI, you should be able to get it done today, right?"
Everyone has a great researcher piped to their desk now. You can have it follow the Twitter zeitgeist for you, you don't have to do it yourself.
The level of appropriate LLM usage in my opinion? Ask Gemini some questions when you need to search the web then read the sources it gives you.
LLM written code still has the problem IBM identified. A computer cannot be heald responsible, and so it cannot be allowed to make decisions. That applies to executive management AND software engineering.
Of course that requires working in an industry, like aerospace, where engineers are (usually, Boeing not counting) heald accountable.
You can ask it to implement the pattern you want. Like we had a bunch of gnarly intertwined logic inside of looping constructs that I could say, hey, create iterators for this, extract out the filtering logic, extract out the transformation logic, etc. I'm making the decision, but it can do the mechanical work. And again, I can tell it to benchmark my refactor to ensure that it actually cleaned up allocations instead of introducing them. It becomes easier for me to actually be more thorough in my work.
You can also use them to quickly build out prototypes for different approaches so that you can make better informed decisions at an architectural level.
You can give them a description of a bug and it will read through your code and figure out where the problem lies, even across multiple repos with complex interactions. You can of course still confirm it, but they've been superhuman at this since at least last December.
LLMs are absolutely a game changer for software development in the hands of somebody who knows what they want. For low-level things, they're an extremely solid junior (they don't ever really make mistakes, they just write tasteless crap). For high-level things, they're an extremely solid discussion partner with wide knowledge and great reasoning skills.
Right, the problem is frontier models can do ~all of that pretty well, a lot better than they could 9 months ago. Luna 6 can do most of it practically free and instantaneously. So, your quote is likely what's increasingly being said in c-suite meetings as they decide to start mass layoffs.
I believe that we're in a scenario where usage is unlikely to go down and neither are frontier AI costs.
I believe we'll see a shift to more organizations building their own harnesses with model routing logic to get central control over who can use what AI and for what.
A marketer doesn't need to default to Opus 5.5 to upload a blog article with MCP, which could be done by a model 10% of the price.
It continues to be an uphill battle to help people at $DAYJOB understand that the relationship between turns and cost is nonlinear. And many still don't understand the idea of a system prompt, that they can control how chatty all responses are.
Yeah, that guy you hired to write the prototype went on a 2 month bender and created 0 usable code. Did you budget for that? Oh, okay. The strategies to deal with spending on tokens are nothing compared to the overhead of managing actual people and their outputs.
If i had to guess upwards of 80 percent in corporate is "useless stuff".
Companies should handhold employees, establish clear SOPs, and train them on responsible and effective AI assisted coding.
Alternate take: it takes a lot of up front effort to ensure when you "max out tokens" the work is valuable.
The same thing goes for token use operating business processes (vs. building software that runs business processes): Let's imagine buying $250k tokens per year to operating some business processes - how much human effort is needed to ensure that level of spend is valuable? I'm using a number that could be "we could hire a skilled human at that level" as that changes the feel of the question.
The whole point of AI is for companies to invest less resources into employees. There's no way they start spending money on training people now, when that was already something that made executives roll their eyes before AI.
At this point, local models have become feasible, and people using them are beginning to get a feel for the tradeoffs vs the frontier models. As the frontier advances, the question becomes “how much will I spend for a given quantity and quality of AI work?” with local hardware providing a pricing anchor point.
When I can price hardware and ops for a given capability level - I have a budget again. From there it’s a question of how much faster/capable/cheaper is a given provider (and, you know, how much do I trust sending them all my IP?).
IMO, human in the loop is the only serious usage of AI (I know, I know, "software factories bro"). Everything else is a hope and a prayer and a big bill.
edit: grammar
I'm glad humans are freaking out when they realize they don't know if spending $10k/month on AI tokens is good or bad.
I'm glad humans are freaking out when an AI pumps out vapid presentations and other humans go ahead and present it to clients.
But.
Too many businesses have tolerated the same mindlessness when humans were in the place of LLMs.
Too many companies telling themselves and investors headcount growth is good without knowing that the new hires are actually doing.
Too many human-slop presentations float around, with the authors and audience just going through the motions.
Did digital photography raise the bar for what constitutes commercially valuable photos? I think so.
I hope AI will similarly raise the bar across all the industries it is touching.
Most of the work i've done in my career has been some random shit no one cared about.
1. choose a good model
2. choose the appropriate reasoning effort
3. choose a prompt to nudge it even further
Then it comes down to understanding the intuition of what kind of task deserves what effort?
I use these models all day long, experiment, and have no clue how to choose that.
It just showed up one day in the interface with no explanation or guidance. Its use is mysterious, its effects unclear except through intensive experimentation, and to this day it mostly seems "how many bad decisions will Claude go forward with when it finally dumps out screenfuls of text instead of getting better guidance early on" though it's certainly not a guarantee on anything.
These models are being released at breakneck speed even before their creators know how to use them. It's a big project of collective discovery to figure out what they are doing and how to use them.
(Couldn't even get this to happen in a 30 person team...)
This summer my team basically shut down when we hit a usage cap because we had recently pushed a decent portion of our devops workflow into agent skills. It had been done in such a way that it was difficult for humans to navigate. Relevant scripts turned out to be buggy and poorly documented, and nobody realized because these harnesses that are tuned to be absurdly tenacious about searching for workarounds had been quietly burning heaps of tokens on muddling through instead of raising any alerts about the horribly broken state of the system.
I do agree that home (or at least independently) grown harnesses might be the next logical step. Harness vendors who charge by the token have an inescapable conflict of interest here. Moderate, well-governed LLM usage isn't good for their revenue.
> I do agree that home (or at least independently) grown harnesses might be the next logical step. Harness vendors who charge by the token have an inescapable conflict of interest here. Moderate, well-governed LLM usage isn't good for their revenue.
It's true that moderate usage harms current harness providers, but I think that's because we currently conceptualize them as AI services. I think in the near future, companies will have harnesses that centrally configure MCPs, CLIs, model routing, etc.
and they'll be much closer to an auth/permission service than to an AI service.
After we started excessively documenting and extracting skills everyone has stopped complaining about running out of tokens, because agents stopped having to reconstruct the full context each time from scratch. Harnesses like claude code also push the model to aggressively keep this documentation in sync so there's little concern about drift.
The worst thing to do from a token usage pov is to give a model a vague open ended prompt because they are so scared to be wrong that they'll waste a ton of tokens "thinking" through the issue and verifying everything. Whereas they almost trust skills blindly and skip all this unnecessary work.
1. Truck drivers are not deciding to take a Ferrari instead of their truck. The company knows roughly how much gas a given delivery is going to consume and how much that is.
With AI, that definitely happens if you default to Claude Opus/GPT Sol.
2. A truck driver's value scales with gas usage. The further they drive, the more valuable.
Not true for tokens.
3. Gas doesn't create extreme outliers. One developer can easily consume 100x the tokens of another, which makes AI economics much more extreme.
Also this feels like a good way to detect runaway jobs. I’ve seen something similar implemented for data plans where you had unlimited data, but every 50 GB had to request another free 50 GB to prevent people from turning their phone plan into their home WiFi
so you basically are rationing your frontier tokens?
> I believe we'll see a shift to more organizations building their own harnesses with model routing logic to get central control over who can use what AI and for what.
This is what happened with cloud computing. People with expertise wrote software around the software creating machine because just giving naked compute to people en-masse didn’t actually achieve anything valuable. Likewise all of the talk of enterprise agents etc, they’re all custom software that wraps around a reasoning model.
Yeah exactly. Plus, most people don't want to configure MCPs and write skills about how to correctly use CLIs. The type of person who uses Hacker News may do that to customize whatever harness they use.
A sales manager at a sheet metal firm in Wisconsin doesn't. They want AI tools that make their lives easier, but they don't know or care whether that happens with an MCP.
I guess the cloud is a great analogy because that same sales manager may have loved Salesforce because it was always available and updated compared to a manual spreadsheet.
I have agents going pretty much nonstop for hobby projects. This is great value for me because I get to see a lot of things I'm interested in come to life. It's very token-inefficient though and if usage limits were reduced much this wouldn't be worth doing.
For my actual job though, I don't let agents go for hours on end and I care about the code. I spend between $1000 and $2000 per month there. Of course we pay API pricing. I wonder how much you're getting for that $20 in API pricing, it could be 10x or even 50x. At least on the max plans it can be that high if you are hitting usage limits, I don't know about the 20.
- full autonomous camp
- developer augmentation camp
- no-ai camp
The trouble I see with all of it is that the future seems unpredictable at the moment. The costs related to AI are low enough at the moment that full autonomous seems to be possible, but we have reasons to believe that costs will rise significantly, which may change that calculus. The no-AI camp is in ostrich mode, and is betting on this all going away once the bubble pops. The developer augmentation camp treats it like just another tool, which is somewhere in the middle.
The trouble is that even if there is a clear advantage today, the ground truth of costs built in is probably not stable.
https://www.reuters.com/business/media-telecom/musk-says-he-...
That will fix adoption of a broken technology and all debt issues!
If you don't have at least some workloads doing that you're missing out on the biggest wins from the current phase of the technology.
We're sorta in the age of alchemy. Lots of cranks out there, but there are real recipes.
Hah, welcome to what it can feel like to manage people.
That's not really possible with frontier models changing every 3-6 months, just like javascript frameworks there's not enough time to learn the ins and outs and they mostly expose the same interface, so unless you have rigorous evals you are working on vibes, and the amount of meetings I have had in the last 3 years of engineers confidently reporting on their vibes is SO TIRING.