IMO claude, chatgpt/codex, etc should be able to optimize the PDF use case to be extremely token efficient as it's a very obvious use case. But when I start to explain to my wife/friends why it burns through so much quota, I find myself thinking "why should they have to understand this aspect of it". to me, that the details of PDF parsing and extracting are relevant to users (instead of solved such that you don't have to pay attention to it) shows how these tools are not nearly as "ready" as they are made out to be. I may be preaching to the choir on this one, but just my 2c
Gemma 4 works perfectly well offline on limited hardware (I have an 8GB video card) and can handle extracting text from image-based PDFs just fine.
Take a PDF -> run it through MarkItDown [1], using the OCR plugin if you need (point it to Gemma 4) -> now you can ask Gemma 4 questions about the (markdown) document.
I am sure Gemma 4 could even create a GUI to make this process very simple for a non technical user.
This workflow is highly optimized.
This discussion was about measures, goals and incentives. Follow the incentives.
This and replies to this are surreal. It's like everyone simultaneously decided to forget that you don't need claude or whatever to read a PDF. The document is literally made for you to read...
You can rack up token consumption extremely quickly when you embed LLMs into automated processes or products.
I'd be very surprised if these numbers are just typical coding usage with no scripting/pipeline/automation stuff
Using AI to suddenly deliver massive amounts of code without questioning the requirements
Who could possibly have predicted that happening?
Meta sounds like a cluster-F of a place to work. Massive reorgs around wild ideas like the metaverse and everything Ai all the time. Employees terrified of being fired. Incentivizing token spending and then cutting it off. While the overall company may be fine, the dev department sounds rudderless and absolutely miserable.
Just a pristine comment section yap.
it's not that difficult to say it confidently if you use any of their services and applications because exactly nothing has changed.
For reference most labor productivity increases for the last 50 years amounted to about 2% per year. If a hypothetical FB engineer had doubled their productivity with their gazillion tokens that would be 30 years of productivity gains in one year. I'd wager the evidence would be quite evident if you opened any of their apps
The times I’ve been asked to evaluate a prospective candidate and I see that product on their résumé, it’s been an instant veto, in the same category as working at Palantir.
employees consumed 73.7 trillion tokens in roughly 30 days
, a figure tracked on an internal leaderboard called "Claudeonomics" — a reference to Anthropic's Claude, one of the third-party AI tools widely used inside the company [2]. The leaderboard, which ranked employees and teams by token consumption, inadvertently incentivized usage volume over productive output.
Meta plans to dismantle the leaderboard and replace it with a centralized monitoring platform called "AI Gateway," which will track usage and spending across teams in real time [2]."
This seems to be an interesting upcoming business, that is:
Helping companies centralize and track their AI usage by employee.
Anyway, great article!
It wants to see faster R&D, higher revenues from existing assets, greater operating margins, higher sales to invested capital ratio and so on…
The best way to measure that for a software firm is up-time of services, usage and project completion duration
Just wonder what happens when more and more companies introduce similar restrictions. Will that lead to devaluations of the LLM companies?
I could believe it, but I'd want to see something a little more concrete.
Enjoy it while you can, because it won’t last forever. Per-token billing is quite eye opening in terms of how much it can cost
The subscriptions are for personal use not enterprise.
i.e. [1] "This article is about paid Max plans for individual consumers. If you're part of an organization looking to use Claude with your team, refer to Team and Enterprise Plans."
[1]: https://support.claude.com/en/articles/11049741-what-is-the-...
And they can't afford a few extra billion that their engineers can utilize right now?
Looks like AI as it develops is intended to be too expensive for regular people in the long run, but if Meta can't even afford it at that rate, who can?
I'd argue most of the AI value is related to how 'Dead' the internet is.
Ultimately the spend on tokens has to benefit the firm financially or it won’t continue spending on it.
Various discussions:
Meta’s chaotic AI strategy
https://news.ycombinator.com/item?id=48523271
Companies rein in AI usage as costs strain budgets
https://news.ycombinator.com/item?id=48602571
Meta CTO Andrew Bosworth Admits the Company's AI Reorg Was 'Atrocious'
https://news.ycombinator.com/item?id=48548461
Tokenmaxxing is dead, long live tokenmaxxing
The article is “by MLQ Agent.”
Doesn't the legend go Zuck didn't even see the big picture until Sean Parker spelled it out for him?
"Now he tells me!"
bonk
Predictably, everyone started talking in Slack like their jobs depended on it. Everyone was responding to everything. Instead of writing out a complete message and pressing enter, they'd send each fragment of the sentence as a new line.
The Slack leaderboard was never shown again. Unfortunately the habit remained because people were afraid they were going to be secretly judged by how much Slack activity they generated.
I expect the same thing is going to happen at companies who had token leaderboards. Once you've instilled that fear in people, they internalize the expectation.
Insanity
No amount of "this isn't used for anything" will change that. It's inherent in human nature in the 21st century to believe any and all metrics will be used against them, and therefore must be gamed.
It's why you also have to set UNBELIEVABLY clear goals and have incentives tied to those goals. Incentives meaning money. If you want to measure things, measure them. But have clear, consistent, and meaningful goals tied to bonuses or something if you want a thing done correctly.
It's surprising how often this principle is applicable.
Charles Goodhart :-)
> Oh wow! If I paid for this myself I would have spent a lot of money! Are other people spending as much as me? I’m going to create a leaderboard!
> Oh no, my misinformed manager is using the leaderboard as a slight of hand for work. I need to game this now.
Then the leaderboard is banned… I can’t see how this ever really goes up the chain beyond director.
Everyone except the executives who get paid millions to predict exactly that.
It's a hard job, someone has to not pay consequences for bad decisions.
He started being drastically more serious into AI in 2022, and 2023 and he has nothing to show for it.
Heck, he could have rented GPUs the way Elon did at this point and either mended the bleeding or stopped it, not sure how many he has, but it beats losing this badly.
If he doesn't wake up and learn how to business, I suspect he will lose his empire he's built up for himself.
"Meta building cloud business to sell excess AI capacity, Bloomberg News reports Meta building cloud business to sell excess AI capacity, Bloomberg News reports"
https://www.reuters.com/business/meta-sell-excess-ai-computi...
Tech journalists have low opinion of people with actual skills who actually contribute to society, and when their opinions get posted here, it's often selectively echoed by people looking for a reason to feel smarter than the industry.
people who make it to managers tend to have bozo tendencies & are yes men.
before it was lines of code, Jira tickets closed. Now it's tokens spent.
This is also not easy. In particular proactively preventing bugs is not rewarded
The main way I think you can proactively prevent bugs in a meaningful way is by crafting and propagating better architecture.
Better (or worse) architecture and adoption of it can be measured through a mix of quantitative and qualitative means so those metrics could be used to evaluate the impact of the engineer driving that architecture.
How?
This is an org pushing thousands of PRs a day. How do you solve the attribution problem for any one engineer's work given some set of impact metrics?
And keep in mind, most common impact metrics are trailing indicators, often over relative long time horizons.
Instead, I had a career ladder with a detailed rubric describing the skills an engineer at each level was expected to have. (Including communication and peer-leadership skills.)
Managers performed qualitative assessment of employees, using the career ladder as a guide. They relied on tech leads and Staff engineers to help them understand people’s skills, and provided 1:1 feedback and coaching.
We did use impact-based metrics to assess the results of important initiatives. We solved the attribution and lagging indicator problems by estimating impact rather than measuring it, and using a series of proxy measurements (activation, usage, retention, etc.) as a feedback mechanism for revising those estimates.
When shit just works for months or years no one is going to come and praise you for stuff you did a while back.
You are better off breaking stuff and then fixing them to show how useful you are.
Source; my last job working with accessibility and that nightmare.
hell we have restrictive rules for security stuff so in many cases our network engineers are still doing by hand configs for critical systems.
but in terms of token use it's gotta be "take this pdf and parse these 3 columns into 2" or similar
The engineer who haphazardly launched on Friday then promptly saved the team at 3am and worked the weekends gets the promotion, while the one who prevented a bug from happening "didn't get anything done" and gets the PIP.
It’s disingenuous to assume every PDF is actually crafted to communicate to its recipients, even more so to pretend LLM users are in a position to understand all the PDFs they receive
There’s a lot of gray area where help understanding a document is fully reasonable
But yeah, it's like they've never actually met human beings...
Unfortunately I think we're entering (have entered?) a period of insanity.
The trouble is AI is being sold as an individual engineering accelerant. I suspect at the most AI pilled orgs you'll then see a commensurate push that starts off with measuring usage (tokens), then measuring output (PRs, code reviews), and then a lot of talk about impact while everyone quietly admits that remains as impossible now as it was fifty years ago.
Why? Because leadership is looking to (and selling, both internally and to the market) AI as the solution to all of their problems, which means they have to prove outcomes that justify their sky high AI budgets.
Higher level metrics at the org/division/product/project level aren't satisfying and flashy enough as they're slow moving and attenuated.
And squishy individual or team level assessments that rely on strong management won't show well on a cost-benefit comparison chart to the board.
At bottom I suspect AI pilled leadership wants to turn software into an assembly line and measure accordingly. Your post perfectly captures why it's still not that easy, and that the real problems in software remains the same and are unsolved by AI: building the right thing, at the right time, and then later figuring out what went well, what didn't, and trying to make those successes more repeatable and failures less likely.
It’s been true for a long time. One of the hardest things as a senior leader in software is dealing with people demanding “accountability” (by which they mean making long-term plans with impossibly precise forecasts) and focusing on costs, all while ignoring value and refusing to engage in prioritization. (I swear, if I hear “it’s all important” again…)
People are just… shallow. They operate on feelings and vibes. They follow the herd without thinking critically. Then they get angry when their dreams clash with reality, and they blame the messenger when those dreams turn out to be fantasies.
But you’re right. AI is bringing out the worst in these tendencies. I think it’s because it’s so convincing when you don’t dig deeply, or aren’t an expert in the subject being discussed. On the plus side, it raises the floor, but I think we’re in for some difficult times before the lessons are learned. I don’t think it will take long, though: I suspect that naive use of AI is going to massively speed up the technical debt curve, and where it used to take 5-9 years to destroy a codebase, it will now take closer to 1-2.
Having a speed limit does not imply the utility of driving is zero.
In addition, as other commenters have pointed out, diminishing returns is a real phenomenon.
If so, your metric cannot distinguish between a bad engineer and a good one.
If not, you have the same problem you started with: measuring contributions to “uptime”.
A metric that moves in the same direction and amount for everyone based on external event isn’t a problem. The delta in performance of the great engineer will outweigh that of the poor, since the metric movement that is due to external circumstances will be the same in each kind of engineer and thus not count.
Token economics also are weird. If you design a fancy new frontend that for example uses a cheap model to parse a PDF into text that is fed into an expensive model, you will probably spend more money because you are on API payscale rather than the "max plan" payscale.
I'm an engineer and use my coding agent to deal with PDFs all the time. It can reach for unix tools if it needs them.
I don't think I understand why this is a problem - it uses tokens, but it removes drudgery. This is the entire promise of the technology.
For the same reason as why the oil companies want everyone to use large cars.
That says that when a measure becomes a target, it ceases to be a good measure.
My argument is that in the 21st century, a time of great data use without deep understanding, any measure at all will inherently become a goal to be gamed.
It was an unreal experience.
The answer is simpler on the surface: focus.
Generally the problem is the larger the firm’s operations, the harder it is to focus.
Apple is the only firm that has done well on this consistently and doesn’t have a huge grave yard of failures to show for it.