Fable and the end of the free lunch(dbreunig.com) |
Fable and the end of the free lunch(dbreunig.com) |
https://ourworldindata.org/data-insights/moores-law-has-accu...
"When Moore’s Law slowed in the mid-2000s (specifically, single-threaded performance stagnated), we suddenly had to think about parallelization, architecture, memory locality, etc."
Your link is talking about transistor count. The article is talking about single-threaded performance. Today's CPUs are faster in large part because they have more and more cores.
I think a lot of people would be very content if they never got smarter, and just kept getting even cheaper/faster. Of course, both things continue to happen on a seemingly monthly basis
It was so amazing to get advices and reflect that it struck me : I could use this model forever - it’s clever enough to help me tons and do lot of work for me - even if ai would stop evolving I would love it
+1 regarding voice usage too, I use it in so many different ways it's hard to enumerate: while driving long distances (think of a custom made, interactive podcast) / as a way to collaboratively build specs or shape an idea / as a way to provide input while vibe coding / just as a normal voice assistant (straight in the ChatGPT app or as OpenClaw input via telegram voice notes). I can't overstate how much my routines have changed over the last couple of years.
And this is inherent to how LLMs work.
I remember thinking the first ChatGPT realtime voice was science fiction, before the limits on its intelligence (particularly as mainline models advanced) became annoying. Perhaps we’ll feel the same way in a year or two - people have been claiming models are plateauing in practical usefulness every year, and they’ve definitely been wrong so far.
especially considering imo most use falls under this instead of those kind of tasks where you'd need the SOTA
The next step would be automatic self-training. A free LLM that could access HN everyday (and the linked sites) for more data would remain current in programming for a really long time.
if deepseek and stuff are 4.6 caliber i literally don't know why im here i should probably just go sign up for openrouter at this point
There's a lot of truth to this. I think we're starting to approach the point where increased intelligence has declining marginal returns, such that it might not even be worthwhile to improve models unless it can be done cheaply.
my experience is mostly frustration and rewrites of anything that requires more than what would take me an hour to do myself, unless it is pure translation / boiler plate work.
the leaps are there at getting to more "shaped" code (code that is correct for linters, static checking, etc), but i don't see the models exhibiting much intelligence. i really can't think of a time using LLMs for building anything where they did something that would make me go, "wow, that is really impressive, i wonder how it came up with that." just brute force search and pattern matching still.
even the interesting results in academic work seem to be more of a function of effort (proofs by exhaustion, fitting puzzle pieces in a search space, etc) than anything else. not to say people aren't using large language models to do impressive things, but the agents themselves do not seem very intelligent to me.
it feels like some engineering teams are aware of this fact and are driving agents using strict rule checks (like hooks on steroids), so they can drive some shape of output that aligns with what they require.
Benchmarks seem gamed at this point, real world experience just doesn't match up.
The world keeps moving on, and so the models need to be retrained so that they can keep up with new information. Otherwise you'll get stuck with a model that only works well with information that existed prior to a dataset horizon that's receding into the past at a constant rate.
At the same time, they have to keep iterating on the training process itself. AI generated text and code is slowly spreading across the internet. Model collapse is a real concern; they wouldn't be spending quite so much energy on buying and scanning rare books if it weren't. But for coding in particular expanding their corpus of old text is not really a good option because of the previous problem - no good training your LLM to write 1980 vintage K&R C that won't even compile on a modern compiler.
[1] https://artificialanalysis.ai/evaluations/omniscience?models...
[2] https://old.reddit.com/r/LocalLLaMA/comments/1vt7l3e/qwen382...
I think lower-cost models will get the largest piece of the pie, as with almost everything that has ever been sold.
Just look at cars: US consumers buy the F-150, EU consumers buy the freaking Dacia Sandero the most :)))
Ferrari/Lambo numbers are microscopic
I imagine it must somehow be possible to update a model's understanding of recent events without training a completely new model from scratch?
I'm definitely not getting smarter. But my tolerance is 1 drink so I'm definitely cheaper. Also as a result, I spend more time training and so I am faster. And yes, I am more content
With how generous subscriptions are, what I actually want is GPT Astra, not cheaper Sol.
All current devices used to run AI are very far from an efficient solution to the problem. What you really want is a pure dataflow architecture, instead of a von Neumann machine. The reason people aren't really making them yet is that when you build one, even if you use SRAM for the weights, you are binding yourself to the dimensions of the model you target -- your chip is only ever going to run variants of that specific model. And SRAM is much more expensive than ROM, so if you want to make a cheap version, you need to design a specific model into silicon.
Once model improvements taper off, the next thing that will happen is everyone will chase speed. There is no physical reason why a mid-sized model could not run at >1 million tokens per second on leading edge silicon, if all computation that can be parallelized, is. No-one will go straight to that, even for a mid-sized model that's like 20 distinct reticle-limited chips. But something like the next version of Taalas HC1 (presumably called HC2?) will probably boost a ~30B parameter model to ten of thousand of tokens+ per second from a single stream within 12 months.
Model on a custom silicon: https://chatjimmy.ai/
1-bit models that run on a CPU: https://github.com/microsoft/BitNet
What evidence?
Well, I'm here to tell you that whatever is going on behind the scenes at Cursor with this Space-X acquisition in the works, the Auto setting is clearly routing all prompts through "Cursor Grok 4.6 High" right now.
This is a degree of subsidy that makes the Microsoft thing look quaint.
I reduced my $200/month subscription to the $20/month level and have proceeded to do what I would have paid about $1500 to do with Opus 4.7 or thereabouts, which is how Grok 4.6 High feels like it compares. I don't have anything remotely like hard evidence to back this estimate up beyond what I'm watching it do and I still somehow have ~10% of my monthly Auto capacity left on my account. It's completely nuts.
Can't say much more because I have more backlog to run before someone comes to their senses.
Maybe Fable can do the same things better than other models, but having to tiptoe around to avoid tripping safeguards makes GPT 5.6 so much easier to work with that I don’t even bother with Fable (or Opus 5) now.
The tendency for absolute inefficiency is effectively unbounded until scarcity is imposed.
Bottom line is the slop used to be manageable, but now there is 100x more code pushed, so that train has departed.
In the end its more bad code for features no one will use.
I fully expect I’ll switch back to Anthropic, or another model in the next 90 days. The fact that we are switching indicates that the models aren’t ready to be baked into silicon.
I wonder if they will ever been that good, or if the lifespan of silicon is longer than the lifespan of a model before it needs to be retrained.
Right now a lot of people have a lot of opinions on which model to use for which task. They get better results for less money by judiciously switching between Fable and Opus and whatever else. Spending my time learning this skill would have an immediate benefit for me.
But on the other hand, maybe the harness vendors will just solve it in 6 months? I'll ask a question, something in Claude Code (or whatever we're using by then) will figure out the most effective model based on the question and the context and my apparent willingness to get it right. I'll get billed X or 10x as appropriate, and I'll be happy with that, because that's what I would have paid if I made my own choice of model every time.
Claude Code already does this a bit, sometimes it will tell me it picked Sonnet for such and such a sub agent, or some other detail I'd rather not care about. The best humans seem to be better at deciding what model to use than any of the tools is, but surely that won't last long.
People say stuff like this a lot, but I have a different take.
The whole "such-and-such model is 90% as good as Fable at 1/10th the price" assumes that the value increase of intelligence is linear. But I think it's exponential: that last 10% makes a massive amount of difference. It can result in a key insight that helps you strategize more effectively, a novel approach that saves a huge amount of time, a feature design that is lot more user-friendly (because top models like Fable also possess substantial non-software domain knowledge that help bridge the gap between user and software), or the depth and breadth of engineering expertise that helps avoid a nasty bug that would otherwise have cost you users and revenue.
Yes, it is totally possible to use Fable as the planner and delegate implementation to lesser models. I do that. But, my theory (which I unfortunately do not have the money to test and prove) is that a codebase designed and implemented by Fable would be substantially better than one that is designed by Fable and implemented by Opus 5, GPT 5.6 Sol, GLM, Qwen, Deepseek, etc. The reason I believe this is because I read the code Fable writes and compare it to code that any other model writes and the difference is night and day. It's not just 10% better. It's mid-level engineer vs. principal/staff-level engineer. And the thing is, even for rote tasks, a more senior engineer is going to be more likely to come up with a clean design than a mid-level engineer. They will also be much more likely to take a step back and ask important questions or propose different approaches.
So if you're using Fable and everyone else is using lesser models, sure they might be saving a lot of money, but there's a higher likelihood that your product will be higher quality, perhaps to a significant extent. And models that are released in the future will benefit from it as well.
Even now, I use Fable as the planner and coordinator, with it farming out to agents. I don't hit my Fable limits either.
Which means I could accomplish more, but these are side projects so I don't need 30x productivity. Still, claude is constantly churning away at something.
Right now, it feels like all of that is today where coding was a year or two ago, and we're on the cusp of some massive improvements outside of coding. It'll be interesting to see what these companies decide to automate next.
The vast majority of people, eg vibecoders, do not need Fable or Sol tier intelligence for their slop To Do app.
One of the primary reasons for this is that computers operate in a vast range of orders of magnitude. There’s several orders of magnitude between cache local cpu operation and dram, then several to disk, then several to network, then several to globally durable guarantees. When your code has literally thirteen orders of magnitude to optimize under, there’s never a free lunch. You always need to understand your stuff.
As 80% of enterprise software is CRUD with a bit of sprinkling of user authorization and tenant customisation. But subtly different for every business domain. It's mainly what properties the models and validations have that are different.
When you add a new module or whatever most of the code you have to write is rote code.
And sonnet can handle that crap just fine, you just point it at a similar example in the code, it picks up your userContext convention, how you're doing i18n, etc. and you're done.
I like saying that enterprise code is often shallow but wide. I must have written at least 4 purchase order systems in my career that are all completely different but almost exactly the same.
The Bitter Lesson says that eventually general approaches which leverage more data and more compute will outperform the handcrafted rules and heuristics that humans add in.
However, it does not say what to do today about the problems of today. We can’t just wait around for 10x faster compute and 10x more data.
For the users, I feel it is more like "free lunch started", with all these awesome open-weight models being thrown around, breaking the monopoly of a few biggies.
I throw everything at claude Opus.
While some people start thinking like OP, A LOT of people just start exploring ai.
And others which are already using it, only understand half of it and just use what they are allowed to use. Claude, GitHub Copilot, Curser, etc.
Besides trying to dogfood my own product I've hit a wall in terms of my patience with a)how slow fable is b)how expensive fable is. Not to mention how often it refuses totally legitimate work.
So yeah- I've moved to DeepSeek and I actually ask the freepi harness to delegate planning to fable but then move back to doing implementation in it's own harness. My current providers are super fast so it's a joy to use.
For starters I don't know if it is an artifact of the model or something by design, but the level of gratuitous cognitive load carried by the complexity of its replies is unbearable.
Yes, it's a beast at coding, and also it's incredible nuanced at improving writing, validating specs, etc.
But when it comes to replying, it's the William Gibson of LLMs [1].
It has this tendency to take extreme detours to say things that could had been said in less, much simpler words. [2]
It really, really like to wrap very simple and atomic ideas on several layers of abstraction, building on unnecessary terms that carry no intrinsic information and assumes this vocabulary as shared and then building on top of it.
By the time I got to the end of the reply I'm bored to death and didn't understand even a third of what it told me.
I think the people at Anthropic should reflect on the maxim "You don't know a subject if you cannot explain it"
If you pardon my french, Fable is an insufferable obnoxious cunt.
---
[1] I apologize on the comparison but, as much as I love his first 2 trilogies, haven't been able to finish any of his last 2 books.
[2] "The residual you're accepting is the one from before: recovery currently rests on beneficial non-compliance, which may erode as models get more literal" == "We already accepted this risk"
" Its observable when it erodes is a stall that survives relaunch — loud at operator level, recoverable from the worklog, and fixable by codifying at that moment" == "When it breaks, it'll break visibly and recoverably"
"That is the iteration model applied exactly as written: resolve on first contact, don't pre-solve " == "So we fix it then, not now"
From the claude settings:
> Switch models when a message is flagged
> When safeguards flag a message, automatically switch to a different model to keep chatting. When off, your session will pause instead. Applies to web and remote sessions.
FWIW I get a ton of usage out of fable and it's only happened to me once.
With Sol we see openai making the model extremely slow and paranoid about process/ceremony. Sure this is a good guardrail against AI going rogue, but it also sets the stage for companies to charge for 2x, 4x, 8x performance, with 1x being barely tolerable and frankly slower than last year's models (though less error prone).
The irony is that the smarter the model, the more it can be trusted to do with less supervision, so one engineer can manage a team of 20 fable subscriptions more effectively than a team of 3 of last year's model subscriptions.
However, I've been using Cursor since it came out. I have a massive amount of institutional knowledge locked into their platform, and vastly prefer it to the other options available even if the switching costs were zero.
Going through what would amount to significant effort/time/cost to switch to a different coding tool as a sort of performative political rebuke is just not how I would recommend anyone protest DOGE.
I usually keep 4+ agents churning, many of them on tasks that take hours or day. I only played with Cursor a bit, but it seemed to want input from me every 10 minutes or so.
I am not saying that you're wrong, just very aware that we appear to use LLMs in radically different ways.
Only the most substantial requests run for ten minutes or more, and I would spend an hour or more writing the prompt for that action.
I am always the bottleneck for how quickly Cursor can do what I want, with the level of outcome that I get.
I would genuinely love to be a fly on the wall while you do what you do, because it just doesn't currently make sense to me.
Like middle school level genetics stuff from a guy who hasn’t been in school for decades.
They need to fix that. It’s just broken. Nobody is making bioweapons if they’re asking the dumb sort of questions I’m asking.
Also, it refused to identify an actor in a popular tv show from a photo. Apparently the policy is it won’t identify ANYONE from a photo, now. Even publicly listed cast members from a very popular show, from a photo of a scene in that show.
It claims that’s a fixed security policy. Nevermind how that makes absolutely no sense… argue about it enough and it terminates the chat.
I don’t know what the Anthropic clown car is even doing anymore, but I won’t be surprised when the others eat their lunch.
There can only be one fix: send Amodei packing and release unguardrailed models.
It happens to me all the time with things that have nothing to do with security, Fable spawns a subagent that then adversarially checks the code Fable just wrote and hits guardrails, with zero prompting from me.
Having not asked a single security question it will write wildly vulnerable code, go back and fix it, and guardrail itself out of existence after charging me a large sum with no refunds for no output and having not fixed it because that might be secuirty adjacents.
And if it doesn't do this you end up with code that has such holes, store xss , no authz ... if it does not go back and notice it has written bad code.
Since they hide thinking and reasoning from the user (who is also paying for those tokens) it is a black box what is triggering it, has the LLM this time thought of "Oh, this has XSS" and used a bad dangerous word such as XSS, while the previous conversation did not ?
In that light, I often go the other way: let Opus (and Haiku subagents) do most of the heavy lifting and then give Fable a shot at finding holes, especially if there are holes or unanswered questions or unearned assertions that I’ve caught on my own in Opus’ output. This, so far, seems like a clean tradeoff that doesn’t burn my Fable credits as hard and still gives solid results.
Really frustrating.
I don't have proof, only my anecdotal experience: I leave plenty of Fable usage on the table because I do not think its implementations of code have been better to Opus 4.8, not even close. It overengineered, obscured and picked awkward constructs all the time over plain, simple, perfectly clean and performant code patterns. Code was smarter AND worse in the kind of way that a brilliant and overeager recent grad often does. (I know I did)
Granted, I laid out a document with coding practices, architecture, and technical design recommendations to steer it towards good engineering. And it's a domain I know super well, so I could give very nuanced feedback on trade-offs + architecture. If it had been left to its own devices, maybe it would have over-engineered the h*ck out of it.
But the code it produced—and the implementations it guided Opus towards—were excellent.
My brother, that's my job.
As a software engineer, I selfishly hope that they spend more effort on non software tasks since I’ve feel like we hit a sweet spot where engineers still have some value and autonomy, but a super charged tool.
Pragmatically, I suspect that “non software” tasks will be a tarpit because most tasks can’t be automated and verified as easily in an RL loop compared to software projects. Especially since most skilled labor is either not nearly as expensive as software engineers (eg biologists), or regulated (eg doctors, lawyers).
This is what the "good enough" people fail to grasp. There's no "good enough" - unless your tasks are genuinely small scope and will stay that way forever. If not, there are always more gains to extract.
i just get fatigued from it, am I holding it wrong or something?
sometimes it's fine but the constant RLHFisms like the constant "worth flagging" and stuff is getting really old
It's highly likely that over the next 10 years we find demand and loss of production further constrains supply.
Designing a specific model into silicon sounds like one of the worst possible ideas. No better way to freeze assumptions and limit growth. Software defined solutions dominate for a reason, because adaptability is key.
Models were never the answer. Eventually we'll get past the nonsense of observationally inefficient neural nets.
BUT
That's a bad model. My opinion is that for exactly this case we need to use RAGs/ APIs/ some retrieval mechanisms.
It's silly to train them on stuff that changes every week/month
I don't learn APIs by heart, I look them up. It's to expensive (my time) for me and (the compute) for the models
But also, LLMs' use of RAG to keep track of API evolution is limited. You can see this if you watch an agent at work using a well-known library that has a high rate of breaking changes such as Polars or Guava. There's a huge amount of churn on repeatedly writing code that works with an older version of the API and then diagnosing and fixing the resulting compile- or run-time errors. It can burn through quite a lot of tokens, which drives up usage costs.
I agree that, all else being equal, using language model training to bake knowledge that's easy to look up into the system is kind of silly and inefficient. That's actually been one of my top complaints about hawking these LLMs as a sort of general-purpose AI. But the fact of the matter is that's fairly fundamental to how they work, and RAG is arguably just a hack on top of the basic design to paper over this limitation. RAG's limits become pretty easy to see when working in knowledge domains that aren't very publicly accessible, and therefore produce little text that would have been incorporated into the models' training corpora. It can be a bit of a, "Ignore that man behind the curtain!" experience.
And no I'm not just talking about local models. I've seen it happen with recent GPT-5 and Claude Opus series models, too.
To get it to the point of being remotely useful, I've had it start to write condensed fact blurbs into the agents.md file. It doubts itself so much and questions its every decision to the point that it'll literally blow the entire context on thinking alone in anything but the most basic CRUD projects otherwise.
What an earlier generation model would just start doing, it went out to research the source code in multiple libraries just to see if what it was thinking would work... then it said "Hey, I should really just do it" then went back and started researching more anyway, on and on (even on medium thinking level).
If there's a better local model for writing code, I'm all ears.
The team that built is working on a better implementation.
Edit: I went back and retested it. It revealed that its data is from July 2021 which explains partially why It couldn't talk about the actress I asked about (she exploded in popularity in 2026 but was still a professional actress in 2021 so idk). I then went and asked questions about a very popular actress and movie in 2010. It got it much better but still hallucinated a ton of details about her.
I guess I didn't fully understand what you were saying. Sorry about that! I look forward to their next releases because upon thinking about what I experienced here, I am super excited to see this progress further!
The biggest generalist models beat the most fine-tuned specialists, as a rule. You can bias an LLM away from literature knowledge and towards coding capabilities, but that buys you very little performance, and for too much effort.
Generality and intelligence seem to be entangled very heavily in LLMs.
It's impressive that it does what it does, don't get me wrong. But if you expect it to replace the likes of GPT 5.6 Luna, let alone Sol? Nah.
Take that away and you'll barely be able to make an app that display a pigeon riding a bicycle (or whatever you ppl are doing these days).
Any given failure is not inherent, they are all dependent failures; what is inherent (due to the SOTA in ML, perhaps or perhaps not the architecture) is how many examples they need to get good at stuff.
When the Industrial Revolution came along it did create 'super farms' relative to the past through increased efficiency and production, but it also created a huge vacuum in the economy that was ultimately filled by industry, to the point that farming, super or not, became a vanishingly small part of the overall economy - even as production continued to increase.
---
LLMs stand to do the same thing for software. If and when we reach the point of 'normal' people being able to reliably compose ultra customized software solutions to their problems, then software is basically done as a problem-solving industry in and of itself. Not 'done' as in dead, but 'done' as in solved. There's just nowhere to really go from there.
And so I think this will do the exact same thing as the Industrial Revolution did to farming and create a vacuum opening the door to all sorts of new interesting expansions in the real world, as opposed to the digital one. I don't know what this means, because it's quite difficult to foresee the impact of the Industrial Revolution when living in agrarian world, but it's not so hard to see that the future will not be agrarian.
---
So it's probably still myopic but my bet would be on the first major manufacturer of cheap customer/enterprise grade generalized robotics hardware shells.
And there I think the winner would be China.
Sorry to cut you off, but have you looked at Nvidia's numbers since the NFT craze? They won.
Sell shovels in a gold rush, make better shovels, repeat on the next rush.
So, my plan would be to invest not in the AI companies, but in the economy as a whole who get to use the AI for their businesses.
Caution though, one thing which AI is already superhuman at is persuasion. Regulatory capture is likely even easier today than one might expect purely from the revenues of the AI companies.
Just like Wikipedia put classic encyclopedias out of business, but wasn't really a financially win for anyone.
What makes you think if one or two AI labs can do this that the rest (including open model providers) won't be able to follow the same path a few weeks/months later?
Even if you believe in the "Singularity", and believe it is coming soon, I still don't see any reason to believe the Singularity will be... singular. There won't be one clear winner, the race doesn't get called as soon as the first person crosses the line.
None of the AI labs are showing any sign of pulling away to a monopoly or duopoly position, to the contrary the early large leads of OpenAI and Anthropic have all been evaporating.
AI has clear economic value. It still isn't clear at all how the providers of AI will capture that value in a moatless environment with the technology becoming rapidly commoditized.
> I will be able to use them forever
Where will you run them when powerful enough GPU and RAM are only sold to hyperscalers?
idk about you but i WANT the second thing, because I like to be alive.
Do you think that fabrication will never progress (in volume) than what we have now? The hyperscalers are already having trouble paying the bills, they can't keep this up forever.
Now, of course I’d prefer no censoring, but I live in the world we live in.
I’m working in the assumption that (like today) there will always be somehow on openrouter, or similar, who will host a model I want to run.
everything would have to be kept under wraps, and you'd need to avoid the scrutiny of the US gov (they already wanna eval SOTA models in advance)
I don't get this idea that "AGI" will just manipulate everyone somehow into destroying the world or something
By censorship I mean things like "nothing happens in 1989". By public safety I mean "no I won't tell you how to build a bioweapon for genocide".
And it would show up in real GDP, not necessarily in nominal GDP.
"nothing happens in 1989" is censorship. "I won't tell you how to build a bioweapon for genocide" is guardrails. I like the second one because I like to be alive.
By my reckoning, there's a significant chance most software engineers will be unemployable within a few years. But I'm not 100% confident that there'll be a utopia waiting for us, as an alternative.
Individuals can still get unlucky. Just like a coal miner might be out of a job, when solar panels become effectively free.
Software engineers are a pretty small part of the general population. And they can move into general white collar work afterwards. Perhaps at a drop in pay compared to software engineering, but still pretty cushy by the standards of ordinary people.
(And if we manage to automate all white collar work to be done cheaply and reliably by machines, well, then we are in utopia.)
So my new pet theory is that this is one of the few times we're seeing a positive effect from the heads of all of these big businesses being part of weird public (e.g. WEF) and private (e.g. Bohemian Club) orgs where they get to together and conspire to conquer the world or whatever. Rapid replacement of labor would be horrifically self defeating, because you'd end up not only tanking your own economy but having a bunch of angry and increasingly desperate people with a whole lot of time on their hands. That doesn't tend to end well for the powers that be.
So I think there's going to be a conscious effort to transition between this era, and whatever comes next, in a more controlled way than $$$ YOLO $$$.
What this tells us is that a 3B LLM can retain enough NLU to understand those word problems. Which isn't particularly surprising?
And also that the same LLM can solve a math or logic problem it understands. Which is a lot more impressive, because early LLMs were already quite good at NLU, but notoriously bad at things like math, logic and iterative problem solving. This 3B model existing tells us we're beginning to figure out how to imbue models with those capabilities reliably.
When LLMs started to get popular, they really were stochastic parrots. I was fully aware that they were completely useless (except perhaps for poets) until they can do math. And I was a bit skeptical that they will ever be able to do math. But they started to do math and recently they got really good at it.
Math is the pinnacle of human achievement. You can't do anything harder with your intelligence than math. And LLMs are now doing it.
The fact that 3B model is capable of doing math on the level that is better than what frontier models trained for millions could do 3 years ago is absolutely stunning.
Math is incredibly hard to humans, but "proving a conjecture" might have a lower intrinsic complexity than "putting together a good joke". It's just that evolution has only ever optimized for one of those things.
Math can easily end up being one of those things that are less "hard" than they are "hard if you're a meat-brained hairless ape" - like chess play did.
Historically? "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division. Then it suddenly didn't require any intelligence whatsoever.
Don't get me wrong - the LLM achievements in math, both as in "solving unformalized problems" like VibeThinker does and in "rolling novel math" like the latest ChatGPT and Fable do are very impressive. We're come a very long way from "formal logic only" systems of the 90s. The AI progress we see now never ceases to impress me.
But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing - so be wary of your intuition when saying things like "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
Might earwax be soon worth more than gold? Experts say: No! What? No.
> "Pre-coded algebra" was thought to require a lot of intelligence too - until someone found a way to make simple logic gates perform addition, multiplication and division.
No. Not intelligence. Diligence. https://en.wikipedia.org/wiki/Computer_(occupation) They didn't hire the smartest to do the calculations. They hired the diligent and cheap. They hired the smartest to do math.
Things that are easy for humans are easy because we have fist sized universal approximator in our skulls, that's just fast enough to keep most of us on two feet, architecturally optimized for very few activities (mostly physical, some virtualized) and trained for years. It doesn't mean things we do are complex.
As for Moravec's paradox ... Guidance system of a missile is not super smart or solving complex problems. It's just brutally optimized for the task and has a fitting form factor. Tasks that are easy for it are hard or impossible for my windows computer and vice versa. Paradox comes from stupidly thinking easy<->hard is one dimensional axis. That kind of thinking is something people are very prone to ... good<->evil, healthy<->sick, young<->old ... while if we go a bit beyond the simplest narratives we can plainly see that everything is a multidimensional landscape. Just because we chose to draw a single line through it, in a semi-random direction we feel is about right, doesn't mean it is relevant for solving anything or even interesting. That's where a lot of paradoxes come from. We just strayed from reality too far and simplified or abstracted something too much.
> But judging intrinsic complexity of a task by whether humans find it hard is the most treacherous thing -
I think we can get a good hang of estimating how hard a thing is. If a thing is hard for a human it's probably pretty hard. We made some of them easy building machines that exceeded human strength and diligence. Now we built first one that exceed human intelligence. On one hand, it's as big as invention of a lever, steam machine or a computer. On the other hand it might be only roughly as important as those things.
... If a thing is easy for human it still might be hard because of hardware optimizations that humans have. Walking on two legs, seems easy. Walking on two arms. Much harder. But truly they are one and the same thing for a robot. So you might easily estimate that walking is not that easy. It's just when it comes to legs humans have a specialized controller, like a missile guidance system. Putting together a good joke? Might seem easy, maybe it's not that easy because humor plays a role in reproductions so we might have some optimization for it, but it's surely not harder than putting together quantum theory. You can see this from whatever the ideas version of cyclomatic complexity is. Some math theories have higher complexity than quantum theory. So a system that's capable of exploring multidimensional landscape of mathematic language, surely has raw capability of doing everything else humans can do with language. And it will once we direct it towards it correctly.
> "you can't do anything harder with your intelligence than math". This kind of statement has an awful track record.
I don't agree it has. And I stand by it.
I doubt that conspiracy theory.
I think the only assumption that's meaningfully debatable is whether the current SOTA are able to supplant labor to a more significant degree or not. And while I suppose that's going to inextricably remain an opinion, I think it's reasonably objective to say that hallucination rates have sharply declined, and overall code quality/coherence is sharply up.
If "human brains become fully irrelevant economically" then that brings into question the entire premise of "share holders" and "financial winners".
What even are money, shares, stocks, and finance in a world where human brains are irrelevant economically? No one knows, but betting that "share holders" will be the winners is a highly questionable bet.
I would much more likely bet that "the armed group who manages to control and benefit from the AI through force" will be the "financial winners" more so than "share holders", who tend to not be terribly military minded at least in America.
If that fails, who knows what things will look like.
If the AI gets as powerful as you think it might, then the group that figures out the answer to that would have the power, I suppose. or maybe the AI does not listen to any of them and does its own thing. Who knows? Personally, I would not bet the share holders are going to come out "on top" whatever that means.
I think a lot of share holders are finance people, not deeply technical AI people and so odds are the share holders will not really understand the AI enough to be the most likely to control the AI.
If I knew, I'd be rich from deploying it onto a substrate for my own AI.
But that doesn't mean that there isn't something there - the current approach seems at odds with how flesh brains work.
I mean, you can power a human brain with 2x bananas for 4 hours, the energy of which might power an H100 for about 20 seconds. It's obvious that there's something different happening.
Having a natural language interface where you need to go out of your way to specify that you want an accurate answer rather than just a plausible one defeats the entire purpose of it being a natural language interface for normal people. In certain professional contexts, it can be useful, but I don't buy it at all that it makes sense to ask everyone in their everyday lives to go out of their way to specify that they actually want correct answers to their questions.
You don't. It goes in the system prompt.
You only need to mention Protected Group Of The Week (I'm one of them and I like to research and read about history, so that makes it extra challenging) or anything resembling negative human emotions (guilty on that front as well), and the model screeches to a halt.
Just because OAI doesn't want another headline like "Chatbot convinces teen to off himself"
I understand this in principle, but I'm not convinced that dulling everyone's knives is better than figuring out how to keep them out of kids' hands
And so you consider that censorship and not guardrails? Huh...
I call that censorship too. I'm curious enough that I want to know about such things. I don't want any limitations on what I'm allowed to understand and know about.
As if your censorship was not going to kill that too. Can't even ask Fable about aminoacids without getting blocked. So much for "learning from first principles".
> but if you want to just ask for the practical steps of making a genocidal bioweapon
Nothing wrong with practical steps. Just because I know how to do something, doesn't mean I'm actually going to do it.
I didn't say "with LLMs". Last time I checked they still teach chemistry at unis and schools.
And yeah, guardrails are not perfect. Honestly I don't think good enough guardrails are possible, it's all eventually defeated or becomes silly. And yes that should be one of the reasons the technology as a whole is banned.
But until then, guardrails are guardrails and not censorship in a pretty obvious way. If you refuse to see that, be my guest. I personally like to live.
> Nothing wrong with practical steps
No thanks from me
(I think it would be a good thing for humanity if they did)
- LLMs have better long term memory (they know more than any human) and more working memory (LLMs have fast, uniform access to their whole context window).
- LLMs are faster than we are.
- Humans have online learning (we can do simultaneous learning and inference), giving us advantages in many novel tasks.
- We can learn concepts from far less data. And we can manage our mental context more smoothly.
- We seem to have better world models than current models. AI video just doesn't look right, somehow.
I expect that these remaining weaknesses can be overcome without resorting to human brain emulation. I see no reason to think that current LLMs are at the limit of what technology is capable of.
For example, it seems that even at Fable scale, simple concepts like the passage of time or (gasp) timezones elude them. I live in UTC+10 and with any RFC8339 data LLMs are constantly confused - is it Sunday the 10th or Sunday the 9th, etc. I have tried many solutions for this and every time it finds a way to get it wrong.
Getting confused about timezones does not place LLMs behind that many humans. (But doing so repeatedly does highlight the lack of online learning).
How do you square that then? They can do amazing things, but they're also not smart? Do you think its possible to solve Erdos problems without any "smarts"? Can you do it without even understanding mathematics?
I find it very hard to hold the idea that LLMs don't understand anything. They can explain concepts, translate them, simplify them and implement them in code. From the outside, LLMs seem to understands most concepts better than most humans do. Do you understand anything? Couldn't I make the same argument? How would you prove that you understand what a for loop is, or that you know what calculus is? I assume you'd demonstrate your knowledge by using a for loop in a program, or explain calculus back to me. But LLMs can do that too.
> For example, it seems that even at Fable scale, simple concepts like the passage of time or (gasp) timezones elude them.
Funny example, because lots of human struggle with this too. The number of meetings I've had with people in the US! "Lets meet on thursday morning australia time!". Only, they actually meant thursday night US time, which is friday morning australia time. "Oooh that's so weird! Its the next day for you!". ...... Yes, I know.
I think LLMs are just a different kind of intelligence than humans. They're better at some things than us, and worse than others. They can find latent security vulnerabilities in the linux kernel, but struggle to count the Rs in strawberry. They're not as smart as humans in many ways. But we're not as smart as LLMs in plenty of ways too. I didn't find those linux bugs.
Are you making a serious argument that it's not?
Because you'll need to explain leading-edge mathematics advances that have come from LLMs, among other things.
Also, you may have noticed in passing that humans aren't getting any smarter, while AI models are.
That's not all of what we are doing for at least a year, possibly few. LLMs are trained increasingly on generated inputs. Soon human sourced material is going to be rounding error in the process of training.
If you add two random numbers and calculate the result and those happened to be numbers noone else ever had idea to add you created a new piece of information. Template is not new, but the piece of information is. And sure, this template might be very simple, too simple, but you can come up with more complex one. And metadata is data. You can create templates in similar manner to how you create new pieces of information using them.
And labs training AI are doing it for years at this point. And it is ever increasing fraction of all training.