GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence(artificialanalysis.ai) |
GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence(artificialanalysis.ai) |
We already know OpenAI has "bel" that is MUCH better than astra and is being used internally
I mean, isn't it almost a guarantee that what we get is a gimped version of what they use internally? They probably already serve themselves next gen level models at 1k+ tps from cerebras machines hosted on perm while we get quantized astra/opus at 50tps on a good day
You're just believing their own bullshit. There's no indication that this is true except from claims from people working at OpenAI.
If they really had a much more powerful model, it would make absolutely no sense to sit on it.
The absolute frontier is largely 2-horse, but the rest of the pack is very close behind, which I'm grateful for. Grok, Facebook, and the Chinese vendors are producing excellent models.
Since Google has their own TPUs, TPS is also pretty high w.r.t. Claude, for example.
Every time a new model comes out, people come out in droves "oh I don't notice anything different".
People have been saying this about <currentModel-1> for 2 years now, and the entire state of AI has changed dramatically.
It cannot be that the next AI model isn't better, but also suddenly what they are capable of is on an entirely different level.
"We tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run"
The only time I remember the newest model making an obvious regression was when the first rolled out MoE. Super speed update but each request had less intelligence at hand. We’re way past that now
x20? 25x? 5x? All in comparison to some other plan that also doesn't have clear limits. "Oh BTW we mean session limits, your weekly is 2x of 5x". "This model eats your limits twice as fast and you can use half of your weekly limits on it". "No this time we mean weekly only".
Using AI to do this, if anything, given the common negative sentiment, is seen as even worse. People think of this as AI using information asymetry to squeeze more out of users, not to cut people a deal.
Taking advantage of information asymmetry is generally looked down upon as well. Look at all the laws we have protecting kids from this. Businesses often don't seem similar protections because those are businesses with big legal teams (and when it is a big legal team vs a small mom and pop store without a single lawyer on payroll, people do start taking issues with it). The power difference between the average company using AI pricing and the average consumer falls pretty solidly in the 'we don't accept this' side of taking advantage of information asymmetry.
I could keep going, but I think these are already plenty enough reasons to why people look at AI price discrimination as not just a bad business practice they don't like, but an immoral/unethical one.
Sure, one can make economical counter arguments, but that's arguing on an orthogonal dimension that simply isn't relevant to where these feelings/thoughts come from.
The results are less buggy, animations are much better.
It can work autonomously for hours and the result is decent most of the time.
That wasn't usually the case with 5.5, which needed more feedback and iterations to get things right.
But they do. See for example the pareto curve they have, try to locate GPT-6 Luna (max), (xhigh), (high), (medium), (low)
Nothing, really. It's like oversampling your data set. You usually get a much better overall baseline performance if you use the default setting.
…but not for all models, which is pretty annoying.
Now I can finally build that half a thing I've had my eye on.
Did they find this improvement within the week? If so, given all their whinging about safety, it seems irresponsible to only test the improved model for less than a week.
Did they find the improvement more than a week ago? If so, why bother releasing GPT 6 if they knew they had a better version essentially ready to go?
If the previous one was as good as worth retiring in 7 days, there's a good chance the new one is also not great.
They probably have 6.2 just about ready, while 7 and 8 are still being cooked up. You can work on multiple releases in parallel.
We're just getting the latest training checkpoints constantly, just to edge out the other lab, while they are trying to come up with something worthy of a new major version number. It's actually worrying, since there's so much pressure to release now.
In a way: We have had "recursive self improvement" (RSI) - that's what genetics is, giving rise to ... us.-
This time around, I think the truly worrying thing is that, having transcended the biological substrate, the pace is unlike anything previously seen.-
Just a thought.-
Gov and Corp will throw their hands up and explain how it's not their fault.
Recent example: on my e-reader, tapping a word I don't now and clicking "Translate" pulls up the possible translations from a dictionary, a local file just a couple of megabytes big, near instantly, on this tiny processor.
Doing the same on my Android phone starts a Gemini-chat with the prompt "Translate the word x into y". Takes forever, internet access needed, results vary, burns who knows who much energy.
Why? Just why?
I neither use Android devices nor Google Search, so the only Gemini thing I see is the Gemini chat interface.
I understand the pain, though.
It's baffling that OpenAI managed to get better at web search than the company literally synonymous with web search.
I don't want it to fill in the gaps though, but make it reference anything and everything it brings, hence it doesn't hallucinate much.
If something feels off, I ask it to back it with concrete data, and if it can't, I don't consider that information correct. That happened once, though, and web doesn't have any information on that thing either. So in that case, not only it had no information on the web, the training data had no information on that thing either short of feeding confidential design documents if they were ever present in the first place.
The question was about an instrument preamp though, so nothing crucial.
Yeah, it does, it might be misaligned, a snapshot of the going on training, bigger than they can serve publicly, not yet completed the full training pipeline.
If you see the knowledge cutoff you can see that sol 5.6 finished the initial main (+stage edited) of the training pipeline on Feb 16, 2026 but it was publicly released on July 9.
The opposite would be weird: if they do NOT have an unreleased in-house model that would be really odd.
Plenty of the tasks that keep a company running can benefit from good-enough (and better than the competition).
Makes total sense if they don't have the compute and can't serve it in an economically viable way. Also it lets you build things no one else in the world can build as fast as you until it's released
"We can generate slop faster than anyone in the world" :evil_emoji:
besides, we know anthropic are sitting on models too
Because clearly they have no problem with releasing newer versions of models even just a week apart.
> besides, we know anthropic are sitting on models too
It is from your crystal ball or from other bullshit you heard from Anthropic employees on Twitter?
We all know Anthropic had Mythos and Fable, and they turned out to be completely normal models, entirely in line with the capability of their predecessors.
All they do is lie, and you're believing their lies.
It's rumored to be over 10T parameters. When released it'll probably be very good at certain tasks, albeit slow and expensive and not necessarily "wiser". You don't have to make this a binary.
Also, Mythos was in fact a significant step-up in several ways. It fits the trend line, but only because the trend line for LLMs is quite steep. Plus Astra is still in many ways less intelligent than Fable/Mythos despite being released much later.
> Because clearly they have no problem with releasing newer versions of models even just a week apart.
You can see how it is pure non-sequitur, right?
> We all know Anthropic had Mythos and Fable, and they turned out to be completely normal models, entirely in line with the capability of their predecessors.
Fable was absolutely not in line with capabilities of other models when released. For cybersec work it was much, much better.