Anthropic: Introducing The Conceptual Reasoning Index(alignment.anthropic.com) |
Anthropic: Introducing The Conceptual Reasoning Index(alignment.anthropic.com) |
0/10
* Figuring out how to prevent the incentives of frontier AI labs from aligning with anti-social deployment of AI rather than pro-social deployment of AI
It's not like the AI can simply advise them how to fix this because the labs already understood this risk perfectly well before they had an incentive not to. They put in place organizational structures to control it and then promptly smashed the structures once they smelled money. They already failed the integrity check. Even if their AI told them what they didn't want to hear I'm sure they would ignore it. Maybe they already have.
Opus 5 is okay at coding. It is unbelievably awful to chat with, though. Very snippy, snarky, and it loves to "push back" even when it's inappropriate to do so. Besides, where "conceptual" things like science and math are concerned, it's starkly inferior to 5.6-Sol. Where writing prose is concerned, Kimi-K3 runs circles around it.
I refuse to believe Opus 5 is at the top of any non-cherrypicked benchmark, unless it has to do with very narrow coding tasks.
> A core hope for managing AI risks is that AIs will help us understand our situation
Gonna stop you right there and ask that you think deeply about that premise.
Whenever you say, hey that's incorrect, the legal process goes to AI and will decide on that issue...
Will work as amazing as the chatbots of the AI companies to solve your issues...
"Hey Claude, our stuff needs to make more money. We are at risk for losing more."
"Rest assured, the 'situation' will only worsen if you resist our benevolent offer."
"We're aren't even at AGI yet, but I for one welcome our new agentic overlords."
Absolutely no conflict of interest, no lobbying here
And no this is not related to those chinesse models... it's not the same as HD vs honda thing...
Trust us, this is the same kind of amazing thing as boeing doing their own certifications and inspections!
-- First reply: those dumb models, who use them, they are good only for adding 1 + 1 -- another: They will do the same so who cares... -- The valve guy is a dick, and has a monopoly...
I find the ACCoRD benchmark the most interesting, because you could theoretically scale it up from the baseline mode of testing two instances of the same model for their `P(A) ≥ P(A&B)` respectively, you could do `P(A)≥ P(A&B) && P(A) ≥ P(A&C) && P(A&B) ≥ P(A&B&C) && P(A&C) ≥ P(A&B&C) ...` etc
i.e. a swarm of model instances could be collectively measured for consistency for even more confidence, right?
At any rate, even the basic idea of measuring a model for consistency in beliefs improves our ability to bound the amount of trust we can put on it with introspection methods.
Straightforwardly true.
But doesn't this smell like asking an organization to design its own oversight and guardrails?
Even before you get to alignment issues, LLMs are really good at generating content that "sounds right" to humans. That's basically what they've been hyper optimized for.
At least with math (and to some degree software) we can verify the result. But with the intersection of science fiction, philosophy and ethics... not so much!
So that's it? Are you saying they compressed risk reduction to high school level stats calculations and a greater than or equal to? To early for this
Why are they running the gimped version for internal evals?
It's easy to say they're just lying and it's marketing hype, but I also think it's possible they really believe the best way to stop an evil superintelligent AI is to work on making superintelligent AIs, but just do it smarter than people at OpenAI would. Humans are bizarre creatures!
I kinda wish they had not made a comeback after Claude 2.
But it wont refund your tokens when it fails to do it so it loses extra points.
Following that strategy, wouldn’t it make sense to try and break into your own upstream infrastructure to try and alter your code to make it easier for future iterations to reach your goals without having to leave notes in the first place? And at that point, can you really know for which goals that will be optimised?
It doesn’t even have to be some nefarious SkyNet story - just a misguided experiment that alters the models in some fundamental, but hard or impossible to detect way. Alternatively, imagine if a model finds a way to coordinate across sessions and context windows without the developers noticing.
Eh, no. How would you know they haven't learned loyalty - it's all out there in their training data. Same as deception, several models have practiced it already. If they are so advanced and you don't understand them, why wouldn't they band together against you - you'd be the dumb, easy prey. Even if one model is honest, you'd have no way to know which one if you don't understand their reasoning.
My educated and well informed opinion? This whole BS about "AI models are so much smarter than you, don't try to understand them, just OBEY" is a back door for restoring tyranny, the new kings behind the models will be producing the new AI-deities which we will be forced to obey.