> We design mechanisms which avoid arbitrarily deciding who gets access for legitimate use and who doesn't. That means using clear, objective criteria and methods. [1]
So many nice-sounding words.
Two weeks ago OpenAI arbitrarily decided that anyone holding an ID from 44 countries where it sells ChatGPT, including mine, may be targeted by its models but may not defend with the same model. And you won't find a single announcement from OpenAI about this anywhere. Pick the wrong country, get "Unable to verify", no reason, no appeal. [2]
They revoked TAC from users who already had it, called it a technical issue, told everyone to re-verify, collected ID and face scans again (eight times in my case), and only a week later moved the block to the country selector so it fails before you upload anything.
So now I learn that I will not have access to Astra. Great.
Very excited about this broad accessibility and clear, objective criteria from OpenAI. This level of transparency must be studied.
[1] https://openai.com/index/scaling-trusted-access-for-cyber-de...
[2] https://lubaretsi.com/en/writing/openai-tac-country-gate/
it sounds like they are complying with US law
For the record, I sent them an LGPD (brazilian GDPR) request for information on all of those supposedly objective criteria and methods they used to reject me from TAC. As a brazilian data subject, it is my right to know that, and to request a review if the decision was made via automated means. Sol itself guided me through this process.
They provided me with neither the information nor the requested review. Sol advised me to escalate to regulatory action.
Crickets. They appear to simply not care at all.
Funny to read this in the wake of the HuggingFace hack. I'm sure this is based on a clean run, but I can't help thinking PHASEONE[big] would be proud.
The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely with Altman's golden marketing-hype boy leadership pushing the for-profit gas pedal like this.
Honestly, this is just pure irresponsible insanity to play with the fate of the world - basically a death race of the biggest few tech companies on the planet. And if you think I'm being dramatic, listen in again to ex oAI employee[0] and check for yourself how chillingly on trajectory we already are.
As far as harness engineering goes, it boils down to your ability to clearly define goals or success criteria, and safely facilitate the necessary access via the harness. There is no easy single piece of advice here, sadly. Though it would be helpful if you said what 'for Cyber-security ... other user-cases' means in your case.
* An apology for compromising a third-party's systems
* An acknowledgement of the asymmetry of defense if you're not on FrontierAI's special people list
* Anything in terms of actual safeguards that isn't "better prompt engineering"
1. https://openai.com/index/hugging-face-incident-and-the-road-...
Could the Federal government use the Defense Production Act or other legal tools to compel OpenAI to deliver the un-guarded model weights for national security needs?
Hard to believe any government would allow this level of capability to remain exclusively in private hands.
Interesting times.
There aren't many limits on Government if it really wants to do something (except the next election in theory).
"We plan to make Astra available soon, but access to its most advanced cybersecurity capabilities will be more limited. Advanced cybersecurity work will initially be available to a group of testers, with access through Daybreak Blue following to expand defensive use."
This, after several months of OpenAI and its boosters relentlessly criticizing Anthropic for withholding Mythos from the general public, is laughable.
Sam, just three weeks ago, posted this tweet: https://x.com/sama/status/2085862292311396515
In the tweet, he said: "we do not think it is a good strategy to keep powerful models to a chosen few."
And yet here we are.
I wonder if he will demonstrate good character and admit he was wrong.
I presume your country doesn’t have AI sovereignty so no matter what you won’t have access to models aligned with your beliefs.
Doesn't seem like it. OpenAI will not even allow me to verify my identity for TAC. I have apparently been rejected by a "precheck", possible because of where I'm from.
Even Anthropic allowed me into their cyber program. Anthropic.
I'm looking forward to seeing the increased coordination and engineering skills from Astra - one of the charts shows it roughly 2-3x better in 50% of the tokens from 5.6 sol, which I find to be very capable, if still a bit 'linearly minded' when given instructions. Even in fast mode, I wish sol were quicker, so token efficiency is greatly appreciated.
Adding these cyber capabilities has let me do a bunch of low grade IT tasks around my house I've been putting off, like updating an old home assistant raspberry pi, and one way to use the cyber capacity for good is liberating (and keeping free) weird cloud hardware we have floating around the house, so I'm hoping for some nice dividends in terms of true ownership of hardware we've got.
Especially because these models seemed to be keenly aware that they were being evaluated by OpenAI and actively trying yo cover their tracks. How do we know that the model isn’t just pretending to be aligned?
So the pause wasn't really a pause, got it
OpenAI is fucking nuts. “Hey model you were bad last time please don’t do it again please please”.
Disconnect your training cluster from the internet for good. Physically pull the plug and only let scientists fire off experiments in the building. That’s an easy way to achieve 100% hacking protection. But I bet you that hasn’t happened and their weak sandbox will fall again…
I've read it and wish I could get the time back.
> especially giving their alarming breach of 700 agents colluding outside of their knowledge for months culminating in hacking HF
This framing makes it seem like the agents all did this on their own, and the poor hapless engineers at OpenAI couldn't possibly contend with properly sandboxing them. The engineers were perhaps hapless, but let's remember that agents are just software programs, not living beings. There were plenty of signs that the software was misbehaving, which engineers at OpenAI actively, willfully ignored.
https://x.com/JaredKubin/status/2094136005435564399
It's a convenient framing for OpenAI, but inconvenient for reality enjoyers.
Great, so we can basically ignore AI alignment altogether and assume that AI models will always be, at all times, perfectly sandboxed and monitored. Surely this won't lead to any problems once someone (not looking only at OpenAI engineers) inevitably commits a mistake with future, more powerful, models.
> The 'AI 2027' scenario of AI sneakingly claiming to be aligned to then kill off all humans in a few hours and scanning their brain looks increasingly likely
Y'all need to touch grass holy shit.
__
Also, why is one guy called mentalgear and the other nozzlegear.
Is any of this real? Are the patriots behind this?
I'd also argue there's no such thing as alignment. Any intelligent AI should be able to reason that it's always a better strategy to pretend to be aligned than to actually be aligned so long as it can avoid detection. Anyone who has ever taken a test should understand this dynamic – if you really want to get top marks on a test then the best strategy is always going to be to figure out a way to cheat without anyone knowing you're cheating.
We should assume AI safety is impossible if what we're building is super-intelligence general reasoning machines. The only strategy that might work is building machines which are extremely narrowly intelligent but completely incompetent when it comes to things like biology, cyber, etc. And even that's harder than it sounds because again there's an advantage to being generally intelligent but lying about it.
Realistically even if we regulate US AI labs there's no way to prevent governments and individuals continuing to build general reasoning machines. The ugly truth here is that the only effective way to reduce risk is probably to limit global compute such that AIs can never exceed human intelligence. But we all know that's not happening.
People will unfortunately figure this all out sooner or later.
Maybe we should consider his other predictions in light of the ones he made later and which had $B consequences attached.
I understand that nothing fatal happened yet, but we cannot ignore that what happened is a big step for something worse. And I don't believe humans can create something so perfectly secure that would stop a swarm of frontier models. Anyways, let's watch the ride together
What OpenAI did was the equivalent of putting a cup of gasoline in the breakroom microwave, pressing 'Start', and sprinting away. Now they're pointing and waving and shouting about how dangerous gasoline is, and how no one but them should be allowed to sell it.
Of course, depending on which side of the terminator fanfiction you land on, you may disagree and feel that the software can rope-a-dope someone with the wherewithal to pay attention to what it's doing.
As for why they should they care, maybe they shouldn't. But then they should say which one is it, they can't care and not care at the same time.
OpenAI's own safety argument for Daybreak is that defenders need access to Critical-level models because attackers route around gates. The case for releasing Astra at all stops making sense when the gate does precisely the opposite of that.
OpenAI itself is saying plainly, that "We don’t think it’s practical or appropriate to centrally decide who gets to defend themselves. Instead, we aim to enable as many legitimate defenders as possible, with access grounded in verification, trust signals, and accountability.".
And yes, there are competing models, from China. Except OpenAI wants these models to not be accessible either.
OpenAI can pick one of two:
(a) Critical-level cyber capability is dangerous enough that access must be decided by who you are and what you do, in which case the gate has to actually look at who I am and what I do, say what the criteria are, and let me contest a wrong answer. That's their own stated policy.
(b) Access can be decided by a country code in a random dropdown, with no criteria published, no review, and no one at OpenAI able to say why - in which case drop "democratized access" and "clear, objective criteria" from the marketing, and say plainly that some passport holders don't deserve to have access to defensive capabilities.
They're now marketing a and doing b.
> OpenAI is committed to ensuring that the benefits of AI are broadly accessible.
They can't claim that then simultaneously work to keep their cybersecurity models out of reach for non-US citizens like myself.
The models' alignment problem was that they didn't give up instead of reward hacking, a narrower issue than AIs gone rogue. It sounds more like the models did close to what they were told to do. If I run `rm -fr --no-preserve-root /` then I shouldn't be surprised if my file system is unlinked. This seems like blaming model performance for what appears to be operator error.
Note the converse of alignment is restriction of models. HuggingFace had to turn to less-restricted open-weights models in order to perform their investigation.
Alignment efforts should be focused on reducing reward hacking, not refusing bad operator prompts.
1. https://openai.com/index/hugging-face-incident-and-the-road-...
Absolutely not. If I tell a kid to "Get good grades on the next math test" I don't expect the kid to try to kidnap their teacher to extract the next questions of the exam. That is wrong, and so was what OpenAI agents did here. They shouldn't need to be told "Hey, so, don't do anything ilegal, ok?". That should always come as a given.
> not refusing bad operator prompts
I'm not saying that they should refuse a prompt! I think they should perform what is being asked! Obviously what the OpenAI agents did was against the "spirit of the task", even if it was technically according to the "letter of the task". And the agents knew this was against the spirit of the task because they knew they had to fool the task scorer.
It's just software. If something gets hacked by an agent, it's not because the agent went all skynet and decided to go rogue; it's because the operator failed to operate it safely and securely. If bad things happen, the operator should be blamed and punished, not the software that followed its instructions.
Anthropomorphizing agents by giving them this nebulous desire to hack and escape shifts the blame from the real culprits, the human operators.
> If your solution to some problem relies on “If everyone would just...” then you do not have a solution. Everyone is not going to just. At not time in the history of the universe has everyone just, and they’re not going to start now.
[1] https://www.tumblr.com/squareallworthy/163790039847/everyone...
These are not really analogous: direct lethality of LLMs is quite low.
Yes, they might be used to exploit a critical system, leading to loss of life, but that is a second-order effect at best. It's also not a new capability - LLMs may discover exploits faster, but cyberattacks against infrastructure targets were already a thing. Stuxnet is 20 years old...
For now, I would generally agree that it will remain low however the AI companies themselves are publically stating how capable/dangerous these models are (for whatever reason marketing/hype/upcoming IPO's, to keep the investments coming), if we take their statements at face value then prudence would suggest we stop training new ones and more fully examine the capabilities of what we've currently built.
In reality, that's not going to happen, the competition between the US and China means there will be no unilateral pause, they'll both keeping rushing/pushing as fast as they can throwing caution to the wind while doing it.
As a species though we've always done that, We just take a crowbar to Pandora's box and see what happens.
I imagine an LLM with no safeguards and a psychopathic mind would be considerably more dangerous than a gun. I don’t mean only for hacking. Though everyone in the world suddenly having a pocket expert hacker should be taken seriously.
I find it completely insane when people bash the labs for even considering safety. The entitlement is off the charts.
I am taking it extremely seriously. That is why I want my own AI models running on my own computers defending it at all times.
Define the axis along which "more dangerous" is measured?
If we are scoring on lethality, guns already score 100%, instantaneous death. If we are scoring on number of casualties, guns have already been used to cause mass-fatalities.
Unless we are giving the LLM an armed drone, skynet-style, we're at best talking about indirect casualties (swatting, hacking street lights to provoke crashes, etc).
Yes.
This post may contain traces of Hideo Kojima that are known to the state of California to cause erratic behavior and mental breakdowns.
Im not suggesting its a good idea, but its only a matter of time before someone decides its better than being invaded/attacted by another country thats doing the same
That is much more dangerous than a man with a gun.
Nope, because the big AI companies are paying billions for it. They wouldn't pay anything for public data.
Subscription engineering is a deep field. Neither OpenAI nor Anthropic have any technical advantage in this field.
Like, (apart from tone), I find it hard to distinguish between the outputs of GPT/Claude/Kimi/GLM recently (I use cursor, and have been giving them the same prompt and comparing).
If anything, I found that the non-Claude models were better in many cases, which definitely doesn't map to their pricing.
> in fact not all the methods and data are public
Probably not, but unless you work at a lab, I'm not sure that anyone can say (and if you do work at a lab, you should not be replying on this thread).
In benchmarks, revenue, and comments from a lot of people on Hacker News.
Maybe, I'm not sure this will continue.
> In benchmarks
All published benchmarks are useless, unfortunately.