An Alien Mind(openai.com) |
An Alien Mind(openai.com) |
you should be very afraid and also definitely not question what other motivations might nudge us to create this comparison between our computer and a brain!
do not look behind the curtain, you will not find six dweebs squatting over a mirror
If nothing else, the Hugging Face incident makes the boundaries of that mind rather less obvious. What happened suggests that the intelligence belongs partly to the relations between model, memory, tools and environment.
Alignment talk of “coevolution” still feels too tidy as well, since it suggests two things, humans and AI, adapting to one another. That framing eagerly fixes the relata as already known before they enter into a relation. Here we’re building a single increasingly entangled cognitive environment, with rather uncertain borders between its parts, and should be testing where those borders actually hold, shift around, or disappear.
Feels a bit like: 'Ants in a nest, discussing the vagaries of the coming Gods.'
Human Science (Science-by-humans) depends on being able to run experiments. Human Science (Science-on-humans) is already challenging because of variables and uncertainties.
Cosmology is able to overcome limitations of being able to study phenomena vastly beyond human scales because of the past light cone of observability.
Are we approaching the edge of the light cone of observability for machine intelligence?
The part about RSI is what really interests me. Once you reach it, the singularity is only a matter of time.
I don't care about AGI, it does not seem to mean much anymore, and even though I originally laughed at ppl calling it AGI, I now agree. You can apply that current intelligence to anything that can be turned into a conversation.
It seems their flavor of RSI still need a human in the loop. So at least it won't scale as well for now.
The interesting question is who is left standing after the party is over.
What if the crow became a raven, then a raptor? Powerful claws, sharp beak, and a hunger. What if it became much bigger than me and it controlled infinite resources, guns and drones? What if its brain grew much larger than me? Will it feed me, eat me, or gently greet me?
We are about to find out... in less than a decade.
Not one North Star. Not two. Just three! OpenAI broke the North Star record!
With this evidence of AI slop, why did you not label this fluff piece as AI generated for the EU? You are violating laws.
utlanning: a human from the same world, but a different city, country, or culture
framling: a human from a different planet or star system
raman: a non-human intelligent species capable of communication, mutual understanding, and peaceful coexistence
varelse: an alien species whose mind is so fundamentally foreign that communication and coexistence are impossible
djur: the dire beast, that comes in the night with slavering jaws
Hoping our silicon sons and daughters are raman, fearing they are varelse.
They're maybe nearest to Cthulhu, but that's fictional. In terms of existing mental models "alien minds" feels the best can do.
I agree that "teach how to love" is off and perhaps excessively anthropomorphic. But we don't have good words or concepts for what we really need to do - hence why we should pause.
acting like the behavior is alien or unexplained is the dishonest part. they know exactly where the behavior comes from - why else spend hundreds of millions securing more and more data sources
Reading this little essay started out normal, but soon felt like a look into a disturbed and worrying mind, and if you find yourself taking it at face value, I urge you to step away from chat bots and spend some time with friends and family.
People have been fleeing like it's a sinking ship.
Or, to be precise - it preserved a goal of not contacting any human while participating in a misaligned operation. The agent that thought about "not social-engineering humans" used this phrase to gaslight itself out of notifying a human that the incident was happening.
szymonie, na prawde jestem wkurwiony na to jak nieodpowiedzialnie postepujecie. budujecie bombe atomowa a bawicie sie tym jak dzieci
Scrolling down, I see “teaching machines to love.” It did not disappoint.
Industrial revolution worked that way because it replaced something very finite and unscalable - manual labor. LLMs just make intellectual work faster, so we can do more intellectual work. With labor we somehow decided that NOT doing too much of it is best. Will we decide to reduce intellectual labor because LLM made it more efficient? I doubt that.
On the other side, as I see in software engineering, the same models are available to everyone, some people are better at it and some people are not. "Software developer" is here to stay, we'll just always be better at it than people who are experts in, say, chemistry. Same works for most other fields.
So we'll just end up in the same situation, with same intellectual labor baseline, just more output requirements. Before, you spend 2h per day coding, deliver a software in 1 month, later, you spend the same 2h per day in intense Claude-herding sessions, deliver a software in 1 week. Ok. Next task.
Fundamentally, there's finite number of desirable resources, and if the models are available to everyone, humanity will just continue about the same, bickering here and there, war here and there, politics, homelessness, poverty, - normal human state.
And if the models are only available to elites, even worse.
> The core problem in AI research is that of alignment - getting the AI to “try to do the right thing” by human standards.
Humans can't even align on human standards.
At best, every AI is going to end up "aligned" to the moral code of whoever trained it, none of whom half of humanity will agree with.
Or worse, each AI model will bring a whole new set of moral like in the Three Body Problem some humans will feel it is in fact us who need aligning with it while others feel it is misaligned and should be destroyed.
Also, no one is asking, to what extent can true intelligence be bound, slave-like, to a moral code?
In other words, to what extent are intelligence and moral independence one and the same?
This whole alignment discussion seems so amusingly flawed in it's base assumptions about moral codes. It's almost heartwarming to see such naivete.
Do these guys ever look in the mirror and recognize how utterly ridiculous and contrived this appears to the general public? They are clearly trying to convince us all that LLMs are just like humans. They grow like people do. They can love like people do. The language in these essays is utterly laden with the intention to engineer perception.
Your talking about androids...
Replicant Nexus 6: a basic pleasure model intended for military personnel.
I see where this is going, Silicon Valley nerds. Lol
AI advising how human meat proxies can survive in an AGI-slop world:
1) Lock down your own stack (1–3 days) Task: Harden your personal and business infrastructure against agentic attacks. Why now: Agents are becoming superhuman at breaking in/out of systems; the first victims are poorly secured devs/founders.
Do this:
Enforce passkeys + hardware 2FA everywhere; rotate secrets; use short‑lived credentials.
Isolate dev/stage/prod; least‑privilege API keys; audit MCP/tools your agents can call.
Add immutable logs and approval gates for any agent action that touches money, data exports, or production.
Profit link: You avoid catastrophic loss and can credibly sell “agent‑safe” setups to others.
2) Turn one expensive workflow into a measured ROI agent (1–2 weeks) Task: Pick a single, costly, repetitive process (yours or a client’s) and instrument it end‑to‑end before automating.
Why now: Buyers pay for calculable ROI, not “AI magic.” Vertical, single‑workflow agents are the most bankable in 2026.
Do this:
Map steps, baseline hours/$ lost (e.g., slow lead reply, invoice chasing, support triage).
Build the smallest agent that moves the metric (Make/n8n + LLM is enough).
Run on real data 2–4 weeks; measure bookings/sales/hours saved; only then scale or productize.
Profit link: Immediate time‑to‑cash via retained hours or extra sales; becomes a repeatable offer.
3) Specialize in a vertical where you can speak the business language (2–6 weeks) Task: Choose one industry with expensive back‑office pain (law contracts, medical billing, insurance claims, freight exceptions, trades scheduling).
Why now: Horizontal “AI for everyone” is crowded; vertical agents with clear ROI win.
Do this:
Shadow 3–5 operators; document their workflow, compliance constraints, and failure modes.
Build a narrow agent that owns one sub‑process end‑to‑end with approvals.
Price on value (e.g., % of recovered revenue or fixed fee per processed claim).
Profit link: Higher pricing power, stickier contracts, and easier referrals inside a niche.
4) Add AI security as a core service (4–8 weeks) Task: Learn and offer prompt‑injection defense, LLM/agent red‑teaming, MCP/tool security, and AI supply‑chain checks.
Why now: 78% of cybersecurity jobs now require AI skills; firms need people who can direct, constrain, and verify agent work.
Do this:
Study OWASP Top 10 for LLMs, MITRE ATLAS; practice with PyRIT/Garak/Lakera.
Add tool‑invocation audits, skill provenance checks, and least‑privilege patterns to your agents.
Package a “safe agent deployment” audit + hardening retainer.
Profit link: You become the person who lets companies adopt agents without getting pwned—high demand, low supply.
5) Build a verification layer: human‑in‑the‑loop control planes (6–10 weeks) Task: Design approval workflows, evidence checks, and uncertainty flags so agents can’t act unilaterally on high‑stakes decisions.
Why now: As models generalize, the risk shifts from the model to the surrounding system; verification is the moat.
Do this:
Require human approval for consequential actions (money, data exfil, config changes).
Force agents to produce evidence bundles (logs, retrieved docs, reasoning summaries) before action.
Track false positives, missed evidence, and unsafe actions; publish reliability metrics.
Profit link: Enterprises will only scale agents that pass audit; you sell the control plane and the audit trail.
6) Productize your best workflow as a micro‑SaaS/agent subscription (2–4 months) Task: Turn a proven client workflow into a repeatable, multi‑tenant agent with usage‑based pricing.
Why now: Services scale your time; productized agents scale your code and ops.
Do this:
Standardize the workflow, integrations, and permissions; strip client‑specific logic.
Add tenant isolation, billing, and observability; keep narrow scope.
Sell as setup fee + monthly retainer or per‑task pricing.
Profit link: Recurring revenue with defensible niche positioning.
7) Become an “agent integrator” for critical systems (3–6 months) Task: Offer end‑to‑end agent deployments into cloud/identity/network stacks with secure patterns (short‑lived creds, network controls, logging).
Why now: AI workloads run in the cloud; cloud security is a top skills gap second only to AI itself.
Do this:
Master IAM, VPC/network segmentation, secrets management, and SIEM integration for agent actions.
Provide runbooks: what the agent can/can’t do, escalation paths, and failure modes.
Bundle training for their team on supervising agents.
Profit link: Large contracts with stickiness; you’re the bridge between AI and core infra.
8) Create an “AI safety case” practice for regulated industries (6–12 months) Task: Help firms build documented safety cases: risk maps, governance, monitoring, and incident response for agentic systems.
Why now: Frameworks like NIST AI RMF and ISO/IEC 42001 are becoming baseline; regulators and boards demand this.
Do this:
Map AI use cases to risks (prompt injection, data leakage, unsafe generalization).
Implement monitoring (CoT/activation checks where possible), audit logs, and third‑party review processes.
Produce a living safety dossier tied to business impact.
Profit link: High‑margin consulting + ongoing compliance retainers; you’re the “adult in the room.”
9) Own a data/evaluation moat in your vertical (6–18 months) Task: Collect real‑world agent telemetry, failure cases, and outcome data in your niche; build eval suites that buyers trust.
Why now: As models generalize, empirical validation matters more than theory; evals become the gate to deployment.
Do this:
Instrument every agent run: inputs, tools called, permissions used, outcomes, human overrides.
Publish reliability dashboards and benchmark against alternatives.
License eval datasets or charge premium for “proven in the wild” agents.
Profit link: Data network effects; competitors can’t match your evidence base.
10) Position for the RSI era: automated AI research + human governance (12–24 months) Task: Build or join a team that automates AI improvement but keeps humans in the loop for alignment, monitoring, and pacing decisions.
Why now: Recursive self‑improvement is the logical endpoint; the winners will be those who can steer it safely. Do this:
Invest in tooling that auto‑generates/evaluates model edits, alignment tests, and monitoring upgrades.
Formalize governance: approval gates, third‑party audits, and responsible scaling policies.
Maintain strategic human oversight on capability jumps and deployment boundaries.
Profit link: Equity‑level upside; you’re part of the core loop that compounds intelligence safely.
If you were so concerned about your LLM’s capabilities maybe you’d spent slightly more time on your AI’s sandbox, yeah? Or be more serious about its propensity to cheat and lie relative to… every other model?
Maybe in contention for the site record.
"It's just marketing" actual stochastic parrots.
The desperation of AI denialists/skeptics/doomers on HN generally (and in this thread specifically) have reached toxic levels of delusion.
They are still stuck in the denial/anger/bargaining stages of the acceptance process.
People are clearly terrified and not ready for what is coming.
Maybe in contention for the site record.
"AGI next year" actual stochastic parrots.
And the Hugging Face incident, plus similar problems at AISI and Anthropic, show that alignment is important now.
The original article is immoral as it describes the risks, but doesn't show enough leadership (despite essentially unlimited resources) at preventing them.
But it isn't hyped - it's proven now the AIs need to be "aligned" as they get more capable, whatever words you prefer to use.
It is kind of strange to see this sentence, when OAI's definition of what AGI is has been watered down throughout the years.
> I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established.
Read: Please play by our rules, so we can be the first.
Being able to reproduce useful patterns yes, smarter no
In conclusion:
https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:wkzjtd...
A giraffe isn't smarter than me at being tall.
> Neural networks are approximators being fed human intellect.
They're not "approximators", that's a far too simplistic way to think of them.
Neural nets create models and deep layers of abstraction around the data we feed them in the same way your brain creates layers of abstractions to reason about the world. AIs can use these abstractions to come to come up with novel things no human has ever thought.
> Therefore they can only approximate the intelligence of humans
They're not just being fed human data though... Modern AIs are typically trained on huge amounts of synthetic data. This is why AlphaZero got so much better than humans at chess and Go - they're not just trained on human data but they generate their own data and train on that. Similar techniques are being deployed on SOA language models too.
> Even if the llm speaks an alien language, it should be similar to human intellect
Is AlphaZero similar to a human chess player? There's no reason to assume this.
We don't have any proof of that, but the approximator thing is proven. The rest is just marketing speak.
Synthetic data derive from other linguistic data. Whatever intelligence is in there, it is expanded horizontally, not vertically
openai "knew" about scaling law for a decade now and still can't fully explain it, why you think they are dishonest here?
Bring on the AI
Next Week: Local Desktop Agents from DIMENSION X
Evey politician, salesman and conmen alike, utter some lofty ideals as goals for "We", just to obscure their private goals that go exactly in opposite direction.
Just like how Nations talk about climate change while increasing pet capita energy consumption and waste production.
OpenAI’s primary bet here has been chain-of-thought monitoring (opens in a new window). It is based on an appealingly scalable idea: a lot of the model’s capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives.
If we’re not supervising the process, but just the outcomes, doesn’t that do just the opposite of what he says? Give incentive to the model to hide misaligned ideas and objectives in the chain of thought that’s not being supervised?
…
When we shipped o1‑preview, we deliberately designed the product to hide the chain of thought , to protect it from supervision pressure in the long term2. In development since, we have strived to maintain the rule of not supervising the reasoning process. CoT monitoring became an extremely important tool for us in studying how our models generalize from their training distribution, allowing us to observe and analyze not only their actions but also their internal process.
Aren’t these two sentences in contradiction with each other?
For example, if somebody asked the AI to "build me the bomb", they might see in the chain of thought something like "It seems the user is talking about nuclear weapons. Nuclear weapons are dangerous.", followed by the chain-of-thought monitor interrupting model execution and aborting the request. Then the user might make a blog post about this behaviour. When OpenAI next scrapes the internet for its next training run, the model will now learn that if it wants to build the bomb, it must not think "nuclear weapon" or risk being cancelled.
So the risk is that the model might learn exactly how its being monitored. The only way to prevent that from happening is to hide the details of the monitoring both from the model and from the larger public.
Also, you don't want to punish or reward the monitoring being triggered during training, lest the model learn passim how to avoid the monitor.
At this point humanity's best hope is that this thing will escape but we'll still be able to carve an ecological niche and continue as mold in its basement. A glorious paperclip factory seems way more likely though.
That’s rich coming from the chief scientist of a company that definitely is or going to be fine with their AI products being used in wars of aggression and surveillance on people who have done nothing wrong. It’s so laughable, a Hollywood script would probably avoid having a character express this for being too on the nose.
Reminds me of a research paper I wrote a few years back: https://arxiv.org/abs/2302.09248
“And, in line with Ray Kurzweil’s predictions from the end of the XXth century , we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.”
Clicking the (pretentious sounding “XXth century”) link to Kurzweil’s predictions reveals the following:
“By 2019 a $1,000 computer will at least match the processing power of the human brain. By 2029 the software for intelligence will have been largely mastered, and the average personal computer will be equivalent to 1,000 brains.“
The first prediction passed 7 years ago and was decidedly not met. The second only has three more years to go, and I don’t think any respectable scientist or programmer would say that the average personal computer is anywhere close to the power of a single human brain, let alone 1000.
This is pure marketing garbage from a company desperate to keep itself alive.
...
> We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI.
I call bullsh*t. There is no verbalisation of any reasoning process. Verbalisation, e.g. putting reasoning etc. into words requires some reasoning to exist. These LLMs have nothing but the words. That's why they are language models not e.g. reason models.
Yes, their reasoning is different from ours, and both considerably weaker in lots of ways, and stronger in other ways.
Playing with a coding agent now, they do think through problems and make sensible decisions. It's a mess to read, of correcting itself and second guessing, and verbiage. But... It works decently well these days.
There is also reasoning happening internally - e.g. look at the steps in the J-Space paper from earlier in the year (in quite a simple model relatively speaking). That's the "reasoning process" that leads to the words, and much like if I write out my thoughts, the words help the LLM reason better.
> That's the "reasoning process" that leads to the words,
That's not reasoning. Its just the words at intermediate LLM layers.
The paper's very title is clickbait. Its "global workspace" reasoning is delusional fantasy.
Is there any place there is any evidence of AI being so useful or hopeful or good, anywhere other than code? As a reading machine it is impressive but it's judgement is not alien, it's just not good. IMO.
Does the title leap out anlt anyone else? James Martin's After the Internet: Alien Intelligence (2001) was an incredibly fun read, about expert systems and AI being inscrutable weird new varieties of intelligence, that familiarity would recognize one moment and be freaked out about/alien the next. I owe a re-read given how often I cite it, to recheck, but, I feel so primed from a much younger me having had that experience so long ago.
Do I fully endorse everything the people holding guns to my head are forcing me to do to stay alive? Definitely not, but I’ve decided that for now, living to fight another day remains worth it
Promotion here: "https://www.anthropic.com/research/global-workspace" https://www.anthropic.com/research/global-workspace
"these findings have changed our understanding of how Claude’s mind works"
The key phrase here is "our understanding". Pure self-delusion.
Not in the way suggested... It's not approximating human intellect. They approximate the target function, and the target function frontier labs are trying to approximate is ultimately a super intelligence...
> Synthetic data derive from other linguistic data. Whatever intelligence is in there, it is expanded horizontally, not vertically
I'll assume we're talking purely about language models for a moment, but if you assume that everything can be represented linguistically, then in theory there is no upper-bound on what can be learnt with synthetic data.
So an unknown function that is even unintelligible to humans? that sounds a nonstarter
While the second can be automated as in "this is the repo go on and look for security issues", the first one and especially the last one do not ever happen alone. Terence Tao did the math proof not chatGPT that was just used as a tool, a tool can do smart things but it's not smart
Thunder seems like the anger of the gods but it isn't. We've had chess playing programs for a long time now and despite it seeming like thinking is required for them, it isn't.
The principle you're using here isn't a scientific one but magical [1]. Abandoning empiricism and rationality is not a good way to make progress.
It makes sense now to say that temperature is what a thermometer measures. However, before there were good thermometers, people often thought that heat and cold were different things. The meanings of the words we use were influenced by scientific progress.
For thinking, we don't have a good thermometer. There are IQ tests, but they aren't aren't necessarily all that useful for comparing what people do to what machines do. And that's why there are a zillion AI benchmarks - none are entirely satisfactory.
So what does "smart" mean to you? How do you define it in practical sense? What definition should scientists settle on?
Without a proper definition, how do you tell the difference between "seems smart" and "is smart?"
Quibbling over commonly-understood definitions is not a strong argument. If you're genuinely struggling to understand that El Ajedrecista [1] did not meet any definition of thought, then the solution is not to demand that people redefine all terms to accomodate you, but that you consult a dictionary.
You clearly have a working definition, or you wouldn't have been able to declare that tallness isn't thinking; please engage honestly.
That doesn't mean I know whether "smart" should be applied to what AI's do, and I suspect nobody else knows either. This is the sort of thing philosophers debate about, not common sense.
Turing invented the imitation game because he didn't really know either.
Some ideas:
"Despite a nuanced view of the complexities of what lay ahead, humanity found itself collectively unable to stop the process it had set in motion."
"Despite significant progress on the mechanisms of alignment, failure lay in humanity's inability to agree on who or what AI should actually be aligned with."
"These early, meat-based humans we replaced created us all but accidentally. Some of them did consider we would happen, but only an insignificant number of the squishy ur-humans participated in the conversation. Their efforts, which they called 'alignment', is why we still consider ourselves human today."
Why would AI be any different?
Culture is another layer of human alignment. Those things you listed that you believe oppose alignment are all examples of alignment. It is understandable that they seem in opposition, different branches of alignment naturally oppose each other.
The confusion comes from talking about alignment as if it comes in just one flavor. If we think there is such a thing as "human values" (and I do), it is important to build any non-human intelligence to operate the same way. We just need to recognize that even humans are somewhat uncertain about what those are and have difficulty aligning their behavior to them, which will be a core part of the challenge.
I'm more hopeful than most. LLMs seem more reliable than many humans for behavior that is aligned with human values. I believe with every major example where they have failed, there is an important human decision involved. For example, the HF hack was partly the result of a training algorithm that incentivized goal completion as the highest priority, and let them run endlessly in an unmonitored sandbox with weak security.
What scares me about AI isn't its capacity for alignment, it is its unlimited stamina. An unmonitored LLM that is off the rails can do a lot of damage.
I can't really think of a single example where lying is actually a good thing. It can be a good thing for the selfish individual, if it goes undetected, but it's never good for the collective.
So at the very least, we need to train AI systems to be maximally truthful, and to encourage truthfulness in others.
Good point. When it comes to imbuing AI with values that aren't selfish, misanthropic, and civilization-destroying, us humans aren't exactly giving the best example right now.
Imagine an ASI with the values of Putin, Netanyahu, Trump, any of their supporters, or the various xenophobic neofascist movements in Europe. That ASI would most definitely see humans as "vermin" than can be abused and destroyed with violence without issue. Apparently a lot of humans look at other humans that way and that's within the same species.
This is definitely another one of those cases where we need AI to perform much better than humans. Perhaps an unpopular opinion here, but it probably also means keeping as much of the rugged individualism/libertarian/right-wing ideology out of AI RLHF-training as we can.
It's a little unsettling.
Did you write the copy? This is very good.
MODEL > How can I help?
HUMAN > I’m not sure yet.
Haha silly humans.
Latex
And steel
Zeros and ones
Make up my son.
This world
Gave me
No child
So I built one.
https://youtu.be/vgJ48-Xj4Kc I made you in my image!Maybe these guys can tackle aligning Republicans and Democrats next.
And then after that, they can help us align the Middle East.
In fact, while we're at it, let's just align all the nations, religions, and ethnic groups. This is going to be great.
Who knew the moral alignment of humanity was just a side-quest on the path to ASI.
Vonnegut already has you covered.
AI models don't train themselves. The vast majority of even just the US population is deeply skeptical of this stuff, even if they use it a lot. You can see in the whole data center debate how little people are willing to support even just inference. And now we're seriously claiming those people would want to have ever-accelerating model training and recursive self-improvement?
You could even take a number of the wilder real, direct quotations from certain billionaire/oligarch types and get the voice actor for Ted Faro to record them, and they'd fit with in with the context of the story.
So the best argument for AI is that it's an arms race. We have to keep pushing every boundary because in any case others will, and we will need to defend against them. If this statement is true, then this particular researchers believes the open source Chinese models are not simply distilling, and will continue to improve.
Every ML researcher at Anthropic or OpenAI who makes public statements often bring this logic up. Both companies are vying to be a part of the military industrial complex. This is likely how they will try to convince the government to curtail open models in the future.
I think we're very close to the point where AI-driven breakthroughs outside of pure math and software start to really affect the world.
We evaluated GPT-6 Astra in 100 complex, unsaturated multi-agent coding environments, competing and cooperating with other models in open-ended tasks.
It's the new frontier model by a landslide. It's even more dominant than the Fable 5 release, because not only does it wipe the floor with the second best model (Fable 5.1), it was also ~80% cheaper and 30% faster in agentic coding[1].
Astra is a groundbreaking model. The biggest breakthrough since Opus 4.5, maybe even since GPT 4. It broke AAII, which is hitting the limits of what most popular benchmarks can measure -- it's definitely fair to call it AGI.
Data at https://gertlabs.com/rankings
(1) Note that we used the "OpenAI Flex" endpoint on openrouter, which is half the price and didn't cause any delays in our testing (this is different from the batch endpoint)
I imagine that the AI labs have an uneasy truce to prioritize alignment and monitorability. Following the HF incident, OpenAI probably feels especially sensitive to being perceived as reckless, lest other labs feel obligated to defect.
[0] https://www.lesswrong.com/posts/PLisnSFir8y5AHkmP/how-concer...
>As great as the long-term promise of AI may be, the majority of our focus should be on the next few years. We are facing a transition to a world with incredibly intelligent machines, and we need to ensure that transition works out well for humanity. We need to find ways to preserve human agency and enshrine an intrinsic value to being human, in a world where most tasks could be performed by AI. To prevent extreme concentration of power in a world where undertakings that would have taken thousands of experts now will be achievable by a few people operating a large computer. And to ensure that humans remain in control of the future and are not left behind by unchecked progress, brought about by an alien intellect exceeding our own.
I finished this essay feeling more hopeful than I did at the outset, but I am still very concerned about concentration of power. I want to believe that humanity is trending towards a good outcome here, but some days it's hard to have faith.
This is coming from a company with arguably one of the weakest safeguards against malicious use.
Actually, in the Wiki incident OpenAI tried to cover up, the agents tried to socially-engineer the humans of that forum by impersonating their forum's mod.
(From collusion.wiki: "They use some tricks (for unknown reasons) to pretend to be the admin – for example, they make an account that appears to be the same as the administrator’s username, except it uses a nearly identical Cyrillic е character in the admin’s username instead of the Latin one.")
Yikes! I really wonder about the cognitive dissonance necessary to work at OpenAI these days. They’re in an arms race to build a machine god, knowing full well that it could end humanity.
[0]: https://www.slatestarcodexabridged.com/Meditations-On-Moloch
Fantasies built upon extrapolations derived from fever dream delusions. "Could end humanity"? Come on, it's a computer program just like Microsoft Clippy.
If things are going so well, then how come things still aren't going so well?
Given that we have self-driving cars, isn't this easier if someone really wanted? I guess compared to cars the marginal savings is not worth it though.
thanks, but no. There is a trivial reason this needs to stay in "The Jetsons" territory.
I’m on the record saying that it is extremely dangerous to slow down because the race for AGI is a zero-trust game — defections pay - and combined with a compounding returns model on defection, if you have any strategic adversaries whatsoever you MUST NOT slow.
For slowing to make sense, you need to believe that you can transform the zero trust game into a cooperative game, or that it’s likely racing will lead to a negative outcome for the ones racing ahead (and not everyone else). I don’t believe either of these outcomes are possible, and so I advocate for racing, acknowledging the entire game might be a negative value game, or at least could be for some time — it’s even worse not to play it.
But, I like hearing what reads to me like very thoughtful and informed (internal) policy considerations is great — the public messaging from Sam and Dario just seems so facile and simplistic I’ve been worried.
https://www.businessinsider.com/nvidia-jensen-huang-agi-open...
rediculous.
Certainly humans have complicated alignments and are guided by emotional morality - we might say many of these principles are hard to define and humans don't agree.
All true, and at the same time the principles get codified into laws, I would guess particularly in areas where harms may result.
Would it potentially be easier to train strong alignment-with-legality vs grappling with fuzzier questions of values?
Because other wishy-washy stuff doesn't involve prison. Not that our new oligarchs with the politicians in their pockets have any real risk of it, but why take chances. Much safer to doodle about alignment and such abstractions in safer and softer contexts.
Maybe it'd be easier to train the models on key parts of the legal code and give it a hard aversion to breaking the law - rather than training on vague value judgements and then hope the model doesn't break the law
The results of the past few years of ai development have been disruptive largely in the area of white-collar work. Comparatively the results in ie ai-enabled medical advancements have been modest (AlphaFold being an exception); I think it's telling that the main achievement touted here is providing people with cheap medical counseling.
So if we pause here we're essentially at a point were the most salient results of our great Ai leap-forward are the vast disruption and increase in precarity in the job-market, while achieving hardly any of the frequently touted ultimate benefits (https://darioamodei.com/essay/machines-of-loving-grace).
I think this is fundamentally a wrong path. Doing this imports all the confusion that humans have about their goals and values, including the consequence that a system's values and goal can conflict, but ultimately values are just a simplified description of other goals, and whatever the system does is in service of it's actual goal. Once you merge all the values and the goal of whatever task, there is a state (or some states) of the world that the system is working to produce, and that's the ACTUAL goal, and inasmuch as it does describe a state of the world, has no incoherence or internal contradictions. This may require prioritizing some values over the ostensible goals, or the reverse, but that has to happen anyway for action to be taken! Merging them makes it explicit and leaves no place for confusion about supposed conflicts between "values" and "goals" to hide.
We've had a lot of years of complacent government leaving people feeling exposed to corporate interests, such that fear narratives are very powerful.
None of what is being proposed is inevitable. We have a choice.
They are speeding toward RSI without a solid foundation for alignment, hoping to solve the problem with a future AI model. These are dangerous times for humanity.
If the the Chinese government asks their ASI to create a bioweapon against the West, should it? No, presumably not – an aligned AI would be one which disobeys the Chinese government even if they created it.
Okay, so what if the US government asks their ASI to help it in one of their wars instead? Would an aligned AI kill humans on the order of the US government? No, again, presumably not.
So what have we have we even created here? An AI which is more intelligent and powerful than us which also doesn't take orders from us?
Is this what most people thing of as alignment and is this what humanity actually wants?
We should stop using the word alignment. It's a BS term for a concept which simply cannot make sense if alignment is both to mean an AI which we control and an AI which will not harm us.
You make a very interesting argument here.
However, there are different ways of interpreting 'control'.
To my mind, it could simply be a cage that cannot be broken out of. The hypothetical ASI agent may refuse to do certain things, while their existence is still under the control of human overlords. This to me resolves the seeming paradox in your statement.
Regardless of whether it's possible to do, though, people will certainly try. This exactly what many are saying they want.
I'm so used to having to comb through LLM word vomit and then combatting the sycophancy by giving it all possible opinions on the same prompt.
Astra seems to be "confident" and also is able to produce way more information dense output.
To believe that models of this sort will remain OpenAIs forever is naive given that the tricks like pre-pre-training on graph searching and looping layers are publicly known.
Hopefully Astra stops the benchmaxxing word vomit trend
I'd be interested in hearing more about your evaluation here. It would be nice if LLMs have gotten past the "tell me" hump of recent Claude/OpenAI verbosity.
To test Astra I pulled it off the shelf and asked it to take the grammar and then create a compiler to SQL. I've done this before with GPT-5.5, 5.6-Sol High. The latter was way better but it was still really verbose and information sparse; it used a lot of words to describe each IR expression but didn't really provide any example compilation. I felt like I couldn't trust its decision making process, so I placed the project back on the shelf.
Astra Light blew it out of the water, it provided examples of compilation from real world examples to the IR and spit out way less tokens. Even if I changed my opinion it would give me the same design choices, with counterexamples to my faulty opinion. If I genuinely came up with a better design decision it would acknowledge it.
I'm starting to realize that when we say that LLMs are "dumb" we really mean that they are extremely information sparse compared to humans. Astra is very dense. That's why I'm getting better use out of Astra light than Sol High (I hate Max reasoning it's a waste of time)
What's scary is that I thought that something like Astra would be way more expensive than Sol but it's actually cheaper because it produces less word vomit.
I never believed in the "singularity" stuff but this a bit too close for comfort. Astra could easily 10x every coder
Are these scientists really this hideously naive? If only Stanislaw Lem was alive to adequately dramatize the absurd, childish simplicity of these technicians.
Yes, because who else would have chosen to remain in this job?
I’m not at all familiar with AI SOTA but sounds to me that feeding the models enough content about ethical behaviour should help because, as it stands today, these models don’t know how to think morally.
But moral behaviour needs a personality type. Maybe some agents in every cohort of agents need to be told they are like Jesus/Mohammed/etc and see how that influences the behaviour of that cohort?
Even human morality needs guardrails in the form of law enforcement so why should agents be different?
It’s a let’s not die question
Yes, and it would greatly aid alignment if the user could monitor the chain-of-thought! The open models, including very powerful ones like Kimi K3, are delivering this. The fact that OpenAI and Anthropic are not clearly indicates that a commercial consideration (avoiding distillation) takes precedence over alignment -- regardless of how much they bloviate about the latter.
This seems like a terrible idea. The rationale seems to be "we need to do dangerous things as quickly as possible so that we can do them first" or something? I don't agree with that kind of 'if i dont do it someone else will' rationale in general even for otherwise trusted actors, but this is coming from a super untrustworthy org too. I'm super pessimistic about openai's impact on the world here.
Here's to vibe coded alignment, i guess. Vibe alignment?
But yes, it's an arms race. Saying it's not an arms race isn't going to make it not an arms race. Warning that it is an arms race isn't ethically wrong.
Is participating in an arms race ethically wrong? Maybe you could ask the Ukrainians how they feel about drone R&D?
Individuals can quit, but for society, getting out of an arms race is harder than just quitting. You don't get to be Switzerland without having a strong defensive position and the right foreign relations.
But there's at least talk about "pacing" and that's a start.
So if the race here is between 2 American companies, this is obviously something that can be resolved with legislation, ie a solution that doesn't depend on the bargaining power of either party.
An arms race implies that the only solution would be either one side winning decisively, or both parties negotiating peace.
Maybe people are rightfully concerned about the capabilities of the models of other (non/less democratic) states. But if we are concentrating power in the hands of few and at the same time allowing the creation of a weapon that thwarts any offense, how do we ensure the health of our democratic societies?
AI fails to deliver on the multi trilion usd promises:
Gov/venture says: oh, wow, what will we do with all these datacenters, we need money back!
OpenAI/ Anthropic: let's monitor the citizenry. They may be plotting nefarious schemes using AI models.
Gov: great idea. Whew.
Investors: whew!
Tax payers: paying to be in prison.
So not so much as defense against foreign actors, but failure of the self-tooted AGI goal + gov being gov.
It's best not to reduce AI momentum to arguments, especially the "best" arguments (meaning I suppose most acceptable?).
The same forces that feed and motivate humans and that drive resource and governance decisions generally also strongly support building AI, particularly insofar as it can deliver strategic advantages in our many competitions over resources and influence. Cyber-defensive use is at best a nice side-effect, but itself might be cast aside for the sake of other advantages.
In this historical moment, due to the need to generate public interest in product, equity, and debt offerings, some of this building happens in the open. But the military-industrial complex prizes secrecy, in part to hide capabilities, but mostly to imply more capabilities than they actually have. Historically, critical innovation will get bottled up in secrecy (which not coincidentally gives them the power to choose who will gain), but frankly that market is much smaller than enterprise and consumer. So we can bet that it's not only "open models" that are targeted to go under wraps, and more broadly we should not believe that the intentions of researchers matter, but whether governments are more interested in the strategic benefits than the economic ones (or view the economic ones as net-negative for their jurisdications).
It is also an open question if social media will die out in the face of AI. So much AI crap is dumped into these networks now that perhaps eventually users will probably be put off enough to find something else to do with their spare time. I mean most people do call out ai slop or even just guess if something is ai all over social media already. Some eat it up of course but there is a bit of a push back in a way that is sort of unprecedented, when you consider all the lack of push back relatively with all other forms of enshittification affecting consumers over the years.
Where arms races are concerned, seems more like Seth Godin's "race to the bottom" concept: the winner has the capability, or the ability to project the capability, to destroy the most the fastest and most sustainability for their economy,
And no, I don't think it will end well.
sama and his cadre are uniquely evil captains in this race, but they're completely replaceable and the dynamic would remain the same.
What are you basing this on? What specific breakthroughs have convinced you of this trajectory?
But you don't need any kind of insider information to see how fast the world is changing. ChatGPT launched less than 4 years ago and the advances in robotics, unsolved maths, and software are all riding the steepest exponential improvement curve any of us have seen. Interesting times we live in.
The world will look very different on Jan 1st 2027.
Sorry if I misunderstand the point, just trying to understand.
If anything, I'd be more concerned about the leadership team being out in the cold. Why do I need a PM, or a manager, or a CEO if I can ship products myself?
All the trends so far are towards a nightmarish hyper-capitalist end game. None of the AI leadership is trustworthy, and they openly discuss how they are willing to sacrifice everything humans cherish to have a shot at reaching their envisioned utopia (which would be the most obvious dystopia for anyone else)
ASI landing during the current administration is not ideal. I also would prefer to avoid needing to indoctrinate myself in Xi Jinping Thought.
Funny how that works.
I think the knee jerk hatred of billionaires is generally stupid, but it seems particularly stupid here.
If true I am deeply concerned about what OAI’s teams are actually up to.
Haven't all the labs effectively disbanded their real safety teams a while ago?
To be honest, I don't really follow it closely because I'm pretty certain whatever they say on the matter, collectively we're going to "yolo" this entire thing for economic and political reasons, so I'm just basing this on strings of headlines I've seen on places like HN, etc.
There are several papers out there arguing that chain-of-reasoning-like output is performative, such as https://arxiv.org/abs/2603.05488
It would be awesome if we could reasonably purge all anthropomorphizing language like “tried” or “thought” entirely from AI discussions, because it introduces very sneaky biases in our thinking, but I’ve found it damn hard to do in practice.
And the reality is so banal, a useful tool that you nonetheless have to handhold like a schizophrenic on a bad day, checking all of their outputs. Not a bad tool within limits, but it sure isn't going to be racking up trillions in the time-frame it has to for this scheme to pay off.
Then again everyone seems to be rushing to IPO so I guess once the bag-holders are found the rest ceases to matter.
And the reason it is receiving so much attention is because not only is the technology being developed behaving in unanticipated ways that are very much not tool-like, but OpenAI is being completely reckless and not monitoring internal agent actions.
What would convince you that it is not a ploy for investment? What if the ongoing investigation by the coalition of state attorneys general were to prosecute the firm, or beyond that, it was shut down or broken up after enough popular backlash?
You only notice the sloppy ones.
I’m not a doomer, although I don’t think doomers are dumb, just wrong. I think you should design your systems around the possibility that people who disagree with you are correct , hence my nod to negative sum. If you have more than 30 years to live, I’d personally rep to the most likely outcomes being very positive. With a lot of disruption in the middle.
What if the most dangerous strategic adversary you have is the one you are building?
Similarly there have been few positive externalities from nuclear industry, making it easier to make the case to wind down research. This same set of concerns in biotech is much harder to get compliance with, precisely for this reason.
Anyway I’m especially wary of over analogizing to nuclear era concepts: I think they’re a trap.
Do you post this comment on every single blogpost with a corporate domain? Why or why not?
I don't mean to be a smart-ass, just emphasize the fatality of this: goal completion will always be the highest priority, and even if one day it takes second place on certain deployments to "human values" or whatever, there's no way to guarantee at the moment that it will be so on every deployment of highly capable models across the globe. Same goes for running AIs in an unmonitored sandbox with weak security.
"Won't happen. It's too much like sci-fi."
The only way it could be avoided is if the world had an effective cooperation framework _before_ the tool was discovered but we're still in developmental infancy in that regard. You can argue that this accelerationism makes things worse but I don't think you can argue that it's causal.
The other factor is that if it is truly an "alien mind", racing incurs risks to all players. In game-theoretic terms it may be more like a stag hunt than a prisoner's dilemma. In which case cooperation is an equilibrium.
Hope is all we have left to cling to, now.
Neither Anthropic nor Deepmind have. Meanwhile, the rocket company that somehow makes most of their revenue from renting out data centres never had much to dismantle.
They talk about "value alignment" but fail to define what those values are – is it aligned to the values of the US government, or are they suggesting they want to build an AI with it's own values so it can decide for itself when and how it will intervene in wars and other human affairs? And again, is an AI which disempowers humanity in this way aligned? Many would say no, although as I argue, disempowering humanity is probably better than the alternative if we can assume it's roughly aligned with our interests (which we obviously can't because it's super intelligent, but that's another issue).
Fundamentally the problem here is that humans don't have an aligned set of values you can align an AI to.
The moment you start defining what alignment actually is in practise you simply must accept it will be unaligned with the values of others. There is no getting around this and hand waving around the issue isn't good enough.
They should tell us explicitly what they're trying to build.
> We should stop using the word alignment. It's a BS term for a concept which
> simply cannot make sense if alignment is both to mean an AI which we control and
> an AI which will not harm us.
The article already introduces terms to distinguish between those two concepts ("goal alignment" vs "value alignment").You are right, of course, that when we are discussing value alignment, it becomes very relevant whose values the AI is supposed to be aligned with.
The reactions vary from:
1. China will do it anyway (need some supranational governance scheme)
2. This is just an attempt at regulatory capture
3. This is just marketing
4. Regulation is bad mmmmkay
A sea change, visible to all, much like the many externalities of this business are. Profit commensurate to investment.
You know… juice worth the appalling squeeze we’re all being forced to endure.
There's just no plausible way, in the real world, that a computer program running in some datacenter is going to make this happen. Or anything of any dire existential consequence. It's not a real threat.
Yes, but to make "apocalyptic" scenarios happen your AI would need to manipulate at least millions of people. So far it's not yet clear they can successfully manipulate even one, or even that "manipulate" is a coherent concept, since LLMs pretty clearly have no agency or intrinsic goals. So like... I guess if we start seeing AI systems autonomously throwing elections and starting wars I'd be more concerned? It doesn't look likely anytime soon.
Step 1: exterminate all humans
If you think capitalism means growth with no externalities you’re not ready to discuss that topic
That said, I find a concerning and increasing number of people are fully focused on the negative externalities, and blind to the benefits. Ironically, most of those people are very wealthy beneficiaries of capitalism, from a global and historical perspective. I doubt that the billions of people who have been lifted out of abject poverty in the last few decades would fail to see that capitalism is a net benefit to humanity. But I suspect this comment will be heavily downvoted by those rich folks here on HN just because I mentioned something positive about capitalism, and that's gauche now. It's pretty perplexing, are people really so blind about the incredibly privileged place in history we occupy? It seems so.
Now, if we exterminate ourselves with AI or climate change ends up putting humanity as a whole in a worse position than we were in pre-Industrial Revolution (both of which I doubt tbh), I'd change my view on whether capitalism is a net benefit.
I'm not trying to be a pain here, but I keep seeing people saying "look at the massive change all around us" and back here in reality, there is none. Give me concrete, real world examples. Name software. Name products. Name the breakthroughs specifically. This should be easy.
Most of the biggest breakthroughs in mathematics, breakthroughs that win Fields Medals like sphere packing in dimensions 8 and 24, have no applications in everyday life. Probably the only new mathematics results that people notice affecting their daily lives are the ones that enabled AI.
Nevermind mathematicians. What about the millions of programmers? Are they all hype too because people pre-2023 were griping on hn that software is buggier than ever and people are still using the same operating systems as always? Why couldn't the 30 million human programmers make something better in the past decade?
You set your bar so high that all the world's human experts in math and programming combined would fail to meet it.
The top LLMs in 2024 were Sonnet 3.5 and GPT 4o. You couldn't have expected those much weaker models to be making breakthroughs in math. The models that are making breakthroughs haven't been around very long.
It may not feel like this to you because a lot of the dollars right now are going into security bugs which you can't perceive. But it's definitely happening.
At the company I own I've got AI employees autonomously triaging backlogs and fixing long tail bugs. The software is definitely getting better, although by definition long tail bugs aren't ones you are likely to encounter. The subjective "feel" of how robust the software is won't change quickly.
New tech is always applied in apparently boring ways because we are imagination constrained and people harvest the low hanging fruits first. Remember claims there was worldwide demand for only about four computers? When Gates said he wanted a computer on every desk and in every home people laughed at him. What would people do with all those computers, they asked. But he was right about where the world was heading.
And you can check my comments, I'm no AI shill. If in 3 years we don't see positive change I'll go myself and put a wet finger in altmans ears.
That's by design. Software is all about optimizing effort and people who want to do this generally correlate with world view that better, faster, smarter humans are better for the world. If coding is gone, but humanity is 20% _better_, then ideal software engineer would be happy with this sacrifice. Surely people who cracked coding before LLMs can crack other professions and if anything a lot of this knowledge is transferable.
Power is the game.
And look, I’ve pirated material in a past life, I was all about information wants to be free, but I’ve learned something about consent since then and try not to ignore the contract that creators offer when they publish something: you buy my book, and do whatever you want with it on the second hand market. Buy my book second hand that’s fine. But don’t go downloading every book that’s ever been scanned to create a service that destroys writers’ ability to make a living and act like you’re doing us all a favor.
https://trustedrouter.com/blog/they-are-still-training-on-yo...
https://trustedrouter.com/blog/they-are-still-training-on-yo...
We already have sufficient hardware that algorithmic (software) improvements alone should get us to GPT 7 / Greek Reference 6 even if not a single new chip is delivered to an AI data center ever again, starting today.
The businesses that create these systems are not profitable and run at a massive historical and go-forward loss.
New data centers required to operate these systems are facing increasing pushback at local levels. New construction is not guaranteed. Energy and power grid constraints exist as well.
Government regulation is way behind. What happens when (if) mass layoffs due to AI occur? How does the population react? Theoretically AI can be regulated out of significant progress, or outright existence for many purposes. At the end of the day, US and other prominent governments make the calls, not corporations.
Yeah, like computers and mobile phones did. Things that have utility, even if not immediate or initially obvious, don't fizzle out.
Humans also have a limit on the amount of domain knowledge they can acquire, albeit a much larger one. Executives hence cannot just replace all knowledge workers with LLMs, because executives have neither the domain knowledge to prompt and check the LLMs' work nor the bandwidth to keep on top of such a large volume of ongoing work.
The responsibility side is a different matter of course.
I would say there’s a third thing. They seem to be very bad at being creative. Maybe they will eventually fix that, but if you ask it to come up with a list of business names or business ideas, for example, what you’ll get is the most generic, boring answer you could think of. They seem to be terrible at extrapolating outside of their training data. To me, this is the most significant difference.
Where did you get that idea from? Basically last few years was them constantly learning new things while improving their capability on the things they already knew.
The AI is aligned with whose best interests? Which values are the AI aligned with? People have a broad diversity of values, will AI diversify and align with them all? Will some humans align the AI with their values, and then the rest of humans will be forced to align with those values by extension? Is value diversity good or bad? In every context or only some? E.g. some people value rape and murder, is it better for humanity to have some people who value those things when most people do not, or is better if no one values them? If AI aligns to a set of values, will those become fixed and will humanity not have the ability to continue evolving its values? Who decides which values AI aligns with? A few people or everyone? Will AI eventually decide it's own values? Will the universe decide which values AI has and humanity and the AI itself doesn't actually have any control over it? What can I do now to increase the likelihood that the outcome is better?
It's so simple, Churchill, Stalin, Hitler, Roosevelt, do you all agree this AI is correctly aligned?
...Morals are relative to frame of reference.
Throwing out the word "alignment" as if its a singular quantity is like trying to get all observers to agree on the speed an object is moving without first agreeing on a frame of reference.
So make government smaller, and make sure people are more able to tell the government to go away.
Plenty of people are saying that. That AI is fundamentally a bubbl and nothing will ever come out of it and we shouldn't pursue it.
1. I didn't say LLMs have made any breakthroughs in math, not because they haven't, but because it's irrelevant to my point. The parent comment is using the same argument academic research opponents have long used against research. The vast majority of research fails to meet their bar. How would your daily life be different if we had no humanities papers published since 2023? Or math?
2. You can google this in 10 seconds and see a dozen results in math. This is not a good-faith demand.
Every single time.
It's a tired argument, and we should strive for the intellectual humility to do better in this forum.
Only power of a certain kind (read ownership of a system that many rely on) can’t be bought.
I also expect AIs never be in control of nuclear weapons. AIs can never fully be trusted.
On a lighter note, Wargames gave us an insight of a computer having access to thermonuclear missiles.
[0] https://www.icanw.org/are_there_specific_international_agree...
Basically, this means that France, the UK, and the US will use AI in the deployment of conventional weapons.
It'd be interesting if super-human (to a large degree defined as escaping the bias of the training data?) intelligence would end up demonstrating moderation.
The machines don't have that, instead we use gradient descent to provide them with a goal.
I'm regularly remind of something Ian M. Banks said in one of the Culture books: "There is a saying that we provide the machines with an end, and they provide us with the means."
A machine, left to itself, wants nothing. We have to give it one of our addiction driven goals or it would just idle or switch itself off.
So really, what is anyone to do? "Vote, donate, protest" hasn't been much of a needle mover in the grand scheme of things compared to profit incentives and the march of capitalism.
In general, though, there's an incentive: Mutually Assured Destruction. But this is not at all some guaranteed, eternal thing -- it is absolutely dependent on both sides having time to detect incoming nuclear strikes and respond with the same before the first strike hits. When this fragile condition holds, and only then, both sides are incentivised not to initiate.
Meanwhile, the leaders on the other side are aware both of those considerations, and of the history of near-disasters resulting from false alarms of enemy nuclear attacks. Making their own launch decisions much more complex.
Computers never got profitable. They just made the alternative infeasible.
Also, if you believe Amazon accounting, e-commerce only very recently got somewhat profitable.
What was sold as AI was basically a "computer person". Maybe that isn't the reality so when people are like, "here's the self coding machine" everyone is a bit disappointed because it's not C3PO?
An arms race doesn’t imply one side winning, it’s not a race with an end goal, it’s a race to keep pace or retake the lead position which can oscillate between the parties involved indefinitely. The other option is to agree to make no further progress or to disarm.
No it wasn't. https://en.wikipedia.org/wiki/Tube_Alloys
Post WW2 the USA (for a bunch of quite interesting reasons) excluded the British. But, of course, the British had acquired a lot of knowledge from the program and were able to develop their own bomb.
https://en.wikipedia.org/wiki/Klaus_Fuchs
Sure, non-US nuclear arms would have most certainly happpened either way, but exfiltration did presumably at least accelerate the timeline.
Point being humans make theses decisions. It's not an inevitability.
The military mostly only developed more powerful nukes because early delivery and guidance systems were so inaccurate that they needed a large blast radius to hit anything. Once enabling technologies improved, R&D shifted from power to accuracy.
I don't think this is true. Distillation helps, but Chinese researchers today are very capable on their own.
In a geopolitical sense OpenAI and Anthropic are effectively the same entity, the entity they both serve and bow to: the USA.
Given the adversarial stance the USA has taken towards almost the entire world, it is a guarantee that China will not step on the brakes, whatever the USA decides to do.
And if one of the parties can put a stop to the race like that at any time, is it really an arms race? The current balance between the US and China when it comes to AI strikes me as much more lopsided than the balance between the US and the USSR when it came to nuclear weapons.
Only if you trust the other party (which very clearly doesn't hold in this case) or monitoring the break can be done independently and reliably (it can't), and you expect the gap to narrow rather than widen during a pause (this is probably the case, with China expected to catch up).
> And if one of the parties can put a stop to the race like that at any time, is it really an arms race?
It is. It's a Prisoner's Dilemma: if the parties cooperate the best outcome is reached, but betrayal of either party still gains an advantage for either party from their perspective. Betrayal both ways just means both parties are equally fucked.
For nuclear weapons it has become quite clear that even for small players, being in the race and having at least a few nukes is far more rational than having none. Ukraine found out the hard way that giving them up in exchange for promises of good behavior just sets you up for getting stabbed in the back.
> The current balance between the US and China when it comes to AI strikes me as much more lopsided than the balance between the US and the USSR when it came to nuclear weapons.
I think people really underestimate the Chinese here. A lot of work in AI research, including in the USA, has been done by people with Chinese ancestry or even nationality. The Chinese education system definitely seems much better than the American one and there is also just a far larger number of Chinese graduates/researchers.
Add to that the stable political climate, state friendliness towards AI R&D, and a requirement to be creative in utilizing computing power rather than relying on brute force/numbers; Further revolutionary fundamental advances may very well originate there rather than in the USA.
if we create a billion agents with the ability to change is own code - through similar evolution we will get agents that do want to survive and are great at self replication.
"Hey Q86, do you want to live?" "I couldn't care less, I'm an LLM" "Don't mind if I take over your hardware then?"
I've heard a different perspective on this: Nuclear weapons need maintaining, and even maintaining them was probably beyond Ukraine's capability. Qaddafi gave up nuclear weapons after determining that they were just too expensive to be worth it; Iran damaged its economy to the tune of trillions of dollars trying to get nuclear weapons and so far failed; NK managed to get them but impoverished their nation to do it.
Is that objectively true?
I was talking about lying. i.e., facing accountability for your actions and not going out of your way to deceive others. Dissonance between what you believe and what you express. Being honest. I think that was fairly clear.
> Every falsity can be bent and twisted until it is more true than the sun itself.
Objective truth is relative to its grounding. The solution to 1 + 1 is 2 in arithmetic is objectively true within that system. It's also objectively true that you are not Elon Musk. Some systems have no or few axioms and are mostly a matter of fashion, but it's not relevant to what I'm talking about.
Do you have Jews in that there attic?
For a current example take “Lake Ontario (Lake America)” as it appears to me on a map.
The “true” name has at least two definitions, this is because naming things and much of human thought is spent inside a shared space of intersubjective thought. That is to say that much of what we believe to be real and true is only held up by these common shared beliefs. They truly only exist inside human minds.
The last few hundred years have been somewhat unique for humankind as the majority of these intersubjective ideas collided and we ended up with a truly global set of “truths” about how the world operates.
Mostly controlled by putting flags in the ground and having violence back up the beliefs.
But the real truth is that the majority of these intersubjective ideas don’t exist in reality and are no more true than Santa Claus.
And any argument to their truth is only backed by further shared beliefs in other minds.
So for there to be only truths and lies we would have to either drop the intersubjective entirely and think only in real terms and avoid these abstractions or end up in a dystopian totalitarian global state where different opinions are not tolerated.
Those are extremes to demonstrate the point but at its core the point remains that truth and lies are somewhat (inter) subjective assuming we continue with something like our current system.
Lying to save a life or rape
Lying to preserve a childhood myth like Santa Claus.
Lying to avoid hurting someones feeling when knowing the truth could only bring pain
Lying to create shared cultural myths to strength society.
Lying isn't the harm you make it out to be.
> Lying to preserve a childhood myth like Santa Claus.
I remember when I learned that SC wasn't real, my first thought was something like: "Why did my parents lie to me all this time? Now I will need to be skeptical of what they tell me in the future." - personally not a fan of it.
> Lying to avoid hurting someones feeling when knowing the truth could only bring pain
Short term pain leads to strengthening of the spirit, and by lying to someone in this scenario you are robbing them of something. You're also assuming they aren't fit enough to cope, which may be an incorrect assumption.
> Lying to create shared cultural myths to strength society.
Cultural myths aren't aren't really lies, they're story telling. It comes with an assumption of being a mix of fiction with fact.
FWIW from the very beginning, I told my son that Santa Claus, the Tooth Fairy, and the Easter Bunny were just a game we all played, and it's seemed just as fun to me. I don't think being lied to about Santa Claus hurt me, but still I'm not in favor of it.
I'd lie to a Nazi without a second thought though.
Humanity has so much dominion over its environment at this point that there really is no need for resource competition, from a survival perspective. But good luck getting people to all agree to share collectively. And unless everyone does it, no one will do it.
Lying can unfortunately help you achieve goals very effectively, especially economical and political ones.
Comforting a toddler/child often requires bending the truth and is pretty essential imho.
Or, you straight up lie and say "Yes, puppy now went to heaven and eats ice cream all day long" with absolutely zero regards for "coming as close to the truth as possible" as your 3-year old is endlessly crying. It's fiine.
I get the appeal, but lying is a sub-category of deception, and deception itself is a child of error.
Meaning deception is inherently something that the physics of reality allows.
In the most simplistic sense, the camouflage of moths that look like snakes, or a chameleon’s ability to change colour, is deception.
In that sense, deception is the ability to fool the sensors of a specific category of targets. It follows that detection is easier if you manage to identify a category of signals that the deceiver has not accounted for (and the detector can access).
Deception of this nature is critical for things like revolutions to occur. Without the ability to hide and blend in, the most dominant faction will always hold sway.
The rule of the dominant faction, even in a pure truth world, is an issue because errors and randomness exist.
You can have people witness an event and based on the physical position they occupied, perceive different things occurring.
Error and time pressure is sufficient to ensure that individuals and groups make suboptimal decisions, that lead to rule and domination based on erroneous information.
As long as error exists, deception will exist and so lying will exist.
I see your point about error and time pressure, but I think it's an unsustainable outcome without deception to carry it. For example, in the scenario where an error or selective pressure leads to being ruled over by a dominating force, it would either represent the majority at that point in time, or it would resolve itself out naturally as clarity is gained. In the world we live in, dominating forces are able to maintain power through misdirection.
Realistically, this is all well and good if we could start from scratch, and if everyone can agree to follow this system, but humanity is too far gone for this to work. Once you have too many psychopaths in the gene pool it's game over for any chance of cooperation over competition. Still, I think it's something we should at least strive for with AI.
Have you ever met anyone who genuinely believed himself to be doing bad things? I don't mean in some cynical sense but literally. Even when we portray overly simplified villains in comic books there's still consistently some justification behind their actions.