We're not racing towards anything. We've been going in circles for years.
We reached the NASCAR-racing equivalent of scientific research.
(Not saying that's how it should be, but that's how it has been.)
I contacted a professor from a university in the UK and he responded since he was working on similar work, then asked me if I wanted to meet with him. We talked for about an hour since we had overlapping results and different methods, specifically different assumptions.
I say all that to say, as a physics student getting my undergrad, simply doing independent research and speaking to experts about it enabled me to network with someone I otherwise likely wouldn’t know. For young people getting into any business, research is a great way to meet new people.
This is what happens to every field as it turns from a science into an industry. Chemists published freely until dyes started being worth money, and then the interesting work moved into company labs and stopped coming out.
I personally wanted an AI that was able to reach out to me about my life before I had to reach out to it. An example, a friend just emailed me asking to meet for at 1pm but I have class at 1:30, so a proactive AI would see that conflict and send me a notification about it, asking if the proposed email it drafted works, then I press send.
My personal setup tracks my mouse movement, keyboard, what’s on my screen, and keeps track of what I’m working on through files on my PC. It can update the backend and then restart it on it’s own, meaning I can develop the thing itself while being away from my PC.
The capabilities are more than what I’ve listed, but I want to avoid being too preachy about something I made. Here’s the repo if you wanted to take a look, it’s open-source and connects to the iPhone app:
The first tried to publish novel results for 3 years in tier 1 journals before finally doing a preprint and telling the tier one publishers to jump in a fire.
The second, and ongoing, isn't publishing anything because of my experience with the first.
That and avoiding openAI and Anthropic copying our results and leaving us with nothing to show for six months of work. The papers only come with the pitch deck.
In the paper, OpenAI is at the top of the chart for cumulative citations. MEGVII, Hugging Face, Waymo, Momenta, Preferred Netowkrs, Anthropic, Owkin, and Databricks, and Aibee follow (in that order). Yes, that is citations, not publications, but they explain that they're trying to use that as a proxy for significance, albeit an imperfect one.
Companies like Google aren't included because they aren't unicorn startups.
We already arrived at that destination a decade ago.
Probably even further back tbh.
There's just a lot more people playing the game now, without the social indoctrination that made it more tolerable in some circles.
It's not that bad, though. September is annoying but you kind of miss the eternal renewal once you're out.
Traditional research publications are no angels. They gatekeep research, and also allow financial incentives to drive them to publish junk with their stamp on it.
But a flood of papers doesn’t actually mean more knowledge. In bypassing this route entirely, AI has swiftly lost the ability to engage with itself as a field. And the cost of that is only beginning to be felt.
Anyway, there’s a repo link https://github.com/q5loisel/unicorn-AI-startup-publications so someone more determined at getting a better answer can poke at it.
50% of startups contributing to public research is actually a crazy good outcome. That’s far more than I had expected.
https://research.google/pubs/distilling-the-knowledge-in-a-n...
# LLM generated summary of the implied irony
Google may consider the standalone frontier-model arms race economically irrational, while still considering frontier-model capability strategically indispensable. Its longer game is probably not to avoid building the biggest models, but to build only enough of them to serve as capability factories—then turn that intelligence into a much larger population of cheap, purpose-built models.
(Human again) If Google knew what they were on to, why wouldn’t they make it their secret weapon from the start? I suspect it’s because they predicted there would be an arms race, and knew how to profit from it. They had a distillation paper published before “Attention is all you need”. In hindsight, is it ironic at all? Or is it obvious?
It was kind of like radiation science before WWII, it would be freely published because it wasn't potentially world changing yet. After it became a government interest, even people doing things unrelated to weapons would be much more apt to hold their work close.
There should be a new ESG (Environmental, Social, and Governance) policy being pushed recognizing the important role this plays. Although ESG and all norms have been set aside in this grim new world it seems.
The process and standards of peer-reviewed publication are tedious. If the objective is efficient communication of knowledge, this is not the process that maximizes that outcome. This process is about credentials and imprimatur, not communication.
Academic publication is largely chasing prestige. The value of prestige accrues more to the individual than the company employing the individual. Incentives for companies reflect this.
Many blogified publications are unqualified slop. This sounds like an indictment but a lot of academic peer-reviewed research is also low-quality slop. In many domains, slop is rewarded, devaluing the contributions of people publishing more serious research, which affects incentives. The signal-to-noise ratio has been poor for decades no matter where you get your knowledge.
Any good research you publish will be exploited against you in myriad ways. It is effectively impossible to assert any IP rights. Research turns into a pure cost center if you treat it this way. Companies recognize this and have adapted to reflect it. In some computer science domains the state-of-the-art has been buried in trade secrets for decades such that the academic literature is embarrassingly obsolete.
Publication costs time and money. Given all of the above, why would you invest productive capacity in public communication of research results? The researchers want to work on interesting problems, the companies want to maximize the leverage from their research. Publishing delivers neither.
A lot of fundamental research is done for the purpose of solving a problem, not publishing a paper. Their motivation and personal reward is solving the problem, not publishing it. Once they've solved it, it is no longer interesting and they are on to the next interesting problem.
Some research is coupled with national security considerations. If you do frontier research this is always sitting in the background.
All of this is a consequence of incentives. A large percentage of all basic research no longer happens in academia. I've greatly enjoyed participating in non-publishing research programs. I also understand why they don't publish. I've done a significant amount of interesting foundational research under my own auspices where it was not worth the effort to publish. I'm not chasing prestige, I had a problem I needed to solve.
No one should be surprised by any of this.
------------
If everyone holds back their publications, the whole sector moves slower. Yes, it's moving alarmingly fast according to many, but is it moving fast enough that the massive investments in data centres will actually pay off?
It's more game theory. Things may go well for one selfish company not publishing in a sea of altruists who publish, but perhaps not if everyone else is selfish too.
It may be moving so fast that they do not pay off, i.e. they get commoditized too fast. This all depends on Ai demand in the end. If you can keep the GPUs saturated (at a profitable rate), then it really shouldn't matter how much a token costs. If you bought too many GPUs, or paid to much for them in a hype cycle, you may need higher token prices to ROI than the market wants to pay given alternatives.
Also, there are simply too many AI papers that make peer-review publishing quite meaningless these days (e.g., AAAI this year got more than 50K submissions).
Think Renaissance Technologies or similar.
I don’t think Dario wants to be a monopolist. But history is full of brilliant people who focused on the social good as priority #1 and lost the game early.
Once you lose the game and go bankrupt you can’t do anything. Same as politics, if you’re not elected you can’t change anything. Of course there are many lines not worth crossing.
How do we fix this.
It's like we know things are destined to go down in a flaming pile, and we're just watching so we can change it AFTER that happens.
Soz. Kinda a "dear diary" response, but am frustrated.
Companies can’t be expected to publish their confidential and proprietary information about their feature development, and academics should consider projects that would have higher impact.
If you’re taking up a seat in a PhD program tinkering with would-be feature ideas for an existing major tech company, you should really just get hired by that tech company, where the resources are abundant and the degree is not required.
Of course they can. Simply eliminate trade secret protections and NDAs, which have literally zero purpose for society and in fact undermine patents' incentive to publish. Give a one year grace period to apply for a patent. Patent protection should be designed to be only as long as needed to try to maximize the development/spread of technology, e.g. maybe 2-3 years for rapidly developing areas like ML.
We should all be expected to contribute whatever knowledge we find to the commons. It makes us all richer.
How we wish to consider public funding and required sharing of gains is an open debate.
...today? (emphasis is mine) And the author was so excited that decided to write up a paper and submit it to Science all in one day? Yeah, right. That's all you need to know who is moonshooting behind this paper.
Even at work, I have come to realize if I simply horde my accumulated custom AI built tools and productivity boosting tricks for myself, I can make myself more competitive as an employee.
I think this is how you get hired now, not by having a good resume, but by making claims of having special processes and personal tooling design that gets massive productivity ROI.
https://arxiv.org/search/?searchtype=author&query=Magarshak%...
So I know that smart people in those companies can definitely publish. In fact, a whole team should probably be publishing like no tomorrow!
More to the point, without determining how much work is “worthy” of a paper it is unclear how much this matters.
Most AI companies are either a product and marketing layer over a model or not meaningfully moving any dimension to be worthy of a paper.
Also, formal papers and blogs and “cards” are all being intertwined.
Cool for Google. Problematic if you are a startup.
Isn't the POINT of publishing research because you want others to copy it?
For startups, no, because the point of that is to make money.
The most common attitude is something closer to "I want to be well known and respected by smart and influential people". In short, prestige.
People using your work is a side effect of success, but simultaneously threatens your ownership of your "brand". If they significantly improve upon the thing you did, your work could be made irrelevant, your prestige ruined.
Moreso if your current employer is a stealth AI startup which hasn't yet produced much notable.
It also can encourage employees to work for fewer wages... 'If you do great work you can publish and then get a $$$$$$ job offer at OpenAI, or you can stay here and get equity in a fast growing startup. If however your work doesn't do well, you walk away from a bankrupt startup with worthless shares and minimum wage'
LLMs are pretty good at this, hilariously. This means a genuinely good paper by an outsider has significantly less chance of getting good reviews than AI slop.
You are not. A disproportionate amount of value in the computing industry was created by <strike>geniuses</strike> decently smart people who worked together and who decided to just tell people how to do things instead of trying to capture the value of being the first person to figure out how to do those things.
This observation pre-dates the current wave of AI hype by a half century or so.
> driven by greed
I can only speak for myself.
For me it's exactly the opposite. If I want to explain how something works, I can just... do that. If I want to share an artifact demonstrating how to solve a particular type of problem, I can just... do that. If I want to mentor/teach, I can just... do that.
Doing those things within the confines of Academia Approved Institutions is exhausting and distracting.
To wit, and the actual point of this post: the term "Publishing Research" in this article doesn't mean "post it on a .html page and share the source code". It means engaging in a very specific and peculiar and extremely political modality of communication.
And it really only makes sense to do that specific and peculiar and political thing you're at a stage in your professional/personal development where you need to play that particular prestige game. (Which there's nothing wrong with, but it is a deeply cargo culted version of the actual scientific process.)
* A disproportionate amount of value in the computing industry was created by g̶e̶n̶i̶u̶s̶e̶s̶ decently smart people who worked together to do things that seemed impossible and who decided to just tell people how to do things instead of trying to capture the value of being the first person to figure out how to do those things.*
There is some excellent publicly-funded research in there, but pivotal papers like Attention Is All You Need and the numerous pivotal OpenAI publications were privately funded.
OpenAI is at the top of the chart in the study.
I think you're bringing some assumptions into this conversation that aren't supported by the evidence.
A distinction should be made between the old nonprofit OpenAI and the current organization. They don't publish technical research anymore.
Example: https://arxiv.org/abs/2607.24653
Exceptions to some like Deepmind, Nvidia, and Thinking Machines
Fortunately
[1] https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
The AI usage has mostly been getting it to explain error messages
The main difference is reproducibility and peer review process.
On the first, I mostly don’t care.
On the second, that’s mostly unavoidable if they want to keep their IP.
They are not perfect, but institutions like peer review or universities do provide a framework and incentives for quality.
Once the business is making money it’s not really a startup any more, it’s a going concern.
OP is an undergrad student. Attracting attention is what will make them stand out. And sharing research at this stage will probably expose them to research others don’t want to publish but do want to share with other researchers.
That would clearly demonstrate it had real substance beyond hot air or puffery.
Better to silo yourself and move in to the penthouse.
To answer your question, sharing the research isn't the important part. These are researchers in the private sector, in a highly competitive field. Arguably, doing the research isn't the important part either. The important part is making money, and the research is a means to that end. Publishing the research doesn't further that goal, and, according to GP, it's so difficult that it may not be worthwhile trying.
edit: edited.
Like I said, I think you're bringing some assumptions to this conversation that aren't based on the how the industry came about.
The discussion has been around, but it's flared significantly recently. Not sure if that's what the person you're responding to is specifically inflamed about.
Regardless, old/rare books are certainly being acquired and destroyed.
Elon tweeted "I’ve asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning"
which is no sense a credible source for what he actually asked SpaceXAI team to do, or if they are doing it, or what they were doing before yesterday. Elon has an extremely long and thorough track record of lying.
The problem remains either way. Small operations around the globe are debinding rare/unique books. I'm involved with an organization considering (not strongly) that exactly.
The scanning process destroys the book.
Internet Archive scans their books one page at a time keeping the book intact and stored in a warehouse.
> "The print original was destroyed. One replaced the other."
(I believe the reasoning is that the author is not being financially disadvantaged by more copies of their book existing).
So the current law prefers them to destroy.
I usually hear that it's extremely hard to track down the copyright owners for many older things. That's the reason the film/video industry often gives for not digitising the backlog online, and I would expect their volume to be much easier to handle than the book industry.
I don't mean to downplay your work, but I think you should come up with a better example use case. Automating away interactions with friends is pretty much the last thing I want AI to do.
The run stalled mid-step around 80M parameters, Orb notified me that it stalled, asked if I wanted to resume at the last checkpoint and kill the stalled version. I simply press “yes” and continue doing whatever I was doing.
For non-technical users and non-antisocial people, remembering things you forgot so you don’t let people down. You told your sister you’d send her some pictures two hours ago but it can see you’re scrolling on reddit and the photos are on your desktop, so it assumes you forgot and reminds you.
The idea was that the biggest issue with the usefulness of an agent is that it has too little context about who I am, it needs more data. So I run all of my data through a smart router, then the local database, then the LLM reviews it and uses reasoning on what’s been collected.
I assumed that's what openclaw basically was, but is Orb different from that? And is it fundamentally a different model from the request/response, or is it just request/response in an autonomous loop?
On a fundamental level the backend was designed to do as little LLM calls as possible, for instance it’ll do scans of my screen every 15 seconds, log what’s on it and what’s going on, and store it in a local database, then Orb reviews the entire database every 6 hours for me. Then it’ll schedule wakeups for itself throughout the day, up to 4 so it doesn’t waste my tokens, and schedule notifications based on the last database dump it made.
I have my Claude Code, Codex, and Grok Build all useable by using the “claude -p; codex -p…etc” so you can also use multiple CLI’s in conjunction at the same time on different projects or the same project.
So your question about a loop is kind of right, but it really just collects your data all day and stores it locally on your PC then calls the LLM of your choice and it reviews all the data and makes those proactive moves we’ve discussed. You could theoretically get it to always be scanning by an LLM but that would be a drastic waste of money from what I’ve seen since most things don’t require a call.
OR
“Event x happened at y time”
There is not an obvious path to making this research public again. That's the reality. For open source to be competitive it needs to be approximately as efficient as the best systems we design.