ArXiv's Updated Rate Limit Policy(blog.arxiv.org) |
ArXiv's Updated Rate Limit Policy(blog.arxiv.org) |
The hosting cost side I can understand, though they received quite some funding recently, but I assume it's not going towards actually running the site, but to who knows what broader impacts and so on.
Moderation for Arxiv is a silly idea. It already shouldn't be taken as a quality signal that something is able to be up on Arxiv. Obviously they should remove illegal stuff, but beyond that, requiring moderation is a misunderstanding of their role and reason for popularity.
If hosting is too expensive, I guess an alternative aggregator could also arise. With just metadata and a hash of the pdf that can be hosted anywhere, and as the sumbitter, if you move the file, you can change the URL.
The main reason for Arxiv's existence is the timestamping and the easy referencing. (Though I admit that the stable hosting is also a pretty important part, but they mention moderation effort as the reason, not the hosting costs.)
Academia is losing sight of the forest for the trees, can't see more than an arm's length ahead of their noses.
I hope something like this could also be adopted for some of our larger conferences - the absolute limits on co-authorship are what seem to cause the most grumbles.
> In September of 2016, arXiv received 9,869 submissions. In September of 2024, arXiv received 20,569 submissions. This September, arXiv received 40,363 submissions, which in turn generated almost 9,000 support tickets for arXiv staff and moderators.
> arXiv now limits submitters to up to two submissions per calendar month, with a limit of three total active submissions at any given time.
I suspect a logical conclusion Arxiv and elsewhere may be an identity management system with an aggressive filter, and shared blacklists. I suspect that classifying people as spam/slop-submitters, then banning them (or whatever identity they used; name, email, name + organization etc), applying incremental rate limits over the general one, may be required.
The situation right now is pretty crazy, for example this established TCS professor [1] has at least 4 arXiv submissions co-authored by him in September [2]. This includes the recent breakthrough on the matroid secretary problem, which has a pretty interesting AI story of its own, if you haven't seen it yet [3].
Of course prolific professors can work around this by working with younger researchers who upload the work, but this just makes the rate limit a solo author bottleneck, which feels a bit weird.
[1]: https://en.wikipedia.org/wiki/Mohammad_Hajiaghayi
[2]: https://arxiv.org/search/cs?query=Hajiaghayi&searchtype=auth...
[3]: Concurrent Discovery Disclosure: The proof of the main result in this manuscript was obtained in a conversation with ChatGPT-6 Astra on Tuesday, September 15, 2026 at 1:02 AM PDT. We then prepared this manuscript for public release, with the intent of uploading it on the morning of Thursday, September 17, 2026. In the early morning hours of September 17, while finalizing the submission, we discovered the manuscript of Abdi, Banihashem, Hajiaghayi, and Mittal, uploaded on September 16, 2026, which contains the same result via an essentially identical approach. We are sharing our manuscript nonetheless in case our exposition is of independent utility to the community, and we hope this experience stimulates broader discussion about concurrent discovery in the AI era. (From https://arxiv.org/abs/2609.20797 .)
And in mathematics, at least, priority on results is determined by the first time it appears out in the wild; for most, this is the arXiv.
Nobody is going to accept waiting a month on something that they fear may be scooped, costing them years of work and thought.
This affectively makes the arXiv obsolete for much of the purpose it has heretofore been put to.
But yes, it'd be nice if arxiv also rate-limited co-author submissions with some higher number to discourage the most flagrant PI co-authorship abuses.