Detecting LLM-Generated Texts with “Classical” Machine Learning(blog.lyc8503.net) |
Detecting LLM-Generated Texts with “Classical” Machine Learning(blog.lyc8503.net) |
Images, absolutely, there are tell-tale artifacts from today's generators that simply aren't emitted by "natural" paths to create them, and you can "detect AI" with high confidence (for now). Words, no, the signal is far too sparse and we are well into undetectable sophistication with today's models, let alone tomorrow's.
> we are well into undetectable sophistication with today's models
The article directly contradicts this, as do you, in your previous paragraph: "Sure you might be able to detect today's tells". The article is literally about a technique that detects today's tells.
Your comment is mostly expressing doubt that this technique will work reliably in the future, but it's framed as opposition to the article, which it's not: the article is about detecting today's AI-written text, at which it seems to be quite successful.
My neighbour is a teacher. She has a really good idea which of her students uses AI to do their homework but 80% accuracy is not good enough. She'd need to be able to prove it with certainty.
As long as LLMs are used to write texts humans wouldn't want to write if they could help it (that's why they're getting an LLM to do it, after all), they'll remain detectable. Even if the reasoning might end up equivalent to "This looks like spam; no human in their right mind would write this spam by hand if they could get an LLM to write it, therefore it's most likely written by an LLM."
I doubt this models reality well at all. If I write the first paragraph, and AI writes the second; a float seems to model that better. If you choose to collapse a float into a bool, I don't think you can make useful conclusions based on that bit?
> since even the shortest text contains more information than that.
I also don't think that's how information theory and bits of information works...
This is word salad, a complete non sequitur.
> to write texts humans wouldn't want to write if they could help it (that's why they're getting an LLM to do it, after all)
Er, that's obviously not true.
Come on, that's circular reasoning.
As an example, you could imagine a giant lookup table that deterministically mapped every text ever written to “human” or “AI”. You would very quickly run into situations where the labels conflict for the same piece of text.
The data is statistically inseparable which makes it impossible to classify from text alone.
Not true at all. Pangram is highly effective and has a very low false positive rate.
The post here is impressive for a small project, it looks like they independently thought of one of the core ideas Pangram uses of creating twins to compare.
You can see how it works here: https://arxiv.org/pdf/2402.14873
If you would not be okay with that, what level of consequence would be acceptable for the output from this tool?
It's really easy to have a false positive and false positives can be very harmful if the person using the detector isn't aware of that risk.
It's also very easy to change the pattern of LLM output. You can provide basic prompting that will significantly change the structure of the output. For example, having it utilize the Wikipedia article on signs of AI writing and avoid everything it describes. https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing
Not really. The false positives for the SOTA detector are very very low.
"It's also very easy to change the pattern of LLM output."
Not in a way that can reliably avoid detection. The problem is the patterns are baked into the distribution itself. It's smoothed over, so it becomes difficult to prompt your way out of that.
Schemes like GPS, CDMA and DSSS are based upon this concept. GPS in particular is quite impressive in its ability to recover information that is received below the thermal noise floor.
Take this sentence: Bob went to the store to buy milk.
Was that AI generated or not? There simply isn't a signal there. The problem isn't noise, the problem is, is there even a signal to begin with.
Sure, you might be able to recognize the quirks of a specific LLM just as you recognize the quirks of a particular person, but as the number of LLMs proliferate, then the signal turns into noise. (The signal isn't buried by noise, it becomes noise. The signal no longer has any discriminating power.)
Largely AI generated books are a vastly different situation than a one paragraph homework assignment. But multiple rounds of homework assignments would change the accuracy.
This does not sit well with personal experience and I wonder if it is just one of these questions of AI people being unaware of the level of skill that exists in domains they think have been automated.
It is of course possible that my tendency to spot LLM-written text has much to do with the way that it sounds like an averaged Californian college student to my British grammar-school-educated ears, as so many of the situations where I am encountering AI text are Brits using it without apparently realising they are giving themselves away.
But I know people who don't have particular technical skills in this sphere or a grammar-school background who also have an uncanny knack for pointing out LLM-written text.
> Words, no, the signal is far too sparse and we are well into undetectable sophistication with today's models, let alone tomorrow's.
I especially don't think this is true. Will they be able to do it in the future? Maybe. Is it possible to prompt a current cloud LLM to write in a way that is obvious? Yeah. (IMO Gemma 4 writes less detectably than most of them!)
But my instinct is that someone with any facility for language is going to be better than chance at spotting LLM-written text once it is three or four paragraphs long. So I think it should be possible in principle to train machine learning systems to detect those patterns.
I do struggle at times with thinking my own writing looks like AI. But I’m an average Californian who went to college half way between SF and LA…
I suspect the difficulty here lies more with your reading of the quoted sentence. British grammar school education, for all the years it devotes to the enterprise, does not always succeed in teaching reading comprehension.
You seem to be treating two rather different propositions as though they were one and the same. If text in general is not sufficiently information dense to support decoding some _arbitrary_ signal of provenance, that hardly establishes that no _specific_ passage can carry distinctive markers of provenance.
For example, you can recognize the unmistakable cadence of the California undergraduate. Impressive. Alas, even in your own example, when your British friends are "giving themselves away", you resort to an external signal, beyond the text, to determine provenance! That is, unless the text itself is claiming that its author is British (like the bots who claim they're John Horsetrader from Arkansas oblast).
When you have to decide whether a 2010s era SAT essay was from a SAT prep book author or an LLM prompted to write such an essay, you will struggle to distinguish one from the other. Not all texts have provenance signals. This is what it means for text to simply not be information dense enough to be able to decode some arbitrary signal of provenance from it.
Hard disagree. LLMs (especially base ones, that only received pre-training) can produce output that is undistinguishable from human writing (because that's what they were trained to do).
But commercial chat models are specifically tuned in a way that maximizes user engagement. It's that specific tuning that is very easy to spot when reading AI slop, and that's not surprising that it's easy to spot automatically either. And I don't think that's going to change anytime soon, unless their incentives change.
(We can say exactly the same thing about man-made stuff optimized for a specific purpose, like stock photography, clickbait titles or industrial food: they aren't stereotypical because their creator lacks the skill to make them otherwise, they are like that because that's what works best).
Did you actually try them? I did.They generated even more "slopey" text than instruction-tuned ones.
There are two problems with this.
The first is that it would still misclassify human-authored text written under the same incentive, and most people have various incentives to "maximize engagement".
And the second is that then people would just make other models that are tuned for defeating that sort of classifier, which would be used whenever the classifier is being used.
This is simply untrue, and completely divorced from reality.
Tarot card readings have literally zero predictive success. Last I checked, LLM-detection had a +90% success.
A better argument is people themselves are just too influenced by reading that they'll sound like LLMs in a couple of years.
You can absolutely still tell.
Most people's issue with AI-generated llmish however is not that it's AI-generated. It's its insufferable tone.
So if we get to a point where we have to read tea leaves (an image you seem to appreciate) to determine if it's llmish or not, we'll have won by then.
Really: it's that full-on asshole tone I (and many others) want to see disappear from blogs, comments, LinkedIn, etc.
It's an aside, but my biggest problem with trying to get up to date and learn about LLMs is how much of the documentation, blog writing, and tutorial material has obviously been written by no-one. It is just so much harder to read (and, like generative AI slop generally, curiously much harder to recall later).
Simply track all citizens' writing patterns throughout their life, from cradle to grave, then diff with any given text's signature--you'll know if it was human written or not.
Better--opt in--install a "personal text signature" on your devices, sign things that you wrote yourself with it.
But I suppose that's just like the image provenance chips on cameras.
Either way father fascism is more with us than ever, praise him!
If the internet is going to drown in LLM text it would be nice to have tools to detect that automatically just like we have adblockers today to avoid wasting time on ads.
(the article was a good read, thanks!)
Check it out: https://slopsieve.com/extension
Accuracy is also much higher than this approach here. 0.9944 AUC, 0.966 acc@.5, 0.971 F1@.5
This is a very startling admission! I checked the Chinese (original?) version of the post, and saw the author uses the word "糊弄" (in the place of "faked"); I'm not a native speaker but I think this may come across more as a self-effacing comment on the low quality and/or effort behind their thesis, whereas the English version implies fraud. May be wise to change this!
E.g. "I faked my way through the interview!" = "I did my best to respond to questions I did not feel fully prepared for, and managed to get through the interview"
From what I understand, your approach is clever, it's like an accent detector. Known models tend toward a specific median approach. Humans have a much richer degree of randomness. Riffing on Anna Karenina... All models are alike in that they present predictable patterns. Humans inevitably write in unique ways.
I gave a lot of thought to the idea that humans will devolve to the median led by volume of AI interactions, but in the end, I think we're still interacting with each other when not at work/on machines, and the fact that we even have a genetic heritage is always going to differentiate us.
He literally demonstrated a working system in this post. Do you mean you'll never get to 100% accuracy? Clearly, but you don't need that.
But how useful is it really to get a result of “This is 80% likely chance of being LLM generated”? Or 75%, or 95%? What if the text is a mix of human written text and LLM text? How would you even begin to test that?
I suppose a text that is half human half LLM would theoretically score in the 50% range, but do you see the problem? You can slap a confidence % score on a test run, but interpreting the results leads to a whole other can of worms.
Point is there are so many variables, and it’s not clear that the result from any of the systems is even valid or applicable to help you make a decision in a real life situation.
... if the assumption is LLMs are being optimized to evade such detection.
PS: I didn't read the actual article.
I think it is instruction-tuning that is having LLMs write differently from humans, and this is not being optimized away.
The link you sent is for generating text which attempts to defeat those classifiers.
Also you have to keep in mind that most AI companies are in fact trying to create and offer legitimate products and services to customers doing actually-useful work. They’re not trying to help fly by night hustlers scam people out of crypto or run spam campaigns, and in fact often voluntarily watermark to prevent misuse of their products.
You could argue that’s “just to avoid bad PR” and maybe you’re right, but that’s just another way of saying that it’s more profitable to prioritize other use cases than the deepfake/spam market. Spammers and fraudsters are shitty customers and a major brand risk.
Done with GLM-OCR, I had to watch text sloooowly crawl out of the llm and still have to live with hallucinations and the model not following the schema
If memory serves, one student objected strenously and ran the professor's own work (published 10 years earlier) into the same tool and it flagged that work as AI-generated.
EDIT: HN item from June 2023 https://news.ycombinator.com/item?id=36215823
If the false positive rate becomes <1% then it's better. The alternative is the world drowning under slop so I'd rather have imperfect detectors and have users aware they may fail in rare cases to avoid witch hunts. The general issue is that people only realize they're reading slop halfway through which is frustrating. If you know it from the start thanks to a detector and move on without commenting, no time waste, no frustration, less negativity towards LLM users.
Of course, you can skew the distribution with some effort and generate text that avoids even the best classifiers out there (like Pangram), but even tech-savvy people aren't usually doing it (see the amount of AI-written posts that end up in HN and get tons of comments complaining about AI mannerisms), so I guess they're successfully avoiding like 99% of the slop using such classifiers.
I don't think it's in the interest of the labs to allow you to generate text that's indistinguishable from human prose. Especially since nobody would pay $1,000/mo just to generate text - but would do so for tasks like coding.
Perplexity was widely considered SOTA in 2022. One part of it is because everyone was evaluating on open models or closed models that were still close (i.e. GPT-2 vs. GPT-3.5). Today, the gap is so much wider between the models you can use to compute perplexity and the frontier models people actually use.
Also so many AI text detection papers used a strawman RoBERTa baseline that was very undertrained for the task.
The synthetic mirrors method for data generation used here is the same as what we use at Pangram. Good blog post, thank you for sharing!
This issue with AI generated stuff is that that it’s sometimes asymmetric: either the author worked very little to produce a lot of slop and now the reader(s) all have to do the heavy effort of reading it OR the author puts a little extra work in once and resolves all future readers’ burden.
If it was possible to boil down an artifact into a prompt + some resources that would be an interesting tool, or at least some way to tell if some artifact is “worth my time to read”
Tomorrow, the LLMs will be training the humans thought patterns that will directly start skewing their natural writing.
Generation alpha is going to have a lot of trouble if we keep perpetuating the myth that you can really interpret text in an ongoing fashion.
https://www.washingtonpost.com/opinions/2025/08/20/chatgpt-c...
For Google Doc users, you can already inspect the edit history over time to verify that text is written by a human.
I wish there was a solution
See https://blog.neurips.cc/2026/06/02/ai-generated-papers-in-th...
> Unfortunately, given the volume of submissions that appear non-compliant, relying on author declarations is insufficient.
...
> Authors whose submissions show significant AI involvement must provide an audit trail that clearly demonstrates that they complied with the policy. We expect that in future years this kind of audit trail will become a default.
The method uses Pangram, but that's somewhat arbitrary. The authors must provide the draft that was used to prompt the AI and show how the changes were made, which should be possible with a chat history. Where it doesn't work is if the authors claim they did not use AI, but otherwise it's quite hard to spoof. Edit history on its own isn't good enough because you could trivially have a program emulate keystrokes to type at a human speed into a document with tracked changes; the proof here accepts that AI was probably used and aims to determine to what extent.
I do think this goes back to if the text is high quality enough that humans don't notice, is that OK? Probably yes. What we don't want is poor quality writing from any source.
If I can do it, an algorithm should be able to do it. Maybe in the future the models will get so good that it is literally impossible to differentiate human vs computer authorship, but that’s obviously not the case today.
- The author is conducting some kind of hustle.
- The author doesn't bother editing.
- The author lacks the taste and awareness enough to see it looks.
- The author thinks you, the reader, lack taste and awareness.
- The author is using it as a kind of smoke bomb to get rid of you.
In such cases, nothing is done about the LLM's distinctive "voice". It dominates the text and it's easy to detect. It stands as a signifier of the above, even if it's otherwise not intrinsically a problem to use AI.
In training all you have to do is take their model as the adversary and then it's useless.
If a signal is weak, it might not even appear in every sentence, but that doesn't mean it doesn't exist. For instance, I don't recall ever consciously using an em dash, but you'll probably need an entire paragraph to find one in LLM-generated text.
My own sense of whether text is generated is partially based on its sheer length - humans typically don't bother writing so much.
But that's the point of corporate speak, you tend not to say thing that may offend your clients and deprive the company of future revenue. Of course there are some companies that make their living being 'counter-culture' and saying what they want, but they are a small percentage of all revenue.
And that's just one particularly egregious case I remember. Posters that are technical writers or use English properly get called bots quite commonly when their post history shows a writing style going back over a decade.
But now that LLMs are causing a language drift in English users our filters of "that's an LLM" will become even more useless.
Blizzard employee once told me anti-botting in WoW was extremely challenging due to the number of real people that acted identically to bots.
Every assumption was invalidated: - unbelievable # of consecutive hours played - consistently repetitive patterns of movement and clicks - farming patterns that aren’t considered fun (“why would anyone do that”) - solo, no external engagement - goes on for months
The problem with botting is many humans ARE bots
Well aren't you a genuine delight?
> Alas, even in your own example, when your British friends are "giving themselves away", you resort to an external signal, beyond the text, to determine provenance!
A possibility I addressed in the actual text you are responding to, where I started the sentence with "It is of course possible" and continued to clarify that "...so many of the situations…" I encounter it are localised.
It's almost like I was expressing just such an awareness of the limits of my assertion, isn't it?
The model appears to be similar to MiniLM-L6 (384 dimension, 6 transformer layers) but uses RoBERTa/GPT-2 style embedding/tokenization (50265 vocab size), used for binary classification, and quantized to INT8.
My guess is that is was distilled from roberta-base, then fine-tuned on freely available pile like artem9k/ai-text-detection-pile or similar.
It's a nice model - I've just used it to create a browser extension which highlights text based on how likely it is to be LLM genereated.
Edit - a quick google search reveals ibm-granite/granite-embedding-30m-english with the same architecture: slap a binary classification head on, fine tune, job done.
> MiniLMv2 L6-H384 (30M params) progressively distilled L24→L12→L6 from RoBERTa-large-v8 teacher (neobert-v2 recipe, min_words=0). Byte-level BPE tokenizer. INT8 quantized.
So if you want to determine whether something was written by a human or by AI, to do better than chance it's enough for there to be a difference in the probabilites of a human writing it and AI writing it, respectively. Whether the resulting accuracy is good enough for a particular use case is another matter. 99% is pretty good odds for love and pretty bad odds for "am I going to survive today?" Hopefully there won't be a death penalty for posting AI-generated content.
On the other hand, how an AI writes will have a big correlation to whether the written text would likely be written by an AI.
The latter is more of a direct relationship.
Maybe A and B are not correlated, and Y and Z are? What pattern are people falling for here?
Their interest is irrelevant in the face of technical impossibility. And that’s before you get into other people who don’t care and will just build adversarial tools to bypass the attempted watermarks. It’s a losing useless battle. Google and OpenAI engage in it to try to catch competitors when there’s a lawsuit or to try to clean their datasets clean.
But it’s absolutely unusable for something like “did someone cheat”.
[1] https://hackerfactor.com/blog/index.php?/categories/1-Image-...
Maybe you meant "many humans have used AI when writing texts"? Your stated reason that they can't be separated because there are many texts of each kind is nonsensical, you clearly need to supply more reasons than "there are many".
I don't know. I mean, it feels like the systems that would detect them are likely qualitatively different to the machines that make them.
One of the things that feels obvious to me is that LLMs are always going to write in a new way, because words do not get all that close to perfectly conveying the inner thoughts of competent writers. Competent writing is always a battle to find the better word, or even to create it.
So sure, you could add another adversary that the generator has to satisfy, but "this sounds like a machine wrote it" is only an observation; it's not a prescription for not writing like a machine.
Maybe it's never going to be possible.
> I do struggle at times with thinking my own writing looks like AI. But I’m an average Californian who went to college half way between SF and LA…
:-)
You guys do just sound a certain way, in the same way Brits sound a certain way to you I expect. But I think the reality is that the final stage of training LLMs was largely done in a Californian voice and with rather Californian communication objectives.
(Though equally I think much of what I am detecting is more Madison Avenue than Palo Alto)
The comment I was responding to, about changing LLM output, referred to prompting, not temp/sampling tricks. I’m not aware of Pangram being beat by clever prompting. There’s some interesting work on creative writing using contrastive prompt techniques, but I haven’t seen it tried as evasion.
Even if you control temp and sampling, they’re not magic. If you raise the temperature too much writing can go to hell, so you may beat the detector but end up with junk. There are some ways to mitigate such a quality drop like raising temperature in conjunction with min-p, but still, I haven’t read any research that shows it getting good results at anything close to 10.
Now you want to get more clever and manipulate logprobs…well ok, you could come up with elaborate strategies designed to evade specific detection methods. But I don’t see that getting done as a weekend project while maintaining writing quality. And if it does happen there’s no guarantee the detector can’t train on its characteristics and start an arms race.
Edit: But I'm also super conflicted about this, because I really want to read what humans think, not what an AI thinks, regardless of the writing quality.
FWIW if I were a student I would definitely be using Track Changes or version control, etc etc, to make clear my work was human-written. Which sucks.
I’d want detectors to be as accurate as possible, false positives of 1 in 10000 seems like a good starting point. I believe their results have been independently tested.
And as a separate matter, any tool for evaluating students should be applied fairly, safely, and with adequate human review and due process.
You need good tools and good oversight.
Plagiarism and cheating sucks for everyone. Worth solving.
Agreed, that's a fair and reasonable stance.
The reason I asked is that I have a hard time understanding the point of these tools. When it comes to education, it can be a matter of learning objectives. But outside that, what's the point?
The prediction from the tool is pointless for deciding on copyright or contract issues, and other text should be judged on its correctness or applicability to the task.
If all the tool is good for is "maybe this student cheated, but only an in-depth investigation would maybe prove it", it isn't a very useful tool, because it's more straightforward to just mandate that evidence is submitted regardless of what the tool says. On top of that, even the lack of evidence of manual work isn't good proof of using LLMs.
But yeah, in general I think you’re right, the actual utility is pretty niche.
>> Not true at all. Pangram is highly effective and has a very low false positive rate.
> So, if the decision from Pangram determined, on every assignment, if you would be expelled from university for plagiarism, would that be acceptable to you regardless of how you actually did the work?
What point are you arguing? Something having a high success rate does not necessarily translate to treating it as a 100% success rate.
You could work for a boss that's a Leo, and he/she believes only Leos deserve to get promoted, or that Leos and Scorpios should never be assigned together on a shared project. Your life and career trajectory under this boss could be totally different, depending on whether you were born a Leo or not.
Certainly "not all bosses" applies, but it's not really that farfetched or uncommon either. The point is that a correlation does exist, but it's a social one and not a physical or astronomical one (and it's also often a self-reinforcing/self-fulfilling one: Leos who read what astrology says Leos should do, may end up choosing to behave more like that).
So in the LLM example, it may not really matter much what physical markers of provenance or physical correlations there are, as social beliefs or perceptions about suspected provenance may be the strongest correlation anyways (in terms of impact and outcomes).
I'm responding to "If the data was separable in this way, you would equally be able to train an AI to mask those signs.": yes, if you wanted to you could, the big names clearly don't consider masking to be a priority.
> But it’s absolutely unusable for something like “did someone cheat”.
This is the one case where I'd most expect it to succeed:
I suspect most of the people who do want to cloak-to-cheat, don't have the skills to do so; I also suspect most of them are so unaware of what they don't know that they won't even ask an LLM to write cloaking software for them.
Other papers have tried to beat it with temperature and it didn’t work, although I haven’t seen anyone try insane levels.
Give it a shot and let me know if you have any success.
That may not last if AI companies start trying to build models that fool it, but for the time being at least, modern models do have strong tells.
And these text didn't train the model in the first place? I just want to ensure clarity on that.
>pangram currently has a false positive rate of about 1 in 10000
Says Panagram.
The problem with just looking at old text is language is a living thing. Say for example I make up the world 'oklambroahaha' right today. Both humans and AI pick up that word and start using it. Now lets say the model says that anything that uses oklambroahaha is 100% AI, you can't just point and say, "well my detection AI is correct on things 20 years old, so it's right skibbidy toilet 6/7".
There is a ton of evidence that use of AI changes the way we speak and write, so it will just turn these AI detectors into bullshit generating classifiers.
And that's before anyone even tries to get the LLM to generate a different style of text. Or for that matter creates a "style model" that rephrases text.
The thing is, humans are significantly worse at maximizing numerical goals than computers.
> And the second is that then people would just make other models that are tuned for defeating that sort of classifier, which would be used whenever the classifier is being used.
Anyone can already do that right now, just grab unsloth studio and fine-tune your local Gemma, but nobody cares. People posting slop content don't care if pangram or I flag their slop with certainty, they are using the easiest option, which is commercial chat models. And given this segment of user doesn't care, the provider have zero incentive to provide a dedicated stealth model for that purpose.
I'm not sure this is even the right premise.
Existing LLMs try to maximize engagement, and they often write in a particular style that has tells, but these two things are not necessarily related. Over-using em-dash or whatever isn't the thing that maximizes engagement.
So the two problems really are, what happens to the actual humans whose writing style is a close match for what a given generation of LLMs output? And, what stops LLMs from using a different style when someone wants to fool the classifier?
> People posting slop content don't care if pangram or I flag their slop with certainty, they are using the easiest option, which is commercial chat models.
They don't care as long as the consequences of identifying it are immaterial, but in that case what's the point of classifying it? Whereas if they need to fool the classifier some threshold percentage of the time in order for enough of their spam to get through, they're going to care.
It's the thing that minimizes the loss during the RLHF phase, and the RLHF phase is the one that's aimed at maximizing engagement (it's literally trained on that).
> what happens to the actual humans whose writing style is a close match for what a given generation of LLMs output?
If a human, for instance because its writing gets polluted by reading too much AI slop, matches the style of an LLM closer than a certain threshold, then his own writing is going to be flagged as well. Whether it's an actual problem or merely a theoretical one is an open question. (unlike OpenAI and Anthropic, humans writers do have an incentive to avoid being flagged as AI).
> And, what stops LLMs from using a different style when someone wants to fool the classifier?
In theory: nothing. In practice if you fine-tune your own model: nothing. In practice with commercial models: the interests of the model making company.
> And, what stops LLMs from using a different style when someone wants to fool the classifier?
Websites have pretty much stopped using ad-blocker-blockers, it seems that it's not a fight worth fighting for them. Does that mean that ad-blockers are useless?
Most people don't even care about ads, I don't think they care about slop either, that's why there's slop posts and obnoxious websites that are unreadable without an ad blocker. A slop blocker used by 10-20% of the internet users wouldn't change the calculation more than ad blockers did.
There is always an incentive to get spam to bypass filters, so as your filters increase in accuracy, those attempting to pass said filters adjust their behaviors.
Spammers/cheaters/whateverers will at least just use a second pass filter that uses one of these 'ai scoring' systems to beat said AI scoring systems. So while it's worthwhile to do it at this moment, this window will rapidly close.
Another example is ad-blocker-blocker. There was a little bit of an arm race between ad blockers and advertisers in the middle of the 2010s, but it didn't last long. Advertisers mostly just decided not to care about ad-blockers.
I don't think RLHF is the biggest reason its style is the way it is.
A lot of it is that it's trained on everything they could get their hands on, which includes domain-specific literature and books that go all the way back to the advent of writing, and then will pick up habits that are common in some specific domain or in 19th century literature etc. that are less common in most modern writing when no attempt is being made to do otherwise.
Do you really think that RLHF humans were requesting more em-dash?
> Websites have pretty much stopped using ad-blocker-blockers, it seems that it's not a fight worth fighting for them. Does that mean that ad-blockers are useless?
Websites have pretty much stopped using them because they realized readers with ad blockers will stop using the site sooner than stop using their ad blocker, and since websites have a network effect, it's better to let a minority of readers block ads when having them makes it more likely they'll distribute links to the site. And because it's the user who controls the browser for web pages, which gives ad blockers a decisive advantage.
> Most people don't even care about ads, I don't think they care about slop either, that's why there's slop posts and obnoxious websites that are unreadable without an ad blocker. A slop blocker used by 10-20% of the internet users wouldn't change the calculation more than ad blockers did.
Sites don't want users to use ad blockers, but having a user with an ad blocker is still better for them than not having the user at all, because of the network effect.
Whereas many sites don't want slop at all, and then if slop detectors work they'll put them in the site itself and block the slop for 100% of users. At which point the slop generators have a 100% incentive to find a workaround instead of a 20% incentive, which is different.
Directly not to care because they lost in court.
And yet the biggest advertizer on Earth (Google) decided to change their browser to make adblocking far more difficult. That or they say "just use an app, oh and turn on notifications". I'm not exactly sure who you think won the arms race there, but it seems like we the user did not.
There is significantly more spam than 20 years ago, just less of it reaches your inbox. This is a very important distinction as the cost of spam filtering is just as high as ever. On top of that most people have given up on their own email servers and instead depend on Google/Microsoft to do it for them. This allows these companies to have an overwhelming influence on email on the internet, to the point they can send spam with near impunity, and where if your system does it will be nuked from orbit by their systems.
And much like now Google supplies both the email spam, and the solution to the spam, they'll gladly supply the LLMs spam and the LLM solution while applying their 'flavor' of what's allowed to the entire internet.
I'm pretty sure the illegal sport streaming websites didn't stop doing that just because it became illegal, otherwise they could have stopped their activity altogether while they were at it…
> I'm not exactly sure who you think won the arms race there, but it seems like we the user did not.
I, at least, won when the webiste showing ads gave up the race (for the past decade at least, only time will tell about the future).
> nd yet the biggest advertizer on Earth (Google) decided
This is actually an argument in my direction! The owners of websites (which are also the ones posting slop today) didn't care enough and the situation only changed because Google moved.
I expect the same thing with slop. Individual websites won't make any effort to make their slop unblockable, and it will only be a problem if OpenAI/Anthropic/Google decide that they care about this market. But unlike Google in the ads market, I don't think the model providers have any reason to care. The web is already dead in their mind anyway.
> There is significantly more spam than 20 years ago, just less of it reaches your inbox.
This goes against your very argument from earlier!
> On top of that most people have given up on their own email servers and instead depend on Google/Microsoft to do it for them.
Out of convenience, but you don't need that to be practically free of spam. Whatever version of SpamAssassin is being run on OVH's mail servers has been enough for that purpose for me.
> they'll gladly supply the LLMs spam and the LLM solution while applying their 'flavor' of what's allowed to the entire internet.
Again, they don't care about the web. They just crawl it for content but they don't want you to read any webpage, they want you to stay in their chatbot. Every other use-case is nonexistent to them (except coding agents, of course, but that's a different product altogether).