Didnt Altman say a few years ago that the next frontier for training would be synthetic data? It’s what everyone turned to (Tiny Stories, etc.). Using existing LLMs to generate more training data for new ones is the obvious step.
LLMs are offered as document generators. This is what people do. They use them to generate documents. And then they train on those documents. The labs do it themselves. They didn’t ask when they trained on the internet. And now the Chinese - and researchers, and hobbyist, and businesses - dont ask when they train on LLM outputs. Especially when they paid for them.
What they call “distillation” is the very knowledge flywheel that we want in society. You buy a book, you may learn from it, and you may write a better one. The author got paid when yiu bought it. Society gets paid when the ideas spread and lead to the creation of more books, products and services, all of which increase the choices that everyone has.
But so if there's distillation, and more distillation, and more distillation, what do we end up with ? A world of ELI5 ? "Grey zones are for losers" ?
Parallels to how China bootstrapped their state manufacturing capacity with Western company joint ventures contributing capital and know how, and then gave them the boot once they no longer needed them.
I've never seen a good blog post or metric establishing what % of a new model's "intelligence" can be captured. I'm not disputing the document's claim. Widespread effort at distillation is clear evidence that it is highly effective. I guess I just don't have an intuitive or technical sense for how powerful distillation is.
"You only leapfrogged us by stealing our tech and copying it" is a very weak line of reasoning. "You did not contribute to the embedded R&D costs to get to this point" is dickering about the price attributed to all three of George Bernard Shaw, Oscar Wilde and Winston Churchill.
The time where even a token justification was necessary has long passed, so I'm surprised they bothered.