I were 17, I'd learn how to build LLMs from scratch(twitter.com) |
I were 17, I'd learn how to build LLMs from scratch(twitter.com) |
I’m sure I’ll get torn to pieces for this but it’s frustrating to continually witness people treat a single person’s prose as the Word of God.
Learn how to make language models from scratch, yes. But learn how to use them, in the context of other machine learning tools, on very small hardware.
When the bubble bursts (and I still tend towards thinking it could burst rather than be deflated in a manageable way), the focus will be on uses of AI that are not like the hyperscalars' products.
People will still be interested in useful AI being added to small things — assistive technologies, home security, garden monitoring, their phones and smartwatches, robotics.
Instead of reductive, reactive make-an-anthropic-competitor advice like this, what about advising 17 year olds to focus on broad, integrated, helpful AI — or on going back through eighty years of history to look at AI projects that failed and reassess them?
I will never understand this mentality, to let yourself be so dependent on something you don't understand at all is to live like a child. But it is very common. I doubt too many 17 year olds will bother even trying to understand what an LLM is, let alone build one from scratch.
Wait no it’s not, that was always happening.
What’s crazy is that people still believe in it.
If the two current bottlenecks, for this LLM madness that could very well be a bubble, are processing capacity and accuracy (a second processing problem) then what comes next? Isn’t that where young people should be looking or are we just giving up on innovation?
If you are 17, go be yourself, whatever that is, in whatever way you want that to be, but do it so authentically and fully. Be unapologetic about what you love and what motivates you, and pursue that with passion and commitment.
Owner of Golf Club Company says I should dedicate my life to golf lmfao.
Possibilities exist now that did not exist before. Those who do not exploit this are fools.
He capitalizes on greed and hype but with a soft, sober and thoughtful voice so as to lull you with rationalism and now 20 years of his “disruption” has mostly ruined modern society and a whole generation of techies have been led astray into trying to “change the world” is the world of today (minus the magic technology really any better than 20 years ago?)
- Good for him and his Tech Bros, bad for the rest of society
Everybody goes to college nowadays and the average white collar has lots of debt and relatively minor financial benefits over a skilled trade worker.
edit: woah, so many people insulted by that. In my bubble and friends, me and another friend are the only people that make very good money compared to non-graduates. Plenty of others opened their shops, went into trades, one learned to tattoo fake eyelashes, one became a (successful) farmer and most make significantly more than the average law/chemist/mathematician/physics/architecture/languages graduates. Sure, the lowest salaries are to be found among the non-graduates too, but I don't see any evidence that graduates make that much more, and that graduating is worth it.
Some answers talking about how "formative college is", but my 25 years old friend with her own shop knows more about real life, business and economy than ivy league MBAs.
Its not good advice for everybody, but it is good advice for a lot of people. What if you want to be a doctor? What if you want to work in R & D? Not everyone enjoys working in a shop or a farm. Also, how old is your friend group? If they are mid twenties you are ignoring the greater scope for advancement in a lot of white collar careers.
> my 25 years old friend with her own shop knows more about real life, business and economy than ivy league MBAs.
Within the narrow limits relevant to her business. How much does she know about macro-economics or financial economics, or scaling up a business? I also suspect you are comparing her to people who went straight on from bachelors to MBA (which is a bad path - study business after having some experience IMO) and lack experience. How will she compare in 10 years time when those people also have real world experience?
If you’re talking about a place to mature, around others who are at a similar phase of life, also yes.
It is where most people meet their cofounders, for example, even if they don’t found anything until much later.
Many people are surprised to learn how affordable elite colleges are if you genuinely need financial aid. I had no idea -- was pleasantly surprised when my alma mater took over 80% off of my tuition.
Just some anecdata, but every single one of my university professors was quite bad, but they think that since they are the professor, that means they are smart and the expert.
Then vote for someone who will make the debts go away?
I mean, have you seen the options for people graduating right now? How people are behaving?
Or forget the data, look at how the story of the new future technology is being told. The people making it recognize that it has the potential to put swathes of white collar workers out of jobs, and they are openly talking/warning/PR-ing about it.
People in tech and SV, the places which have a underlying culture of near delusional optimism, are talking about trying to avoid being part of "the permanent underclass".
Gambling is up, and prediction markets are being treated as financial investments. Wall street bets is a thing, and outright speculative investments are the hope people have to get ahead.
This is happening in the USA, forget the weaker or smaller economies.
When people see the future as one massive zero sum game, with no way to win by building, then they are going to change how they plan their future.
As opposed to vote for someone who gives even more power and money to the oligarchs? You bet your last dollar that people will choose the former over the latter.
I spent a days reviewing the lecture notes for CS336: Language Modeling from Scratch - and then trained a nanoGPT-esque model in PyTorch.
I'd recommend trying it for those who are curious. Computational bottlenecks become much more intuitive when you've looked at the overall process.
He's seeing a future for models running on everyday hardware just capable enough to do what the use case requires.
Agi is the academia solution Software is the practical solution
learning about llms is not useful so that you can make llms later, you want to learn about it so that you can work on next generation architectures. llms before long i imagine will be left in the dust by ebm / physics oriented models especially that can have an embodied understanding of the world. but a lot of things you learn about them are transferable by doing something like this
That's why. That future is uncertain. So why gamble your future on something that's popular at the moment for something that could change completely a year later?
if i were 17, i'd do these things:
1) LC until mediums
2) Calculus & linear algebra even if i understood nothing i'd just stare at vectors and derivatives.
The core technoology is pretty basic, developing a rudimentary understanding for why the individual parts work as well as they do is tricky.
I don’t see why learning how LLMs work is a bad project for a 17 year old.
Optimizing your entire career and the next decade+ of your life on LLMs? Yeah, probably not ideal. It’s almost always a bad idea to make long term decisions based on current trendy things.
And since everyone is using this topic to give their ideal advice to 17 year olds, my advice as a mid-30s guy: seriously consider becoming highly skilled at a specific thing, and don’t be scared off by the idea that it’ll take 5-10-15 years to get there.
When you’re 17-25, the timescale of a decade seems infinite. But it’s really not, and a decade spent “exploring and keeping your options open” sometimes just ends up with you being pretty decent but not amazing at a lot of random things.
Sometimes I wish I had just become a carpenter, chef, electrician, etc. – a specific skill set that leads to mastery over time, rather than the endless exciting-new-thing hamster wheel of working in tech.
And this is a problem with modern society today: expecting 17 year-olds to know what they want to do professionally for the rest of their lives, and to focus heavily on professional development aligned to that.
By that age, I think it's not uncommon for individuals to have interests and perhaps even dreams, but a well-defined career focus that serves as the foundation of an actionable skills development plan? Nah. That just isn't common and I'd argue not desirable. 17 year-olds should be exploring their interests, enjoying early adulthood, learning valuable lessons in the social realm, etc. Not training themselves to become compliant little worker bees.
(Besides the obvious nano gpt)
Wealth inequality is a huge factor in our ability to excel.
Where I live it's basically free and I still sometimes regret not going into the trades. But I suspect this feeling might mostly be a "grass is greener" thing.
Plenty of blue collar workers make more than white collar ones, and have a huge debt free head start in life.
A plumber or electrician will make significantly more than the average bank employee or translator or nurse or teacher.
And they will also have an easier time starting a business as many trade workers are self employed and make much more than hired ones.
Financially, intellectually, socially, it’s a good idea - college is a formative period of life in American culture. That is more and more true the better the college gets. At the upper tier (Ivy League, etc.) you don’t really have to pay anything if your family isn’t already wealthy, and the connections and degree you’ll make more than pay for themselves.
Even worse when they ask themselves.
As to the substance of your comment, there's nothing stopping a 17-year-old from writing poetry, learning guitar, and also learning about how LLMs work, if they're so inspired. Indeed I'm sure pg would encourage it, and that he did the equivalent of all those things himself when he was young. There's plenty of evidence of that in his essays: he wrote short stories, published a scandalous school newspaper, played soccer, and studied fine art and philosophy. See:
I assume the people at HN are intelligent enough to understand this without me having to explain the most obvious banalities to them. The point is why you would do anything at 17:
> if they're so inspired
You said it yourself. The Twitter post, to me, reads like career advice or something in that vein. One has plenty of time to run the wheel like a lab rat later, but at that very point in life you have possibility to find the thing that truly inspires you, that is its own motivation.
"Why would you reinvent the wheel rather than making it better?"
"To learn how wheels are made"
Banger reply.
I have my opinion on this but I'd like to hear the HN opinion, I will just say one thing:
If you are starting with little knowledge, like a 17 year old would, letting an LLM explain it to you is a terrible idea.
My first read was "if you could redo your life from age 17 what would you do". To which "figure out how to make an LLM" would be an insane answer.
I think his advice here is maybe... a year or two too late to be good advice. I can't foresee the job market for machine learning experts being better than it is now in 5-6 years time when said hypothetical 17 year old would be most ready to start career hunting. Either the bubble is gonna pop and the market will be flooded with laid off AI talent.
Even if I'm wrong about there being a bubble at all, I still think that in 5 years time the tech will just have matured to the point of diminishing returns on refining existing architectures. Plateauing until some PHD comes up with something as ground breaking as attention.
I was incredibly lucky to have had a father who worked that out for himself then impressed it on me. The people I feel sad for are those who don’t have people around to make them feel it’s ok to try things that seem impossibly advanced for your age or social status.
the lessons have svg diagrams, ai chat inline (google docs), etc
apologies if this comes off as slop, but its been working great for me.
Sometimes if you want to be heard you must go to the public square, whatever the flags there.
The rest of the post is unintelligible: add some verbs at least.
I'd rather simply write another mnist implementation and check if I really like all that AI stuff at first place. Even then, before going into mature-on-the-way-to-dying tech (LLMs) I'd rather focus on fundamentals - good ols linear models, regressions, stat etc.
If I'd get a buck every time someone said something like this to me when I was in the 13-18 range, I wouldn't have a ton of money, but it's so very annoying when people tell you this.
Regardless if they're "gifted" or not, regardless if you believe in myths like that or not, let children explore what they want to explore, even if you don't understand what it is or why they want to explore that, just let people explore, regardless of age.
It was such a terrible experience being a young kid growing up, with so many adults spending hours trying to convince me to stop sitting in front of the computer so much doing whatever; "why are you even trying to learn that stuff, you have to go to school to understand anything of this" and so much other similar trash.
Sorry, not your fault and I'm borderline trauma-dumping now, but really sad to see this sort of gatekeeping on HN of all places, age is irrelevant to learning ANYTHING, in my humble opinion at least.
Kids, find anything interesting? Jump into it, ignore what adults tell you, and do whatever you feel like, you'll find your place eventually.
There still will be varyy small number of outliers among youngsters who'd be able to extract tremensous value from such an excercise, but for most that'd be _IMO_ waste of of time, with illusion of understanding w/o actually having any.
This also meant that by the time I was actually offered to take a programming class in school (junior year of HS), I had already been able to self-teach myself well beyond what that class was covering, thanks to just working on random projects that scratched an itch I had at the time, looking up anything I didn't know or understand, and internalizing those concepts over time.
In short though, I definitely agree, young kids and teens (and also, frankly, adults too!) should be encouraged to explore things that they have a passion for, without being told 'you need to go to school for this' or 'you cant understand this at your age'
Because I can?. JK. Because that was my experience, of someone who is 2.5 older than 17?
> For some (many?) people, a 'proper' understanding develops _after_ the exploration.
I am afraid you have a too confrontational attitude here, but I'll answer anyway: because I do not believe you can simply "explore" such complex topics like building an LLMs. You'd simply be unable to build LLM drom scratch, unless you'd call cargo-cult chaining magic numpy incantations you've taken from Karpathy's tutorials "exploring".
If I were in "exploratory" state of mins, I'd rather go from entirely different side - I'd try playing with LoRA-ing existing small LLMs, such as venerable 2 y.o. Mistral Nemo, to get "feeling" for what training is and how hyperameters influence the process.
Can you provide an example?
I attempted many projects at a young age that I was absolutely not equipped for. The result of the attempts more often than not left me equipped, every time it left me better off. This is terrible advice.
Transformers are difficult to understand even to people with strong ML background, let alone a teenager.
But sure, make the kids even more depressed by telling them they need to learn how to build an LLM so they can get a job working themselves to death to make Paul and friends rich.
I told my much younger brother when he was 12 what programming was and it'd be a great career. He looked into it and within months was writing CLI games. Eventually releasing his own unity 3d game on steam as a teen.
Eventually he got into CS and did really well because none of it was scary and new. He parlayed that into role at Meta out of university.
My point being, 17 year olds have time to learn new skills and guidance can go a long way.
>Whoa. I’m 19 and I trained a 100M language model from scratch. Did a v2 now with a new SFT experiment to see if I can get better results on same size.
Either way, this isn't really advice for 17 year olds. Pg is thinking out loud about the pathways for founders.
I don't think you can really call yourself a developer unless you at least have an idea how to build a more complex software project like a compiler, and maybe have built a toy one either at uni or for fun.
It's not clear how long this LLM age of AI will last (to be replaced by something better), but nowadays any developer should at least understand the basics of ANNs, and more than just the "hello world" of a cat vs dog CNN. An LLM/Transformer is maybe the equivalent of a compiler in that regard - something that we all use and is complex enough to present a bit of a challenge. You should at least understand the basics of how an LLM is built, and maybe building a toy LLM will/should become the new Comp. Sci. degree toy compiler replacement.
That 17yo would have already built many uncommon bases, and would build further.
That is, a 17yo with proper mentality.
Like, right now especially in the US you’d probably sound crazy saying that one day nobody’s going to buy a gasoline car, and that EVs are going to completely replace them. But that’s basically an inevitability based on the direction we know technology is going to go, it’s only a matter of when.
The point being that the way we heavily rely on LLMs specifically right now is likely to cease to be the case, and that in the future they’ll be almost like an irrelevant component.
Same! I just happened to disagree with your opinion, and frankly, I'd say trying to gatekeep what people learn is closer to "borderline irresponsible" compared to asking people to build/learn/do X.
> youngsters who'd be able to extract tremensous value from such an excercise
But they're youngsters, who are about "extracting value"? Life is about fun, not extraction, not value, not avoiding waste of time but literally enjoy what you do, nothing is more important (IMO).
Then who knows, doing fun stuff sometimes lead to useful stuff, like in my life. But if you only think about "extracting most value for time spent" or similar "optimization strategies", then you'd never discover this part of life.
This is, pardon, demagoguery. There is always "future fun" and "present fun" which a normal person would assign different nonzero weights (https://en.wikipedia.org/wiki/Discounted_utility). Besides, building a LLM _truly_ from the scratch, just using the famous 2017 paper and numpy manuals is not fun at all, esp. for a high schooler.
To you it isn't, is my entire point here. But why extrapolate what you think is fun, to others? Sure, I don't find that fun either (although useful), but who am I to say it isn't fun for others?
Previously they had a class where they'd build apps for other teachers. Like tracking when clubs are, etc. But now that's become easy for teachers to vibe code themselves.
I suggested they shouldn't prereq this class on CS. Heck invite anyone in interested in "building things" and they can get practice at building apps for other people / themselves. Maybe that becomes a gateway TO CS - people who want to learn how things work under the hood.
I suggested post CS class for the CS people should probably be building an LLM or something :)
Is it something that's going to be fundamental for future technologies? I always plan to learn, but end up in a death spiral feeling like I'll invest so much time and energy only for the puck to have moved somewhere completely different.
Would gladly accept any advice :)
I do not see any past constructive experience as a waste of time.
If the goal is to understand LLMs deeply, one would be better served by either joining one of the big AI companies or doing a PhD. And to be honest, I think this journey should have been started 5 years ago, because right now there's too much competition.
Whether that makes a good long-term investment is quite debatable. IMO, LeCun’s take (at the end) is far more forward-looking since it aims to gain more insight into what could come next based on what we know about the current state of the art.
All that being said, the extent to which these people capitalize on our tendency to be blinded by the halo effect is incredible. A constant stream of bite-sized aphorisms…
This is especially at a different level for early startup figures who happened to be in the right place at the right time and usually did little more than digitizing mundane, traditional day-to-day processes. Yet they’re treated as geniuses and prophets, with people hanging on their every word as though everything they say contains some deeper wisdom. PG and the like often strike me as broken clocks and they’re still profiting from having been very early players in the game, who had good instincts for commercialization and capitalism.
————
LeCun’s reply:
> I would try to figure out why LLMs can write my essays but not clean my bedroom. Then I would study topics in college and grad school that could help solve that problem. I'll figure out a set of methods and architectures beyond LLMs that can quickly learn to perform physical tasks as efficiently as humans and animals. That last item is also what I would if I were 30, 40, 50, or 66 years old
Better to start working with harnesses, evals, statistical analysis, etc. - where you don't need the huge hardware for pre-training etc.
I'm sure they'll get right on powering up their computer from the hamster wheel, Paul.
My heart goes to all the kids out there that didn't get the fair shake let alone fair access to tech that gets these condescending "learn to code/learn to LLM" bootstrappy talks from rich pricks that don't know what life really can be like for a lot of American kids out there.
I mean first that is already what plenty of 17yo are actually doing, because that is what they do at school or in parascholar activities. There are already countless of such tutorials where you can do that in an afternoon.
The pointless part though is precisely why Amazon and others are hunting for rare books, all the low hanging fruits have been picked already so just training a bigger model will simply mean burning more energy and money. Sure training a small one for the basic principle is a great pedagogical thing, training another one, medium, then maybe a large one, is also good in term of learning the process and architecture, but one should not expect it to be useful out of that context.
Pure players are precisely doing everything they can to corner the market by making their own scale unreachable by others. Smaller players with access to lesser infrastructure are thus betting on different market, e.g. embedded systems.
17yos should definitely build their (L)LMs from scratch and whatever bigger model they can train for free, or for cheap, but they should not expect that to bring them any riches.
You should want to train a LLM from scratch as an intellectual curiosity itch that needs to be scratched.
The idea of learning one hot skill that has a pot of gold waiting at the end of it was a brief moment in time that came and went.
When I was 17, we would have said obviously support vector machines are the future. Neural networks overfit and don't work.
2. LLM from 0 to Hero, and nanoGPT by Andrej Karpathy
The real value is to unlock meaningful insights and directions from existing data that is there inside companies.
I live far outside any tech city, so maybe I do not understand. But working with LLMs full-time, building for clients and tons of own experiments, I see no value in building on LLMs.
He said, if he is 17 _now_, with no family, no commitments, no pressure to start a career, he'd invest his time to learn to build LLMs from scratch instead of trying to start a company.
The big companies stole the data. The average person can’t do that
That's what I would call exploring.
I too started exploring programming as a teen by cargo culting. Fooling around and getting results is what made it fun. Understanding came later.
Then it is not "building llm from scratch" in my book. Just mindlees following instructions. Could be educational yes, but only trivially useful, if you have no bloody idea what you are doing.
> Fooling around and getting results is what made it fun. Understanding came later.
It is not how LLMs are "built from scratch".
But you're learning.
Maybe just accept that not everyone starts from fundamental theory, and there are lots of people who start learning by fooling around.
It is sold by PG as something special though.
> Maybe just accept that not everyone starts from fundamental theory, and there are lots of people who start learning by fooling around.
Even then LLMs are strange thing to advice to play with, when there are so much more interesting and theoretically accesible for a 17 y.o. so I wonder why would you'd particularly single it out.
(Fully free, of course.)
You should check the code of both world models (JEPA class for example) and compare to GPT. Many of the tricks stay the same, representation is still embeddings, there is a loss function, etc.
The exact architecture will change, but unless there's a new discovery in that area, we've cracked the text component already. We're hitting the limits of LLMs because of the intrinsic limits of text as a medium. But the way we work with text is pretty much settled, fundamentally.
Yeah, if you see it as pointless, then I prefer not to. Have a enjoyable day!
Also many people/kids don't have access to proper "productive" systems anymore, since the whole computing and electronics industry shifted to make "consumer"-devices like smartphones or laptops made for netflix, gaming and spotify.
Breaking the barrier to build a custom system, install linux (or developer tools for Windows, MacOS) is already a complex AND costly task. It was just way simpler in the late 90s and 00s to get something working.
There is very little reason for humans to get all too engrossed in this type of work now, today, with the hope of being good enough at it to command a high salary in 3-5 years. AI can already do it incredibly well, and they can do it persistently and doggedly 24 hours a day.
And younger people I meet don't even own laptops. I had a genz/millennial cusp friend who wrote all her college papers on her iphone.
There really isn’t a very strong correlation between tech industry hiring strength and AI as of yet. Various studies that are out there haven’t even witnessed AI workflows contributing more than modest gains in software engineering efficiency. I.e., being able to write code 20-40% faster isn’t a seismic shift in the industry where everyone is getting laid off tomorrow and we’re all replaced by software.
Even with the questions surrounding the current job market, it’s still an incredibly good ROI career compared to so many other jobs out there.
For example, in my local area you can get a job as a registered nurse working nights in the emergency room and only make ~$115k.
I make almost double that telling an LLM what to do from my house in my pajamas during the day with less time spent in university.
Even if tech roles lose half their salary to automation pressure it’s still a really good gig.
This right here is why nobody is shedding tears for the massive employment crisis in tech.
You make double that shilling ai slopware while they work nights saving lives.
I would tell a 17 year old that the world will always need nurses, same can’t be said for guys sitting in their pajamas burning tokens.
What is happening now in tech has been a long time coming, and it can’t happen fast enough.
For the AI part, I recently started working through this CMU course on Modern AI which teaches you how to build an LLM from scratch. It was posted on HN a few months ago and has been great: https://modernaicourse.org/
If anyone has other recommendations for beginners, please share them!
The reality is that an incredibly small minority of companies in the world do any real training or optimisation. It's unnecessary and inefficient for most purposes unless you are fully dedicated to being an LLM company, and still then it's a struggle. Those few that do train, they spend most of their budget on compute and have relatively small teams.
Getting experience in this field requires having access to very expensive hardware to begin with. And the skills will be quite hard to convert into any real value for someone, leading to a decent income, unless you have a ton of funding from patient investors, or you have decent contacts in Bay Area networks to get hired at the right place.
With all due respect, paulg is in somewhat of a bubble, this is not congruent with the global situation.
Building an LLM from scratch has a hard split between a tutorial project you can complete in a weekend (that's useless for actual usage) and then a solid 1km high brick wall if you want to create anything actually useful from scratch.
Modified open models have a very active community around them, without the need to look much further than Hugging Face.
If you are really 17, my advice is to identify use-case for AI (ideally relevant for businesses) that work most of the time and find ways to make them work pretty much every time. AI reliability is the scarcity right now.
Benjamin: Yes, sir.
Mr. McGuire: Are you listening?
Benjamin: Yes, I am.
Mr. McGuire: Plastics.
Benjamin: Exactly how do you mean?
Mr. McGuire: There's a great future in plastics. Think about it. Will you think about it?
---
I love this scene because it so perfectly captures what it's like to be young and given advice, however well-meaning, by an older generation living in a world that no longer exists for the young. And it's ambiguous and trite enough to be essentially useless even if the underlying idea isn't terrible.
And it's basically a weekend project to put transformers together in a ML library and train it.
The follow up comment,train it to play a game also doesn't make sense? Llms Sony really play games and there are better ml approaches to do that?
Good times.
I wonder if you can dig that out of the historical reddit database. I'd like to see that again. I love how everything is recorded now.
I mean, except when it's censored. That part is a shame. HEY! Is there a tool that tracks the censored comments? I bet there is. And if there's not... that will be a delicious new project for my agents.
Long-term there's probably a lot more potential, and much bigger markets, in robotics than in LLMs, but it will really take the long term to get there.
There's also this little-known concept called learning things for learning's sake and not always trying to capitalize on it.
I think this kind of mindset should be taught way more in school so that people really appreciate learning a subject deeply.
Second in line, build your own agent, that's more in our ballpark, then customize it to your needs, both virtual and physical
As a platform engineer being based mainly out of Australia/Hong Kong, opportunities seem to be getting less unless targeting high frequency trading or banking.
It seems like building a startup with the help of some AI tools might be the best bet.
I'd only recommend taking that path if you already have your first paying customers standing by or are extraordinarily good at marketing.
Incidentally, the skills for the lowest levels of LLMs aren't that far removed from those needed for mobile telephony, in that both are based on maths, computation and information theory.
What most people want from mobile technology is for it to work, not too expensively, and for it to get out of their way.
What most people want out of AI is for no leader to emerge and wield supremacy against the rest of us. People are afraid of it in ways they weren't afraid of mobile, so they're more willing to work together against whoever is in the lead.
Its more like an arms race and less like a utility. The disadvantage I face when my competition has better mobile coverage and bandwidth is minor. The disadvantage I face when my competition has better intelligence on tap is much more significant.
We don't really have the demand for as many telecom companies as actually exist in the world. There's a reason we just have one Whatsapp and one Instagram, not three or four almost-but-not-quite clones in every single country that mostly differ in branding. The reason for the current situation has mostly to do with regulation and traditional, enterprise, "obviously every country needs a separate local branch, because that's what mcDonalds does" thinking. Technology has very little to do with it.
This is why the telecom world now consist of equipment manufacturers, who do most of the hard tech stuff, and actual telecom companies, who operate the equipment, rig towers in their local country, and maybe write some glue code to integrate a core from vendor A, a billing system from vendor B and a CRM / corporate invoicing system from government-approved local vendor C.
Banking also works similarly, though modern Neobanks / Fintechs and bank consolidation are slowly dissolving the concept of national bank branches.
It is viable as a toy project, but there are vanishingly few career opportunities.
The same is not so different for AI. A few people design novel AI, but there are a lot of people training AI (especially if you include fine tuning) and implementing AI, even as a hobby.
It seems to me like he’s saying that doing this thing would be 1) fun and 2) a great way to become employable in the future. I don’t believe he’s saying that this project would be some kind of job training exercise.
But that was way back in the early 1970's and all I had to work with was a mainframe.
Well the mainframe itself wasn't bad, the real show-stopper was that I didn't own the computer outright, no strings attached, no debt, etc.
>I'd probably try to make an LLM that I could use on some specific problem.
I thought so too back then, still do so I guess this is one of those things that could stand the test of time. I always wanted to start with something a lot simpler than a Moon mission myself. At 17 I already had a significant breakthrough in the chem labs and it was from alternatives to a single processing step plus everything that descended from that, rather than trying to tackle a much more complex detailed multi-step synthesis. I was only 17 but I was not trying to be a slouch, I don't think pg was either at that age but his advice is not for just anybody. I couldn't have done it if I hadn't made major progress since being 16, and it really emphasized at the time how much maturity can make a difference. My imagination ran wild as I extrapolated :)
In a reply from LeCun to pg:
>>I'll figure out a set of methods and architectures beyond LLMs that can quickly learn to perform physical tasks as efficiently as humans and animals. That last item is also what I would if I were 30, 40, 50, or 66 years old
I see no reason to stop at 66 either ;)
But I figured that people owning more computer power than I could ever afford were going to be doing something like this as soon as they could, without having to wait for something like an LLM to arrive before getting peoples' attention.
It did seem like things were going to take longer than you expect, so it's pretty good to have a lifetime of concentrating on the specialized natural science domain expertise, focused now for 50 full years on how it would combine if AI ever got good enough.
Both the natural science and the AI need to be a major cut above, I still see dramatic room for improvement in my own work. If I'm going to have to rely on "other peoples' AI" then that natural science component is going to have to pull a lot of weight to keep up with the kind of computers that only rich-as-hell high-rollers have access to.
For anyone who wants to dork around there is https://github.com/rasbt/LLMs-from-scratch which is something amazing that I think anyone who wants to engineer things around LLMs should at least blast through and read.
And the job postings are often ridiculous. I recently was an AMD job advert in Germany for an ML Kernel Engineer, not Senior mind you. The requirements went something like
> Masters Degree required with strong preference for a PhD with peer reviewed articles in {journals_list} > 10+ years of experience in C/C++ > GPU programming experience required > 10 more ridiculous lines
No idea how a teenager self teaching himself LLMs is supposed to even get a shot...
1. I can make turn a stone and water into a delicious soup
"17 year olds, learn to build an LLM from scratch"
2. This soup would be more delicious if we add a few carrots. Does anyone have carrots
"Increase your chance of success by getting a Masters degree"
3. How about potatoes?
"And get a PHD"
4. What about some salt?
"And publish some peer reviewed articles in {journals_list}
5. We should also add beef
"Now work in the industry for 10 years"
6. See, this soup is delicious, and I made it all with a stone
"See, you're rich, and it's all because you learned LLMs as a 17 year old"
The only jobs that he found he was highly over qualified or paid very little.
In any case, it doesn't look like there's this crazy rush to hire all ML talent, even the one that understand the math and technology deeply.
Maybe people simply don't want math PhDs but something else? Since 1-2 years ago I started doing consulting/freelancing in the ML space, but more on the infrastructure, deployments and similar stuff, as a general purpose developer, and I have a waiting list of clients interested in more work, some of them even trying to recruit me to work for them full-time as well. I'm based in continental Europe, fwiw.
There are high school students competing in contests that cover parts of the (Math) theory behind AI. A lot of high school research programs are integrating AI with other things and complex mathematical models…
To me, this is bizarre as Calculus is barely taught in high schools (and likely poorly).
Don’t get me wrong, these kids certainly aren’t the usual lot.
Yet, I really wonder if they know the fundamentals. Do they even understand derivatives or just memorized the rule for polynomials? Can they even explain what a transistor is?
Normal curriculum takes 5 years to go from Algebra I to Calculus. Real Analysis, Linear Systems, etc. are fundamentals taught only in college…
Feels like too many are trying sprint before even learning to walk.
Sure you can gradually climb the ladder by demonstrating your skills bit by bit and getting access to more resources. It has very good prospects if you do manage to push through. But it's a hard and risky path, and you will not be able to get any interesting results for the longest time.
For a young middle-class student, it just doesn't make much sense. You can do much more impressive and impactful things with your time without getting into that black hole.
I know how to build an LLM, I know plenty of fellow young engineers that do too. It's really not that complex. But they can't do much with it without capital or access.
Good engineering has never been a bottleneck in this field, it's been all about having access to capital and taking smart but dangerous risks burning it on compute, without much idea of how long you need to keep burning for. There's still no end in sight, some are still managing to convince investors and keep burning, and we are seeing progress, but the business case is still unclear. If you want to get in that game, go ahead, but it's not something I would advice the average young engineer.
Learning should not be done only as a direct path to getting paid.
Learn to create pattern matching and intuition to solve future problems.
When you are 17 it is a good time to understand how the world works so you can build on top of it in the future. If we assume most tech is going to have an LLM as part the stack, a solid basis in how LLMs work is likely to help you in future endeavors the same way a solid basis in how the web works helps you today.
Maybe a 17 year old should learn both. As a small anecdote when I was 17 I learned a lot about load balancers, failover, and building self-healing systems running small hosting company that had to be fault tolerant when I was attending high school. This wasn't at state of the art levels (e.g. I wasn't configuring gigabit routers or global CDNs -- but it was useful pattern matching for future problems)
I currently don't touch any of that tech, but I have working knowledge that still serves me today.
Think long term.
Paul, I think, is talking about achieving outsized outcomes in relatively shorter timeframes (as the timing is just right to be investing in learning this tech) for high agency folks who can also afford the ordeal in wanting to maximize for impact & ambition. Of course, there's real risk one may get no where, but even in failure, given you were building the LLM yourself, you might end up with other adjacent, high reward opportunities.
what kinds of startups ?
And so I think the idea is more to understand tomorrow ... from first principles.
In the late 80's, as a teenager, I learned x86 assembly and C because that was the only way to squeeze out enough juice from my shitty CGA (and later VGA) card to programm the games/graphics that interested me.
I haven't written assembly in years.
But whatever I did in my career: it helped me and gave me an edge over my peers to have a foundation that is very close to the metal.
We don't know this. So many people are simply claiming this confidently, and a lot of them are betting their careers on it, but nobody has a crystal ball. Whenever someone tells you confidently, and without any doubt or qualifications, that something "is the future," be skeptical.
I remember when the Segway was definitely going to change urban planning worldwide.
Current AI can automate significant amounts of grunt work in programming and math. It's good at running web searches and writing summaries. There are a few other niches where it is currently successful. But other than that, many corporate AI projects are spectacular failures.
So just given what we have in hand, assuming no further breakthroughs, then we're maybe looking at AI being somewhat bigger than the Internet. Which would make it a revolutionary technology, sure.
But to get from "a revolutionary technology" to "the substrate the future runs on", then you need to assume more breakthroughs: long-context operation over weeks or months, displacing human workers 100% instead of 75%, and the ability to directly economically compete with actual humans. And people are investing literal trillions of dollars to make that future come true, without really thinking through what truly competitive-with-human AI would actually mean. We might be looking at massive job loss, centralization of power, fully automated "companies" with no humans dominating markets, and other dystopian scenarios.
And in those worlds, it's unclear that being good at CUDA and matrix math will be all that helpful, careerwise. The AIs are already pretty good at that stuff. Data scientists get paid OK when they actually get hired, but it's not everything college students were promised in the 2010s, either.
We can't yet build a fully-general competitor for the human mind. But we're getting closer. And if we ever do build one, the consequences will be really weird in any number of ways. So I worry about visions of the future that assume AI keeps improving significantly, but that also assume it still somehow remains a "normal" technology that doesn't, for example, render most humans fundamentally uncompetitive.
AI is a great solution in niche areas but generally doesn't make much money. All the large companies are in the negative.
The steam engine was less of a bubble and was much more revolutionary and had a greater impact.
There are plenty of areas were we need people to do this for insurances, banks etc.
AI/ML exists on many levels.
Paul likely assumes there will be a sequence of additional papers with the same impact as Attention is all you need, which will spawn a lot of opportunity for a larger group of experts who are conversant enough to advance the field even if they do not themselves create such a major innovation. Not only is this deeply exciting, it is also highly meritocratic as there is still scarcity of the kind of intellect and creativity necessary to swim there.
Machine intelligence might soon surpass it, though, and Deepseek is 100% Chinese mainland educated. Paul's description of building an LLM from scratch is meant as a vague starting point for being an innovator of the highest value aspect of modern AI innovation, not as a specific prescription.
-- 3Blue 1Brown's Neural Network Series: https://www.youtube.com/playlist?list=PLZHQObOWTQDNU6R1_6700...
-- Karpathy on LLMs: https://www.youtube.com/watch?v=7xTGNNLPyMI
-- Stanford CS336: https://www.youtube.com/watch?v=JuoVZkPBiKk
Then do this hands-on:
-- Karpathy's zero-to-hero: https://karpathy.ai/zero-to-hero.html
I think this is disingenuous. One could say that drones are nowhere the efficiency limit either: a bee can fly for hours on the energy contained in just a few milligrams of honey, while our best battery-powered drones can't stay airborne for more than 30 minutes. But comparing energy efficiency of electric/mechanical devices to their biological counterparts is not an apples-to-apples comparison. There's a world of difference between the energy storage and delivery mechanisms.
And as many have pointed out already in the siblings, it's not just about the compute but the access to petabytes of training data.
I can see the point. It's unlikely that a 2.4T LLM will be integrated into, say, a pesticide drone. You'll still need some kind of LLMs to achieve maneuvers that "normal" programmings can't achieve.
But what if everything basically turn into that? Essentially, instead of build me a web app to solve X and do Y, build me an LLM to serve X and do Y. (unless the current LLMs are able to do it end-to-end but then they can hardly write coherent software/personal opinion).
We had scores of students study how microprocessors work and compilers work over decades, yet we have 3 or 4 major processor companies and a handful of programming languages. Yet, what they learned was probably crucial in their development as engineers.
We are also so early right now that even 2-3 years from now who knows how many LLMs and model firms survive (esp. given the "snake eating its tail" venture/investor funding situation)
But LLMs are tools. Does a great engineer need to know how vscode works? Might be helpful to understand how extensions work, LSPs, and project configurations.
Usually when working with any tools, you need to understand how to get the most out of your tool for your needs and that's about it. Core fundamentals about how software and hardware works in general seems like it would be MUCH more useful than LLM core knowledge.
Finetuning model is cheap and incredibly useful for deployment. You don't need to pre-train a frontier llm from scratch to make useful models.
There is tons of domains where you and fine-tune llms and deploy them for value in companies and for your own entrepreneurship ambitions. I have made this a big part of my career for the last few years and now I'm working on finetuning models for starting my own companies.
End of the day they're all customized data stores and protocols to interact with them. May as well stick to a uniform toolkit with fine-tunes.
Not that other tools aren't useful. But reaching straight for a bunch of infrastructure reliant services is like jumping in with k8s when you're still at a stage where basic mocks in code are sufficient.
I won't roll my own encryption or UI lib but want to stay focused on the incompleteness of the project I have to ship not all the buttons and knobs of some dependency or framework. Same old manage context switch problem.
But there are also a lot of prerequisites, namely does the enterprise have its sh*t together on a technical level. Does it have the processes and data pipelines available to train and benefit from these models? Probably not!
Applied ML is at the crown of a tech pyramid whereas most enterprises are still struggling at ground level. Being able to build from be ground is likely a safer skillset than only knowing how to work at the (non-existent) apex.
In reality the problem is that it gets blasted out of the water by a much worse architecture trained on 10000x the infrastructure. And while I'm sure the freshly brought in ML student came up with a 10%, even 30% better architecture, it just doesn't matter. (and never mind that even OpenAI hasn't really solved a voice model yet. Try it. It can probably match 2026-quality call centers, but it's no substitute for an actually empowered human)
... and yet, if you look at what hyperscalers are getting paid for ... comfortably more than half the income is training. Which makes no sense on so many levels.
e.g. https://valueaddvc.com/blog/inference-chips-vs-training-chip... (I get it, not great first source, but st
For some very niche cases I think this is probably the case but for the vast majority, the company's data isn't as useful as they think it is or anywhere near the size needed.
I don't know where you are located, but in EU, in China, and yes even in Silicon Valley, the vast majority of companies do not do any real AI engineering. There's nothing wrong with it, it's just not a smart path for most purposes. You can do amazing things without training, and if you try to train, you cannot get anything amazing unless you burn millions.
Very few people can afford to play the long game and cross that dessert. And, sure, you will not get far without good engineering, but good engineering is definitely not sufficient and is not the primary bottleneck.
However, I do learn stuff about models that takes it from “magic” to “useful tool I understand the limitations of.”
Do I do it for that reason? No not really, I’ve never had luck learning something because it would be good for my career. I do it because at my core I’m a bored teenager who wants to make the computer do cool shit.
A 17yo who trains their own LLM will have a much richer understanding of what AI is, how it works, what its potential capabilities and pitfalls are, versus someone who spends the same time doing something else.
Now with LLMs, people write native apps in Rust, and I'd like to think some of them found that there isn't such a huge jump in difficulty they assumed there would be.
That's the same case finding companies that will actually pay for hand made LLM instead of using something from big providers will be hard because most companies won't be able to afford it.
Yeah if you find a company that will do that stuff directly, good for you, but you will have to be very lucky and you will have to compete with other people who followed PG advice.
So I would rather learn all there is about properly using LLMs and integrating them with existing systems, that will most likely by useful for 90% of companies out there.
Building business niche harnesses is in my opinion much better direction. Knowing what will work best in specific cases is it FTS or vector search, optimisation of usage, getting best results while using cheaper models, knowing how to use tools to run models on the servers, and all the tooling around that like various MCP or just tooling that will be provided to models.
That is what I am currently busy with and I already have customers for that knowledge.
But that’s the tip of the iceberg. If you have any ambitions of doing this professionally, it quickly becomes clear that all it’s all about knowing how to deal with problems that are only present at massive scale, when an LLM is actually L and becomes AI.
The mundane details about how to build a tiny autocomplete model and the maths behind it you can learn in a couple weeks easily. It’s not black magic, there are much harder areas of computers science.
Not that expect to make it big as a LLM researcher but building something from scratch gives a much deeper understanding than what you can get from simply using something.
Much in the same way as implementing and designing your own programming language makes you a much better programmer.
At the scale you are probably imagining, this is true - but take the hype out of the OP and what you have is just someone saying the field of data science exists and is growing.
I would think not, but when I started to look into OCR options recently - assuming that obviously a dedicated tool would do a better job than an LLM - I was wrong (apparently).
I think a lot of people assume that only the big AI labs can do cutting edge research, but there's a strong argument you can do it as part of little tech as well.
I just finished fine tuning Gemma e2b for local code completion on my local machine.
This comment just reinforces what the post actually means. We need people that are LLM natives, computing solves itself with time and with scale adjustments
I didn't write a LLM from scratch but it is on my "wishlist" so to speak. From what it seems, a GPT-1 class LLM can be done from scratch in a few days and tens of dollars of cloud compute or a high-end gaming GPU.
It is an exercise not unlike building a compiler, a school classic. You are unlikely to ever work on a compiler, but at least, now, you know your tools a little better. It is not about becoming an expert, that takes years, it is about knowing what you are doing.
If you intend to make software engineering your career, you will want more than surface knowledge. And that part is entirely on you, or on your school if you are a student. Companies will not pay for you to learn the fundamentals, they want short term returns, because you may leave at any time. But you as a software engineer may have 40+ years left, so it is worth thinking long term. Claude code may become obsolete a few years, but linear algebra is not going anywhere.
Probably this also is too cynical and simplistic, but: really what’s special is the ability to get this kind of capital, with the freedom to burn it on mad moonshots, with long enough leeway to actually get to see a few of the moonshots come true.
No wonder that the head of YC made this happen, this is exactly what they are world-leading at.
It similar to understanding how a very basic CPU works. Just because I'm not going to work at intel or nvidia or whatever optimizing the hell out of a chip, it doesn't mean I just throw my hands up and think "magic" - the basic architecture isn't that difficult, and the value of knowing it is astronomical for anyone writing software.
I feel like the "ALWAYS HAS BEEN" meme is apropos here.
1. Both training and optimisation will get significantly cheaper and easier quickly.
2. Politics will probably get even more insane before a potential reprieve on the 20th of Jan 2029.
3. The big AI firms will become part of the surveillance capitalism network, if they're not already.
So I think for self-protection a lot of companies will be looking near to medium term AI independence.
For the time being, unless you truly have millions, the outcome from training will be very net negative, while focusing on building on top of existing AI will yield amazing things if you apply the same talent and effort.
When it does get cheaper, then it will be easier to acquire the skills and experience too, and the struggle you went through by trying to do it now will be somewhat wasted.
Besides, I am well versed in this field, and it is not rocket science. There are plenty of software engineering domains that are a lot more challenging, like high-end graphics, large-scale data engineering or kernel programming. People will learn to train LLMs when people want them to.
In reality, enterprises are happy to offload even risky tasks to others as long as they get some contractual guarantees about their data. Would they like more choice in who to buy from? Yes, but not enough to in-house such a specific discipline.
Fine-tuning a model or LoRA based on the companies data set is more feasible but you're likely going to need several runs as you test/try out different base models, parameters, etc. This is why there are a lot of fine-tuned models on huggingface based on base or instruction-trained models from the larger AI companies that have released open weight models (Microsoft, Google, IBM, Mistral, DeepSeek, Qwen, etc.).
Training is limited on memory first (storing training data and weights) and computation second. Realistically you need to own or rent 2-8 H100/B100 devices or Google's TPUs.
The majority of workflows for a company providing AI capabilities are likely best solved by tailoring a system prompt for the chosen model, evaluating the prompt and model with tools like promptfoo, and then running it on a compute cloud provider (including AWS Bedrock). If the company is big/financially well off enough they could look at buying the hardware needed to run it on their own servers.
For other uses like agentic software development you'd need to spin up a suitable model on a compute cloud provider (or local hardware if the model is small enough) and then tell your IDE/editor to use that model. You would need some way of benchmarking and evaluating the models to see if they are capable of doing the tasks you need. -- There have been some tests done by people on YouTube that suggests that Qwen 3.8 27B is a decent model, but your needs may vary.
You’re missing the point. Understanding how Unity works fundamentally makes you a better Unity dev.
Writing your own game engine makes you realize that the Unity engine is not really that well written....
No.
That’s what people like us on HN want. The people out in “Greater Userland” just want the black box to answer their questions. They could care less who is behind it. They don’t yet attach their black box to Amazon or Microsoft etc. And most won’t care enough to be inconvenienced even when they do make the connection. (As your competition argument implies.)
Heck, a lot haven’t even made the connection between the black box that gives them answers and data centers. They think, “ ChatGPT good” and at the same time think “data centers bad”.
So I'd say that understanding LLMs to the level that you get to by training your own baby one would be a solid foundation for pretty much anything built on top of the "real" ones.
At 75% replacement of a worker we would already have huge job losses as each individual would be doing what several before did.
The only alleviation would be the creation of new equivalently paid jobs, which is no better than a hypothesis right now.
Citation needed
I don’t make the rules for how much each profession is able to make in compensation.
The world will always need nurses, but that doesn’t mean that it’s a fantastic career to get into if you have a neutral career preference and your primary consideration is university tuition ROI, expected compensation, work schedule, and day-to-day physical exertion.
My point isn’t to debate the virtues of each career, I am intending to stick to objective aspects of them.
In that sense, telling today’s kids that there’s no future in tech careers just because there’s a short term hiring slump is extremely premature. I certainly wouldn’t tell a kid who is passionate about tech to avoid the field just because the unemployment rate is currently 7%.
https://www.investopedia.com/bachelor-s-degrees-with-the-bes...
2. That's why I have an agentic agent as well installed, Qwen 27B, outrageously good, better than sonnet 2 years ago. And it's mine, I can give it confidential info to work with since I own the whole chain. See where I'm going with this?
When I was a pre-teen I stumbled upon CD-rom hacking guide to bypass disc requirements on games, I remember opening up the file and the screen being filled with HEX code. I was so overwhelmed I just closed it and never touched programming after that for 15 years. My life would have been totally different if I had embraced the unknown instead of retreating.
Citation needed. I'm guessing you're not counting children.
I "guess" you could put all your effort into moving up to the next level in computer science, or you could put the same progressive effort into reaching new levels in video games on the same hardware.
Alternatively you could even put all your effort into social activities and leave the technology to other people entirely.
It might even be possible to find a balance between things that are widely "understood" "socially", and those that are not ;)
But it was generally seen as a gimmick instead of desired before Apple made it look good. Even when the iPhone came out, one of the jokes was how the grid of icons looks like how a Windows user's desktop would look like when they didn't understand the filesystem.
If you look at model training jobs a lot of the work at this point is creating RL gyms (normal programming work), but most people still think the work is all neural architecture research. Doing the former is fine but won't teach you much about how to build LLMs, whatever that means now. Doing the latter is a very hard market to get into: not many jobs and requirements are often like, "you must have published at one of the following conferences". Prior experience is assumed. Most of them seem to treat Google as ML university and source of new recruits. It's understandable given the cost of training runs.
I'm not sure why it's like this. If you look at the real world, you have stuff like ggml, which is about as hardcore as it gets in the LLM space, and it was made buy just a guy. Same for this like ComfyUI
If you get enough academics in a place, they tend to close rank, and not let anyone in without the same credentials. Data science used to be like this, they were constantly on about how you need a Math Phd to even apply, yet when I met these guys IRL, most of them were just running Python math libraries.
These previous examples show that if you understand at least a part of the problem space, you can 100% contribute without academic credentials.
I don't think the "incredibly small minority of companies in the world do any real training or optimisation" part is necessarily as true, as some parts of the work I do get is about helping them optimize training and infrastructure around training. Mind you, none of this is for building LLMs from scratch, it's 99% fine-tuning existing checkpoints.
I'd also agree with "paulg is in somewhat of a bubble" regardless of this, which is worth remembering whenever you read his content. Same goes for any person living in SF, and dare I say the US. But also, YMMV, I live and work in Europe, probably why I have this perspective.
Also bunch of past workplaces who've adopted AI in various ways who reach out once they find out what my current focus lies, but that's harder for others to replicate unless you've already had a career as a developer.
Is it possible to see some of your old works? Personal research?
The question is more one of opportunity cost. At 17 you need to start finding your way in the world. It's best to learn skills lots of people need.
Worst of both worlds - no casual gossip feed in the Bay Area, no big dog meetings in the UK. (Which mostly has no idea he exists.)
As for the question - what are the odds LLMs will be anywhere near the top of the tech tree five years from now?
The trend seems pretty clear to me - local/offshore models are snapping at the heels of the big names in the US, and the current investment arc is insane.
I wouldn't bet on Anthropic or OpenAI being leaders five years from now. Longer term, I especially wouldn't bet on the US build-yourself-a-monopoly corporate model surviving AI at all.
I’m 40, and I don’t.I took that abstraction for granted and “left it to the big labs”. However I want to build my own LLM for learning purposes.
On needing big expensive hardware.. necessity is the mother of great innovation. Perhaps 18year olds trying to build their own LLMs in constrained resources environments will result in ground breaking ideas of achieving better intelligence than the one we currently have….
The world needs pragmatic folks who work at a higher abstraction and make LLMs useful, AND also folks who think why not “this other way”? And build newer ways to do fundamental things.
Given the usefulness of current LLMs, I would certainly encourage anybody to try and build their own LLMs, and see what they come up with…
Heck if they build a rack full of old laptops and run something with it that could be done “better” with modern servers, I’d still appreciate the learning running things on those little machines bring.
Maybe with a decent consumer GPU like a 4090, you could do experiments like distilling and fine tuning a small image model for edge deployment for specific tasks.
Even there, many use cases might require renting compute for $10/hour and investing a few hundred.
A LLM from scratch? Forget it. You can do theoretical experiments, but not build anything remotely useful with that kind of budget.
If you're talented enough to come up with revolutionary methods, maybe an university or AI lab would be the place to be.
10000%.
Even more so when things are not just expensive by nature, but truly overpriced beyond that point.
>10000%
Once in a while you do get somebody who only spends a dollar and gets more out of it than a seasoned high-roller spending $10000. Most of the time the waste is borne by those who can afford to throw away $10000 more easily than an economizer can afford to lose one dollar, so nobody is crying about it.
With how ridiculously large the language models have gotten though, a 10000x improvement in actual intelligence does seem like it could be lurking unrecognized at a different point on the compass.
You can however learn everything you need to know to get on the career ladder as a software engineer on a regular home PC.
Do you take the first step or rule it out because you don’t yet see the complete picture.
As a teenager I never hesitated to try things out. As a young adult I wanted the whole picture. Now I’m back to playing / trying things out. I kinda wish I’d not given it up. PG being a bit older and reminiscing - I bet he’s in that bucket too, whereas someone trying to establish themselves professionally probably (aka me early 20s) wants to see the path.
The “large” qualifier dates back to pre-transformer language models, where even training a multi-million model was hard due to how poorly it scaled. GPT-2 was a large language model, despite being only 124 millions parameters.
Due to how much high quality data is readily available, anyone can now train a sub-billion (L?)LM on commodity hardware.
And I'm personally convinced that pretty much any enterprise use-case of an LLM (except coding) is better served by a fine-tuned small (<2B) model that is trained specifically on the task, rather than a generalist frontier model, so learning the engineering around fine-tuning is a key skill that companies will realize they need sooner than later.
I saw my first language model in action in 2014, I was writing blog posts about them back in 2016. In recent years, like many of us, I spent some time learning ML frameworks to see if it'd be a fun career pivot.
But:
1. It doesn't seem especially creative. All the frontier labs have nearly identical model personalities, capabilities and even app designs. We're not seeing them differentiate from each other, implying that the design space might not be that large. In which case the opinions and unique approaches of specific engineers aren't that important, they are interchangeable at the right level of skill, and what to do next is usually obvious to everyone.
2. It doesn't seem like a big job market. A lot of ML jobs were wiped out in recent years by the rise of LLMs. Lots of NLP specialists etc were suddenly replaceable with a cheap API. The jobs that remain have compacted into a small number of companies. It's a small community which greatly increases career risk, especially as so many are unprofitable and/or have strong ideological requirements.
3. It's unclear how much demand for better models there actually is. Do we actually need smarter models? In robotics clearly yes and robotics is interesting and high potential, but for pure LLMs/image models, most users are already incapable of setting tasks that stress the best models and are happy with the cheaper smaller ones.
Using the models on the other hand is a very large design space, and has a lot of scope for creativity. I see use cases for AI everywhere, but most companies seem to stop at putting a chatbot on their website or asking Copilot to rewrite an email before they send it. A lot of companies have hollowed out their IT departments over the past twenty years. It feels like a new golden age of consulting work could be upon us.
I really don't think I'd tell a 17 year old to learn how to train LLMs. Learn how they work and how to use them, sure, absolutely.
There is the more likely reason they are not differentiating. They use almost exactly the same class of model. Everything is linear, parallelizable. It's incredible path dependence that's now invisible enough we think it's a natural law. Nature is not linear.
Good luck relying on in-context learning for a 600M LLM.
> The actual adapter training is automated and put behind simple APIs.
That's like saying it's worthless to learn infra because you can use serverless instead…
> All the frontier labs have nearly identical model personalities, capabilities and even app designs. We're not seeing them differentiate from each other, implying that the design space might not be that large
The design space for a generalist model isn't large, by definition. But the design space for specialized smaller models is much larger. If you can train a 200M model that, for your use-case, is competitive with a frontier one, then you'll make your company save a lot of money in tokens.
> 2. It doesn't seem like a big job market. A lot of ML jobs were wiped out in recent years by the rise of LLMs. Lots of NLP specialists etc were suddenly replaceable with a cheap API. The jobs that remain have compacted into a small number of companies.
We are in a strange place where a few companies are collectively burning a hundreds of billions a year to sell things a few pennies for the dollar. Of course it's going to be cheap and concentrated. How is it supposed to end though?
> 3. It's unclear how much demand for better models there actually is. Do we actually need smarter models?
That's the thing actually: I don't think we need better models this much, and if we don't need better models we need the cheapest possible model for a given use-case.
I'm with you both 10000% on this :)
I just figure the next level would be if a brilliant innovator comes up with something more than 10000x as effective as an LLM, for one reason because it has nothing to do with an LLM.
Probably by somebody who has gotten good at making every dollar count for quite some time.
Learning to hack something together in high school using the latest technology (vacuum tubes, radios, microprocessors, web/javascript) has been a common theme in the tech world for generations. With LLMs and online tutorials, this isn't even a difficult suggestion. Do people think learning new tech is somehow wasted effort?
The above is, after all, the whole genesis of the word 'hacker'. We should celebrate that.
I tried to modify the embedding output of bert to make it generate box embeddings instead of point ones. At the time I had access to university provided A100 gpus but even with all that a training run took half a day. Models these days I don't think I can train it in any reasonable time with that much compute.
I can't recall or point out exactly when, but there is a stark before/after moment where the opinions of anything pg went from "Interesting and maybe true in some ways" to what we see today, lots of knee-jerk reactions and hardly any comments about the actual content.
Hazarding a guess, I think the moment Altman became the CEO and later during COVID, the sentiment seemed to have been shifting towards what we see today. But this is all based on hazy memory, rather than looking at the data. I'm sure there is a blog post waiting to be written about analyzing the sentiment of comments to PGs articles on HN, and you'll see a shift somewhere.
The sycophancy on HN is starting to break down because there is a higher proportion of users sceptical towards the outputs of the VC and wider investment world than ones who believe they're potential beneficiaries of it.
Tech industry people are becoming less interested in HN as a warm handshake into the startup world because, frequently, they're disgusted by it. And this reflects on the sentiments people post on PG's articles.
Increasingly if those at YC want the same kind of low-bar praise they got before, they will need to get it from machines.
This might be hard to disentangle from the tech correction layoffs of 2023 and 2024, but my perception is that most of the spite comes from Tech Outsiders opposed to jaded Developers.
Hard disagree. This submission is still being highly upvoted, while another recent post[1] on the harms caused by Graham’s fellows[2], with a fairly tame comment section, has been flagged. That is a constant on HN. It’s not a fluke, it’s as predictable as the sunrise and getting more pronounced.
I’m sure we’re both biased in our perceptions. Mine is that HN in general (certainly more than any other website) used to worship[3] everything he wrote, together with others like Musk, until things started to really go to shit and many eyes have been opened to the effects of the unfettered greed of rich tech guys out of touch with reality.[4]
[1]: https://news.ycombinator.com/item?id=49411762
[2]: A better English word is escaping me.
[3]: That word I choose hyperbolically but deliberately. It definitely was not “interesting and maybe true in some ways”, it was much more hardcore than that.
[4]: That is not “knee-jerk” but a slow realisation still ongoing.
Don't many of the commercial ones prevent you from using them to build LLMs?
I would say the reason for the negativity is not because it's a bad idea for a project, or that doing projects in general is a bad idea (it's not!), it's because it's a very specific thing that is not for everyone. The best thing about computing is the low barriers to entry. You can basically work on anything that takes your fancy. So those who are interested in ML will be drawn to learn about LLMs. They don't need anyone to tell them to do it. Telling everyone to do it reminds me of the "just learn to code" stuff of a decade ago. No, please don't, please find something you enjoy.
Looking back, it was an ideal start to a 40 year career as a software engineer. And of course currently I'm reading "Why Machines Learn: The Elegant Maths Behind Modern AI" because the urge to understand the machines I want to master and control hasn't gone away.
Good. It took a long while for ICs to be paid more than 'management' in the US. It's still a hard-fought battle here.
// a flood of coders in the '10s-early '20s whose only creed was FYPM and "what's the minimum
I wonder what possibly could have happened in that timeframe?
https://www.theguardian.com/technology/2014/apr/24/apple-goo... https://en.wikipedia.org/wiki/High-Tech_Employee_Antitrust_L...
Sadly, I attribute this to a decline in having “good will” towards others, and an increase in cynicism, prejudice, and meanness.
"Barely?"
It's insane to lump vacuum tubes together with servers with 1 MB RAM. My first PC, which I used for a decade, had 640KB RAM. And that was an upgrade from 512 KB RAM. My other PC had only 128KB RAM. None of these were considered the equivalent of (by then long dead) vacuum tubes in their day.
You can get a lot done in 1 MB.
No. But funnily enough that is a promise by some of the AI cretins and their boosters. Oh yeah best case scenario you learn how to build LLMs for us. We’ll employ you. And then ultimately that just becomes training data for the LLMs to do it themselves.
But why are people cynical? they ask.
It would be a good idea for young people to deeply know how these programs work. Not so that they can spend their career building them, but so that they can approach the next class of problems we'll all start trying to solve, with intuition all the way down to the weights and underlying mathematics. And also, to develop a healthy intuition of when "Just LLM it" will not be the right choice.
"Build an OS" wasn't a common university project because we were all expected to go out and work on Windows, but because understanding the bare-metal firmware for a computer helps you deeply understand how to intuit building for a whole class of problems.
I didn't do it because it was useful to me in a practical sense. It's because LLMs are fascinating and I want to know how they work. From that perspective it's been a great experience. I have afirm grasp of the basics. This makes it much easier to understand frontier concepts like compressed latent attention. I can follow the field and understand it.
Not sure I would have got as much out of it at seventeen. I have a lot of background and experience which made it much easier to learn. I wasn't struggling with the linear algebra or with python. I already knew pytorch and neural networks. That helped a lot and I covered these tutorials fast and could skip over large sections. A few evenings and the odd weekend day over a couple of months was enough for me.
For seventeen year olds the tutorials are good enough to make it possible to learn this but it would have taken a lot longer to understand. On the other hand I would have learned a lot more. I think I would have learned a lot of valuable stuff.
However I also think 17 year old me was studying for his A levels and probably this was right choice in terms of maximising future opportunities. I'm not sure I think learning about LLMs instead is sensible. Indeed it might be bad advice. But I can absolutely agree with the sentiment.I think 17 year old me would have wanted to do this too.
http://languagemodelbuilder.com teaches you (in a few hours to days) how to build an LLM from scratch. It's entirely free, without accounts, and without data collection.
Thanks a ton for building this.
Putting this in contrast with programming, I learned coding when I was 8, and it was incredibly stimulating to learn because you can quickly iterate and there were thousands of books and YouTube tutorials that dumb everything down and teach you fundamentals. All you needed was a $300 computer, and you can learn nearly anything you want, without being gatekept from this or that because you don't have enough vRAM / an sm_100 GPU.
There are other types of models like diffusion models right now that are showing more efficiency and have a higher ceiling for improvement. Understanding math and fundamentals are more important.
But of course, 10 years ago this wasn't obvious.
That said, I a good starting point for a 17yo is reading about perceptrons[0], then the basics of neural networks[1] (ex. 3-layer perceptron) then writing a program to train a 3-layer perceptron and classifying the MNIST dataset[2] - a dataset of characters.
This can anywhere between a day and a week and you will demystify the basics of neural networks and work your way forward with more advanced contemporary concepts.
Fun fact: any multi-layer perceptron neural net can be reduced to a 3-layer perceptron network.
[0] https://en.wikipedia.org/wiki/Perceptron [1] http://geeksforgeeks.org/deep-learning/neural-networks-a-beg... [2] https://www.kaggle.com/datasets/hojjatk/mnist-dataset/data
I’d probably say something like: do something you enjoy and seems like it might be useful, but accept that the pace of change may mean that whatever you study ends up being irrelevant.
Whatever solution there ends up being to this, it’s not going to be one that an individual 17 year old can implement. We’re past the point where individual good and bad choices matter that much to economic outcomes.
Telling other people's children what to do is easy and basically doesn't have any downside to being wrong. With your own children things are a bit different.
So: what are people here with school age children telling their own kids about the future? If their kids ask, what kind of careers would they encourage them to pursue, assuming they have the skills and interest?
When I was last in the Bay Area, maybe about a decade ago the bookshops were full of titles like "Python for Preschoolers" (I exaggerate, but only slightly). Clearly at the time a lot of people working in tech thought that cultivating an interest in programming was going to be the path to being a successful (by some metric) adult. Is that still the case?
In this regard, Healthcare is a polar opposite. It's pretty hard as a male nurse to not "accidentally" become a home wrecker.
Being a super rich and an unhappy workaholic, or a super-impressive engineer who wakes up one day at 45 and realizes they regret wasting half their life (I ran into way too many of these) is a much worse fate than "not being rich from your startup" and working a relatively regular job while feeling fulfilled and happy by more than just work.
Especially in the US, which is uniquely bad at this and encourages people to work themselves to death, mental health and work life balance are much more valuable things for 17 year olds to focus on than finding good startup ideas.
In case you think i'm being a bit dramatic, let's look at the state of 17 year old mental health in the heart of Silicon Valley:
"The City of Palo Alto and the Palo Alto Unified School District approved a funded contract to place 24/7 human security guards and monitors at all four local Caltrain grade crossings, including the Churchill Avenue crossing directly adjacent to Palo Alto High School."
(in case it's not obvious, it's because of suicides by high school students)
The 17 year olds do not need advice on better startups, and this situation will never get better if we focus our advice on how to be better at work instead of how to be better at life. This will require redirecting the conversations.
I'd move to the middle of nowhere and work multiple jobs on a farm and in construction. Learn how to grow food, and build things. Meet the farmer's daughter, and marry her. Then, buy my own land, grow my own food, and build my own things.
I started investing in farms, have 50 pigs and 100+ chickens now. We are planning to grow to 100 pigs and 2000 chickens in a year. We will start growing Shiitake mushrooms in a few months too.
Nothing wrong with learning the theory and understanding the papers. Getting to that point you’ll have to get your fundamentals down. Might be an interesting exercise.
But as a future? I guess we’ll see. I suspect the next financial apocalypse will determine if there is one. Another AI Winter that may outlast all others so far.
The few exceptional individuals will innovate what the millions will benefit from.
At 17, the mother of Isaac Newton removed Isaac from school and tried to make him a farmer. We all know that wasn't his destiny.
You could also walk out your front door and get hit be a meteor tomorrow.
On an Amiga, I took various public domain text documents from cover disks and counted the probability of the next word given the previous word. Then spat out random sequences of words from it and printed them out. It was called "Splurge". Basically a very very simple single layer statistical language model.
Some of the sentences were randomly not bad sentences, which seemed amazing at the time!
That kind of thing (and Core Wars and Tierra etc) did lead me to getting a job at an artificial life startup at the end of the decade. But that was in turn about 10/15 years too early (no GPUs).
There's some lesson from this about timing, but honestly I've gained the most as a person when I did something that was fun, ethical and gained an audience. A tricky combination.
None of us know what the future of work, education, or AI is going to look like. But your best bet is to become a life-long learner. Be it LLMs, musical instruments, physics, or business.
Learning AI isn't like learning HTML in the 90s then expecting to get a job at a tech company building websites. You can't just "learn how to build LLMs" and expect a frontier lab to hire you so I'd argue this is rather bad advise.
Additionally, unlike web development in the 90s you cant really do anything interesting yourself... All of the interesting/useful stuff will require huge amounts of compute and data so there isn't even much point in learning to start your own thing either.
As someone whose built many of NNs from scratch (hand written code, long before the days of LLMs), it's more or less useless knowledge if I wanted to work in a frontier lab or do anything interesting in the field.
I also think anyone thinking about going into a field which is basically a crossover of CompSci and Maths is absolutely insane right now. Even if you think there is a place for CompSci and Maths post LLMs, there's almost no chance anything you learn today will be relevant to the skills required in say 5-10 years.
(Edit: And learn how honest business works)
I think the last one was seeing a skilled electronics repairman do surgery on a CT machine controller.
Moreover I am not sure it is even good advice? Would you advise a 17 y.o. to learn how transistors work or how to code (i.e. is LLM training the right level in the stack)? LLM training, a discipline where relevant work is already out of reach for 99.999% of budgets really as essential as this post implies?
>"Someone asked what I'd do if I were 17. I'd learn how to build LLMs from scratch..."
That's funny, Paul Graham, because if I were 17 again,
I'd learn how to program in LISP.
https://www.paulgraham.com/rootsoflisp.html
https://www.paulgraham.com/iflisp.html
https://www.paulgraham.com/hundred.html
(And/or other LISP derived languages... Clojure, Scheme, Racket, TinyScheme, etc.)
I guess "the grass is always greener..." as that old expression, that old "chestnut", goes... :-)
I first studied them in 2011, and I was like what? Just a bunch of partial derivatives?
I keep looking at AI to check if now it's something else but it keeps being gradient descent.
Okay, it's great that you can perform miracles using gradient descent but that doesn't make it captivating in any way.
I'm much older and less wise now, but I still afforded myself the opportunity to follow karpathy's tutorials to build a LLM from scratch. Got to play with a few ideas. Seen similar ideas turn up in frontier model work, which is quite gratifying.
There are so many ideas to try.
Currently playing with autoencoders that takes A and B and produce latents A', B', and C'. Reconstruction of A is from A' and C', B is from B' and C'
The idea is if C' can be made to improve both outputs, it must store as much information as it can about what is common to both inputs.
The abilities of LLMs are emergent. You can experiment with LLMs and know as much about their observable behavior as the experts. But there’s no way to “crack open” an LLM and see precisely where each skill or tendency lives; as far as we currently understand, it’s all mashed together.
Lot's of neat stuff to learn.
Most software engineers do not have a deep understanding of CPU architectures. In fact they probably don’t even have a shallow understanding and get around just fine. How many of them are looking up the instruction set for the CPUs they deploy their CRUD app to in EC2?
But the load-bearing assumption is that intuition has to be human-shaped intuition. Humans can't intuit thousands of orthogonal directions because we project everything down into a 3D metaphor and hope it holds. That's a fact about our hardware, not about the systems.
And the reason why is the most interesting part: nothing requires the compression step. A model or an agent can operate over the actual objects, holding thousands of runs and ablations in context and noticing regularities in the native dimensionality, without translating them into a picture of a ball rolling down a hill. No bottleneck at "can you visualize it."
So the narrower claim: it's not that intuition here is impossible full-stop, it's that human intuition is unreliable. Your post-hoc rationalization point is evidence for that, not against it. The story exists because a person needs something to hold in their head. Drop that requirement and the failure mode goes with it.
Well it is framed as quite specific advice.
(I'm done with mining PG tweets for meaning)
IMO this is still relevant, everything surrounding the LLMs needs such a vast infrastructure that I don't know if I would find it more useful to learn the maths behind ML than CS
Indeed, if credulous folks look to the world expecting people to bestow success upon them... than the disillusionment with reality will hit their savings harder.
The Shrek movie market correction correlations are undeniably funny, and a new film is due July 2027. OpenAI may be going public in the next few months while still losing $2.25 for every $1 of customer revenue, and with 6 other firms sharing over $4Tn in debt disclosed to investors in a footnote.
There is only one direction things can go at the Peak of inflated expectations. Popcorn ready. =3
Some people will be delighted. Most are going to have a very hard time adjusting.
YC and adjacent isn't it.
Sure, but it was for a specific degree with a syllabus that taught you the foundational knowledge. It was not expected from the law students to learn how to build one.
It would be a good idea for _everyone in the industry_ to deeply know how LLM training, inference and "agents" work, not least because it removes the ability of shysters to bamboozle with bullshit.
But, as much as a good idea it is for the young to understand this, it's the elderly who will be really taken advantage of if they do not keep up - just look at Facebook for good examples of why.
I would suggest starting with Andrej Karpathy's YouTube video: https://youtu.be/kCc8FmEb1nY?is=oiDsrBYJg_MUUmoD
This video is excellent. I'm a huge fan. Also the video is zero commitment and instantly available which makes it a good way to check you are interested.
The book by Sebastian Raschka is slightly less accessible but very reasonably priced and the experience of working through a book is a lot nicer than skipping back and forth in a video (for me). Sebastian's blog posts on recent architectures are absolutely great too.
you can now play the inventor and start with question "i want to build next token prediction software" and go as far as you can with your current knowledge while brainstroming with chatgpt as a rubber duck.
dont not start with a course on probabality , linear algebra or calculus . do not watch 3 hr videos or 3blue animations .
there is agood video on this way to learn.
Watching Karpathy's "Zero to Hero" won't make you AI specialist. It will improve your understanding of the basics.
Or more reasonably, the same old thing : use Linux, hack a little, why not learn programming basics. But learn to own your technology, fight against centralization of technology. The same old RMS story.
IDK where the tech industry is going, if there will be jobs anymore or not, but what I'm sure (and what have been the case for the last 10-15 years anyway) is that for most tech jobs, having good technical knowledge beyond the basics is pretty useless and will probably not be recognized.
If you can, stay a computer geek if that's your thing, but don't make it your career choice, the Eldorado is behind us.
What? You want me to tell kids not to be ambitious?
Personally I don't see the problem, as long as you're aware there is survivorship bias involved here.
What's the alternative really, seek advice from unsuccessful people? That seems worse :)
Personally I do both, read about what worked for people, also read about what didn't work for people, then ignore both and do whatever the fuck I want.
Seek advice from the averagely successful people, since that is statistically what you're most likely to be.
Intuitively, I would guess that they have a better grasp of what made them fail than successful people have of what made them succeed.
I know who the 17 years old is closest to.
Not sure why would you think so.
Inverse reasoning is very powerful, and unsuccessful people can give you plenty of "don't do this mistake", which the survivors would not even think about.
Do both. Get advice from successful and unsuccessful people and take the diff
Pretty tall ask, especially if this is targeted at 17 year olds...
Learning from their mistakes (where it did go wrong) can be as valuable as.
If you can't do that - not sure anyone can help you.
That’s what you get from listening to “successful people”. You get to learn about all the things they tried that failed, then the things that did work on that 24th try, which was successful.
The “survivorship bias” people always seem to assume that the “survivor” lucked into his fortune on his first try ever, so he can’t have learned anything, so we don’t have to listen to him. But that’s seldom the case.
I’ve written about this before:
In 2000 (his era), it would have been really smart to study the source of Linux or Apache. Would have paid dividends over decades. Cuz that knowledge was so rare. The number of people hacking on LLMs now dwarfs the number of people hacking on web servers 30 years ago, by several orders of magnitude.
And if you turn back the clock even more, I mean just even having access to a computer, let alone owning one, would have put you at a massive advantage.
I don't know what to call it. The pioneers should be respected obviously, but at the same time you need to understand that for them, the game wasn't nearly as played out as it is now.
I just don't think you can afford to be dicking around with LLMs like you could afford to dick around with random Linux distros 20 years ago. Too many people willing to do it for free these days.
You don't wanna end up being the 2030 equivalent of a certain SNES emulator developer, or maintainer of a package manager for jailbroken iPhones, I mean the list goes on and on. Being a hacker doesn't automatically give you a path to being rich, or even making a decent living. It hasn't been that way for a while.
You're right that it's probably a little too in the weeds, but it's also a nice clear and fun objective that teaches you the basics. Like building a TODO list in JavaScript to learn webdev or a Gameboy emulator in C++ to learn how a CPU works.
I think it is. He isn't saying to learn how to train a LLM so that you can go on to train LLMs. He's saying to learn it so that you gain a deep understanding of how LLMs work. Ordinary startups can still benefit from things like training or fine tuning highly specialised smaller models, knowing how to select and configure an appropriate model for the task at hand, knowing what software to use and why, understanding what's going on behind the scenes instead of treating everything like a black box, having a higher level of intuition about LLMs generally, etc.
Most computer science courses do in fact teach things which are lower level than coding, such as how transistors work.
how many of us out here are doing work directly in what we got a degree in? I majored in economics and now I'm a CTO.
I would absolutely advise a 17 yo to learn how to code, understand how transitors work and how to code an llm. even if he never works on llms, you basically end up with a kid with applied knowlege of statistics, math, physics hardware, logic and a whole lot of practice in critical thinking.
Since he has not given any reference point it is equally as good advice as „just learn everything slightly valuable“.
So the real question is: „What should I not learn in favor of learning this.“
Or in other words: His comment is not usable because it can mean anything or nothing.
Hell yeah. Transistors are pretty awesome.
But it feels like it sort of backs up my point about there being good models at every size class. Fine tuning Needle looks automated. Yes, you need to know basics like what validation loss means and how to use Python, but otherwise it's all about creating the dataset.
Why the everliving fuck has everyone stopped building tools?
Vacuum tube computer: https://www.youtube.com/watch?v=KAlnJnYt5do
Making vacuum tubes in a garage: https://www.youtube.com/watch?v=-UEfqAWb3fE
(As a TML person, I'm obviously biased, but I couldn't resist because of "tinkering").
TBF it's hard to imagine a real architecture change that wouldn't require a ton of compute, but you could certainly fine tune and play with different recipes, loss functions, etc. And Claude can carry you a lot of the way through doing this.
One fun task is to invent a tool and then train a small model to use it. You could export that small model and run it locally for free forever to do your thing. I think this is what a lot of Software Engineering will look like later.
There are a lot of other high level abstractions here to look at. Prime Intellect has one.
The other thing to play with is self-hosting small models, but IMO most of the interesting stuff is actually related to multi-gpu or multi-node inference so there's not necessarily a ton to learn here.
https://ravinkumar.com/GenAiGuidebook/book_intro.html
This guidebook covers pretraining, post training (SFT, RL) and a couple other topics. And others authors have also written books that fit on single node reasonable hardware.
If you want to start with a pretrained base I built Gemma 270m and released it last year. This fits on a raspberry pi.
https://developers.googleblog.com/en/introducing-gemma-3-270...
The fundamentals of AI don't require industrial amounts of large scale. Think of it like this, when I was learning how a plane worked when I was a kid I didn't build a 747 at home, I started with scale sized model planes. Same idea here.
And FWIW I'm a staff researcher at Deepmind (opinions here are my own) so I want to specifically encourage all people out there, you can learn a lot about how these LLMs work at home, for (mostly free), using resources like colabs or spot pricing on accelerator providers. There's many great resources out there and I encourage anyone willing to learn to go for it!
The best analogy is strip mining (big labs) vs cave exploration (solitary/small teams). I think this is how science progresses at the boundaries by smart/curious/hardworking individuals because depth is a requisite for finding the right questions and then the answer. It is not for everyone and it does not always work. But you learn a ton even if it doesn't pan out to be a big breakthrough.
So, you build one from scratch.
I fail to see how PG suggesting that a 17 year old, today, should study something... is in any way a parallel to formalizing it as part of the curriculum.
I would, in fact, expect PG to recommend something completely different in 10 years.
The whole point of hacking is to understand the world around you, today
What I meant was that if you're rich you don't need any trade school or to run a business, you can just sit on your ass living off interest from capital. Only "poor" people need to work for a living.
Might be that these people are from outside the US as well, where things like "honest business" is very much possible today, probably most businesses I interact with AFK on a daily business are "honest businesses".
Mechanistic Interpretability has entered the chat.
For a classic example, see https://www.anthropic.com/research/tracing-thoughts-language...
The spirit of your point stands, though. This kind of research is interesting to read about, but it's very hard, more like neuroscience or biology than computer science ("LLMs are grown, not made"). You're dealing with a lot of extremely _messy_ complexity, for which organic life is really the only good point of comparison. Most of us here are't really equipped for that kind of work; it's not at all like, say, reverse-engineering a piece of software written by humans. And of course the only people who can do it on frontier models from Anthropic and OpenAI are people within the labs themselves. (But I'm optimistic we'll see more of this work on open weights models...)
If stochastic gradient descent isn't taught in whatever CS Theory 101 is now, it really should be these days.
> why not learn programming basics
Do you think implementing an LLM from scratch would not require programming!?
A generic project based goal like this is the best approach to self learning, because you fight your way to it, picking up what you need to know on the way. And, once it's all working, and you get your first meaningful sign of success (in this case, some tokens out), after your long path, you fucking celebrate!
Absolutely none of that mattered in the long run because capital markets think environmental destruction is a small price to pay as long as some people get rich.
What kids learn about personal cooling tech and what not won't matter when the supply of the raw materials required to make them is reserved for some VC-funded trillion-dollar 'startup' trying to synthesize an undetectable chemical weapon to get past the bureaucratic red tape of international treaties.
There are successful people that failed many times before being successful. Somehow these days, previous failures and persistence seems to be ignored and they focus just on the luck you got on the 20th try.
I guess it depends on what submission you look at, previous comment of mine solely based on memory. Now I went to https://news.ycombinator.com/from?site=twitter.com/paulg, clicked "More" a bunch of times, and seems my memory was more or less correct, none of the submissions I clicked on are "pg worship" (hyperbolic or not). Just one example: https://news.ycombinator.com/item?id=19418701
Maybe you need to enable "Show Dead" or something? Pgs articles on HN definitely never was free of any critique in the HN comments, just like any article. Although I do agree with you that it used to be different than it is today, and same with Musk too, and Altman, and probably more individuals, where they were lauded before but now pretty much just mentioning them poisons the conversation.
That post has barely any points and comments. It’s not a good indicator of general sentiment, it’s just an indicator of people who were on HN at that time.
> Maybe you need to enable "Show Dead" or something?
I have it enabled.
> Pgs articles on HN definitely never was free of any critique in the HN comments, just like any article.
Of course. I very explicitly wrote “in general (certainly more than any other website)”. That does not mean “always”, or “never”. HN is not a hive mind, there’s never going to be 100% agreement. The general trend is what’s being discussed, and we both agree that in general the sentiment on Graham used to be higher. We’re just disagreeing (we may be able to find ourselves agreeing through tough thorough thought, though[1]) on where exactly it is now and how to interpret it.
[1]: Sorry, can’t believe there was a real organic opportunity to use that sentence, had to take it.
From a biology perspective - we have a good understanding of the carbon atom and how forces influence it. We really don't know much about a cell.
That said, I think traditional resources are a better way to get breadth and to frame the topic. Just chatting about something can be a little disorienting imo.
But how can I know for sure that that particular mistake is actually why they were unsuccessful? Has exactly the same issue as listening only to successful people as they hardly know what actually made them successful most of the time, but they still compose large blog posts with their reasoning for why.
Again, I still think my approach of reading both but then regardless go my own way is the preferable approach, at least for me, ymmv.
What bucket should I put this advice in?
But some advice is less specific than other advice though. Some things are always stupid, and some things are always smart, if you look at the context of our world and society. I find myself pursuing these "fundamental truths" with great interest lately, especially now that the world is changing so quickly.
Agreed, my previous stated "ignore both and do whatever the fuck I want" approach has worked out very well for me in life, people should probably focus on identifying better what their gut tells them, rather than what randoms on the internet thinks and writes.
Building a browser in the early web was actually a very achievable goal for exactly the reasons it isn’t now. There was not JS. No CSS. No SVG. In fact very few widely supported image formats (and graphical browsers weren’t around in the earliest days of the web anyway). TLS didn’t exist. HTML only had a subset of methods. And even POST was usually just managed by CGI/BIN calling an external process, often written in C++ or Perl.
It was a simpler time.
That’s like saying someone who fabricates cars doesn’t have the skills to drive them. Perhaps not, but they’re very well placed to pick it up quickly. They’ve also shown they can do something far more challenging, which is actually better than hiring for narrow immediate skills.
In other words: I’d hire that candidate in a heartbeat.
If anyone reading this has this un-marketable skill, Carnegie robotics in Pittsburgh is hiring
carnegie-robotics.breezy.hr/p/2d85f5321cc7-software-engineer
[Blake Ross] worked as an intern at Netscape at the age of 16 ... Ross became disenchanted with the browser he was working on and the direction given to it by America Online, which had recently purchased Netscape. Ross and Hyatt envisioned a smaller, easy-to-use browser that could have mass appeal, and Firefox was born from that idea ... in 2003 all of Mozilla's resources were devoted to the Firefox and Thunderbird projects. Released in November 2004, when Ross was 19, Firefox quickly grabbed market share ... with 100 million downloads in less than a year
https://en.wikipedia.org/wiki/Blake_RossNot a lot of demand, but also probably not a lot of supply.
You don't need your impressive product to succeed to land a good career.
If Alice is doing LLM-from-scratch work today, and Bob is doing agent harness work today, Bob's project is far more likely than Alice's to become useful/popular/profitable.
But if neither project survives, in 5 years, Alice will be more employable/at a higher market rate than Bob.
In general C++ work and similar, if not in Chrome development.
Oh hey thank you for that. It really helps. Hope you have your rug pulled from under you today too.
— signed, a career changer trying his best.
Good luck with that approach when trying to toy around with models and their training/inference.
There is also a lot of math basics missing that a 12 year old may be able to grasp, but I would bet they are at least 13 by the time the knowledge is deep enough to understand what operations are happening.
17 is an interesting age. There are way too many comments here saying things like, 17 year olds should just do whatever seems interesting or bum around the world or focus on getting into university. But historically most kids were expected to be productive adults at 16 or 18. 17 is about the right time to be thinking seriously about what kind of work you'll do, how you'll make a living. University won't help and will just delay this decision.
I contributed to a browser engine around that age (KHTML, which later became WebKit and Blink), and while I don't work in browsers right now, much of that knowledge, mindset and of course the professional network have done much to shape my life. And a fairly successful career, for that matter.
A lot of people will make better decisions with a few yeas more maturity, and spending a few years developing themselves.
University will help a lot of people, and for some it will help.
There is a lot more to life than making a living.
The early Internet (which the GP mentioned) didn’t even have the web. But that’s nitpicking.
Paradoxically I coded way more between ages 12-14, I regret my wasted late teens.
Which is paulg's point really
Jeff Bezos did similar, but he did his own fulfillment and hired out the coding.
Basically every part of the original transformer was replaced with something more efficient or better:
LayerNorm -> RMSNorm
Sinusoidal position encoding -> RoPE
MHA -> GQA
ReLU -> GELU
What this means is that there is ample opportunity to improve on what we’ve done thus far.
Yes you might progress the field, but will you really understand why? You can make up an explanation and anthropomorphise it with a few contrived diagrams and everyone will cheer!
https://www.youtube.com/watch?v=3LopI4YeC4I
An advantage that is not "advisable", like being born in january, in a rich country, in an above average family, or just having luck, might have more influence on the outcome than any conscious action. It is almost sure that one-in-a-million level people only edge over the other 999,999 they competed with is just "have more luck".
Of the 10 most talented and hardworking people above, I bet they have wildy different advise other than "do hardwork": use more LLM, use less LLM, keep away from computers, etc. And we are talking 1 in 100, not 1 in 1,000,000!
It comes to mind someone like Kary Mullis. He was no dumb, but I wont put him either at the peak of talented or even hard working people in the field. Yet, he was "lucky" enough to invent the PCR, which earned him deservedly the Nobel Price. Now that he was made to believe that he is some kind of illuminated and well above the top human, he started giving (at the bare minimum wrong, probably even dangerous) "advise" left, right and center: take LSD, astrology works, extraterrestial glowing racoons exist, etc. Downtune the guy a notch or two, and you can find thousands of Kary Mullis around giving useless advise as they were messiahs.
Now, for most software careers (including my own), that knowledge isn't directly practical. BUT, the most important things for a successful, fulfilling career are curiosity and a willingness to dive into subjects _without_ necessarily knowing how (or even if) that knowledge will yield practical applications. That drive is what's going to lead to insights and breakthroughs over your career that wouldn't happen if you just let abstractions be abstractions.
I’d find some high level tutorials online that maybe use one of the free circuit simulator tools.
falstad.com/circuit/
Nice animation on this site
If you are really curious there are many simple circuits you can build.
It’s quite a thrill to get an LED to blink with a self built circuit.
I’m also a crusty old sw eng. I did spend sometime being a young electronics engineer.
Good luck with the journey.
Hard agree! Though the discussion was short, I thank you for it. Good start of the week, I wish you a good one.
Browsers from scratch are multi-year projects for multiple people. Even just skinning and minor tweaks to modern browsers is a deep well for one person.
FWIW I built browsers from 1999 for a long time. (But I was never a wunderkind, just somewhat tenaciously curious).
And I guess I am doing fine, but not amazingly rich or so.
Browsers were always a project closer to research/charity. I think Marc Andreessen said something similar - that he would never do that again. B2B is where you can make money.
For me, I know that one major reason for putting an upper limit on my success is my inability to form effective professional relationships. I know I should go to events and talk to people and use these relationships to my advantage, but that's just not something I've ever been good at, and I find it so incredibly unpleasant that I also just don't want to do it.
Do you think the most successful people got that way by attending events? I thought they were hacking in their garages, building great products.
Nonetheless, there are many successful people I would gladly listen to for advice, though they are often successful in a different meaning than what venture capitalists would use (e.g. parents with great kids, managing to keep a healthy work-life balance, happiness, and maybe even having time to spend on some cool hobby project -- you are heros!)
Of course, there are degrees and exceptions on every side, but the coping mechanism is very strong among people.
ages 8-15: lots of great 8-bit computer fun.
ages 15-20: girls, booze, motorcycles.
20 onwards: get a PC, back to computers, realize how much I've been missing.
To be honest, judging by my own kids and their friends, late teens seem to generally be an era of hard to avoid stupidity.
If teenagers could make small contributions to LLMs via open source then sure, go for it. Optimizing llama.cpp or similar would be a good learning project that might later get you good work via social networks. Contributing to open source is how I got started too.
Unfortunately, training LLMs isn't something that fits well to open source open collaboration. Inferencing codebases are better.
In that sense it's more a "seek out the open source community and real projects when young" rather than "do web browsers", with a bit of "look for ambitious types of projects few get to work on".
I really hope we get another hiring boom like in 2020 when I decided to study CS. Otherwise my career will be very rough. I love it and can't imagine doing anything else.
You should really read up on disruptive innovation. Many of those who you call "serfs" who'll "take 1/3rd of my pay" will tomorrow establish themselves, justify their presence, and move up the wage/income scale, while you'll be left complaining.
I am from and live in a third-world country and have spent some time in a first-world one. I've seen the ETH Zurich tag open doors that I didn't even know existed. If you aren't leveraging your passport, your education and your network to reach a place where you don't need to worry about "serfs", then you're doing something seriously wrong. What you have, compared to the competition you denigrate, is invaluable, making best use of it is your job.
who would have guessed when first world were enjoying their heyday as they were colonizing the world, and teaching everyone and their mother to learn English.
Kinda sad now you have to compete with serfs from thirls world. chu chu chu. so cute.
shouldn’t have robbed others then. a richer world for you comes at poorer world for someone. we are playing infinite game in finite world.
Young people are not responsible for the crimes of their ancestors centuries ago.
Bad advice. You’re going to need family connections and loyal people you can bet your life on to survive if the rest of what you’re saying is remotely true.
We never imagined that we'd live to see Snow Crash depict a utopia by comparison to what we got.
People keep talking about a browser in the modern context but the GP specifically said “early days of the internet” (which, in fairness, would mean pre-web. But I think it’s safe to assume they meant “web” not Internet).
In those days, it was actually a much simpler exercise to write a browser than it is today. I even wrote one! And writing a browser absolutely teaches you how HTTP and HTML worked. Plus a lot of backed development was forms data sent to CGI and thus written in languages we wouldn’t even dream of using for web development nowadays, including C++.
So in the early web, writing a browser absolutely was a transferable skill. It might not be now, but in the context defined by the GP, it was.
I recently switched roles, and among the seven places I interviewed, none of them seemed to see my then-current browser job as a problem, even though they were not related to browsers. (The closest one was a company implementing a HTTP reverse proxy, and I did not work on the browser's HTTP stack.)
I know there were opportunities I could have benefited from if I'd had better relationships with specific people.
> Do you think the most successful people got that way by attending events?
That's not what I said.
As you say, there's a lot of value to unlock with understanding the generic principles, and I would add specific application.
Lots of people are building fantastic Tools or pulling down million dollar salaries without groking the precise representation of a single weight.
I'll borrow my response from a far greater applied mathematician than myself: https://www.youtube.com/watch?v=_oNgyUAEv0Q
Now with llms we need way more that 6 dimensions so we can start thinking of assigning matrices to each 3d point for example. That allows us to increase the dimension from 3 to 3 + whatever the matrix dimension is.
We can visualize the matrices instead of having numbers as having colors for each entry, so they can be a sort of cube with each vowel being a different color.
Now you can start to visually intuit about how these massively high dimensional spaces can be formed of these colored matrices that can react to some input training data.
That's a start of an idea for intuiting things that might seem impossible to have an intuition about. I think visualization is a great way to start.
visualizign 100 dim matrix will not tell you how llm work. so what you even talking about.
Linking to a Twitter thread about training a small LLM to be good at Wordle is not an example of what I'm talking about. It might well be a useful task but it doesn't allow us to understand deeply what's happening.
Like a cat who knows how to teleport so rapidly around the room that he becames a blur. All the while taking into account stuff falling around he knocked down.
The reason remote worker can out compete you is because of this currency arbitrage you designed. Your pennies goes a long way in this so called third world. Your one day dinner at nice place is a months salary for family of 4.
Seriously, if you can’t compete globally, then you don’t deserve your privilege of being born in rich world.
Stop labelling other humans as serf when your entire wealth was built on shaky foundations.
You == not you the person, but the government and laws.
The fix would be obviously to control and limit foreign workers by regulating companies.
Depending on your knowledge of math, I recommend starting with linear algebra, building an understanding of the equations and try to visualize more and more complex systems, then study llms to see how you can apply your linear algebra intuition to your understanding of llms.
VTK is a great toolkit for visualizing complex systems. 3 blue one brown on YouTube has other visuals that might help you.
It takes time but it's possible. Good luck.
> Visualizing large dimensional matrices is a way of starting to develop intuition about llms.
one last time before i disengage. how do you know this and what intuitions have you personally developed.
Is this the guy you’re talking about?
> I’d imagine those teenagers ended up with pretty good careers in technology?
The guy you mentioned doesn't seem to have ended up with a bad career in technology?