Old and new apps, via modern coding agents(terrytao.wordpress.com) |
Old and new apps, via modern coding agents(terrytao.wordpress.com) |
https://www.reddit.com/r/mathematics/comments/1tryyw7/terenc...
Every time.
The difference to me is one of directionality - maths research is seeing a far off island and getting there by hook or by crook; bridge, draining the swamp, inventing an airplane or boat, whatever it takes. Software engineering is like covering a plain with tiles - every feature is ultimately filled in and the underlying beauty is obscured by a fractal of complexity required by the ever growing requirements.
Ookay.
Back in the day i was confused by 'Linear programming', which is optimisation and has nothing to do with coding.
> every feature is ultimately filled in and the underlying beauty is obscured by a fractal of complexity required by the ever growing requirements.
Right. I would say Mathematics tries to unobscure (patterns in) nature. Engineering is creating tools, sometimes leveraging natural patterns. But yeah, Fourier or Laplace definitely created tools, too.
Yes, functional programming feels a lot more mathy than procedural or OOP.
https://htmx.org/essays/universities-and-ai/#demos-visualiza...
Many visualizations that I have always wanted but just didn't have the time to build, I now have.
To give an example, I wanted a simplified 8-bit computer to complement the 16-bit teaching computer I use and designed this in a few days with the help of claude:
It helps me digest the content faster and allows me to read more articles than I otherwise would.
When I did my microcontroller class with lecturer hand drawing an 8-bit computer, the registers, memory, instructions on the white board, it was v cool to understand how things worked under the hood.
Wondered if someone could make more simulations for what was being taught. Teaching is about deciphering a thing into it's components and seeing how they interact. Vibe coded simulations are a great tool for that.
Sounds like 50/50 for the distribution? That means you are okay with a student getting a 40% across all your quizzes and then passing the class with a C-?
My experience is that no students get a C- except for students who blow the first part of the class and try to work back. I usually work a deal out with them anyway.
And then realizing they put together something that would have taken you a few days to do.
The supply of software is about to go way up, and that's going to massively impact demand unless every firm on earth is clamoring for more.
We're going to see if Jevons paradox holds true, or if wages get impacted drastically.
"as such [LLM-coded interactive] supplements are not mission-critical to the core of the paper, I again feel that the downside risk of using guided interaction with LLM agents to generate such visualizations is acceptable."
It's a tool. Good for some things but not others and generally not to be trusted.
I am not sure how to feel about agents solving the problem via proper modernization. It's certainly positive that students will be able to interact with this content in a modern and more accessible way, but the educational use case for our product, although not commercially important, has always been a source of pride.
https://chromewebstore.google.com/detail/cheerpj-applet-runn...
https://github.com/bradfitz/koffer#der-verloren-koffe
Play online at https://bradfitz.github.io/koffer/js/
So neat seeing ~30 year old code come back alive.
I have been interested in machine-assisted ways to do and teach mathematics from as far back as 1999, when I started coding several applets in Java 1.0, both for my complex analysis and linear algebra courses, to visualize various mathematical objects I was interested in (such as honeycombs or Besicovitch sets).
Really bullish on LLMs expanding code development by a very large group of people who are really smart in some domain but could not get into 'coding'.
As for profit, there's a reason why governments and AI companies are hiring philosophers and mathematicians. It's not to make the world a better place for everyone, or to encourage the progress of human knowledge; but to gain cutting-edge advantages over their competitors. Same reason why theoretical physicists were prized before/during the Second World War.
By famous I mean someone whose biography is in the training data. All models know a lot more about Terrance Tao than they know about me, when he's working on his projects do the models know they don't need to explain "Besicovitch sets".
Since the system prompt likely includes something about not insulting the user, does the LLM modify it's responses if it realizes it's talking to famous politician, like "dont mention the time $politician was cancelled".
I agree completely you always need to check the work of LLM agents, but it does strike me as a tiny bit funny to anthropomorphize AI by using ‘trust’ while warning against anthropomorphizing the AI by using unchecked output. ;) Generally speaking, “trust” in AI has been going up very quickly as the models & harnesses improve, and as people figure out effective workflows.
I trust my hammer with nails but not screws… does that mean the hammer should generally not be trusted? The problem with AI is we don’t know the difference between nails and screws. (This may be where my analogy breaks down. :P) But I feel like saying don’t trust it isn’t as helpful as saying something like you should expect to spend more time planning and iterating than before, and you should expect tot spend more time reviewing and checking output than before, and learn how to use skills and context and subagents, and learn to use AI on some non-production low-consequence projects first. Saying ‘generally not to be trusted’ implicitly suggests not using AI, and doesn’t leave the reader with how to use AI. The goal is to build trust by building good workflows and by understanding what works well and what doesn’t, right?
I trust a hammer to be able to hit a nail, without breaking. But if the hammer is old and the wood brittle, I don't trust it anymore.
Using it for anything else (screws) has nothing to do with trust, but using the wrong tool.
- nice to have but clearly not important enough to invest time in
- definitely not worthwhile paying someone to do
- not going to hurt anything or anyone if it wasn’t implemented as close to what was envisioned.
If you read between the lines then trust may also mean privacy, if you’re furthering the goal of some company that may be stealing your data for training even when they say they are not because there are legal loopholes that allows them to get away with it, etc.
Your example of hiring Donald Knuth to write your code doesn’t fit into what’s being said about trust either. If you were never going to trust anyone to write your code, it doesn’t matter who it is anyway.
For most of us, even if we are engineers, chances are hiring someone legendarily good at writing code to requirement will produce far better results than what we knew we could achieve.
I would trust someone who is that good to write the code — more than myself — to do things I want to do it better than I can while being able to catch all the things that didn’t realize I needed.
Donald Knuth's code would be much more likely to meet my standards during a code review than some LLM output.
There are many AI bulls who adamantly disagree and cite Tao’s statements about LLMs for mathematical proofs as an example of how advanced and autonomous these systems already are
> Marco Pierre White passionately defends chefs using microwaves. White dubbed microwaves “sensational things” and revealed he thinks they’re far better at preparing kippers than any other technique, like boiling or grilling
https://www.independent.co.uk/life-style/marco-pierre-white-...
And another one:
> José Andrés, a renowned Michelin-starred chef, New York Times bestselling author and internationally recognized humanitarian. He listed the microwave omelet as his number one foolproof dish and called it the “best fluffy omelet in the history of mankind!”
https://www.tasteofhome.com/article/jose-andres-microwave-om...
What Tao and other artists of his caliber are demonstrating is that the tech is capable of building the rig. And the machine makers are incrementally demonstrating that the machine can make not only the jewelry box rig, but rigs to build rig-making machines.
Are there any documented essays or reactions from the great chefs of back in the day reacting to the first microwave dinners?
Nov 2025: https://terrytao.wordpress.com/tag/artificial-intelligence/
https://academy.openai.com/public/blogs/terence-tao-ai-is-re...
I’m old. If I had to, I could retire tomorrow, albeit on a restricted budget. But I worry about the younger folks (like my 25-year-old nephew) who haven’t built up the resources to survive without working who are in the field right now. There’s going to be a mega disruption and writing code is going to go the way of calculating square roots by hand or hot metal typesetting. There will still people doing it, but it will very much be a niche endeavor.
Whether reviewing agentic code rather than writing it is a job he wants to do... Different question.
Its older people who can't or won't retool that are going to find that in the game of musical layoff chairs they won't have a chair left when the music stops (and some I've met haven't really internalized that they're even part of the game...)
I don’t know what you’re reading, but always and never are strong words. I’ll predict by this time next year you’ll have seen some pretty serious AI uses, and can no longer say always/never. Widespread use of AI coding is brand new, and the models only just barely got good enough to do serious things. It’s way too early to be using words like always and never, but FWIW I’ve already seen some serious uses. There are good reasons personal blog posts rarely talk about ‘serious’ production code; it may be against organizational policy, it may involve code that isn’t’ public, it may reveal proprietary information, and more…
Teaching, research and publication are the core activities of his job as a math professor. How does it get more serious than this?
Replace "vibe coding" with "outsource to India" and it's the same situation we were in 20 years ago: I remember around the time when YouTube had well-established itself (2009-ish?) and so a cottage-industry mushroomed all selling visual-lookalike YouTube clones, but implemented as monolithic single PHP project or WordPress plugin, using the local filesystem for video file storage (or worse: Base64-encoded text in MySQL, lovely).
...to someone with no prior experience or knowledge how highly-scalable, Google-tier web applications (like the real YouTube) work or are built, these "Me-Too-Tube" sites were indistinguishable from the real thing and as far as they were concerned they were happy with it.
Never mind that even if those $25 scripted websites were ever actually scalable, it isn't enough to simply have a video-sharing site; you need to get millions of people to decide to start using it - and that's something which will elude Claude's abilities... possibly forever, I feel.
Mermaid, Graphviz and friends but in HTML pages.
Sometimes it is turning other things into perfetto.dev format for multi-machine tracking (like turn a build process into the same format as Chrome traces).
If you need more flexibility, you end up reaching for p5.js and three js (rather, tell the model to use it).
Once you're touching distance from WebGL, the equivalent of something you make can start looking like something from ciechanow.ski over a single weekend.
What is your thought on D2 compared to those? GIven the WebGL mention (The one I'm familiar with), I suspect this is a matter of interactive vs static/diagrams. I assume the latter due to my own project which falls in that category!
Just to nitpick - will it actually? I don't know what your relationship with Donald Knuth is, but to the extent I'm familiar with him and his approach to work, I would expect that even if I could afford him, I would not be able to get him to agree to accept my coding standards, whereas I found that LLMs (although flawed in many ways) can be cajoled relatively effectively to adopt my particular standards.
⸻
1. A Mac menu-bar app to make it easier to know when I need to either do a PR or look at someone’s feedback on my own PR. Clicking the icon shows a list of all open PRs that I’ve created or that my review has been requested on (and clicking an item there will open the PR so I can act on it). If there’s a PR that I’ve not left comments on or approved, or if there are unreplied/unclosed comments on my PR or it’s been approved, a badge is added to the menu bar icon.
But, no, it's not "any day now." The required size and structure of the ANN is to be determined.
Or, at least, I don't tend to postulate existence of such mechanisms in the brain after hearing about every empirical ANN deficiency.
BTW, I'm not talking about supernatural. The "magic" could be unknown physical processes, unknown quantum algorithms, unknown physics.
/s
Every time someone paid me to write software it was some combination of 1. not that interesting of a problem 2. no real utility i could see or touch, useful in some abstract way of making a number go up 3. involved a constant, painful maintenance burden 4. involved incident management of one kind or another 5. involved a long tail of details with no unifying principle other than a lot of implicit legacy constraints and stakeholders whose involvement waxed and waned with no seeming rhythm..
I'm a big fan of the new capability, it opens up new regimes of performance and correctness and capability for what I can achieve, that in turn grinds me up against math and theory that I had thus far been able to avoid, it's pushing me up the ambition ladder hard and that's a good thing.
But the change is a change in degree not in kind at least in the vibecode regime: it was always relatively fun and relatively easy to do one small program with modest requirements around defect rigor that had a big legible "oh cool!" surface that I didn't have to maintain. Fable doesn't seem any better than Opus at grinding detail work in the bowels of a compiler, but it sure can make an iPhone-scoped platform game with a bunch of bugs in it in a single shot?
If there's a job where you get paid for doing fun, high defect, "oh wow!" factor one-off software that you can immediately disavow any responsibility for? Fuck man, I should have had that job before Fable got that job.
It’s kind of like saying wedding photographers will be impacted because of the people posting on r/iPhoneography. Seems kind of silly doesn’t it.
I don’t know how it is all around the world but where I live, the photography is part of the whole packaged wedding plan and there’s no choice/option BUT to use one, guaranteeing their (expensive) service and job, regardless of what people can produce on their own.
OTOH, it’s probably better to have someone neutral there who can take pics rather than eat/drink/cry and whatever else people do at weddings ;)
The last one I probably took a thousand photos and it took about 10-12 hours to process. Researching, I found out that I could have outsourced the processing for pennies a photo, which is what I'm assuming a lot of wedding photographers do now. But based on what I was doing (delete blurry/bad, color balance, pleasing crop, select the best photos for a book, put the rest on a drive), I suspect I could get that down to a few hours with an AI workflow after I worked out the kinks. I'm not a professional, but I suspect I'd be ahead of the average wedding photographer after I worked that out.
People always argue "well, Slack and Notion have distribution and the product isn't everything." Ok and? The person making it for themselves doesn't necessarily need distribution for it to be valuable. In fact, it's even more attractive that way.
I don't anticipate replacing my payroll company, for instance.
Capitalism will capitalism.
I built an AllTrails equivalent (iOS app and all). It allows me to scan my backpacking gear to load it into a library (with metadata attached), build a weight profile for my trip, has offline maps, and more. I wanted to make something that was more suited to long backpacking trips and will eventually add itinerary planning. I was originally planning on trying to monetize it but I don't know how to sell software and honestly the QA process is so time consuming (hike, fix bugs, hike more, fix those bugs) that I just gave up. I plan on redeploying it to my homelab since I previously ran it on Hetzner.
The idea would be that you would scan all of your gear in and then create a map (or select one) and then Gemini would tell you whether or not you were prepared for that trip - does your sleeping bag insulate well enough, do you have the right shoes, etc. You could even do neat tricks like calculate expected route completing time based on assumed weight. I was also planning a "ask a human" feature that would snapshot all that data and post it to an internal forum.
Ultimately though, outdoor people hate AI (understandably). Selling it as "AI powered" is a no-go.
And I genuinely do not understand how to distribute software. Getting people to use your software is extremely hard unless it is very polished and has no bugs. The feature set was too similar to AllTrails so people will expect AllTrails quality. And something that I'm struggling with for all software that I make is I don't know how to get beta testers to get a better feedback loop. Going "hey my software isn't perfect but can you try it out and let me know what you think" is genuinely met with disdain for bothering the person, apathy, blank stares, or whatever. Doesn't really matter if it's offered for free or not - or even if the person expressed a desire for a solution. A person can say "wow that is a great idea I would love to try that" and I would follow up with them with practically no success.
At the end of the day I think I just want to give up and go back to big tech and optimize ad distribution pipelines to deliver weight loss ads to children or whatever tech is up to these days.