AI is breaking our proxies for expertise(seangoedecke.com) |
AI is breaking our proxies for expertise(seangoedecke.com) |
Likewise, they point to merit (which I think we consider very much like expertise) as one such fundamental concept in our society, which is very much undermined by unchecked use of LLMs. This coming not just via "cheating," but by the way in which we so quickly are willing to claim, and ourselves believe, that we deserve praise for that which the machine has created. On a wide scale, their use will not just compete with those who may not use the machine, but will destroy even our shared idea of personal merit.
Beyond merit alone, AI might lead generally to our "own cultural values becom[ing] not just decadent or debatable but unintelligible." At the end, Berg and Baskin basically say that hope is not sufficient (hope that the old concepts will be replaced with new ones); the proper attitude is to do everything in our power to preserve them in the present.
I prefer a world that's equalocracy, with a focus on specific personal-freedoms (live and let live principle).
These people are trying to be John Henry, but forget about the part of that story where he dies trying to out compete the machine.
imo, The author of this essay does not understand the concrete problem that mathematicians are upset about. There is an idea that math [1] and coding [2] are human activities whose purpose is to achieve a certain kind of insight or mental clarity of things. The simplest description of this is by Feyman [3]. AI generated proofs short-circuit human understanding and therefore goes against the primary purpose. The declaration is calling this out loudly to reiterate that the purpose of the endaevor is not the generation and rewarding of proofs.
[1] "On proof and progress in math" https://arxiv.org/pdf/math/9404236
[2] "Programming as theory building" https://pages.cs.wisc.edu/~remzi/Naur.pdf
[3] "What I cannot create, I do not understand"
If you think the declaration is saying AI is not useful it's you that "does not understand the concrete problem that mathematicians are upset about". Terence Tao uses AI heavily and has been writing extensively about how useful it is ever since it became useful in maths.
there's a level of obfuscation for any knowledge work where you rely on existing but incomprehensible-to-you systems that you just trust to work. are you unable to do any kind of mathematical work if you don't understand every single layer of proof that exists under-the-sun that touches your subject matter - or can you trust that some of these antecedents have been battle-tested and are functionally true for your purpose?
you could make an effective argument about the state of modern general-purpose LLMs that's founded on the idea that they are fundamentally untrustworthy and all results need to be validated but the larger categorical narrative, that the only true way to understand something is to know the logic from the most base principles, seems faulty
I think that AI should enhance said proxy.
For example: I have Crohn's so crohns.ai has the entire AGA [gastro.org] and each member is an agent that can participate in my program / protocol.
Same for MNT and dietmanager.com
this is NOT a promo, its a model I am trying to prove; AI can enhance the support that domain experts provide if we remove the barriers.
It's really a matter of AI-native Governance and how we handle that.
My hunch is that this ultimately doubles back to those that excel at story telling and human coordination. As the AI systems “offload” not just production but I believe some initiation of what to build, the “why” and how to rally groups for any appreciably complex work matters more.
I also hope to see a plenty of solo shops succeeding in spaces that used to take entire teams, but (for now) remain convicted human coordination remains a key need for most endeavors.
The "why" behind most products is of little interest to the majority of workers. While it may be of the upmost importance (on the surface at least) for the leaders of a business, I see no reason why Bob from accounting is going to give a fuck about your company's grand vision.
This is especially true if your company is in a mundane lane like B2B SaaS. You could argue that workers at SpaceX care about the "why", but 99.99% of companies aren't SpaceX.
People skills imo will always trump other skills. The best ideas never live because people don't know how to sell. And knowing and being very successful at selling requires that excellent story telling and human coordination.
Football analogy; AI is the wide receiver and the human is the quarterback. Even if you're the best WR in the game you're still not producing touchdowns unless you have a decent QB.
It's super easy to smell vibe coded projects.
All of these proofs and vibe code are impossible without human work. Call me when GPT whatever writes gcc from scratch
If AI is the wide receiver, it’s perceived currently by many to be an absolutely elite top-tier receiver.
Especially at lower leagues, QBs who have these clearly elite WRs are discounted and considered more-or-less replaceable all the time, because “anyone could throw the ball to Megatron 2.0”.
See, for example, Graham Harrell at Texas Tech: Paired with Michael Crabtree, put up insane numbers, went undrafted. Shedeur Sanders is another more controversial recent example.
It’s definitely not a guarantee that people will see the QB as valuable if the WR is that good.
Your point also extends the analogy, the AI might seem enough at lower levels due to that reason
The bar just took a very noticeable jump that people are still adjusting to.
Anything the WR can do on it's own is often seen as 'AI slop' - so now what matters is the QB. Because any random junk QB that doesn't meaningfully improve the WR's output is indistinguishable from AI slop. The distinguishing factor is solely the QB.
LLMs are computer programs, so there are math problems which they cannot solve. AKA, ideas which are not possible for them to generate.
The argument for this is that Busy Beaver function is uncomputable. More specifically, some N-state Turing machine requires a proof that it doesn't halt. At some point N is too large and LLM being a computer program, it cannot generate the required proof.
See the Busy Beaver Frontier [1]
This is VERY DIFFERENT from the Halting Problem. In the Halting Problem, we see that no computer can decide whether an arbitrary given input program halts. With the argument above, for a fixed LLM, there is specific math problem which is beyond the capability of proof by the LLM (though other LLMs or humans could perhaps prove it).
Humans are not bound by the argument since we aren't finite computer programs (no proof for this anyways). LLMs which "evolve" over time with input from the natural world also aren't bound by this, since their code is effectively infinite. The argument only applies to a static program with fixed input, no dynamic information sources.
Some people believe in divine inspiration. Maybe you could believe that humans incorporate information from the natural world which LLMs don't have access to. Either of these beliefs would imply that humans have an edge.
I also believe there are limitations of LLMs, but not necessarily where people think. I won't expect LLMs to be creative solvers until they can tell a novel funny joke with any recognition/consistency.
A few years ago they couldn't do basic arithmetic. Is there any reason to think that capabilities flatline from now?
> proving by negation, not proving for all cases
I'm finding 2.2k "∀" symbols in the Navier-Stokes proof repo:
https://github.com/search?q=repo%3Aopenai%2FNavierStokesAndE...
It's not like can't proof universal properties.
4 years ago ChatGPT created better slop than current Claude Opus 5. It's a regression: things good objectively worse. You can find multiple threads here or on Reddit about that. Models won't necessarily get better.
Just like how social media has possessed many people with cultivating an outward facing image that often diverges with reality, AI posses people to portray themselves as an artist/developer/musician/etc. without having put in any of the work.
The article talks about how many new ideas are relatively worthless and the real goal is to find the "concepts that 'carve nature at its joints.'" I think this is the crux of the whole thing and I haven't seen a satisfying discussion of it anywhere.
I mean, FLT is mentioned. Is that an accessible proof to humans? Is it full of these high value, refined concepts or is it more like a bunch of little hacks that at least dozens if not hundreds of people randomly stumbled upon, in an all out attempt to solve one of the most famous math problems?
I'm not totally convinced what value math concepts have beyond "you can use them to solve even more math problems." I really want to believe there is. But if not, it's just a pure benefit to have faster ways to solve them, no?
But I think FLT is about as far as one can get from that situation! Wiles's work was the culmination of centuries of theory-building work, and the concepts that were developed over that time are far more important than FLT; the thing Wiles actually proved (a special case of something called the "Modularity Theorem", the full version of which was proved a bit later) is itself much more valuable to human understanding of mathematics than FLT. It's certainly very cool that it can be used to answer such a simple question that was open for so long, and it makes for a great headline, but I think if you asked number theorists working in the area they would almost all tell you that they're much more grateful for the theory that came out of this quest than for the mere fact that the quest was completed.
The value of these math concepts is in their explanatory power. We want to understand! Insofar as we want to maximize the diffusion of powerful thought-tools, ironing out the wrinkles at the frontier is needed to be able to fold everything up neatly. The benefit we should seek is wide diffusion of powerful theories, and the folding and the neatness has been long undervalued.
I do not really care what other people think. Build something for the purity of building it for yourself or for a purpose.
Thanks to AI, I feel like I have been writing the best code in my life. Yes, writing, not vibe-coding. I mainly just ask questions and ask for hints and clues. I do not use LLMs to do the fun parts for me.
I am currently working on a game. If and when I ever finish it, I want to be able to say that I wrote every single line by hand. Will it make me better than anyone? No, not at all. I want to do it for myself.
This misses the point entirely. If you know proposition X is true or false, you won’t bother spending years trying to prove or disprove it, developing deep understanding and potentially even developing entire new fields of math in the process. (See FLT.)
That’s why it’s so destructive to the discovery process to have an LLM just generate a proof or counterexample without the side effect of generating useful explanatory or generative structures that we can build on.
That said, there are examples like Ramanujan where someone did an info dump of unexplained theorems that people try to mine useful structures from, but that’s not at all the mainstream of mathematical progress.
The LLM itself is finite, the axioms it knows are fixed, there is an N where BB(N) is independent of those axioms, so the LLM cannot solve it.
(actually I am wrong. You would introduce a new proof, and then step the verifier on all ongoing proofs, so non-termination isn't a driving concern)
The proof verifier uses fixed math axioms. The busy beaver function at high enough N cannot be proven with those axioms.
you like typing out the code, but i see that as the drudgery that I have to slog through after the fun of drawing architecture diagrams and writing out the api specs.
It's equally likely that in a few years we see everyone who doesn't have skill the same way we view people who follow the Kardashians and other influencers: vacuous and incapable of non-augmented thought.
In an ideal world, we would all be able to do such a thing all the time, and not worry about our position in the howling ape hierarchy.
I have two projects running concurrently. One is my hand-written project. The other is one "demolition zone." In the latter, I can vibe-code and do whatever I want just to see what LLMs can hack away at in comparison.
Besides the enjoyment I get from programming by hand, I have a more overarching concern when it comes to a building a full game. I am using Monogame, and I realized that is probably not the library to use if I am going to have no understanding of how things work under the hood. I tried to review LLM code, but without knowing enough of Monogame's API, I really couldn't make a good judgement call on the quality.
Looking at the LLM code, for example, were the keyboard and game pad implementations correct? I am not sure at first glance. I haven't implemented one since I used C++ in a game dev class in college over 12 years ago. The LLM's tests passed, which is fine and dandy, but how do I know the tests are truly valid if I do not understand what is fundamentally being tested?
I am not building this game entirely for fun either (though I am having a blast). I am working with a few non-technical friends on making this game. So, I do not want to disappoint them by creating something with bugs that I cannot fix because I have no earthly clue how anything works. I believe that would be embarrassing for me and disrespectful to them.
Do what I do sometimes: use opaque pointers (in C, or C++), and don't give the LLM the implementation, give them only the header, when asking it for tests.
I've lost count of the number of times I'll give only the header, but the LLM insists it needs to see the implementation as well in order to write the tests.
If you're not using C or C++, well, then write interfaces, give that to the LLM.
To put another way, if it were true that some fixed LLM could solve every math problem, it would implement Halt, which is impossible.
In my understanding of history (pre-enlightenment), people didn’t gather around ideas that were “correct more often” but gathered around ideas and personalities that reinforced their own ignorance, superstitions and biases and were openly hostile to “correct” ideas and truth, especially if those ideas upset tradition or dogma.
I also don’t understand how you’re getting to “violence was not a default until populations grew” since, perhaps, “default” violence was only properly recorded after populations grew and education/writing became more widespread. It’s very likely everyday violence in small groups was simply accepted as “the way things are”and simply ignored, as it often is even to this day.
Systemic violence came with the agricultural revolution, food and population stores, the ages of Bronze and Iron. Life was too precious in small bands for intense violence, even against rivals.
[Edit: 100 men armed with wooden clubs is an extremely bloody rugby match. Things generally did not start getting very-very bad until metal weapons were combined with boats or horses.]
Hobbes was a bit wrong it turns out.
This has not been my experience inside of a FAANG, in fact rather the opposite: People have been getting praised for AI slop because they can produce it quickly and it’s “good enough”.
The pendulum is slowly starting to swing back a little bit, but mostly because people have realized there’s effectively infinite noise now and doing anything, regardless of how skilled of a QB you are, is not being rewarded.
We're specifically talking about the context of 'LLM-aided human output' not about general capability of these systems.
Some buttons on the front page don't even work.
And then you have Joachim Breitner using Claude Code to produce Soundness proofs in Lean4, mostly vibe coded.
It really is the QB in charge methinks
Somebody will now respond that it doesn't matter that it doesn't benefit the field if it can produce the product without the mathematicians, but the whole point is that an LLM-generated Lean proof isn't the product. The proof of a singularity in Navier-Stokes is of zero commercial value (except bragging rights) in itself. If it ever results in something of commercial value it will be far downstream, after any new, valuable ideas in it have been digested, explored, explained, etc. This is work for mathematicians that LLMs cannot (currently) do.
And next someone will respond that LLMs will be able to do that work, however that is a prediction that is yet to come about. It might happen, it might not. To tear down something valuable now on the assumption that it will become obsolete in the future is wildly irresponsible.
The million dollars is an award from the Clay Institute. It doesn't represent value in the proof itself, but rather is offered as a reward for an activity that, when done the human way, is expected to result in valuable ideas. As the Clay Institute says on their website, on the Navier-Stokes page:
> Why ask for a proof? Because a proof gives not only certitude, but also understanding.
It's also a fraction of the amount OpenAI invested in generating the proof, and its value to them is dwarfed by that of "bragging rights".
"Everything that can be proven" is relative. PA can prove some things, ZF more things. In 200 years we could develop more powerful math foundations which can prove more things. Today's proof verifiers could never prove them, but tomorrow's proof verifiers could. And the cycle repeats.
I have updated the neomacs's landing page, make the "work in progress" indicator more prominent.
Also, I'm working on using neomacs's wasm build as neomacs's landing page:
I believe that, after continuous multiple iterations of development/test, neomacs will undoubtedly become excellent software。
This is quite simple. f(p) = C implies p does the job quite elegantly.
Interestingly it's harder to do the opposite, to simulate ZF in ZFC, because there is no way to express "forget that you know about C". Such a construction cannot be possible in general because if a contradictory axiom is added, then everything is true, and a theory where everything is true is useless and can't simulate anything. However for C in particular I believe it should be possible to make such a construction but I can't immediately think of how I would do it.