Truth is not a direction: a Tarski attack on LLM probes(abeljansma.nl) |
Truth is not a direction: a Tarski attack on LLM probes(abeljansma.nl) |
The best anyone expects from an LLM "truth vector" is that it would encode the model's belief about whether the statement is true. Of course a perfect truth oracle is impossible.
I'm admittedly not a logician and don't formally study paradoxes, but I never quite understood the whole category of "this sentence is false" paradoxes, I feel that the sentence is paradox and therefore it's neither false nor true.
I do agree that it's a truth vector sounds like a silly panacea fantasy, though. But more logical formality is not the counter argument that would convince me of it, rather I believe that there's less formal and rigorous ways to get closer to truth.
If terms we use are not well-defined, then sentences using such terms can not have meaning.
But for the sake of argument let's explore, what could the "this" in (so called) "self-referential" sentences refer to?
Do they refer to a specific encoding of the sentence you are reading, as some bits in computer memory perhaps?
That would require that those bits somehow have a unique "identity" and the "this" in a self-referential sentence would have to refer to those bits in specific addresses of a specific memory-chip.
But of course the "this" does not specify which memory chip, which specific (concrete) encoding of its (purported) meaning it is referring to. And if it did, then it would be talking about that specific set of bits in that specific memory-chip, not of "itself".
The fallacy is that what we perceive as a "sentence" we read, is somehow "speaking" of something. But no, the sentence is not a subject, a sentecne can not speak, and THEREFORE it specifically can not speak of itself.
A sentence can not speak only actors, only subjects, like humans and AI, can "speak". And their speech must be encoded in some physical mediusm. A written sentence like "This sentence is ..." gives us the false impressions that somehow the SENTENCE IS SPEAKING of itself!
But speech can not speak, speech is the product of speaking.
Hence, in my view, "self-referential sentences" do not have any meaning and whatever paradoxes they might seem to create are results of confusion between ontological levels of "Subject" vs. "Speech".
Paraconsistent Logic
Intuitionistic Logic
Dialetheism
The same analysis applies: the probe tells us what the LLM thinks about the truth value if the sentence, not the truth value of the sentence. I don't think anyone claimed that these probes were truth oracles.