Using Opus 5.5 to discover a new eyewitness record of the dodo(resobscura.substack.com) |
Using Opus 5.5 to discover a new eyewitness record of the dodo(resobscura.substack.com) |
If the introduction HP-65 was the first time an individual was capable of fully replicating the logical and cryptanalytic functionality of the WWII Turing-Welchman Bombe, then the roll-out of Agentic AI could be said to represent the first time an individual was capable of replicating the functionality and deliverables of the WRENs of Bletchley Park. Mindboggling.
https://en.wikipedia.org/wiki/WRENS
Even more impressive is the Astra-6 report recently. Given access to the Crypto Cellar Research archive of unbroken intercepts, the AI wrote its own custom Python and C++ software to simulate an Enigma machine and a mechanical Enigma Bombe.
https://www.tomshardware.com/tech-industry/artificial-intell...
In any case I stopped writing it up at that point and tried to focus on the messages. But I have the mostly-done writeup and I'll post it soon.
Do you work in university/research/history & co, or how did you find this specific niche? :-D
> They are also notably bad at judging the historical significance of what they find.
I use LLMs for some things that are outside the more common use-cases (in my case 3D design for 3D printing) and one thing I've noticed is that the errors it makes are so completely unlike human errors that they are hard to anticipate.
It will do things like build perfect snap catches but put them so the the pieces they are connecting are rotated 90 degrees from how they should be. It's "dumb" error, but hard to say the model itself if dumb because it does other very hard things so perfectly.
> seven chord groups
This sounds a lot more like Opus 5.0 than Opus 5.5 TBH. I wonder if that was an earlier investigation because 5.5 has improved that kind of language a lot.
If they were the total I would say it was arguably more impressive the author was able to narrow it down to just 3000 pages than it was to find the dodo mention amongst those!
This is a fantastic concept. This is where the blacksmith and the baker alone will never come up with the idea for a new furnace that seems rather obvious if you knew both trades.
A person capable of making the rest of the video would never make those mistakes, but an AI does. Perhaps our intelligence is also spiky, and we're just used to the general shape and variance within humans.
You could image aliens coming to earth being shocked that we were able to discover general relativity but can't remember 100 digit numbers in our heads.
some of the metaphors make no sense
To be fair some famous rappers are guilty of this, tooYou say that’s an error a human couldn’t do, but imagine if the human has never seen or touched the kind of item you were making and relied entirely on text descriptions to build its ontology. Off by 90 seems like such a believable mistake.
(even if it's through some silly means like explicitly talking to themselves like "if I were a human doing this, what [... 5 million tokens in 2 seconds ...]" but also of course if they crack ASI and get something more efficient and intelligent than a human brain by then)
However I can image a part, see it in my mind, rotate it, place it in context, and have an intuition about it.
Opus 4.6 and 4.7 were bad, but GPT 5.2 and above were very usable. Opus 4.8 was usable, but the GPT 5.x series was better.
Fable is great.
Opus 5.0 was interesting. It could solve some problems that Sol 5.x couldn't solve (applying a G2 curve on a 3 way corner where one face was a Bezier curve) but you had to be super prescriptive ("only answer the question"/"only do what I tell you and stop when done") or it would go on a hugely involved validation journey that didn't really achieve a lot.
Opus 5.5 is better than that was in that respect.
Sol 6.x is great, and my daily driver for this (I use Opus for coding though)
Astra can solve problems that Sol can't but for some reason on easy stuff makes uglier solutions.
For all models it's very interactive though - we aren't at the "agentic design" phase for most things yet.
Here's a sample of what I've been able to get them to design with me: https://x.com/nlothian/status/2099023496794018067
When I need to use a program I seldomly use I'm far more likely to remember where I need to click to open it than the word I need to search for to open it.
I honestly don't know if an image model would help, or if it might analyze the output and go "13 fingers? ship it!" anyway.
Through at least Opus 4, LLMs were practically useless for authoring any sort of coherent procedural closed-curve geometry (I know this with strong confidence because of the little animated guys at https://letterspractice.com).
Opus 5.5 can bang it all out. Possibly a deliberate RL sort of thing or maybe another surprise emergent capability.
Maybe the model isn't intelligence in any form, except perhaps as an imperfect reflection of the intelligence of its training data.
I scanned quite a few documents many years ago and it's been fun trying new tools every couple of years to see how good they are getting. I would say it's still not 100% there, but if you have them scanned you can basically just just keep trying and compare the results.
I think the only disappointment in recent years has been that storage has gotten more expensive instead of cheaper. I was waiting for SSDs to get cheap enough to justify moving all my documents to a fast flash array to process and search through them faster, but that doesn't seem like it will happen anytime soon.
How much data are you working with? lol
Seems like you're doing something professional-grade if this is the case. I imagine most people, like the OP, can basically use whatever machine they have laying around and never be concerned with storage size/speed.
I'm also curious if you've tapped into cloud computing. Not that using the cloud is cheap, but I suppose if I was in a position where I'm concerned with the cost of SSDs for a processing task, then I'd be exploring all of my options and I'd be surprised if the cloud wouldn't be an "easy" solution.
I would say it's about 80% accurate, which means it's missing enough key words to make a lot of it uselessly unintelligible. I can easily compare the images against text I turn up in a grep which is nice if I'm looking for something.
Allegedly Claude set up a system for retraining for my handwriting, but it would require me to manually revise several hundred pages by hand so I don't think I'll ever do it.
GPT 6.1 and Gemini Flash 3.8 both do pretty well, their OCR of your sample image is only "wrong" in the sense that the original has typos and they corrected some inadvertently and/or filled in gaps where you had "unintelligible" in the canonical text.
If you have the budget and want the best possible results, you need to run each image through multiple models and then combine the outputs into a final "merge these" prompt. Better scanning helps too, your sample image is rotated and you used a phone in low light. Try a DSLR or a flatbed scanner and process only one page at a time instead of two at once.
Something about the way some AIs are trained to write causes them to go for metaphors aggressively - and they don't always come up with good ones off the cuff. But they don't double back and get rid of the failures.
So yeah, they are very weak at spatial reasoning.
I haven't tested this recently but my possibly dated experience is frontier LLMs can't figure out which model is correct or incorrect if there's disagreement on vision recognition.
Have you found otherwise?
(Edit: I see you gave an anecdote about merging terrible results. My experience is with merging overall accurate results).
I don't think this is viable for critical record OCR. I think the only way to do that is one pass with a frontier model and then a mechanical turk manual review passthrough with good compensation that allows for a slow and methodical approach. Plus of course much better scanning than a phone camera.
I basically kept rabbit holing this problem and finally settled on "80% and done is better than sitting on this problem for 4 years waiting to have time and equipment for a 99.99% solution." Crank the pictures between Claude code sessions, run the local LLMs when I'm asleep, done, now I can free text search years of journals plus I have photo backups now finally of them.
Roughly:
Get API keys for multiple vendors or just use OpenRouter (but availability of frontier models tends to be limited). Alternatively, Azure Foundry has everything except Google models, so just two subscriptions is enough.
Run the same prompt and same input image through each of your chosen models.
Then feed the smartest model the original image together with the collected output texts. Use a prompt along the lines of "Merge these attempts to OCR together into an corrected and improved combined version, taking special care to exactly preserve the original's typos, etc, etc..."
You can do this manually, it's just fiddly. It's not hard to automate, most of the "code" is English instructions!
The downside of this approach is the cost: even the "light" frontier models are a few cents per page, which is not so bad until you're doing this 5x or 10x times per page and suddenly scanning a notebook can set you back tens of dollars, more than buying a good novel at a book store.
I picked up on this technique back when GPT 4 was released. People noticed that it could translate ancient Akkadian, but only if you ran the prompt through 4x times and merged. I tried this with a few random samples I found online and the merged translations were generally better than the "official" ones, even thought the individual attempts were unreadable gibberish.
There are already scripts/tools floating around for this!
Look into OpenRouter Fusion, Consensus AI, Multi-Model Debate, etc... or just whip up something yourself.
Take anything "intelligent", alter it to be spikier and spikier, and eventually *poof* somehow the intelligence vanishes. You can do the same with the phrases "flawed intelligence" or "specialized intelligence."
> Calculators [have] no intelligence. [...] I mean there's no adaptability
To short-circuit a long discussion, I submit that "adaptability" will (once the Scooby Doo gang catches it) turn out to be "intelligence" in a tautological mask, both equally undefinable except in relation to one-another.
Something will be intelligent because you perceive adaptability, and it'll be adaptable because you infer intelligence. If it doesn't seem adaptable, it can't be intelligent, and if you don't want it to be intelligent, it won't have "real" adaptability.
> A Chess bot has [adaptability] (in its narrow domain of Chess). A calculator does not.
My calculator solves equations with unknown variables, what makes that insufficiently adaptable? What determines the cutoff-point?
All intelligence is 'spiky' or 'jagged' or whatever, even human intelligence. We evolved in certain environments and situations and sometimes there's a mismatch and it causes all sorts of wonky things. We just call them funny names like optical illusions and cognitive biases. But it's the same thing. I agree there's no point in call LLMs a 'spiky' intelligence, but for probably the oppoosite reasons as you.
>My calculator solves equations with unknown variables, what makes that insufficiently adaptable? What determines the cutoff-point?
A chess engine can be dropped into a board position it has never encounterd and search over possible continuations, evaluating and selecting actions based on the state it finds itself in.
A calculator solving x+3=7 is doing something quite different. The fact that x can take arbtrary values doesn't make the calculator adaptive; it just means the fixed procedure operates over a range of inputs. Every problem a calculator can solve was effectively anticipated when it was built, and anything outside that grammar produces an error with no partial credit. A chess engine can face positions nobody enumerated, in situations nobody could even dream of and still produce sensible moves.
The core of being intelligent is being able to make decisions independently. That's why we hire smart people - to make better decisions. You can't make decisions if everything is spelt out for you. And naturally, if you can't make decisions, you can't adapt.
I thought you meant calculation and output, in which case “hey calculator, what’s the capital of North Dakota” demonstrates incorrect output.
I am obviously hoping you don't mean that in the trivial sense, otherwise the command 'cat' would also count.
Sounds pretty similar to a calculator with a numerical root-finding algorithm, if you only substitute board position it has never encounterd with a polynomial it has never encountered.
>The core of being intelligent is being able to make decisions independently
What does it mean for a deterministic algorithm to make decisions?
>What does it mean for a deterministic algorithm to make decisions?
Newton-Raphson isn't choosing among possible actions. Given (x_n), its next step is mechanically specified by the update rule: compute the derivative, take the tangent intercept, repeat. The intermediate result changes the next input, but that's not by-itself decision making.
A chess engine, again, does something different. From a position, there are many legal actions it could take. It considers alternatives, estimates their consequences according to some objective, and selects one. The engine has to work out which available move best advances its objective.
You can make both algorithms determinstic, but determinism isn't the distinction i'm drawing. 'Decision' here doesn't mean some metaphysical excercise of free will. It's more about selecting an action from alternatives based on an evaluation of their expected consequences. Determinism is orthogonal to decision making. Deterministic doesn't mean predictable, nor does it make its choices any less it own computation.
That aspect at least is not an issue, because: what does it mean to say that a dice roll is a "decision"? That's just probability, and it's as mindless as determinism. Some people associate free will with randomness, for no reason other than that it's an escape from the constraint of determinism, but it isn't any more meaningful. Yet just because meaningful thought is pre-determined by physics doesn't stop it from being thought, and hence being a decision.
Exhaustive search is impossible in chess, so again, chess engines do something different. A chess engine has to stop well before terminal positions and make judgements about positions it cannot fully resolve.
Reducing the search depth would make it more interesting because it too has to evaluate unresolved positions. But then the interesting part becomes the evaluation function is. For tic tac toe, it's going to be very easy to be written in such a manner where most of the judgement is supplied by the designer and not the system.