> The professional human team begins from a legal eight-to-ten-minute game state in which it has decisively won the laning stage: roughly a 6,000–8,000 team-net-worth lead, a 4,000–6,000 team-XP lead, two enemy Tier 1 towers destroyed, the third badly damaged, and all three friendly Tier 1 towers standing.
In go ELO like scoring he’s something like 120 points over the next strongest player. No other player has ever broken a 3800 rating let alone 3850. Ke Jie (the previous long time champion) peaked at 3755. Shin Jinseo’s strength graph is the most absurd straight line.
2 stones is historically the gap between a 9P ranked and a 1P ranked professional player (very roughly the gap between super grandmasters and an almost grandmaster)
That is to say it’s shocking that Katago (almost certainly significantly stronger than AlphaGo) is a mere 2 stones stronger than Shin Jinseo. I suspect it would be 3-4 stones vs any other human pro.
And of course you would need to take into consideration the scale of go ratings and chess ratings when making that comparison. With top chess ratings being around 2800, being 1000 less than the top go ratings, one would have to apply a factor of roughly 3/4.
It's not a comparison of the worth of the games (I play both, though I'm better at Go, and prefer it), but the dynamic range of Go is larger.
That said, any cross-game/sport comparisons of this kind are pretty tough to do properly.
English is not my first language but for clarity perhaps the title should be:
Go grandmaster Shin with a two-stone handicap defeats AI KataGo
However, you can't compare goratings over time, the top ranks are not nearly stable enough. https://www.goratings.org/en/history/ (I think it's believable Shin Jinseo is better than Lee Changho, but not that there has been steady progress since the days of Lee Changho, so that there are now 20 players stronger than him).
Ratings drift over time, based on the total population of people competing. I think the most accurate way to view Elo ratings is as a measure of skill vs. the average rated player.
If you want to compare Magnus Carlsen's peak rating of 2882 in the year 2014 to Garry Kasparov's peak rating of 2851 in the year 1999, you have to know how strong the total pool of players (including all the amateurs who compete at lower levels) was in the year 2014 vs. 1999.
The only way to actually anchor the Elo system over long periods of time would be to have rated humans occasionally play against a set of unchanging computer players, which could then serve as static rating calibrators. You could use those games to then calibrate Elo ratings from different time periods to a common scale, by asserting that the computer ratings don't change.
Not without flaws of course, but probably interesting
Interesting KataGo is an open sourced Go program written primarily by David Wu in C++ and recently heavily vibe coded by Claude. It's running on four Nvidia RTX-3090 GPUs with 96GB VRAM. [1]
Based on the youtube video, it looks like katago was only using 16 seconds per move, is that right? https://www.youtube.com/watch?v=-86zF4mTWOY
Is 20 seconds on that hardware really overkill and well into the diminishing-returns curve, as a top-level comment suggested, or is it plausible katago could have played better if given 40 seconds per move?
match details: https://gostonebase.com/blog/shin-jinseo-vs-katago-kishin-ma...
It can't be just play like AI.
Any other Korean on the Korean Go program could have done the same.
In fact, many did when AlphaGo was the pinnacle of AI.
I wonder if this means the best Go play is closer to theoretically perfect play or if it just happened the current computer methods didn't manage to get much farther than humans. Go has vastly more valid games but also a simpler ruleset, so I'm not sure if there is really a good way to tell beyond "keep trying and find out"?
I am going to go out on a limb and say part of it is a matter of respect, and part "lets not discourage and scare the shit out of anyone who knows even a little about Go"
A long time ago I ported a Go game to the iPhone for a job, the only way to control difficulty was to limit the amount of time it could spend monte carloing. On the iPhone 4, "Hard" would take 20-30 seconds a turn and drain your battery. I learned the game as I was making the app, got into single digits vs humans years after that. When the iPhone 5S came out it was capable of doing so many more iterations a second that even the easiest difficulty destroyed me until I spent enough time figuring out which moves created the worst spaces for it to search, which isn't really playing go anymore it's more like mining bitcoin by hand.
E: 4x3090 tops out at ~1.4k nnEvals/s. 4x5090 is 3.2x. 8xRTX 6000 something like 7x.
For example, we don’t even know whether perfect White play can possibly overcome two correctly placed Black stones against perfect play.
I recall a while back someone came up with a set of "anti-computer" strategies that allowed even an amateur to defeat a strong go program. These moves weren't anything like ordinary go moves (and perhaps the "loophole" has been closed now) but imo, their existence suggests that a study of programs may reveal other unexpected weakness.
Until a more average player confuses the AI with an unseen behaviour (pulling creeps between the towers etc) to get an advantage.
This in no way detracts from how absurd and remarkable it is that Shin Jinseo can beat KataGo (it gets a LOT of training and architecture refinements https://katagotraining.org/#eloGraphButtons) with 2 stones of handicap.
Shin's genius was to play out a complex variation of the flying knife joseki that was, in essence, a one-way path to reach an equal board position that occupied about 1/4 of the board. Due to the 2-stone handicap, the position favoured black with the game ~25% complete. KataGo could not have played any other way, where a human may have tried to foil the plan by introducing further complications.
What was truly incredible was how Shin held the advantage from that point on.
Shin took a 2-stone handicap from KataGo which means that Shin is the weaker of the two. But to give that more context, Shin is also the strongest human player to have ever lived in raw strength terms by a good margin, and is known as replicating AI move-for-move more closely than anyone else.
If they were to play even then there’s no chance any human could win (and pretty much all pros agree with that). Lee Sedol beating AlphaGo in game 4 of that series is largely considered the last time a human beat a modern AI in an even game, which is why it was so amazing.
RE the game, Katago was set to use the strongest available model and ran on a 3x 3090 GPU system, which is a lot for KataGo. 20 seconds might sound like a handicap, but that’s over 100,000 play out variations which is essentially infinite for modern KataGo models (anything over 10,000 is overkill).
Shin played well in all games, but his strategy was to avoid complexity. KataGo reads out complex fighting like an absolute monster, so Shin was trying to play very very solid and very very calm so as to not give KataGo an in.
The 2-stone handicap could be thought of as roughly 10-15 points of ‘buffer’. That’s massive in professional games, and that’s what Shin used to win. He played so overly solid that it sometimes cost a point or two, but it removed an opening for Katago to fight. He did this at the key opening and middle-game sections and never burnt through the full buffer of handicap points in the last two games. That’s why those games look kinda ‘boring’, it’s because Shin wanted them to be that way.
Also note that KataGo probably could have won if its ’variance’ was tuned higher (basically it taking risks). Standard KataGo won’t take risks, it just wins with brute force. For handicap games though you can tune its willingness to start fights higher to prevent people from just playing ultra solid (like Shin did).
Shin did an absolutely amazing job and he deserves all the recognition. Katago routinely beats professionals giving them 3-4 stones of handicap, so the win by Shin highlights to me how strong he is, but also just how well he understands how the AI ‘thinks’.
I watched every game years ago between Lee Sedol, even though it was late at night. I love AI and I love Go.
Off topic, but I wrote the first commercial Go program for the Apple II in the late 1970s.
Also, the human played a strategy tailored to that huge initial advantage. He said that the AI did not handle this particularly well, and played high probability moves instead of trying to lure him into a mistake.
Also, even though this was the best Go engine, it was not running on a supercomputer, and had a relatively limited amount of time per move.
So, this was an important victory for a human, but not a sign that humans are now stronger than AIs at Go.
• Deep Blue played Be4, declining a pawn capture to achieve a long-term positional advantage and fueling Kasparov’s suspicions of human intervention.
• AlphaGo played P10, a brilliant 5th-line shoulder hit that live commentators initially dismissed as a blunder or misclick.
https://franky07724-57962.medium.com/amazing-coincidence-in-...
It's similar in chess but there's a model specifically trained to play a knight down, and it's pretty cool to see the insane tactics it'll try. Though a knight down is considerably more than a 2 stone handicap in Go.
You can choose which piece it's missing. It's humbling to lose starting up a queen
It's never made sense to me that so many go players study AI go play in the hope of emulating it in a human game. We're not machines. We can't do thousands of Monte Carlo tree searches per second.
> We're not machines. We can't do thousands of Monte Carlo tree searches per second.
You sound so sure of that, but I have seen people catch a ball, and I am not so sure that any artificial person should be any more aware of the tremendous maths they are "solving"
From the perspective of a game with far fewer rules, "trying to imitate AI" might not mean anything like what you are thinking.
from what i remember, dogs sometimes run in curves so that the perceived trajectory of the object they try to catch is more linear
so not everything might be in-brain math but also good trickery
Correction: a very conservative strategy, so as not to lose the advantage he started with.
I predict the same thing is going to happen in Go once the engines catch up.
(You can play those chess bots on Lichess for free. Challenge LeelaQueenOdds or LeelaRookOdds if you are master level or stronger)
Within our lifetimes (unless you're quite young) it was doubtful if a go ai would ever beat a decent human. Same was true for chess a generation or two earlier. It's news because it's the handoff of man to machine being the best at a particular thing.
This current news is news from the other way, this human did _exceptionally_ well.
That's a way to see it. Another way to see it is that for any given event in a zero sum game, one of two outcomes, where two players are involved, will occur. Man only had to beat machines for what, 200-400 years only for it to beat man how many times for it to be "best"?
Obviously its an impressive human intellectual feat but the hubris is instructive perhaps for our wider interactions with AI as its abilities accelerate around us.
I started playing I believe in March of 2025. I had the very good luck that one of my high school best friends is a 4 dan EGF player, and was in the French team 20 years ago. We played constantly against each other with 9 stones handicap and him teaching me a lot.
But when I read that book, in a few weeks/months we went from 9 stones to only 3 stones by August 2025. and we have stayed at 3 stones since, playing very even games, with only a handful or couple of points differences. Only very recently did I win by 9.5 points. Which made me hope I had made new progress, but I had to resign in the very next game.
After you get stuck, you should pick one thing at a time to focus on improving. But the one thing should probably be related to tesuji or life and death.
One follow-up: Any recommended go servers to play on? OGS, Fox, Pandanet?
For a double digit kyu there are two simple things than can give you 2-3 stones boost overnight if you understand them.
- don't play aji keshi.
- endgame starts earlier than you think. If nothing is urgent, play the largest endgame in the corner / side of the board.
What a powerful story. Humans have a chance of remaining superior because emotions are the fuel for our intellect and wisdom.
As with chess, roughly equal matches between computers and human players are probably happening for the very last few times.
Blue spot is an adversarial ai and it's managed to beat average professionals on five handicaps, which is absolutely insane.
Second, if it is a fixed sequence, how position independent is it? In chess, tactical sequences end up depending on the entire board state to work when they get sufficiently long. I guess what I'm asking is, could a player with some capacity for stategic thinking recognise this idea and take steps to make the flying knife impossible?
I really should spend some more time learning go, it's such a fascinating game.
However, when Shin executed the 50 move flying knife, the board was pretty much empty. So there is really no need for calculation, both Shin and the AI know it’s locally optimal. But getting to play a very long locally optimal sequence is good for the weaker player, so they have less “real” moves to lose EV on. Notably Shin probably can’t open with the flying knife in one corner past a certain point in the game, even if that corner were completely empty - the rest of the board positions would change the end values of the variants.
If the AI could know this, they might play a variant that ends 30 moves sooner but is 0.01 pts worse. Then they would have more time to mess Shin up through organic new moves (which the AI will be better at of course).
(disclaimer: only ranked 1 dan)
To account for this, KataGo uses a "playout doubling factor". When the AI plays against itself to learn, the developers set one instance of the AI to use fewer playouts compared to the other one, but gives the weaker AI some handicap. This allows the AI with more playouts to learn that although it may be in a losing position, if it makes the board position chaotic enough, it may still win.
The flying knife is objectively an extremely complicated position, so the AI played it assuming that the opponent would be forced into a very complicated reading battle where they could make some mistakes. Unfortunately, Shin has memorized the flying knife joseki more thoroughly than any other human on the planet, so he could play exactly like a very strong AI. It would probably be possible to train an adversarial network specifically to beat players like Shin, but that would take a substantial amount of effort, and Shin is strong enough that it probably wouldn't make too much of a difference -- Shin won by 11.5 points in game 3 without a flying knife shenanigans, only losing 7 points of value throughout the entire game.
Yeah, katago's training is not really focused at all on handicap games, because it's by nature learning from even games against similar-strength opponents.
It doesn't have specific training from playing in a way to exploit a weaker player. In a handicap game you have to give your opponent opportunities to fuck up if you want to play optimally.
If a move loses 0.0005 points if the opponent plays optimally, katago won't play it even if there's ~zero chance a weaker player would play it right.
There have been go AIs that tried to train more directly on uneven opponents, one called "sai" comes to mind, but katago has huge advantages otherwise and won out over the others (for very good reason, it's a great project).
Just stating this off the top of my head so I could be misremembering, but I heard that the KataGo settings used were tweaked to favor complexity. This was most apparent in Game 1 which Shin Jinseo lost, where the AI had an unusual opening. However, the last game was quite plain leading me to wonder whether that setting was present in the last game (or at all).
In any intellectual contest between human and machine, all the machine winning implies is that the endeavor is algorithmic.
The machine can be given practically unlimited memory and compute; we consider it cheating if the human would use memory aids. The machine could be implemented as many agents cooperating; we'd think it's not right if thousands of humans collaborated to face the machine, etc.
So statements like "not a sign that humans are now stronger than AIs at Go" are pretty meaningless, IMO
Totally wrongheaded, actually, since the computer gave the human a 2 stone advantage from the start.
Another way to look at this: Go's handicap system gives us a genuinely interesting metric for the distance between a human and a machine at this specific game. Instead of just "computers beat humans" we get a quantified gap.
(This is especially true if they don't get briefed "this is time-traveling Lee Changho, he doesn't know contemporary joseki, play a trap variation").
A 2 stone handicap, however, does not scale linearly with strength. It becomes relatively more impactful at higher levels. For pro players, 2 stones are gigantic, and for beginner players, they make zero difference. Same for a point advantage (komi adjustment). So there might come a point where it is physically impossible for a computer to beat a top human under some handicap.
If you have an LLM add 2+2, it is doing a tremendous number of additions and multiplications it doesn't have to, same as us, and I am not sure the LLM can answer any more questions about that process than we can about ourselves.
It can take a few minutes to get a match on OGS. When that happens, I take the opportunity to warm up with a few tsumego drills on goproblems.com
If you know anything about go or computer go from last 10 years then it is obvious which direction the handicap goes.
Anyway. Yes if you throw examples into training it will be able to handle the situation - but handling unseen things for me is a key goal.
Is this also the case for Go's joseki? Or is any deviation easy to punish?
The feeling among the go players I know is that the match had been kind of a setup for the human to win.
It's an equivalent of playing in chess with a bishop handicap and a computer playing deterministic chess so that a human can prepare a forced sequence into a winning endgame.
Shin is an absolutely amazing player and not anyone could have achieved that, but nobody really feels like he won vs a computer team that tried its best.
One thing you'll hear people say sometimes -- I've said it myself -- is that this shows that in some sense go is a "deeper" game than chess; there's more to know and understand, more variety of possible human skill.
That might well be true. It certainly feels a more elegant game, and involves longer tactical sequences, and so forth. But this may be misleading.
Consider the game of treblechess. To play a game of treblechess, you play three games of ordinary chess and look at the overall result.
Suppose that when we play ordinary chess, I win with probability W, lose with probability L, and draw with probability D = 1-(W+L). And suppose separate games are independent of one another (which might not be true in reality, but never mind). What happens when we play treblechess?
I win 3/0 with probability W^3. I win 2.5/0.5 with probability 3W^2D, because that happens when I draw any one of our three games and win both of the other two. I win 2/1 with probability 3(W^2L+WD^2), because that happens when I win two and lose one or win one and draw two, and for each of those there are three choices for which game is which. So I win at treblechess with probability W^3 + 3(W^2(1-W)+WD^2).
I can draw by getting one each of WDL (probability 6WDL) or by drawing all three (probability D^3).
Suppose that when we play chess I win 40% of the time, draw 50% of the time, and lose 10% of the time. Then our Elo difference is about 107 points. In triplechess, I will win 65.2% of the time, draw 24.5% of the time, and lose 10.3% of the time. Our Elo difference is about 214 points.
If in ordinary chess I win 65% of the time, draw 25% of the time and lose 10% of the time -- about the same odds as for treblechess in the last example -- then our chess Elo difference is about 215 points. At treblechess I will win 84% of the time, draw 11% of the time, and lose 5% of the time, and our Elo difference will be about 375 points.
If in ordinary chess I win 15%, draw 75%, lose 10%, then our Elo difference is about 17 points; in treblechess I will win 31%, draw 49%, lose 20% and our Elo difference will be about 41 points.
Treblechess Elo differences are on the order of double ordinary chess Elo differences! Clearly treblechess is a game with twice the depth of ordinary chess!
But it isn't. It's just longer and gives more opportunities for the better player to come out ahead overall.
Go is also a longer game than chess, though of course not in the same way as treblechess is. Perhaps the larger Elo range of go is more because of that than it is because of actual deeper strategy and tactics?
So I think there's truth to what you're saying.
If you make a Backgammon board bigger it would just be a slog and no real increase in tactical challenge.
Chess rules dictate a set size of board, which I guess has been refined over time.
If a Go board is made bigger or smaller it just adjusts the problem space, the rules and core of the game remain the same.
1.00 +800
0.99 +677
0.9 +366
0.8 +240
0.7 +149
0.6 +72
0.5 0
0.4 −72
0.3 −149
0.2 −240
0.1 −366
0.01 −677
0.00 −800
The second column is your rating minus your opponent's, and the left is your predicted result. So if you are rating 1849 and your opponent is rated 1700 then you'd be expected to score about 70%. To have a 1% expected score against Magnus, you'd need a rating of about 2150.Two major federations are American Go Association (AGA) and European Go Federation (EGF). EGF uses an Elo-inspired update rule since 2021. AGA uses a quite-different Bayesian system without pairwise update; they provide a paper and a C++ reference impl.
Asian countries don't bother with such numeric ratings. Instead, rankings are titles which are won through tournament promotion structures (sounds similar to Sumo to me).
Interesting, because I always thought that it was more "apples to apples", and that the higher upper limits of Go rankings was somehow indicative of the higher "dynamic range" of the game compared to chess. For example, if Elo were applied to basketball, what would the Elo of Lebron James be compared to a playground hooper (leaving aside that 1-on-1 isn't the best part of Lebron's game)... would it be higher or lower than Magnus Carlsen in chess? I don't have an intuition.
And you can run a huuuuuuuge number of sims in 16 seconds.
If it is "only" 4x3090 at 16s, you will definitely get a drastic playing strength boost from doubling the thinking time. This is still clearly within the interval of a linear relationship between thinking time and playing strength, i.e. elo ~ log time. The relationship, to my knowledge, becomes less clear only starting at about 10-20x the number of playouts.
Source: Wrote a paper on this. https://ieeexplore.ieee.org/document/10645535
[1] From wikipedia, citing deepmind's paper: 4858 Elo vs 3739 Elo
You can kind of tweak towards play this metric or that, but it's not the same.
I agree that you cannot do it reliably, in principle.
However, Bobby Fischer peaked at 2785, and Gary Kasparov peaked at 2851. These are not far from what informed observers suggest--maybe 50 or 100 points off. They are well into super-grandmaster territory. Kasparov would be 1st today, Fischer would be 4th.
But on goratings.org, the top player of the 90s would be roughly 30th today.
My point is that the go ratings are much more unstable than the chess ratings. With chess ratings, you'll be wrong in the details. With go ratings, you'll be catastrophically wrong.
I imagine that FIDE has all of this data somewhere, right? They don't publish more in-depth distributional analyses of players somewhere?
Elo ratings measure differences in skill between different players. A player rated 400 points above another player will win 90% of the time. Only rating differences are meaningful. Absolute ratings aren't.
However, since the population maintains some continuity over time (players gradually enter and then leave over time), would it not be possible to reconstruct relative ratings between players that didn't play during the same era?
Or
Go grandmaster Shin, playing with a two-stone handicap, defeats AI KataGo.
Since the AI was playing scratch (the normal, unmodified play style), it seems odd to say it had a handicap. Rather, the human had a positive handicap of two stones.
X defeats (Y with Z)
or
(X defeats Y) with Z
with the latter usually understood as meaning X has Z, as far as I know.
Or that at least is my guess as to why it sounds wrong.
Also pretty bad at spelling.
Chess requires you to be really sharp. Experience increases with age, but there's some cross-over point at which the decline in calculating speed is more important than increasing experience. Heck, I'm measurably better at chess tactics in the morning after a good night's sleep than in the evening after work.
Edit: 10 years since i made the HN account and I only now notice you can't see the scores of other people's comments. Here's hoping the brilliance of the above comment is appreciated.