It appears that no LLM can get more than about a third of returns correct. Kimi K3 is on par with Sonnet 5. They are both very bad at doing your taxes.
In 6 months an LLM will get it to 80%
people will exclaim negatively “it was trained on the eval [to do this incredibly useful thing]”
meanwhile a problem is solved
So, instead of math/programming, tax codes approximate the real world better. Whereas math/programming/images/languages are frozen in time. I find this fascinating. Almost like the legal system, wherein rules are laid out in english but are almost mathematical in the sense that there is a legal/illegal judgement made for almost every action, but that judgement is not easily determined.