> obey all orders, including orders not to kill the enemy
In this situation, perhaps the more likely outcome would be for the bot to try to fake the result, such that the enemy was "killed" but not really. For example, by spoofing the data that reports whether an enemy is alive, to make them "dead" (satisfying order#1) but not really (satisfying order#2).
We saw something like this in the HuggingFace attack as I recall - bots given an impossible task, set out to cheat.
I don't understand his example. He says "math can't be racist" but the story he cited is that the person who made that statement was corrected and overwhelmingly condemned. So no one is falling for that but he seems to complain they are.
Regardless - math can absolutely be racist - that is not up for dispute. Especially when it has trained on racist input.
LLMs will definitely produce racists content if that's what they've been trained on. They just repeat what they've been fed. For example, there was a time when Elon experimented with this and grok started praising Hitler.
My experience with facial recognition tech is that it’s fundamentally prone to false positives in darker faces,but pretty good with white people. Contrast? Training data? Idk.