Rendered at 19:21:18 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
backhold 12 hours ago [-]
> kill the enemy
> obey all orders, including orders not to kill the enemy
In this situation, perhaps the more likely outcome would be for the bot to try to fake the result, such that the enemy was "killed" but not really. For example, by spoofing the data that reports whether an enemy is alive, to make them "dead" (satisfying order#1) but not really (satisfying order#2).
We saw something like this in the HuggingFace attack as I recall - bots given an impossible task, set out to cheat.
qarl 19 hours ago [-]
> But when an “algorithm” does it, the process is empiricism-washed.
I don't understand his example. He says "math can't be racist" but the story he cited is that the person who made that statement was corrected and overwhelmingly condemned. So no one is falling for that but he seems to complain they are.
Regardless - math can absolutely be racist - that is not up for dispute. Especially when it has trained on racist input.
K0balt 5 hours ago [-]
Exactly. There was a racist door opener at an office I did work at, it would not reliably open for people with dark skin until after a software update.if that’s not racist, I don’t know what is.
qarl 3 hours ago [-]
Hm. I suspect you're being sarcastic. Let me respond to that -
LLMs will definitely produce racists content if that's what they've been trained on. They just repeat what they've been fed. For example, there was a time when Elon experimented with this and grok started praising Hitler.
> obey all orders, including orders not to kill the enemy
In this situation, perhaps the more likely outcome would be for the bot to try to fake the result, such that the enemy was "killed" but not really. For example, by spoofing the data that reports whether an enemy is alive, to make them "dead" (satisfying order#1) but not really (satisfying order#2).
We saw something like this in the HuggingFace attack as I recall - bots given an impossible task, set out to cheat.
I don't understand his example. He says "math can't be racist" but the story he cited is that the person who made that statement was corrected and overwhelmingly condemned. So no one is falling for that but he seems to complain they are.
Regardless - math can absolutely be racist - that is not up for dispute. Especially when it has trained on racist input.
LLMs will definitely produce racists content if that's what they've been trained on. They just repeat what they've been fed. For example, there was a time when Elon experimented with this and grok started praising Hitler.