Adversarial Policies Beat Superhuman Go AIs
2022/11/01 by Tony T. Wang, Tony Tong Wang, Adam Gleave +21 · 18 voices · 7 citations
Computer Science · Social Sciences · #Adversarial Robustness in Machine Learning #Advanced Malware Detection Techniques #Ethics and Social Impacts of AI
paper · pdf · doi:10.48550/arxiv.2211.00241
Abstract
We attack the state-of-the-art Go-playing AI system KataGo by training adversarial policies against it, achieving a >97% win rate against KataGo running at superhuman settings. Our adversaries do not win by playing Go well. Instead, they trick KataGo into making serious blunders. Our attack transfers zero-shot to other superhuman Go-playing AIs, and is comprehensible to the extent that human experts can implement it without algorithmic assistance to consistently beat superhuman AIs. The core vulnerability uncovered by our attack persists even in KataGo agents adversarially trained to defend against our attack. Our results demonstrate that even superhuman AI systems may harbor surprising failure modes. Example games are available https://goattack.far.ai/.
Cited by
Discussions
- Adversarial policies beat superhuman Go AIs (2023) [hn, 306 points, 139 comments]
- Adversarial Policies Beat Professional-Level Go AIs [hn, 195 points, 46 comments]
- These are very odd failures! arxiv.org/abs/2211.00241 [bsky, 102 points, 3 comments]
- There is hope for humans in the AI wars, it seems that they have surprising failure modes, and we can beat them at games like Go, after all. arxiv.org/abs/2211.00241 [bsky, 4 points, 0 comments]
- Adversarial Policies Beat Superhuman Go AIs arxiv.org/abs/2211.00241 [bsky, 4 points, 0 comments]
- "en même temps" : arxiv.org/abs/2211.00241 [bsky, 2 points, 1 comments]
- "Adversarial policies beat superhuman Go AIs (2023)" Adversarial policies show that challenges exist even for the best AI in Go. Understanding these findings could help improve future AI. Simpler exp [bsky, 1 points, 0 comments]
- ORIGIN OF THE MENTATS arxiv.org/abs/2211.00241 [bsky, 0 points, 0 comments]
- “Adversarial Policies Beat Superhuman Go AIs” AI systems are bridle in the weirdest ways. arxiv.org/abs/2211.00241 [bsky, 0 points, 0 comments]
- Adversarial policies beat superhuman Go AIs (2023) https://arxiv.org/abs/2211.00241 (https://news.ycombinator.com/item?id=42494127) [bsky, 0 points, 0 comments]
- https://bsky.app/profile/news.ycombinator.com.web.brid.gy/post/3le2rxxwomzi2 [bsky, 0 points, 0 comments]
- Adversarial Policies Beat Superhuman Go AIs (arxiv.org) Main Link | Discussion [bsky, 0 points, 0 comments]
- Adversarial Policies Beat Superhuman Go AIs view on hacker news [bsky, 0 points, 0 comments]
- I think the fact that even now its possible to find adversarial strategies against alphaGo-esque ai go bots suggests that something as open ended as ‘maff’ is going to have some weird holes when rl is [bsky, 0 points, 0 comments]
- Superhuman Go AIs are vulnerable to adversarial patterns. Can be beaten with relative ease (including by humans). https://arxiv.org/abs/2211.00241 [bsky, 0 points, 1 comments]
- Adversarial Policies Beat Superhuman Go AIs https://arxiv.org/abs/2211.00241 https://news.ycombinator.com/item?id=42494127 [bsky, 0 points, 0 comments]
- Adversarial policies beat superhuman Go AIs (2023) https://arxiv.org/abs/2211.00241 arxiv.org [bsky, 0 points, 0 comments]
- Adversarial policies beat superhuman Go AIs (2023) https://arxiv.org/abs/2211.00241 [comments] [220 points] [bsky, 0 points, 0 comments]
Related