Machines that halt resolve the undecidability of artificial intelligence alignment
Abstract The inner alignment problem, which asserts whether an arbitrary artificial intelligence (AI) model satisfices a non-trivial alignment function of its outputs given its inputs, is undecidable. This is rigorously proved by Rice’s theorem, which is also equivalent to a reduction to Turing’s Ha...
Uloženo v:
| Hlavní autoři: | , , , |
|---|---|
| Médium: | Artigo |
| Jazyk: | Inglês |
| Vydáno: |
Nature Portfolio
2025-05-01
|
| Edice: | Scientific Reports |
| Témata: | |
| On-line přístup: | https://doi.org/10.1038/s41598-025-99060-2 |
| Tagy: |
Žádné tagy, Buďte první, kdo vytvoří štítek k tomuto záznamu!
|
