Despite the recognition that the problem of AI control may be one of the most important problems facing humanity, it remains poorly understood, poorly defined, and poorly researched says Dr Roman V. Yampolskiy in his upcoming book, AI: Unexplainable, Unpredictable, Uncontrollable.
In the book, AI Safety expert Dr Yampolskiy looks at the ways that AI has the potential to dramatically reshape society, not always to our advantage. He explains: “We are facing an almost guaranteed event with potential to cause an existential catastrophe. No wonder many consider this to be the most important problem humanity has ever faced. The outcome could be prosperity or extinction, and the fate of the universe hangs in the balance.”
The risks of uncontrollable superintelligence
Dr Yampolskiy has carried out an extensive review of AI scientific literature and states he has found no proof that AI can be safely controlled – and even if there are some partial controls, they would not be enough.
He explains: “Why do so many researchers assume that AI control problem is solvable? To the best of our knowledge, there is no evidence for that, no proof. Before embarking on a quest to build a controlled AI, it is important to show that the problem is solvable. This, combined with statistics that show the development of AI superintelligence is an almost guaranteed event, show we should be supporting a significant AI safety effort.”
He argues that our ability to produce intelligent software far outstrips our ability to control or even verify it. After a comprehensive literature review, he suggests advanced intelligent systems can never be fully controllable and so will always present certain level of risk regardless of benefit they provide. He believes it should be the goal of the AI community to minimize such risk while maximizing potential benefit.
AI (and superintelligence), differ from other programs by its ability to learn new behaviors, adjust its performance and act semi-autonomously in novel situations. One issue with making AI ‘safe’ is that the possible decisions and failures by a superintelligent being as it becomes more capable is infinite, so there are an infinite number of safety issues.
At the same time, Yampolskiy explains, AI cannot explain what it has decided, and/or we cannot understand the explanation given as humans are not smart enough to understand the concepts implemented. If we do not understand AI’s decisions and we only have a ‘black box’, we cannot understand the problem and reduce likelihood of future accidents.
Minimizing the future AI risk
To minimize the risk of AI, Yampolskiy says it needs it to be modifiable with ‘undo’ options, limitable, transparent and easy to understand in human language.
He suggests that all AI should be categorised as controllable or uncontrollable, and nothing should be taken off the table. Limited moratoriums, and even partial bans on certain types of AI technology should be considered.
Instead of being discouraged, he says: “Rather it is a reason, for more people, to dig deeper and to increase effort, and funding for AI Safety and Security research. We may not ever get to 100% safe AI, but we can make AI safer in proportion to our efforts, which is a lot better than doing nothing. We need to use this opportunity wisely.”
Resilience Forward comment
We thought it would be interesting to ask an AI system itself for ideas on the best ways to stop artificial intelligence from leaving human control. This was Chat GPT’s response:
Preventing artificial intelligence (AI) from leaving human control is a complex challenge that requires a multi-faceted approach involving technical, ethical, and regulatory measures. Here are some strategies that can help mitigate the risks associated with AI systems becoming uncontrollable:
- Robust Governance and Regulation: Governments and international bodies can establish comprehensive regulatory frameworks to govern the development, deployment, and use of AI systems. These regulations can include guidelines for transparency, accountability, safety, and adherence to ethical principles.
- Ethical Design and Development Practices: AI developers should prioritize ethical considerations throughout the entire lifecycle of AI systems. This includes designing AI systems with built-in safety mechanisms, adhering to ethical guidelines such as fairness, transparency, and accountability, and conducting thorough risk assessments during development.
- Human Oversight and Control: Maintaining human oversight and control over AI systems is crucial. Developers should design AI systems in a way that ensures humans can intervene, monitor, and override automated decisions when necessary. This may involve implementing mechanisms such as kill switches, human-in-the-loop systems, and decision explanation capabilities.
- Transparency and Explainability: AI systems should be transparent and explainable to enable humans to understand their behavior and decision-making processes. This enhances accountability and helps identify and address potential issues before they escalate beyond human control.
- Continuous Monitoring and Evaluation: Implementing robust monitoring and evaluation mechanisms can help detect anomalies, biases, or unintended consequences in AI systems. Regular audits, testing, and feedback loops are essential for identifying and addressing emerging risks.
- Education and Awareness: Educating policymakers, developers, and the general public about the risks associated with AI and the importance of responsible AI development and use can help foster a culture of accountability and responsible innovation.
- International Collaboration: Given the global nature of AI development and deployment, international collaboration is essential to develop consistent standards, share best practices, and address common challenges related to AI safety and control.
- Research and Development: Investing in research and development aimed at advancing AI safety techniques, such as robustness, verifiability, and alignment with human values, is critical for developing AI systems that remain under human control.
While these strategies can help mitigate the risks associated with AI systems leaving human control, it’s essential to recognize that achieving complete control may be challenging due to the inherent complexity and unpredictability of AI systems. Therefore, ongoing vigilance and adaptation of strategies are necessary to address emerging risks and ensure that AI continues to serve humanity’s best interests.






