Amount Awarded: $25,000
AI is critical in today's rapidly changing world. New Artificial Intelligence (AI) technologies are poised to drastically transform how we address societal challenges and build a more sustainable future. Generative AI (GAI) technologies, particularly large language models (LLMs), have known weaknesses such as hallucinations, misinformation, and disinformation. Additionally, LLMs are vulnerable to various attacks and abuses, raising important questions about AI risk management.
The EU AI Act regulates these risks in Europe, but different approaches are taken elsewhere. The Center for Long-Term Cybersecurity at UC Berkeley has published a white paper titled "AI Risk-Management Standards Profile for General-Purpose AI (GPAI) and Foundation Models," which outlines risk-management practices and controls for identifying, analyzing, and mitigating risks associated with GPAI and foundation models.The recent development of integrating AI as agents and components in a system of systems has led to more unpredictable behavior due to the indeterministic and unexplainable outputs of LLMs. This development calls for emergent research to defend against possible attacks caused by malicious or abusive use of GAIs.
The main objective of AntiGAIAbuse is to combine the data engineering and policy expertise of UC Berkeley with the software engineering and technical security expertise of the Norwegian University of Science and Technology (NTNU) to tackle the risks of foundation model abuse in software systems that integrate LLM models as components. Through collaborative research between the Principal Investigators (PIs) and researchers at the Center for Long-Term Cybersecurity at UC Berkeley and the Department of Computer Science at NTNU, the project aims to strengthen existing and establish new, long-term collaborations between UC Berkeley and NTNU. This collaboration will contribute to managing GAI risks from both US and European perspectives and enhancing UC Berkeley's AI risk management initiatives from policy and technological approach perspectives.