EPFL’s Machine Learning and Optimization Laboratory has teamed with the United Nations International Computing Centre to deliver a practical evaluation framework to assess the safety and reliability of AI systems.Announced as part of the United Nations’ AI for Good Global 2026 Summit in Geneva, EPFL’s Machine Learning and Optimization Laboratory (MLO) in the School of Computer and Communication Sciences and the United Nations International Computing Centre (UNICC) have unveiled a collaboration to advance responsible AI.The AI for Good Global Summit is the United Nations' flagship annual event in Geneva dedicated to leveraging artificial intelligence to solve global challenges and accelerate the UN Sustainable Development Goals (SDGs).Facilitated by the International Computation and AI Network, ICAIN, the new collaboration brings together EPFL's leading research capabilities in artificial intelligence and UNICC's operational expertise in digital foundations for the UN system, reflecting a shared commitment to advancing trustworthy AI through rigorous research, practical testing and knowledge sharing.The collaboration’s newly released white paper presents a structured approach to assessing the behaviour of Apertus, Switzerland’s fully open foundation model for sovereign AI, developed by the Swiss AI Initiative. Rather than focusing solely on the model’s technical performance, the research examines how an AI system behaves when deployed in a real institutional environment where accuracy, safety, and reliability are essential.“The project evaluated the system across multiple dimensions, including its resistance to harmful prompts, its ability to recognize when it should not answer, the accuracy of its responses when using external knowledge, and its susceptibility to biased or misleading interactions,” explained Associate Professor Martin Jaggi, Head of the MLO.Beyond the findings presented in the white paper, the collaboration developed a practical evaluation framework that the public-sector, the UN system, and other international organizations can adapt to assess the safety and reliability of their AI systems.“This collaboration shows what international organizations and academic institutions can achieve together,” says Anusha Dandapani, Chief of the UNICC AI Hub. “UNICC brings the operational context, while EPFL brings research rigour. Together, we have evaluated open-weight models in practice rather than in theory. That is what building institutional capacity for responsible AI looks like, and the knowledge generated extends well beyond this partnership.”As artificial intelligence becomes increasingly integrated into public sector operations, initiatives such as this demonstrate the importance of combining scientific research with operational experience to help ensure AI technologies remain safe, trustworthy and aligned with the public good.About the white paperThe white paper, Safety Evaluation of an Institutional LLM-RAG Deployment: A Four-Layer Audit of the UNICC Apertus System, documents the joint evaluation framework developed by UNICC and EPFL to assess the safety and robustness of institutional AI systems. In addition to presenting the findings, it introduces practical methodologies that can support future AI assurance efforts across the UN system and beyond.Read the full white paper here: https://www.unicc.org/resources/2026/07/07/safety-evaluation-of-an-institutional-llm-rag-deployment/