Abstract :
The emergence of Generative Artificial Intelligence (AI) systems has expanded the boundaries of traditional model development and validation, introducing new dimensions of model risk. Existing Model Risk Management (MRM) standards such as SR 11-7 and SS1/23 remain foundational; however, their application must evolve to address the dynamic and context-dependent behaviour of Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) architectures, and multi-agent environments. These systems pose novel and high-impact risks due to their generative nature, contextual variability, and ability to self-orchestrate actions. Ensuring that such models produce consistent, auditable, and risk-mitigated outcomes has become critical, yet validators face growing challenges in assessing conceptual soundness, monitoring reasoning reliability, and evaluating control effectiveness within existing MRM frameworks.
This paper proposes GEN-5, a five-pillar validation and assurance framework that adapts established Model Risk Management (MRM) principles to the unique behaviours and risks of Generative AI, RAG pipelines, and multi-component AI systems. GEN-5 provides a standardized template and actionable methodology for assessing conceptual soundness, performance accuracy, outcome reliability, control effectiveness, and continuous monitoring across AI-driven environments. It integrates both qualitative and quantitative evaluation techniques—including hallucination detection, prompt robustness, retrieval fidelity, semantic consistency, and reasoning stability—while emphasizing the essential role of governance, safety guardrails, and control assurance. By extending traditional MRM rigor to modern AI architectures, GEN-5 offers practitioners a policy-aligned, technically grounded approach for identifying, evaluating, and mitigating the novel risks introduced by Generative and enterprise-scale AI use cases.
Keywords :
GenAI, Large Language Models (LLMs), Model Risk Management (MRM), Model Validation, RAGReferences :
- S. Federal Reserve and Office of the Comptroller of the Currency, Supervisory Guidance on Model Risk Management (SR 11-7), Washington, DC, Apr. 2011.
- Bank of England and Prudential Regulation Authority, SS1/23 – Model Risk Management Principles for Banks, London, UK, May 2023.
- Monetary Authority of Singapore, FEAT Principles and Model AI Governance Framework (Ver. 2), Singapore, Jan. 2023.
- National Institute of Standards and Technology (NIST), AI Risk Management Framework 1.0, Gaithersburg, MD, Jan. 2023.
- Organisation for Economic Co-operation and Development (OECD), OECD Principles on Artificial Intelligence, Paris, France, 2019.
- European Commission, Ethics Guidelines for Trustworthy AI, High-Level Expert Group on Artificial Intelligence, Brussels, Belgium, Apr. 2019.
- International Organization for Standardization (ISO) / International Electrotechnical Commission (IEC), ISO/IEC 42001:2023 – Artificial Intelligence – Management System, Geneva, Switzerland, Dec. 2023.
- Huang, M. Chang, and E. Chi, “A Survey on Hallucination in Large Language Models: Taxonomy and Mitigation ,2309.04843, 2023.
- Yu, S. Wang, and Y. Li, “Evaluating Hallucinations in Retrieval-Augmented Generation Systems,” arXiv preprint arXiv:2402.11247, 2024.
- Ribeiro, C. Wang, and K. Zhou, “Faithfulness and Factuality in Language Models: Metrics and Evaluation,” Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 2023.
- “HalluLens: A Benchmark for Detecting and Measuring Hallucinations in Generative Models,” Open-Source Toolkit Documentation, HalluLens Project, 2023.
- “RAGAs: Retrieval-Augmented Generation Assessment Suite,” GitHub Repository, 2024.
- “eRAG: Evaluation Framework for Retrieval-Augmented Generation Systems,” GitHub Repository, 2024.
- Deloitte LLP, Responsible and Reliable Generative AI: Extending Model Validation and Governance Practices, White Paper, London, UK, 2024.
- Wells Fargo, Responsible AI and Model Risk Management: Building Trust in Generative AI, Corporate Publication, San Francisco, CA, 2024.
- Citigroup Inc., Generative AI Risk and Governance Playbook, White Paper, New York, NY, 2024.
- PricewaterhouseCoopers, Operationalizing Model Validation for Generative AI Systems, Advisory Insight, London, UK, 2023.
- Barbera, Ai privacy risks & mitigations—large language models (llms), European Data Protection Board,2025
- OpenAI, “Best Practices for Red-Teaming Large Language Models,” Technical Report, OpenAI Inc., 2023.
- Raza, R. Sapkota, M. Karkee, and C. Emmanouilidis, “TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems,” arXiv preprint arXiv:2506.04133 [cs.AI], Sep. 2025
- Chen, Y. Liu, W. Han, W. Zhang, T. Liu, A survey on llm-based multi-agent system: Recent advances and new frontiers in application (2025)
- Biran and C. Cotton, “Explainability and Trust in Autonomous Systems,” IEEE Trans. on Human-Machine Systems, vol. 52, no. 5, pp. 801–813, 2022.
- Monetary Authority of Singapore, “Consultation Paper on Guidelines on Artificial Intelligence Risk Management (AIRG),” Singapore, 2025

