<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1d1 20130915//EN" "http://jats.nlm.nih.gov/publishing/1.1d1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" article-type="research-article" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">SAJEMS</journal-id>
<journal-title-group>
<journal-title>South African Journal of Economic and Management Sciences</journal-title>
</journal-title-group>
<issn pub-type="ppub">1015-8812</issn>
<issn pub-type="epub">2222-3436</issn>
<publisher>
<publisher-name>AOSIS</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">SAJEMS-27-5786</article-id>
<article-id pub-id-type="doi">10.4102/sajems.v27i1.5786</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Original Research</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Integrating traditional and non-traditional model risk frameworks in credit scoring</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-2730-4891</contrib-id>
<name>
<surname>du Toit</surname>
<given-names>Hendrik A.</given-names>
</name>
<xref ref-type="aff" rid="AF0001">1</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0001-7551-2009</contrib-id>
<name>
<surname>Schutte</surname>
<given-names>Willem D.</given-names>
</name>
<xref ref-type="aff" rid="AF0001">1</xref>
<xref ref-type="aff" rid="AF0002">2</xref>
</contrib>
<contrib contrib-type="author">
<contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-4126-1739</contrib-id>
<name>
<surname>Raubenheimer</surname>
<given-names>Helgard</given-names>
</name>
<xref ref-type="aff" rid="AF0001">1</xref>
<xref ref-type="aff" rid="AF0002">2</xref>
</contrib>
<aff id="AF0001"><label>1</label>Centre for Business Mathematics and Informatics, Faculty of Natural and Agricultural Sciences, North-West University, Potchefstroom, South Africa</aff>
<aff id="AF0002"><label>2</label>National Institute for Theoretical and Computational Sciences (NITheCS), Pretoria, South Africa</aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><bold>Corresponding author:</bold> Hendrik du Toit, <email xlink:href="drikus0329dutoit@gmail.com">drikus0329dutoit@gmail.com</email></corresp>
</author-notes>
<pub-date pub-type="epub"><day>08</day><month>10</month><year>2024</year></pub-date>
<pub-date pub-type="collection"><year>2024</year></pub-date>
<volume>27</volume>
<issue>1</issue>
<elocation-id>5786</elocation-id>
<history>
<date date-type="received"><day>10</day><month>06</month><year>2024</year></date>
<date date-type="accepted"><day>20</day><month>08</month><year>2024</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024. The Authors</copyright-statement>
<copyright-year>2024</copyright-year>
<license license-type="open-access" xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>Licensee: AOSIS. This work is licensed under the Creative Commons Attribution License.</license-p>
</license>
</permissions>
<abstract>
<sec id="st1">
<title>Background</title>
<p>An improved understanding of the reasoning behind model decisions can enhance the use of machine learning (ML) models in credit scoring. Although ML models are widely regarded as highly accurate, the use of these models in settings that require explanation of model decisions has been limited because of the lack of transparency. Especially in the banking sector, model risk frameworks frequently require a significant level of model interpretability.</p>
</sec>
<sec id="st2">
<title>Aim</title>
<p>The aim of the article is to evaluate traditional model risk frameworks to determine their appropriateness when validating ML models in credit scoring and enhance the use of ML models in regulated environments by introducing a ML interpretability technique in model validation frameworks.</p>
</sec>
<sec id="st3">
<title>Setting</title>
<p>The research considers model risk frameworks and regulatory guidelines from various international institutions.</p>
</sec>
<sec id="st4">
<title>Method</title>
<p>The research is qualitative in nature and shows how through integrating traditional and non-traditional model risk frameworks, the practitioner can leverage trusted techniques and extend traditional frameworks to address key principles such as transparency.</p>
</sec>
<sec id="st5">
<title>Results</title>
<p>The article proposes a model risk framework that utilises Shapley values to improve the explainability of ML models in credit scoring. Practical validation tests are proposed to enable transparency of model input variables in the validation process of ML models.</p>
</sec>
<sec id="st6">
<title>Conclusion</title>
<p>Our results show that one can formulate a comprehensive validation process by integrating traditional and non-traditional frameworks.</p>
</sec>
<sec id="st7">
<title>Contribution</title>
<p>This study contributes to existing model risk literature by proposing a new model validation framework that utilises Shapley values to explain ML model predictions in credit scoring.</p>
</sec>
</abstract>
<kwd-group>
<kwd>machine learning models</kwd>
<kwd>credit scoring</kwd>
<kwd>model risk frameworks</kwd>
<kwd>model interpretability</kwd>
<kwd>model validation</kwd>
<kwd>Shapley values</kwd>
<kwd>model transparency</kwd>
</kwd-group>
<funding-group>
<funding-statement><bold>Funding information</bold> The authors received no financial support for the research, authorship, and/or publication of this article.</funding-statement>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s0001">
<title>Introduction</title>
<p>The field of machine learning (ML) has gained a lot of popularity in recent years. Therefore, the importance of interpretable ML models in regulated environments such as the banking sector has increased significantly over the last decade. Machine learning<xref ref-type="fn" rid="FN0001"><sup>1</sup></xref> is a name for a group of models that are typically classified under the umbrella of artificial intelligence (AI). Although some of the techniques are not new, they have been supported by the advancements in computational power (Bertsimas, King &#x0026; Mazumder <xref ref-type="bibr" rid="CIT0005">2016</xref>). For this study, we do not aim to define ML. Machine learning models in this study refer to supervised classification algorithms such as tree-based models. Many banks and other financial institutions have seen the benefit of these so-called ML models and are implementing the necessary infrastructure to productionalise these models. Banks have benefited from traditional model methodologies such as logistic regression for over 30 years and developed a body of knowledge and rules that are used to judge the appropriateness of these models when used in practical applications.</p>
<p>The South African Reserve Bank (SARB) adopted the definition of model risk as defined in the Basel II market risk framework (SARB <xref ref-type="bibr" rid="CIT0028">2015</xref>). In this framework, two forms of model risk are identified. The first form of model risk has to do with an incorrect valuation methodology, and the second is unobservable (and possibly incorrect) calibration parameters in the valuation model.</p>
<p>The model risk as identified by SARB (<xref ref-type="bibr" rid="CIT0028">2015</xref>), is managed by banks using a function called model risk management. De Jongh et al. (<xref ref-type="bibr" rid="CIT0007">2017</xref>) state that this function usually comprises robust and sensible model development, sound implementation procedures, appropriate use of models and consistent model validation. These measures exist to ensure that model risk can be measured and mitigated appropriately. Traditional model risk frameworks refer to model risk frameworks that typically already exist in model development teams and are being used to evaluate traditional models such as logistic regression. Non-traditional model risk frameworks are model risk frameworks that have been proposed in recent literature to evaluate non-traditional models such as ML algorithms.</p>
<p>Model validation is the set of processes and activities intended to verify that models perform as expected, in line with their design objectives and business uses (OCC <xref ref-type="bibr" rid="CIT0025">2011</xref>). The Office of the Comptroller of the Currency (OCC) (<xref ref-type="bibr" rid="CIT0025">2011</xref>) report further describes that effective model validation techniques should be able to test model soundness, identify potential limitations and test certain model assumptions. The model validation process includes testing the model&#x2019;s accuracy, testing if the model data is stable over time, determining if the model can differentiate between events and non-events in the data and evaluating if the variables used in the model are intuitive. When referring to the validation of ML models, it relates to best practice in the application of ML models in real-world production environments.</p>
<p>Van der Burgt (<xref ref-type="bibr" rid="CIT0033">2019</xref>) notes that the financial sector is particularly important from an ML development perspective and needs an adequate regulatory and supervisory response. The reasons given are:</p>
<list list-type="bullet">
<list-item><p>The financial sector is commonly held to a higher social standard than many other industries and AI-related incidents can have serious reputation effects.</p></list-item>
<list-item><p>Incidents could seriously impact financial stability, given that the financial system is interconnected in many ways (i.e., systemic risk).</p></list-item>
<list-item><p>The progress of AI and the increase in the importance of these models in the financial sector directs us to rethink traditional supervisory frameworks.</p></list-item>
</list>
<p><xref ref-type="fig" rid="F0001">Figure 1</xref> shows the layers pertaining to model risk as discussed so far. It highlights the role of validation as an important role in the model risk framework and proposes ML interpretability techniques as an additional layer to model risk management.</p>
<fig id="F0001">
<label>FIGURE 1</label>
<caption><p>The layers of model risk.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="SAJEMS-27-5786-g001.tif"/>
</fig>
<p>The use of ML in credit scoring has been proven to be very efficient and financial institutions are exploring different ways of leveraging the accuracy of ML models (Lessmann et al. <xref ref-type="bibr" rid="CIT0018">2015</xref>). With this new domain of models entering specialist areas such as credit scoring, certain questions come to mind about how these models can be validated and proven sound. These questions, among others, have led to a new field of research called ML interpretability techniques. Machine learning interpretability can be described as the set of techniques that can be used to explain a so-called &#x2018;black-box&#x2019; model. There are many different techniques, and each of them aims to describe a piece of the model, or in some cases, it attempts to describe the model entirely.</p>
<p>The European Banking Authority (<xref ref-type="bibr" rid="CIT0010">2021</xref>) notes that the main challenges regarding ML models come from the complexity of the model, which leads to difficulties in interpreting the results, ensuring management functions adequately understand them, and, lastly, justifying their results to supervisors. The <italic>EU Artificial Intelligence Act (EU AI Act)</italic> was recently published (European Commission <xref ref-type="bibr" rid="CIT0011">2024</xref>). This act includes a classification of AI applications in terms of risk. According to the act, there are four categories of AI applications:</p>
<list list-type="bullet">
<list-item><p>prohibited applications</p></list-item>
<list-item><p>high-risk applications</p></list-item>
<list-item><p>applications with special requirements</p></list-item>
<list-item><p>low risk applications.</p></list-item>
</list>
<p>Credit scoring is classified as a high risk application and must fulfil comprehensive requirements in areas such as transparency. The Institute of International Finance (IIF) and Ernst &#x0026; Young (EY) (<xref ref-type="bibr" rid="CIT0014">2022</xref>) survey report on ML uses in credit risk and anti-money laundering applications notes that some of the key challenges in the adoption of ML included data quality, explainability and IT-infrastructure. The survey further found that existing model risk management frameworks often govern ML applications. Because of the challenge of approval, frameworks need to be created to enable a smoother road towards the implementation of ML models. As such, traditional frameworks are first inspected in the &#x2018;Overview of traditional model validation process&#x2019; section. The section titled Machine Learning Model Risk Management consults recent literature on how model risk management is conducted for ML models. Because of a lack of practical examples, this section further describes the principles for using ML models in the financial sector as proposed by various authors. The section concludes by gathering recent literature on how existing model risk frameworks can be expanded to incorporate ML models. Shapley values are introduced and explained as an ML interpretability technique in the &#x2018;Methods&#x2019; section. This technique is proposed as a model validation technique for ML models. The section titled &#x2018;Case study&#x2019; explains the integration of Shapley values into model validation frameworks to validate ML models. Practical tests are proposed within the &#x2018;Results&#x2019; section to illustrate how Shapley values can be used to validate ML models. The study concludes in the last section and proposes areas for further research.</p>
<sec id="s20002">
<title>Overview of traditional model validation process</title>
<p>The model validation process is considered a critical step in model risk management. Therefore, regulatory authorities such as the Basel Accord have attempted to set the standard for the model validation process. However, according to our knowledge, research has not been able to establish a definite set of global standards for this process, and not a lot of focus has been placed on providing examples of how these standards could be achieved. In this section, we will focus on a few key resources that encapsulate the core aspects of a traditional model validation framework.</p>
<p>Quell et al. (<xref ref-type="bibr" rid="CIT0027">2021</xref>) list the following typical aspects of a model that should be validated and how this should be achieved:</p>
<list list-type="bullet">
<list-item><p>Model data:
<list list-type="simple">
<list-item><label>&#x25A0;</label><p>data representativeness</p></list-item>
<list-item><label>&#x25A0;</label><p>data traceability and data quality</p></list-item>
<list-item><label>&#x25A0;</label><p>feature engineering</p></list-item>
<list-item><label>&#x25A0;</label><p>other exploratory data analysis techniques.</p></list-item>
</list></p></list-item>
<list-item><p>Conceptual soundness:
<list list-type="simple">
<list-item><label>&#x25A0;</label><p>model design and algorithm selection</p></list-item>
<list-item><label>&#x25A0;</label><p>model assumptions and limitations</p></list-item>
<list-item><label>&#x25A0;</label><p>dynamic learning</p></list-item>
<list-item><label>&#x25A0;</label><p>explainability and interpretability of the model</p></list-item>
<list-item><label>&#x25A0;</label><p>overfitting and bias.</p></list-item>
</list></p></list-item>
<list-item><p>Model implementation and ongoing validation.</p></list-item>
<list-item><p>Model documentation and use.</p></list-item>
</list>
<p>De Jongh et al. (<xref ref-type="bibr" rid="CIT0007">2017</xref>) propose a similar model validation framework (see <xref ref-type="fig" rid="F0002">Figure 2</xref>). This framework consists of validation governance, validation process and validation policy. Specifically, the model validation process consists of:</p>
<list list-type="bullet">
<list-item><p>conceptual soundness and developmental evidence</p></list-item>
<list-item><p>process verification and ongoing monitoring</p></list-item>
<list-item><p>outcomes analysis.</p></list-item>
</list>
<fig id="F0002">
<label>FIGURE 2</label>
<caption><p>Traditional model validation process as proposed by De Jongh et al. (<xref ref-type="bibr" rid="CIT0007">2017</xref>).</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="SAJEMS-27-5786-g002.tif"/>
</fig>
<p>Both the model validation frameworks of Quell et al. (<xref ref-type="bibr" rid="CIT0027">2021</xref>) and De Jongh et al. (<xref ref-type="bibr" rid="CIT0007">2017</xref>) highlight conceptual soundness as a critical aspect of the model validation framework. Within conceptual soundness, explainability and interpretability of the model prediction are considered important validation checks. Other authors, such as Abrahams and Zhang (<xref ref-type="bibr" rid="CIT0001">2008</xref>), describe model validation by identifying areas and components of importance (see <xref ref-type="table" rid="T0001">Table 1</xref>).</p>
<table-wrap id="T0001">
<label>TABLE 1</label>
<caption><p>Model validation areas and components.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Validation area</th>
<th valign="top" align="left">Validation components</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Inputs</td>
<td align="left"><list list-type="bullet">
<list-item><p>Input assumptions</p></list-item>
<list-item><p>Input data</p></list-item>
<list-item><p>Lending policies and practices</p></list-item>
</list></td>
</tr>
<tr>
<td align="left">Process</td>
<td align="left"><list list-type="bullet">
<list-item><p>Model development</p></list-item>
<list-item><p>Model selection</p></list-item>
<list-item><p>Model implementation</p></list-item>
</list></td>
</tr>
<tr>
<td align="left">Output</td>
<td align="left"><list list-type="bullet">
<list-item><p>Model results interpretation</p></list-item>
<list-item><p>Holdout sample testing</p></list-item>
<list-item><p>Performance monitoring and testing</p></list-item>
</list></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>Source:</italic> Abrahams, C.R. &#x0026; Zhang, M., 2008, <italic>Fair lending compliance: Intelligence and implications for credit risk management</italic>, Wiley, Hoboken, NJ</p></fn>
</table-wrap-foot>
</table-wrap>
<p>Baesens, Roesch and Scheule (<xref ref-type="bibr" rid="CIT0003">2016</xref>) explain that one can quantitatively validate a model by comparing realised numbers to predicted numbers. These numbers will rarely be identical, and therefore, appropriate performance metrics and test statistics should be specified to conduct the comparison. The authors describe the process of back testing models to check data stability. This process determines if the population on which the model has been developed is representative of the population being observed. Back testing is also used to determine if the ranking of the model predictions is comparable to the event that the model is predicting, in other words, comparing predicted events to actual events. The second type of quantitative validation method described by the authors is benchmarking. The method involves comparing the output and performance of the model that is being validated with a reference model, otherwise known as a benchmark. Examples of benchmark models are credit bureaus, rating agencies, existing models and even expert models. Where a relevant benchmark is not available, other regulatory authorities, such as Hong Kong Monetary Authority (HKMA <xref ref-type="bibr" rid="CIT0013">2006</xref>), have suggested the development of an internal benchmark or an expert based benchmark.</p>
<p>Authors such as Abrahams and Zhang (<xref ref-type="bibr" rid="CIT0001">2008</xref>) have proposed typical statistical measures and their applications in the model validation process. These include measures such as the Kolmogorov-Smirnov (K-S) test, Receiver Operating Characteristic (ROC) curve, Gini coefficient, cumulative gains chart and the Chi-square statistic. These measures are used to measure model performance and population shift, and to analyse model input.</p>
<p>Siddiqi (<xref ref-type="bibr" rid="CIT0032">2017</xref>) lists resources that are used as guidelines for model validation under the umbrella of model risk management. These include resources such as supervisory guidance on model risk management (OCC <xref ref-type="bibr" rid="CIT0025">2011</xref>), GL-44 guidelines on internal governance issues (EBA <xref ref-type="bibr" rid="CIT0009">2011</xref>) and the Basel Committee on Banking Supervision Working Paper 14 (BCBS <xref ref-type="bibr" rid="CIT0004">2005</xref>), to name a few. Although there are many suggested methods to validate models, both Siddiqi (<xref ref-type="bibr" rid="CIT0032">2017</xref>) and De Jongh et al. (<xref ref-type="bibr" rid="CIT0007">2017</xref>) specifically note that there is no single set of global standards to validate models.</p>
</sec>
<sec id="s20003">
<title>Machine learning model risk management</title>
<p>To define the need for a revised model risk management framework, it is important to understand the dangers of ML models and what risks need to be mitigated. To this end, Quell et al. (<xref ref-type="bibr" rid="CIT0027">2021</xref>) listed a few common dangers of ML models. These dangers include explainability, where they note that models such as &#x2018;black-box&#x2019; models are difficult to interpret. This is significant because interpretation is often required for financial models, especially if they must adhere to external regulatory requirements. Many researchers have attempted to provide a comprehensive definition of interpretability. Miller (<xref ref-type="bibr" rid="CIT0021">2017</xref>) explains that interpretability can be understood as the degree to which a human can understand the reason for a certain decision. Furthermore, Kim, Khanna and Koyejo (<xref ref-type="bibr" rid="CIT0015">2016</xref>) define it as the degree to which a human can consistently predict a model&#x2019;s result.</p>
<p>Although this research will focus on interpretability as a key area of concern, the interested reader can consult Quell et al. (<xref ref-type="bibr" rid="CIT0027">2021</xref>) for a complete list of the typical dangers of ML models. Other dangers of the application of ML models include:</p>
<list list-type="bullet">
<list-item><p>overfitting</p></list-item>
<list-item><p>robustness and population drift</p></list-item>
<list-item><p>bias, adversarial attacks, and brittleness</p></list-item>
<list-item><p>development bias</p></list-item>
<list-item><p><italic>p</italic>-value arbitrage.</p></list-item>
</list>
<p>Recent literature on model risk management for ML models proposes principles that model risk frameworks need to adhere to, with less focus on proposing practical techniques that can be considered. This article aims to bridge this gap and provide practical techniques that a typical credit scoring team can implement within a credit scoring model validation framework. As such, this section aims to investigate these principles and gain an understanding of how they can be used as validation criteria within a model risk framework. To narrow the scope of this research, a clear focus is placed on a ML governance principle, namely transparency. The below-stated review solidifies what the concept of transparency means and how non-traditional model risk frameworks intend to ensure that the principle of transparency is met.</p>
<p>General principles for using AI in the financial sector have been proposed in recent literature. The principles provide a broad overview of what model risk management teams should keep in mind for the use of ML models in finance. H&#x00E4;rle et al. (<xref ref-type="bibr" rid="CIT0012">2015</xref>) identified six structural trends that will transform bank risk management in the future. One of these trends is the continuous expansion of the breadth and depth of regulation. Furthermore, the same authors note that compliance with existing rules will likely not be sufficient in the future and that banks will need to comply with broad principles to protect themselves against potential future rules and interpretations of existing rules.</p>
<p>Van der Burgt (<xref ref-type="bibr" rid="CIT0033">2019</xref>) introduces general principles for a responsible application of AI in the financial sector. The author notes the following six key principles: soundness, accountability, fairness, ethics, skills and transparency. These are known as the &#x2018;SAFEST&#x2019; principles. The principle of transparency is particularly interesting as it states that firms should be able to explain how and why they use AI in their business processes and how these applications function.</p>
<p>The Monetary Authority of Singapore (MAS <xref ref-type="bibr" rid="CIT0024">2018</xref>) outlines similar principles to aid the use of AI and data analytics in Singapore&#x2019;s financial sector. The MAS (<xref ref-type="bibr" rid="CIT0024">2018</xref>) aims to provide the financial sector with a set of foundational principles to consider when using AI and data analytics in decision making. It also aims to assist companies in contextualising and operationalising governance of the use of AI and data analytics. Lastly, it aims to promote public confidence and trust in the use of AI and data analytics. The article introduces the principles of Fairness, Ethics, Accountability and Transparency (FEAT). The MAS (<xref ref-type="bibr" rid="CIT0024">2018</xref>) clearly states that these principles are not intended to replace existing relevant internal governance frameworks and that companies should continue to comply with all applicable laws and requirements. Although these principles are not intended to be prescriptive, the authors believe that through industry engagement, there might be areas where more specific or technical guidance would benefit the industry, with the FEAT principles serving as a foundational framework. In summarising the transparency principle, the article explicitly recommends that data subjects are provided, upon request, with clear explanations on what data is used to make an AI and data analytics decision about the data subject and how the data affects the decision.</p>
<p>PricewaterhouseCoopers (PWC) released a report describing the practical implications of the FEAT principles on industries such as banking and insurance (PWC <xref ref-type="bibr" rid="CIT0026">2018</xref>). It advises banks to prepare processes and tools to deal with client requests for explanations. More specifically, the report requests banks to provide information on the input factors and the potential output scenarios of an ML model using non-technical language without revealing the underlying intellectual property. The report further encourages banks to evaluate explanatory techniques, especially for deep learning models.</p>
<p>Shifting the focus slightly to models relevant to the regulatory capital space. The European Banking Authority (EBA) (<xref ref-type="bibr" rid="CIT0010">2021</xref>) released a discussion article that aims to understand the challenges and opportunities of applying ML models in the context of internal ratings-based (IRB) models to calculate regulatory capital for credit risk. The article proposes a set of principle-based recommendations based on trust elements:</p>
<list list-type="bullet">
<list-item><p>ethics</p></list-item>
<list-item><p>explainability and interpretability</p></list-item>
<list-item><p>traceability and auditability</p></list-item>
<list-item><p>fairness and bias prevention</p></list-item>
<list-item><p>data protection and quality</p></list-item>
<list-item><p>consumer protection aspects and security.</p></list-item>
</list>
<p>The EBA (<xref ref-type="bibr" rid="CIT0010">2021</xref>) lists interpretability as one of the concerns of using ML models. Together with the concerns, the article also lists some of the techniques frequently used to obtain insight into the internal logic of an ML model. More specifically, for interpretability techniques, the following are listed by the EBA (<xref ref-type="bibr" rid="CIT0010">2021</xref>):</p>
<list list-type="bullet">
<list-item><p>Graphical tools such as partial dependence plots (PDP), and individual conditional expectation plots (ICE). These plots are designed to show the effect of an explanatory variable on the model.</p></list-item>
<list-item><p>Feature importance measures the relevance of variables in the overall model.</p></list-item>
<list-item><p>Shapley values quantify the impact of a variable on the final prediction.</p></list-item>
<list-item><p>Local explanations such as Local Interpretable Model-Agnostic Explanations (LIME) and anchors give a simplified explanation of the model from a local point of view.</p></list-item>
<list-item><p>Counterfactual explanations show how changing the input variables can influence a model&#x2019;s prediction.</p></list-item>
</list>
<p>Based on the overview of the general principles for using AI in the financial sector, the next section will focus on seven key pillars that outline the changes required to integrate AI and ML models into existing model risk management frameworks. These changes are a first step towards identifying the practical modifications needed to integrate AI/ML models in existing model risk management frameworks.</p>
</sec>
<sec id="s20004">
<title>Integrating machine learning models in existing model risk management frameworks</title>
<p>KPMG (<xref ref-type="bibr" rid="CIT0016">2022</xref>) suggests that model risk management for AI/ML models can be integrated into existing (traditional) model risk management frameworks. In doing this, industries can benefit from synergies that arise from using proven processes and methods. Integrating the new types of models into existing frameworks addresses many regulatory requirements as listed in the <italic>EU AI Act</italic>, but minor changes need to be made to address AI/ML models specifically. Considering this, the white article suggests seven key pillars that outline the changes needed to integrate AI/ML models into existing frameworks. These seven pillars are listed in <xref ref-type="table" rid="T0002">Table 2</xref>.</p>
<table-wrap id="T0002">
<label>TABLE 2</label>
<caption><p>Key model risk management pillars as proposed by KPMG (<xref ref-type="bibr" rid="CIT0016">2022</xref>).</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Key pillar</th>
<th valign="top" align="left">Description</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1. Establish a definition of AI/ML models.</td>
<td align="left"><list list-type="bullet">
<list-item><p>Banks establish an enterprise-wide definition of what AI/ML models comprise. Expand model inventory to include these models.</p></list-item>
</list></td>
</tr>
<tr>
<td align="left">2. Updating the model tiering definition.</td>
<td align="left"><list list-type="bullet">
<list-item><p>Model tiering parameters, specifically around materiality, criticality and uncertainty, need to be revised to correctly incorporate the risk of AI and ML models.</p></list-item>
</list></td>
</tr>
<tr>
<td align="left">3. Establish an appropriate risk appetite.</td>
<td align="left"><list list-type="bullet">
<list-item><p>Banks need to leverage peer networks to establish a first draft of a risk appetite statement and appropriate thresholds. This is necessary because traditional risk appetite statements are not designed for ML models and regulatory guidance on this matter is still being developed.</p></list-item>
</list></td>
</tr>
<tr>
<td align="left">4. Identify accountability.</td>
<td align="left"><list list-type="bullet">
<list-item><p>Establish clear definitions of roles and accountabilities within all the risk management functions.</p></list-item>
</list></td>
</tr>
<tr>
<td align="left">5. Invest in skill enhancements.</td>
<td align="left"><list list-type="bullet">
<list-item><p>Develop a skill set inhouse or involve external parties. External subject matter experts (SMEs) can help benchmark banks with the latest risk management, controls and techniques for model validation using AI/ML models.</p></list-item>
</list></td>
</tr>
<tr>
<td align="left">6. Enhance the compensatory control framework.</td>
<td align="left"><list list-type="bullet">
<list-item><p>Designing additional compensatory controls around areas such as benchmarking, feature selection, bias elimination, among others, to account for the lack of transparency.</p></list-item>
<list-item><p>Enhancing existing data management framework.</p></list-item>
<list-item><p>Building control frameworks around compliance and operational risk.</p></list-item>
<list-item><p>Conducting enterprise-wide training programmes.</p></list-item>
</list></td>
</tr>
<tr>
<td align="left">7. Develop additional tests and procedures for AI/ML models.</td>
<td align="left"><list list-type="bullet">
<list-item><p>Interpretability</p></list-item>
<list-item><p>Bias elimination</p></list-item>
<list-item><p>Dynamic calibration of models</p></list-item>
<list-item><p>Implementation</p></list-item>
<list-item><p>Ongoing monitoring.</p></list-item>
</list></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p>AI, artificial intelligence; ML, machine learning; SME, subject matter expert.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>The general principles of using ML in finance establish a clear directory of what regulatory bodies consider important and relevant when using ML models. Finally, this section introduces seven pillars that need to be adopted to incorporate ML models into existing model risk management frameworks. The following section builds on these pillars by proposing an additional test that can be linked to pillar 7 (see <xref ref-type="table" rid="T0002">Table 2</xref>). This test specifically focusses on interpretability as a requirement in the model validation process. Furthermore, because Shapley values will form an integral part of the tests proposed, the next section will uncover some details about Shapley values.</p>
</sec>
</sec>
<sec id="s0005">
<title>Methods</title>
<sec id="s20006">
<title>A brief overview of the Shapley value</title>
<p>This section introduces the Shapley value as an ML interpretability technique. It provides an overview of the technique and explains the interpretation of the Shapley value with a simple example. The section concludes with a case study that illustrates how a model development team can use the technique to validate the interpretability of ML models.</p>
</sec>
<sec id="s20007">
<title>Shapley value</title>
<p>Considering the different ML interpretability techniques listed in the section titled: Machine learning model risk management, the best method of selecting ML interpretability techniques is considering the characteristics of each available technique. Arrieta et al. (<xref ref-type="bibr" rid="CIT0002">2020</xref>) provide an extensive overview of concepts, taxonomies, opportunities and challenges with respect to the proper use of AI. In light of this, the article proposes the use of model-agnostic post-hoc explainability techniques. Model-agnostic techniques can be applied to any ML model, while post-hoc explainability techniques are used to explain the inner workings of an already developed model which is not intrinsically interpretable. Model-agnostic post-hoc interpretability techniques contain feature relevance explanation techniques. These techniques measure the influence, relevance or importance of variables. Shapley additive explanations form part of this group of interpretability techniques and have been proposed by authors such as Lundberg and Lee (<xref ref-type="bibr" rid="CIT0020">2017</xref>) and Chen, Lundberg and Lee (<xref ref-type="bibr" rid="CIT0006">2021</xref>).</p>
<p>Considering the more recent integration of ML models into traditional model validation frameworks, Shapley values have been mentioned as an interpretability technique to be considered. See, for example, EBA (<xref ref-type="bibr" rid="CIT0010">2021</xref>) and Scheda and Diciotti (<xref ref-type="bibr" rid="CIT0029">2022</xref>). Although the technique is not without its weaknesses (as is evidenced by the shortcomings listed in Molnar et al. <xref ref-type="bibr" rid="CIT0023">2020</xref> and Woznica et al. <xref ref-type="bibr" rid="CIT0034">2021</xref>), many researchers, such as Du Toit et al. (<xref ref-type="bibr" rid="CIT0008">2023</xref>) consider this technique to be especially applicable for interpreting ML models in a model validation process.</p>
<p>Shapley values originate from game theory (Shapley <xref ref-type="bibr" rid="CIT0030">1953</xref>). Shapley values show the impact of a specific predictor variable on the model outcome. The interested reader can consult sources such as Du Toit et al. (<xref ref-type="bibr" rid="CIT0008">2023</xref>) and Kumar et al. (<xref ref-type="bibr" rid="CIT0017">2020</xref>) for a detailed explanation of how the Shapley value is calculated. SHapley Additive exPlanation (SHAP) is also proposed by Lundberg and Lee (<xref ref-type="bibr" rid="CIT0020">2017</xref>). Considering how expensive the calculation of the Shapley value is, two estimation approaches to calculate SHAP values were introduced by Lundberg and Lee (<xref ref-type="bibr" rid="CIT0020">2017</xref>). The first is KernelSHAP, a kernel-based estimation method inspired by local surrogate models. The second is TreeSHAP, an estimation method for tree-based models. The interested reader can consult Molnar (<xref ref-type="bibr" rid="CIT0022">2020</xref>) and Lundberg, Erion and Lee (<xref ref-type="bibr" rid="CIT0019">2018</xref>) for more details on the estimation techniques.</p>
</sec>
<sec id="s20008">
<title>Shapley value explanation</title>
<p>The following simplified example, as used by Du Toit et al. (<xref ref-type="bibr" rid="CIT0008">2023</xref>) illustrates how the Shapley value is applied. In <xref ref-type="fig" rid="F0003">Figure 3</xref>, the Shapley value ranges from &#x2013;0.2 to 0.25 and for this explanation, the value is based on a model that predicts the probability of default. The Shapley value is the marginal increase or decrease in the probability of default contributed by a certain variable entering the model for a specific application. This contribution is added or subtracted from the average model prediction.</p>
<fig id="F0003">
<label>FIGURE 3</label>
<caption><p>Shapley value explanation.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="SAJEMS-27-5786-g003.tif"/>
</fig>
<p>To illustrate, assume the average probability of default for the model is 25&#x0025;. For the variable depicted in <xref ref-type="fig" rid="F0003">Figure 3</xref>, if the specific applicant had a variable value in bin 1, then the prediction, for instance, would be calculated as 25&#x0025; + (-20&#x0025;) = 5&#x0025;. Conversely, if the variable value was in bin 5, we would not expect the variable to influence the average prediction (i.e., 25&#x0025;) because the Shapley value in bin 5 is equal to 0. From this example, we can infer that the higher the variable bin, the higher the Shapley value for the specific instance and thus, the probability of default increases.</p>
<p>The research conducted by Du Toit et al. (<xref ref-type="bibr" rid="CIT0008">2023</xref>) evaluates Shapley values as an ML interpretability technique for credit scoring models. The authors test the technique on simulated data generated from various underlying distributions representing real world credit variables.</p>
<p>The Shapley values are generated after fitting both traditional models and ML models. The authors compared the resulting Shapley value to a well-known measure called the Weights of Evidence (WOE) value to evaluate the technique. The interested reader can consult Siddiqi (<xref ref-type="bibr" rid="CIT0031">2006</xref>) for more information on this measure.</p>
<p>These two metrics are compared using Spearman&#x2019;s correlation and the mean squared error between the standardised values of the two metrics. The results show that the Shapley value can explain ML models similarly to the WOE value. The study shows that the Shapley values represent the relationships and interactions simulated in the data. The study encourages model validation teams to use the technique to explore acceptable thresholds for the Shapley value explanation.</p>
</sec>
<sec id="s20009">
<title>Integration of Shapley values into model validation frameworks for machine learning models</title>
<p>To illustrate that a new technique can be integrated into an existing model validation framework, we refer to the model validation process as proposed by De Jongh et al. (<xref ref-type="bibr" rid="CIT0007">2017</xref>). Recall that the three distinct elements in the model validation process were:</p>
<list list-type="bullet">
<list-item><p>conceptual soundness and developmental evidence</p></list-item>
<list-item><p>process verification and ongoing monitoring</p></list-item>
<list-item><p>outcomes analysis.</p></list-item>
</list>
<p>The authors propose a model validation process scorecard to determine if the best practice model validation framework has been adequately assembled and implemented. The idea is to rate the overall scorecard a point out of 4, where 1 represents no evidence and 4 indicates full evidence. The authors note that the validation process scorecard consists of seven elements, as <xref ref-type="table" rid="T0001">Table 1</xref>-A1 indicates. This generic scorecard will require changes depending on the product, the institution and the phase at which the process is being performed, that is development, implementation or monitoring. To this end, the following case study shows how a traditional model validation framework can be retained and enhanced to cater for ML models. Note that the case study is aimed to provide model validation teams with a methodology of validating ML models in credit scoring. The data used in this section are hypothetical and assume a model has been selected and trained on a set of data. The tests proposed are based on hypothetical data to demonstrate how this methodology can be followed. It is proposed that a practitioner uses their own model development data or model training data and validation/testing data set to conduct the proposed tests.</p>
</sec>
</sec>
<sec id="s0010">
<title>Case study</title>
<p>The model validation team of Bank ABC is tasked to validate a credit scoring model that will be used to determine the probability of default of prospective clients. The model selection and development process has been completed, and the model development team decided to implement a Random Forest<xref ref-type="fn" rid="FN0002"><sup>2</sup></xref> model based on an extensive list of selection criteria. The team identified transparency as one of the key principles that need to be adhered to when implementing ML models in the financial sector. For the purpose of this case study, the focus of the validation will be on formulating tests that will prove the transparency of the ML model. All other typical validation steps, such as testing population stability and testing accuracy, are considered out of scope for the particular task at hand. The team can use the existing model validation framework where applicable.</p>
<sec id="s20011">
<title>Proposed steps to integrate transparency tests into existing model validation framework</title>
<p>A possible solution to the following case study could be to utilise an existing model validation framework and integrate new research on the best practices for model validation processes in ML models. To illustrate this concept, the traditional model validation process proposed by De Jongh et al. (<xref ref-type="bibr" rid="CIT0007">2017</xref>) will be used as a starting point.</p>
<p>The following steps outline the proposed process of incorporating new model validation steps into existing model validation frameworks in a practical manner:</p>
<list list-type="bullet">
<list-item><p>Identify the general principle(s) for the use of AI in the financial sector, which needs to be incorporated into the model validation process: This could be one or more principles depending on the state of the existing model validation process. In some cases, there might not be an existing model validation process available. For this specific use case, transparency is selected as the principle to incorporate.</p></list-item>
<list-item><p>Identify and describe the existing model validation process:
<list list-type="simple">
<list-item><label>&#x25A0;</label><p>Identify all the elements of the current model validation process and the typical criteria that relate to the principle as identified in step 1.</p></list-item>
<list-item><label>&#x25A0;</label><p>In some cases, there might not be an existing model validation process available. In these cases, a high level process should be established that specifies the most notable validation elements that need to be tested.</p></list-item>
<list-item><label>&#x25A0;</label><p>Typical validation elements within a model validation process include (De Jongh et al. <xref ref-type="bibr" rid="CIT0007">2017</xref>):
<list list-type="simple">
<list-item><label>&#x25E6;</label><p>understanding and evaluating the model paradigm</p></list-item>
<list-item><label>&#x25E6;</label><p>ensuring model methods/theory is based on sound assumptions</p></list-item>
<list-item><label>&#x25E6;</label><p>determining if the model design is appropriate</p></list-item>
<list-item><label>&#x25E6;</label><p>testing data/variables used in the model</p></list-item>
<list-item><label>&#x25E6;</label><p>evaluating algorithms and codes used to develop the model</p></list-item>
<list-item><label>&#x25E6;</label><p>understanding output generated by the model</p></list-item>
<list-item><label>&#x25E6;</label><p>assessing how the model will be monitored.</p></list-item>
</list></p></list-item>
</list></p></list-item>
<list-item><p>Determine which criteria within the existing model validation process can be used to evaluate the chosen ML model risk principle. Expand on this criteria where necessary:
<list list-type="simple">
<list-item><label>&#x25A0;</label><p>As highlighted in the &#x2018;Overview of traditional model validation process&#x2019;, Miller (<xref ref-type="bibr" rid="CIT0021">2017</xref>) explains that interpretability can be understood as the degree to which a human can understand the reason for a certain decision. Furthermore, Kim et al. (<xref ref-type="bibr" rid="CIT0015">2016</xref>) define it as the degree to which a human can consistently predict a model&#x2019;s result. To showcase Shapley values, the focus will be on developing criteria for understanding model output, one of the main elements of a typical model validation process.</p></list-item>
<list-item><label>&#x25A0;</label><p>The following criteria, as proposed by De Jongh et al. (<xref ref-type="bibr" rid="CIT0007">2017</xref>), can be used to evaluate the transparency of a model within the model output element:
<list list-type="simple">
<list-item><label>&#x25E6;</label><p>was model output benchmarked against best practice models (e.g., against a vendor model using the same input data set)?</p></list-item>
<list-item><label>&#x25E6;</label><p>was the reasonableness and validity of model outputs assessed?</p></list-item>
<list-item><label>&#x25E6;</label><p>has a comparison of model outputs against actual realisations been performed? (Commonly referred to as &#x2018;back testing&#x2019;.)</p></list-item>
<list-item><label>&#x25E6;</label><p>has a range of outputs been examined versus a range of inputs &#x2013; are solutions continuous or jagged? What is the behaviour of hedging quantities and/or derived quantities over the same range?</p></list-item>
<list-item><label>&#x25E6;</label><p>are all results repeatable? (e.g., Monte Carlo simulations)</p></list-item>
</list></p></list-item>
<list-item><label>&#x25A0;</label><p>Additional criteria that can be added to this list are:
<list list-type="simple">
<list-item><label>&#x25E6;</label><p>can the model prediction be explained to stakeholders on a global level?</p></list-item>
<list-item><label>&#x25E6;</label><p>can the model prediction be explained on a local level for a specific instance?</p></list-item>
<list-item><label>&#x25E6;</label><p>what is the certainty to which the model prediction can be interpreted?</p></list-item>
<list-item><label>&#x25E6;</label><p>how much variability is present in the model output over time?</p></list-item>
<list-item><label>&#x25E6;</label><p>does the model output make logical sense when compared to assumed outcomes gained from business expertise and experience developing similar models?</p></list-item>
</list></p></list-item>
</list></p></list-item>
</list>
<p>In this section, certain key criteria used to test model transparency are obtained from an existing model validation process. The focus is placed on model output as a key validation element to evaluate the transparency of a model. The next section proposes practical tests using Shapley values that model practitioners can perform to evaluate the transparency of ML models in credit scoring.</p>
</sec>
<sec id="s20012">
<title>Ethical considerations</title>
<p>Ethical approval to conduct this study was obtained from the Faculty of Natural and Agricultural Sciences Ethics Committee (FNASREC), North-West University (NWU&#x2013;01251&#x2013;23&#x2013;A9).</p>
</sec>
</sec>
<sec id="s0013">
<title>Results</title>
<sec id="s20014">
<title>Shapley value tests</title>
<p>The previous section assessed the existing model validation process and identified validation elements that need to be considered when evaluating transparency as a key ML model risk principle. The next step of the process suggests tests that will provide evidence to meet the highlighted criteria of transparency. Before we propose the various tests, it is important to note that these tests are performed in the model validation process. Although Shapley values can be used at various stages in the model development process, these tests aim to provide transparency to model predictions by explaining the model output. These tests assume that the variable selection, model selection and training, model parameter tuning and basic performance checks have been completed. The trained model is used to generate the Shapley values, which are used to perform the various tests proposed in the next section.</p>
</sec>
<sec id="s20015">
<title>Proposed structure of Shapley value test</title>
<p>The following tests are designed in the format of a report to ensure that they can be included in the model documentation. The tests illustrate through certain calculations and graphs how Shapley values can be used to prove that the model output is logical, accurate, stable and, therefore, transparent and trustworthy. The first three tests (Tests 1&#x2013;3) are variable-specific and are designed to analyse variable characteristics. This is similar to the typical scorecard development steps proposed by Siddiqi (<xref ref-type="bibr" rid="CIT0031">2006</xref>), which is called variable characteristics analysis. The next two tests (Tests 4&#x2013;5) are model prediction analyses and focus on providing validation on a model level, thus focussing on the global explanation of the final model prediction.</p>
<p>To perform the tests, certain prerequisite steps have to be completed. These steps are listed below:</p>
<list list-type="bullet">
<list-item><p>Fit the final<xref ref-type="fn" rid="FN0003"><sup>3</sup></xref> model with the development train sample and include all variables selected through variable selection techniques. Note that the development train sample is a subset of the development sample and is used to train the model. Similarly, the development test sample is also a subset of the development sample used to test the model&#x2019;s performance. Additionally, the out-of-time sample is an independent sample that either precedes or succeeds the development sample. This sample serves as another sample that the model developer can use to test accuracy and stability over a different period of time.</p></list-item>
<list-item><p>Generate the Shapley values for the model based on the development train sample. This value is based on all observations in the development sample and will give a Shapley value per observation/row for the specified variable.</p></list-item>
<list-item><p>In the case of continuous variables, the variable can be binned and the average Shapley value can be calculated to create a summarised view of the Shapley value for a range (bin) of values. Binning is a technique that is commonly used in scorecard development, see Siddiqi (<xref ref-type="bibr" rid="CIT0031">2006</xref>). Although the binned results are not used as input in the model, it creates a convenient way of summarising model output for visual reports.</p></list-item>
</list>
</sec>
<sec id="s20016">
<title>Variable characteristics analysis</title>
<sec id="s30017">
<title>Test 1: Shapley value versus default rate (per variable bin)</title>
<p>The first test visualises the Shapley value and default rate per variable bin. The Shapley value is calculated from the development sample as mentioned in the prerequisite steps. These values are grouped per variable bin (in the case of continuous variables), and the average of the binned group is reported on. The same steps are followed to obtain the default rate per bin. This test is shown with an example in <xref ref-type="fig" rid="F0004">Figure 4a</xref>.</p>
<fig id="F0004">
<label>FIGURE 4</label>
<caption><p>Shapley value tests 1&#x2013;5: (a) Test 1: Shapley value (RHS) versus default rate (per variable bin) (b) Test 2: Shapley value stability over time (per variable bin) (c) Test 3: Shapley value (RHS) versus percentage of population (d) Test 4: Top x positive and negative contributing variables (e) Test 5 Gini rank versus absolute Shapley value rank.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="SAJEMS-27-5786-g004.tif"/>
</fig>
<p>Test 1 shows how the average Shapley value generated with the development sample compares to the default rate of the development train, development test and validation samples. It is expected that the average Shapley value follows a similar trend to the default rate, considering the Shapley value explanation provided in the section titled &#x2018;Shapley value explanation&#x2019;. The test explains the variable&#x2019;s expected impact on the final prediction depending on the bin from where the observation originates. Furthermore, the tests enable us to inspect if a logical trend is present for the variable under consideration. This is a very important step, as noted by Siddiqi (<xref ref-type="bibr" rid="CIT0031">2006</xref>).</p>
<p>An additional test that accompanies Test 1 is a correlation report as shown in <xref ref-type="table" rid="T0003">Table 3</xref>.</p>
<table-wrap id="T0003">
<label>TABLE 3</label>
<caption><p>Test 1: Correlation report.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Results</th>
<th valign="top" align="center">Development-train</th>
<th valign="top" align="center">Development-test</th>
<th valign="top" align="center">Validation</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Pearson Correlation</td>
<td align="center">0.86</td>
<td align="center">0.82</td>
<td align="center">0.87</td>
</tr>
<tr>
<td align="left">Spearman Correlation</td>
<td align="center">0.86</td>
<td align="center">0.94</td>
<td align="center">0.89</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="T0003">Table 3</xref> shows the Spearman and Pearson<xref ref-type="fn" rid="FN0004"><sup>4</sup></xref> correlation between the average Shapley value and the development train, development test and validation default rates. This illustrates how closely aligned the average Shapley value trend is to the three<xref ref-type="fn" rid="FN0005"><sup>5</sup></xref> default rate trends.</p>
</sec>
<sec id="s30018">
<title>Test 2: Shapley value stability over time (per variable bin)</title>
<p>The second test measures the Shapley Value Stability Over Time (per variable bin). The Shapley value is calculated from the development train sample. These values are grouped per variable bin (in the case of continuous variables), and the average of the binned group is reported. The data is grouped by month for the development train sample, and the results can be depicted in a stacked bar graph, as seen in <xref ref-type="fig" rid="F0004">Figure 4b</xref>.</p>
<p>Test 2 highlights an important step in assessing the predictions&#x2019; continued transparency, namely the predictions&#x2019; stability over time. The test shows how stable the Shapley values are and determines how much the value fluctuates throughout the year per variable bin. This can point out certain seasonal trends that could be present in the data, which the model development team might want to cater for.</p>
</sec>
<sec id="s30019">
<title>Test 3: Shapley value versus percentage of population (per variable bin)</title>
<p>The third test shown in <xref ref-type="fig" rid="F0004">Figure 4c</xref> compares the Shapley Value to the population distribution, expressed as the proportion of the population per variable bin. The Shapley value is calculated from the development train sample and grouped per variable bin (in the case of continuous variables), and the average of the binned group is reported. The average Shapley value is plotted against the volume of observations in each variable bin for the development train, development test and validation samples.</p>
<p>Test 3 shows the distribution of population within the 10 variable bins and overlays the average Shapley value per bin. This test shows the expected impact of a certain Shapley value prediction on the overall population, considering the size/proportion of the group it relates to. This test could point out illogical trends in the Shapley value and its impact on the overall population. This test is aimed at providing transparency by assigning the size of the impact on the Shapley value contribution for a certain subgroup in the population.</p>
</sec>
</sec>
<sec id="s20020">
<title>Model prediction analysis</title>
<p>The variable characteristics analysis focussed on tests that could be performed per variable. This section introduces a test that can be used on a model prediction level, that is understanding what influences the final model prediction.</p>
<sec id="s30021">
<title>Test 4: Top x positive and negative contributors</title>
<p>The fourth test shown in <xref ref-type="fig" rid="F0004">Figure 4d</xref> shows the top <italic>x</italic> variables that positively (decrease default prediction) and negatively impact (increase default prediction) the final model prediction. The Shapley value is calculated from the development train sample. These values are grouped as before. The average Shapley value is plotted per variable bin for the top x positive and negative contributors.</p>
<p>Test 4 gives a high-level explanation of the final model prediction by visualising the top <italic>x</italic> contributing variables that positively and negatively impact the model prediction. This test can add significant value when a challenger model is being developed. Using this plot, the model developer can compare the contributions of the same subset of variables for the challenger model and investigate differences between the variable contributions of the models. This will highlight key differences between the champion and challenger models, and guide the development team in choosing the best model for the specific purpose. While not an extensive explanation, these visuals can be used to explain core drivers in the final prediction to business stakeholders.</p>
</sec>
<sec id="s30022">
<title>Test 5: Gini rank versus absolute Shapley value rank</title>
<p>Test 5 shown in <xref ref-type="fig" rid="F0004">Figure 4e</xref> is used to determine if the Shapley value contribution per variable ranks similarly to the Gini value calculated for each variable. To calculate this test, the Gini value is calculated per variable. To calculate the Gini value, the model prediction and actual default rate are required. Subsequently, the Shapley value contribution for each variable is calculated. Note that both these calculations are performed on the development train dataset.</p>
<p>Finally, the rank of the Gini and the absolute Shapley value is determined with respect to the subset of variables selected for the model. The absolute Shapley value is used because the value can be both positive and negative. For this test, we are not interested in the direction (positive or negative) of the impact but rather the magnitude of the impact. The rank of the Gini versus the rank of the absolute Shapley value is compared and illustrated visually.</p>
<p>This test evaluates the assumption that high contributing Shapley values are good in discriminating between default and non-default events. If a strong correlation between the rank of the Gini and the rank of the absolute Shapley does not exist, further investigation is warranted to understand why the assumption is not evident in the data.</p>
<p>A common misperception when explaining ML model predictions is that the model outcome can be explained through one metric. To our knowledge, such a golden standard does not exist, at least not one that is model-agnostic. Although the tests proposed above offer practical examples for the model development team, it&#x2019;s important to acknowledge their limitations. For instance, the process can be time-consuming as each variable needs to be evaluated individually. Additionally, certain variable assumptions that are not met may be difficult to explain and may take some time to investigate. The test can be modified and improved to meet the requirements of the relevant stakeholders. The final test will be different for specific use cases and different audiences because the relevant stakeholders will be responsible for approving the final model.</p>
</sec>
</sec>
</sec>
<sec id="s0023">
<title>Conclusion</title>
<p>The consensus has been that introducing effective methods for interpreting ML models is widely regarded as a crucial step required for the validation process in credit scoring. This article aims to contribute to this research by comparing traditional model validation processes to more recent proposed frameworks that include ML models.</p>
<p>Furthermore, the article aims to compare these two methodologies and identifies transparency as one of the key elements to instil trust in ML models. The comparative study suggests that many of the same criteria still apply to ML models, the only difference being the methods and techniques to evaluate the criteria that need to be adjusted and/or extended for ML models.</p>
<p>This study motivates, through prior research such as Du Toit et al. (<xref ref-type="bibr" rid="CIT0008">2023</xref>), that Shapley values hold immense potential in explaining ML models on the level of detail comparable to that provided by well-known scorecard metrics, such as the WOE metric. A list of validation elements related to ML transparency is identified, and our research guides practitioners on how Shapley values can be used to evaluate the criteria within the model validation elements practically.</p>
<p>This article illustrates how the typical model validation practitioner can integrate existing validation frameworks with the principle based guidelines proposed by recent research. It showcases the usability of techniques such as Shapley values and illustrates the importance of maintaining model validation processes that have proven very successful. The novelty in this research is that an interpretability technique is proposed specifically for credit scoring model validation in the banking sector.</p>
</sec>
</body>
<back>
<ack>
<title>Acknowledgements</title>
<sec id="s20024" sec-type="COI-statement">
<title>Competing interests</title>
<p>The authors declare that they have no financial or personal relationships that may have inappropriately influenced them in writing this article.</p>
</sec>
<sec id="s20025">
<title>Authors&#x2019; contributions</title>
<p>H.A.d.T., W.D.S. and H.R. contributed to the design and implementation of the research, to the analysis of the results and to the writing of the manuscript.</p>
</sec>
<sec id="s20026" sec-type="data-availability">
<title>Data availability</title>
<p>Data sharing is not applicable to this article, as no new data were created or analysed in this study. The data supporting the findings of this study are available within the article.</p>
</sec>
<sec id="s20027">
<title>Disclaimer</title>
<p>The views and opinions expressed in this article are those of the authors and are the product of professional research. The article does not necessarily reflect the official policy or position of any affiliated institution, funder, agency or that of the publisher. The authors are responsible for this article&#x2019;s results, findings and content.</p>
</sec>
</ack>
<ref-list id="references">
<title>References</title>
<ref id="CIT0001"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Abrahams</surname>, <given-names>C.R</given-names></string-name>. &#x0026; <string-name><surname>Zhang</surname>, <given-names>M</given-names></string-name></person-group>., <year>2008</year>, <source><italic>Fair lending compliance: Intelligence and implications for credit risk management</italic></source>, <publisher-name>Wiley</publisher-name>, <publisher-loc>Hoboken, NJ</publisher-loc>.</mixed-citation></ref>
<ref id="CIT0002"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Arrieta</surname>, <given-names>A</given-names></string-name>., <string-name><surname>D&#x00ED;az-Rodr&#x00ED;guez</surname>, <given-names>N</given-names></string-name>., <string-name><surname>Del Ser</surname>, <given-names>J</given-names></string-name>., <string-name><surname>Bennetot</surname>, <given-names>A</given-names></string-name>., <string-name><surname>Tabik</surname>, <given-names>S</given-names></string-name>., <string-name><surname>Barbado</surname>, <given-names>A</given-names></string-name>. <etal>et al.</etal></person-group>, <year>2020</year>, &#x2018;<article-title>Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI</article-title>&#x2019;, <source><italic>Information Fusion</italic></source> <volume>58</volume>, <fpage>82</fpage>&#x2013;<lpage>115</lpage>. <comment><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.inffus.2019.12.012">https://doi.org/10.1016/j.inffus.2019.12.012</ext-link></comment></mixed-citation></ref>
<ref id="CIT0003"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Baesens</surname>, <given-names>B</given-names></string-name>., <string-name><surname>Roesch</surname>, <given-names>D</given-names></string-name>. &#x0026; <string-name><surname>Scheule</surname>, <given-names>H</given-names></string-name></person-group>., <year>2016</year>, <source><italic>Credit risk analytics: Measurement techniques, applications, and examples in sas</italic></source>, <publisher-name>Wiley</publisher-name>, <publisher-loc>Hoboken, NJ</publisher-loc>.</mixed-citation></ref>
<ref id="CIT0004"><mixed-citation publication-type="book"><person-group person-group-type="author"><collab>BCBS</collab></person-group>, <year>2005</year>, <source><italic>Studies on the validation of internal rating systems (revised)</italic></source>, <comment>Working paper 14</comment>, <publisher-name>Bank for International Settlements</publisher-name>, <comment>viewed 13 November 2023, from <ext-link ext-link-type="uri" xlink:href="https://www.bis.org/publ/bcbs_wp14.htm">https://www.bis.org/publ/bcbs_wp14.htm</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0005"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bertsimas</surname>, <given-names>D</given-names></string-name>., <string-name><surname>King</surname>, <given-names>A</given-names></string-name>. &#x0026; <string-name><surname>Mazumder</surname>, <given-names>R</given-names></string-name></person-group>., <year>2016</year>, &#x2018;<article-title>Best subset selection via a modern optimization lens</article-title>&#x2019;, <source><italic>The Annals of Statistics</italic></source> <volume>44</volume>(<issue>2</issue>), <fpage>813</fpage>&#x2013;<lpage>852</lpage>. <comment><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1214/15-AOS1388">https://doi.org/10.1214/15-AOS1388</ext-link></comment></mixed-citation></ref>
<ref id="CIT0006"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Chen</surname>, <given-names>H</given-names></string-name>., <string-name><surname>Lundberg</surname>, <given-names>S</given-names></string-name>., <string-name><surname>and Lee</surname>, <given-names>S.-I</given-names></string-name></person-group>. (<year>2021</year>). &#x2018;<chapter-title>Explaining models by propagating Shapley values of local components</chapter-title>&#x2019;, in <person-group person-group-type="editor"><string-name><given-names>A.</given-names> <surname>Shaban-Nejad</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Michalowksi</surname></string-name>, &#x0026; <string-name><given-names>D.L.</given-names> <surname>Buckeridge</surname></string-name> (eds.)</person-group>, <source><italic>Explainable AI in Healthcare and Medicine: Building a Culture of Transparency and Accountability</italic></source>, pp. <fpage>261</fpage>&#x2013;<lpage>270</lpage>. <comment>Studies in Computational Intelligence, Volume 914</comment>. <publisher-name>Springer</publisher-name>.</mixed-citation></ref>
<ref id="CIT0007"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>De Jongh</surname>, <given-names>P.J</given-names></string-name>., <string-name><surname>Larney</surname>, <given-names>J</given-names></string-name>., <string-name><surname>Mare</surname>, <given-names>E</given-names></string-name>., <string-name><surname>Van Vuuren</surname>, <given-names>G.W</given-names></string-name>. &#x0026; <string-name><surname>Verster</surname>, <given-names>T</given-names></string-name></person-group>., <year>2017</year>, &#x2018;<article-title>A proposed best practice model validation framework for banks</article-title>&#x2019;, <source><italic>South African Journal of Economic and Management Sciences</italic></source> <volume>20</volume>(<issue>1</issue>), <fpage>a1490</fpage>. <comment><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4102/sajems.v20i1.1490">https://doi.org/10.4102/sajems.v20i1.1490</ext-link></comment></mixed-citation></ref>
<ref id="CIT0008"><mixed-citation publication-type="thesis"><person-group person-group-type="author"><string-name><surname>Du Toit</surname>, <given-names>H.A</given-names></string-name>., <string-name><surname>Schutte</surname>, <given-names>W.D</given-names></string-name>., <string-name><surname>Raubenheimer</surname>, <given-names>H</given-names></string-name></person-group>., <year>2023</year>, &#x2018;<article-title>Shapley values as an interpretability technique in credit scoring</article-title>&#x2019;, <comment>PhD thesis</comment>, <publisher-name>Centre for Business Mathematics and Informatics, North-West University</publisher-name>, <publisher-loc>Potchefstroom, South Africa</publisher-loc>.</mixed-citation></ref>
<ref id="CIT0009"><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>European Banking Authority (EBA)</collab></person-group>, <year>2011</year>, <source><italic>Guidelines on Internal Governance (GL 44)</italic></source>, <comment>viewed 13 November 2023, from <ext-link ext-link-type="uri" xlink:href="https://www.eba.europa.eu/regulation-and-policy/internal-governance/guidelines-on-internal-governance">https://www.eba.europa.eu/regulation-and-policy/internal-governance/guidelines-on-internal-governance</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0010"><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>European Banking Authority (EBA)</collab></person-group>, <year>2021</year>, <source><italic>Discussion paper on machine learning for IRB models &#x2013; European Banking Authority</italic></source>, <comment>viewed 01 May 2023, from <ext-link ext-link-type="uri" xlink:href="https://www.eba.europa.eu/regulation-and-policy/model-validation/discussion-paper-machine-learning-irb-models">https://www.eba.europa.eu/regulation-and-policy/model-validation/discussion-paper-machine-learning-irb-models</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0011"><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>European Commission</collab></person-group>, <year>2024</year>, &#x2018;<article-title>Artificial Intelligence Act</article-title>&#x2019;, <source><italic>Official Journal of the European Union</italic></source>, <comment>viewed n.d., from <ext-link ext-link-type="uri" xlink:href="https://eur-lex.europa.eu/eli/reg/2024/1689/oj">https://eur-lex.europa.eu/eli/reg/2024/1689/oj</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0012"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>H&#x00E4;rle</surname>, <given-names>P</given-names></string-name>., <string-name><surname>Havas</surname>, <given-names>A</given-names></string-name>., <string-name><surname>Kremer</surname>, <given-names>A</given-names></string-name>., <string-name><surname>Rona</surname>, <given-names>D</given-names></string-name>. &#x0026; <string-name><surname>Samandari</surname>, <given-names>H</given-names></string-name></person-group>., <year>2015</year>, <source><italic>The future of bank risk management</italic></source>, <comment>viewed 01 May 2023, from <ext-link ext-link-type="uri" xlink:href="https://www.mckinsey.com/~/media/mckinsey/dotcom/client_service/risk/pdfs/the_future_of_bank_risk_management.pdf">https://www.mckinsey.com/~/media/mckinsey/dotcom/client_service/risk/pdfs/the_future_of_bank_risk_management.pdf</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0013"><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>Hong Kong Monetary Authority (HKMA)</collab></person-group>, <year>2006</year>, <source><italic>Validating risk rating systems under the IRB approaches</italic></source>, <comment>viewed 13 November 2023, from <ext-link ext-link-type="uri" xlink:href="https://www.hkma.gov.hk/media/eng/doc/key-functions/banking-stability/supervisory-policy-manual/CA-G-4.pdf">https://www.hkma.gov.hk/media/eng/doc/key-functions/banking-stability/supervisory-policy-manual/CA-G-4.pdf</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0014"><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>IIF &#x0026; EY</collab></person-group>, <year>2022</year>, <source><italic>IIF and EY survey report on machine learning white paper</italic></source>, <comment>viewed 13 November 2023, from <ext-link ext-link-type="uri" xlink:href="https://www.iif.com/portals/0/Files/content/32370132_iif_and_ey_survey_report_on_machine_learning_-_uses_in_credit_risk_and_aml_applications_-_public_summary.pdf">https://www.iif.com/portals/0/Files/content/32370132_iif_and_ey_survey_report_on_machine_learning_-_uses_in_credit_risk_and_aml_applications_-_public_summary.pdf</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0015"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kim</surname>, <given-names>B</given-names></string-name>., <string-name><surname>Khanna</surname>, <given-names>R</given-names></string-name>. &#x0026; <string-name><surname>Koyejo</surname>, <given-names>O</given-names></string-name></person-group>., <year>2016</year>, &#x2018;<article-title>Examples are not enough, learn to criticize! Criticism for interpretability</article-title>&#x2019;, in <source><italic>Advances in neural informatio&#x2019;n processing systems 29 (NIPS 2016)</italic></source>, <comment>viewed 15 January 2024, from <ext-link ext-link-type="uri" xlink:href="https://papers.nips.cc/paper_files/paper/2016/hash/5680522b8e2bb01943234bce7bf84534-Abstract.html">https://papers.nips.cc/paper_files/paper/2016/hash/5680522b8e2bb01943234bce7bf84534-Abstract.html</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0016"><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>KPMG</collab></person-group>, <year>2022</year>, <source><italic>Modern risk management for AI models &#x2013; Re-imagining the model risk management function for artificial intelligence</italic></source>, <comment>KPMG white paper, viewed 19 March 2023, from <ext-link ext-link-type="uri" xlink:href="https://assets.kpmg.com/content/dam/kpmg/xx/pdf/2022/07/modern-risk-management-for-ai-models.pdf">https://assets.kpmg.com/content/dam/kpmg/xx/pdf/2022/07/modern-risk-management-for-ai-models.pdf</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0017"><mixed-citation publication-type="conference"><person-group person-group-type="author"><string-name><surname>Kumar</surname>, <given-names>E</given-names></string-name>., <string-name><surname>Venkatasubramanian</surname>, <given-names>S</given-names></string-name>., <string-name><surname>Scheidegger</surname>, <given-names>C</given-names></string-name>. &#x0026; <string-name><surname>Friedler</surname>, <given-names>S</given-names></string-name></person-group>., <year>2020</year>, &#x2018;<article-title>Problems with shapley-value-based explanations as feature importance measures</article-title>&#x2019;, in <person-group person-group-type="editor"><string-name><given-names>H.</given-names> <surname>Daum&#x00E9;</surname> <prefix>III</prefix></string-name> &#x0026; <string-name><given-names>A.</given-names> <surname>Singh</surname></string-name> (eds.)</person-group>, <conf-name>37th International Conference on Machine Learning, online</conf-name>, <conf-date>July 13&#x2013;18, 2020</conf-date>, pp. <fpage>5491</fpage>&#x2013;<lpage>5500</lpage>.</mixed-citation></ref>
<ref id="CIT0018"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lessmann</surname>, <given-names>S</given-names></string-name>., <string-name><surname>Baesens</surname>, <given-names>B</given-names></string-name>., <string-name><surname>Seow</surname>, <given-names>H</given-names></string-name>. &#x0026; <string-name><surname>Thomas</surname>, <given-names>L</given-names></string-name></person-group>., <year>2015</year>, &#x2018;<article-title>Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research</article-title>&#x2019;, <source><italic>European Journal of Operational Research</italic></source> <volume>247</volume>(<issue>1</issue>), <fpage>124</fpage>&#x2013;<lpage>136</lpage>. <comment><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.ejor.2015.05.030">https://doi.org/10.1016/j.ejor.2015.05.030</ext-link></comment></mixed-citation></ref>
<ref id="CIT0019"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lundberg</surname>, <given-names>S</given-names></string-name>., <string-name><surname>Erion</surname>, <given-names>G</given-names></string-name>. &#x0026; <string-name><surname>Lee</surname>, <given-names>S</given-names></string-name></person-group>., <year>2018</year>, <source><italic>Consistent individualized feature attribution for tree ensembles</italic></source>, <comment>Working paper, viewed 13 November 2023, from <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/pdf/1802.03888.pdf">https://arxiv.org/pdf/1802.03888.pdf</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0020"><mixed-citation publication-type="conference"><person-group person-group-type="author"><string-name><surname>Lundberg</surname>, <given-names>S</given-names></string-name>. &#x0026; <string-name><surname>Lee</surname>, <given-names>S</given-names></string-name></person-group>., <year>2017</year>, &#x2018;<article-title>A unified approach to interpreting model predictions</article-title>&#x2019;, in <person-group person-group-type="editor"><string-name><given-names> I.</given-names> <surname>Guyon</surname></string-name>, <string-name><given-names>U.V.</given-names> <surname>Luxburg</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Bengio</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Wallach</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Fergus</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Vishwanathan</surname></string-name>, <etal>et al.</etal> (eds.)</person-group>, <conf-name>31st International Conference on Neural Information Processing Systems (NIPS&#x2019;17) proceedings</conf-name>, <conf-loc>Curran Associates Inc.</conf-loc>, <conf-date>4 December 2017</conf-date>, pp. <fpage>4768</fpage>&#x2013;<lpage>4777</lpage>.</mixed-citation></ref>
<ref id="CIT0021"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Miller</surname>, <given-names>T</given-names></string-name></person-group>., <year>2017</year>, <source><italic>Explanation in artificial intelligence: Insights from the social sciences</italic></source>, <comment>Working paper, viewed 13 November 2023, from <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/pdf/1706.07269.pdf">https://arxiv.org/pdf/1706.07269.pdf</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0022"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Molnar</surname>, <given-names>C</given-names></string-name></person-group>., <year>2020</year>, <source><italic>Interpretable machine learning. A guide for making Black Box models explainable</italic></source>, <comment>viewed 12 November 2023, from <ext-link ext-link-type="uri" xlink:href="https://christophm.github.io/interpretable-ml-book/">https://christophm.github.io/interpretable-ml-book/</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0023"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Molnar</surname>, <given-names>C</given-names></string-name>., <string-name><surname>K&#x00F6;nig</surname>, <given-names>G</given-names></string-name>., <string-name><surname>Herbinger</surname>, <given-names>J</given-names></string-name>., <string-name><surname>Freiesleben</surname>, <given-names>T</given-names></string-name>., <string-name><surname>Dandl</surname>, <given-names>S</given-names></string-name>., <string-name><surname>Scholbeck</surname>, <given-names>C.A</given-names></string-name>. <etal>et al.</etal></person-group>, <year>2020</year>, <source><italic>General pitfalls of model-agnostic interpretation methods for machine learning models</italic></source>, <comment>Working paper, viewed 13 November 2023, from <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/pdf/2007.04131.pdf">https://arxiv.org/pdf/2007.04131.pdf</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0024"><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>Monetary Authority of Singapore (MAS)</collab></person-group>, <year>2018</year>, <source><italic>Principles to promote Fairness, Ethics, Accountability and Transparency (FEAT) in the use of artificial intelligence and data analytics in Singapore&#x2019;s financial sector</italic></source>, <comment>MAS white paper, viewed 13 November 2023, from <ext-link ext-link-type="uri" xlink:href="https://www.mas.gov.sg/publications/monographs-or-information-paper/2018/FEAT">https://www.mas.gov.sg/publications/monographs-or-information-paper/2018/FEAT</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0025"><mixed-citation publication-type="book"><person-group person-group-type="author"><collab>Office of the Comptroller of the Currency (OCC)</collab></person-group>, <year>2011</year>, <source><italic>Supervisory guidance on model risk management</italic></source>, <comment>Supervisory document OCC 2011&#x2013;12</comment>, <publisher-name>Board of Governors of the Federal Reserve System</publisher-name>, <publisher-loc>Washington, DC</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>21</lpage>.</mixed-citation></ref>
<ref id="CIT0026"><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>PWC</collab></person-group>, <year>2018</year>, <source><italic>Building trust in AI and data analytics</italic></source>, <comment>viewed 20 April 2021, from <ext-link ext-link-type="uri" xlink:href="https://www.pwc.com/sg/en/publications/assets/building-trust-ai-data-analytics-122018.pdf">https://www.pwc.com/sg/en/publications/assets/building-trust-ai-data-analytics-122018.pdf</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0027"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Quell</surname>, <given-names>P</given-names></string-name>., <string-name><surname>Bellotti</surname>, <given-names>AG</given-names></string-name>., <string-name><surname>Breeden</surname>, <given-names>J.L</given-names></string-name>. &#x0026; <string-name><surname>Martin</surname>, <given-names>J.C</given-names></string-name></person-group>., <year>2021</year>, <source><italic>Machine learning and model risk management</italic></source>, <comment>Model Risk Managers&#x2019; international association (MRMIA) white paper, viewed 18 March 2023, from <ext-link ext-link-type="uri" xlink:href="https://mrmia.org/wp-content/uploads/2021/03/Machine-Learning-and-Model-Risk-Management.pdf">https://mrmia.org/wp-content/uploads/2021/03/Machine-Learning-and-Model-Risk-Management.pdf</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0028"><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>SARB</collab></person-group>, <year>2015</year>, <source><italic>Directive 4/2015: Amendments to the regulations relating to banks, and matters related thereto</italic></source>, <comment>viewed 10 August 2016, from <ext-link ext-link-type="uri" xlink:href="https://www.resbank.co.za/Lists/News&#x0025;20and&#x0025;20Publications/Attachments/6664/D4&#x0025;20of&#x0025;202015.pdf">https://www.resbank.co.za/Lists/News&#x0025;20and&#x0025;20Publications/Attachments/6664/D4&#x0025;20of&#x0025;202015.pdf</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0029"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Scheda</surname>, <given-names>R</given-names></string-name>. &#x0026; <string-name><surname>Diciotti</surname>, <given-names>S</given-names></string-name></person-group>., <year>2022</year>, &#x2018;<article-title>Explanations of machine learning models in repeated nested cross-validation: An application in age prediction using brain complexity features</article-title>&#x2019;, <source><italic>Applied Sciences</italic></source> <volume>12</volume>(<issue>13</issue>), <fpage>6681</fpage>. <comment><ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3390/app12136681">https://doi.org/10.3390/app12136681</ext-link></comment></mixed-citation></ref>
<ref id="CIT0030"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Shapley</surname>, <given-names>L</given-names></string-name></person-group>., <year>1953</year>, &#x2018;<chapter-title>A value for n-person games</chapter-title>&#x2019;, in <person-group person-group-type="editor"><string-name><given-names>K.</given-names> <surname>Harold William</surname></string-name> (ed.)</person-group>, <source><italic>Contributions to the theory of games</italic></source>, vol. <volume>2</volume>, pp. <fpage>307</fpage>&#x2013;<lpage>317</lpage>, <publisher-name>Princeton University Press</publisher-name>, <publisher-loc>Princeton</publisher-loc>.</mixed-citation></ref>
<ref id="CIT0031"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Siddiqi</surname>, <given-names>N</given-names></string-name></person-group>., <year>2006</year>, <source><italic>Credit risk scorecards: Developing and implementing intelligent credit scoring</italic></source>, <publisher-name>John Wiley &#x0026; Sons</publisher-name>, <publisher-loc>Hoboken, NJ</publisher-loc>.</mixed-citation></ref>
<ref id="CIT0032"><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Siddiqi</surname>, <given-names>N</given-names></string-name></person-group>., <year>2017</year>, <source><italic>Intelligent credit scoring : Building and implementing better credit risk scorecards</italic></source>, <publisher-name>John Wiley &#x0026; Sons</publisher-name>, <publisher-loc>Hoboken, NJ</publisher-loc>.</mixed-citation></ref>
<ref id="CIT0033"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Van der Burgt</surname>, <given-names>J</given-names></string-name></person-group>., <year>2019</year>, <source><italic>General principles for the use of AI in the financial sector</italic></source>, <comment>DeNederlandscheBank white paper, viewed 13 November 2023, from <ext-link ext-link-type="uri" xlink:href="https://www.dnb.nl/media/voffsric/general-principles-for-the-use-of-artificial-intelligence-in-the-financial-sector.pdf">https://www.dnb.nl/media/voffsric/general-principles-for-the-use-of-artificial-intelligence-in-the-financial-sector.pdf</ext-link>.</comment></mixed-citation></ref>
<ref id="CIT0034"><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wo&#x017A;nica</surname>, <given-names>K</given-names></string-name>., <string-name><surname>P&#x0119;kala</surname>, <given-names>K</given-names></string-name>., <string-name><surname>Baniecki</surname>, <given-names>H</given-names></string-name>., <string-name><surname>Kretowicz</surname>, <given-names>W</given-names></string-name>., <string-name><surname>Sienkiewicz</surname>, <given-names>E</given-names></string-name>. &#x0026; <string-name><surname>Biecek</surname>, <given-names>P</given-names></string-name></person-group>., <year>2021</year>, <source><italic>Do not explain without context: Addressing the blind spot of model explanations</italic></source>, <comment>Working paper, viewed 13 November 2023, from <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/pdf/2105.13787.pdf">https://arxiv.org/pdf/2105.13787.pdf</ext-link>.</comment></mixed-citation></ref>
</ref-list>
<app-group>
<app id="app001">
<title>Appendix 1</title>
<sec id="s0029">
<title></title>
<p><xref ref-type="table" rid="T0004">Table 1-A1</xref> describes the seven elements contained in the model validation process scorecard as proposed by De Jongh et al. (<xref ref-type="bibr" rid="CIT0007">2017</xref>) in more detail.</p>
<table-wrap id="T0004">
<label>TABLE 1-A1</label>
<caption><p>Model validation process scorecard as proposed by De Jongh et al. (<xref ref-type="bibr" rid="CIT0007">2017</xref>).</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left" rowspan="2">Validation process</th>
<th valign="top" align="center" colspan="4">Score<hr/></th>
</tr>
<tr>
<th valign="top" align="center">1: No evidence</th>
<th valign="top" align="center">2: Due consideration lacking</th>
<th valign="top" align="center">3: Some consideration</th>
<th valign="top" align="center">4: Fully evident</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left"><bold>Paradigm</bold></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was the conceptual soundness of paradigm checked?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was the review performed by suitably skilled experts?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left"><bold>Methods or theory</bold></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent is the underlying model theory consistent with published research and sound industry practice?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent were research publications considered of appropriate quality/standing?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was the methodology benchmarked against appropriate industry practice?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent are approximations made within agreed tolerance levels?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left"><bold>Design</bold></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was it ascertained that assumptions are clearly formulated?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was the appropriateness and the completeness of assumptions checked?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was it checked that all variables employed have been clearly defined and listed?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent have the causal relationships between variables been noted?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent have input data been assessed in terms of reasonableness, validity and understanding?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent has it been ascertained that outputs are clearly defined?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent has the design been evaluated in terms of over-complexity/over-simplification?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent has the model builder benchmarked the design against existing best practice models?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was the design independently benchmarked against existing best practice models?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent have special cases been dealt with appropriately? (e.g. terminal conditions or products with path-dependent pay-off)</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left"><bold>Data or variables</bold></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent have input data been checked to gauge reliability/suitability/validity/completeness?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent has it been checked that data involving subjective assessment of expert opinion been appropriately incorporated?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was the procedure for the collation of expert opinion scrutinised?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent has expert opinion been validated in terms of logical considerations?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent has the expert selection process been assessed as sound?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was it verified that data are representative of relevant (general and stressed) market conditions?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was it verified that data are representative of the company&#x2019;s portfolio?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent have inadequate or missing data been re-assessed and reviewed for model feasibility?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left"><bold>Algorithms or code</bold></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was the algorithms/code checked against the model formulation and underlying theory?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent were key assumptions and variables analysed with respect to their impact on model outputs?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was an independent construction of an identical model undertaken?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was the code rigorously tested against a benchmark model?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was technical proofreading of the code performed?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left"><bold>Outputs</bold></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was model output benchmarked against best practice models (e.g. against a vendor model using the same input data set)?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent was the reasonableness and validity of model outputs assessed?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent has a comparison of model outputs against actual realisations been performed? (backtesting)</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent has a range of outputs been examined vs. a range of inputs (e.g. are solutions continuous or jagged? What is the behaviour of hedging quantities and/or derived quantities over the same range?)</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent are all results repeatable? (e.g. Monte Carlo simulations)</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left"><bold>Monitoring</bold></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
<tr>
<td align="left">To what extent has the model been monitored for appropriate implementation and use?</td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
<td align="left"></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</app>
</app-group>
<fn-group>
<fn><p><bold>How to cite this article:</bold> Du Toit, H.A., Schutte, W.D. &#x0026; Raubenheimer, H., 2024, &#x2018;Integrating traditional and non-traditional model risk frameworks in credit scoring&#x2019;, <italic>South African Journal of Economic and Management Sciences</italic> 27(1), a5786. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.4102/sajems.v27i1.5786">https://doi.org/10.4102/sajems.v27i1.5786</ext-link></p></fn>
<fn id="FN0001"><label>1</label><p>Note that some literature refers to AI, other refer to ML, but for this article, these terms are used interchangeably.</p></fn>
<fn id="FN0002"><label>2</label><p>Any ML model could have been used as the framework, and is not limited to random forests only.</p></fn>
<fn id="FN0003"><label>3</label><p>As per the case study above, the final model is a Random Forest model and it is assumed that the software used to fit the model can be used to generate Shapley values.</p></fn>
<fn id="FN0004"><label>4</label><p>The normality assumption should be tested before the correlation is calculated. If both the Shapley values and the default rates are normally distributed, then the Pearson correlation is sufficient. Otherwise, Spearman rank correlation is recommended.</p></fn>
<fn id="FN0005"><label>5</label><p>Note that the default rates are based on actual default and thus give a good indication of how accurate the Shapley value trend is tracking the actual default rate.</p></fn>
</fn-group>
</back>
</article>