Skip to main content

Journal of Risk Model Validation

Risk.net

Beyond mirror validation: the cost of ineffective challenge in model risk management

Gustavo Coleho Haase, Paulo Henrique Dourado, João Paulo Vieira Costa and Eduardo Medeiros Rubik

  • Mirror validation misses approximately 32% average PD underestimation (500 replications).
  • Five tests, via a corroboration rule, catch failures discrimination metrics miss.
  • Undetected degradation causes an 8.5% regulatory capital shortfall (loan-level).
  • The framework operationalizes the SR 11-7 / SR 26-2 effective challenge mandate.

A common failing in model validation is to confirm that the developer’s model has been correctly reimplemented and that discrimination metrics remain acceptable while leaving calibration under changing conditions unchallenged: a practice we term “mirror validation”. A controlled simulation of a retail credit probability of default model under regime change, replicated over 500 draws, shows that this practice fails to detect calibration degradation: standard discrimination metrics (the Kolmogorov–Smirnov statistic, Gini coefficient and area under the receiver operating characteristic curve) remain above acceptance thresholds in 100% of replications even as the model underestimates the probability of default by approximately 32% (with a 95% confidence interval of 21%–42%). We assemble five complementary tests (population stability index analysis, calibration testing, sensitivity analysis, challenger model benchmarking and cumulative sum control chart structural break detection) into an adversarial testing framework and show that under a corroboration decision rule, this framework detects the degradation with full power while holding the false positive rate on stable data near to zero. Computing capital loan-by-loan with the Basel II internal ratings-based formula (avoiding the aggregation bias of plugging the portfolio-mean probability of default into a nonlinear capital function), we find a capital shortfall of approximately 8.5% on a US$1 billion portfolio. Effective challenge, per Supervisory Guidance SR 11-7 and its 2026 successor, SR 26-2, requires testing that goes beyond replication and historical backtesting.

Sorry, our subscription options are not loading right now

Please try again later. Get in touch with our customer services team if this issue persists.

New to Risk.net? View our subscription options

You need to sign in to use this feature. If you don’t have a Risk.net account, please register for a trial.

Sign in
You are currently on corporate access.

To use this feature you will need an individual account. If you have one already please sign in.

Sign in.

Alternatively you can request an individual account here