Few banks formally evaluate GenAI human-in-the-loop controls
Risk Benchmarking: G-Sibs and challengers use tools to test controls efficacy; others rely on judgement
This piece is part of a series benchmarking bank model risk management practices. Risk Management subscribers can view selected cuts of the underlying data here. Sign up for Risk Benchmarking emails here.
Only one in five banks use tools to formally monitor the effectiveness of the human-in-the-loop (HITL) controls they apply to generative AI or large language models (LLMs), Risk Benchmarking’s
Only users who have a paid subscription or are part of a corporate subscription are able to print or copy content.
To access these options, along with all other subscription benefits, please contact info@risk.net or view our subscription options here: http://subscriptions.risk.net/subscribe
You are currently unable to print this content. Please contact info@risk.net to find out more.
You are currently unable to copy this content. Please contact info@risk.net to find out more.
Copyright Infopro Digital Limited. All rights reserved.
As outlined in our terms and conditions, https://www.infopro-digital.com/terms-and-conditions/subscriptions/ (point 2.4), printing is limited to a single copy.
If you would like to purchase additional rights please email info@risk.net
Copyright Infopro Digital Limited. All rights reserved.
You may share this content using our article tools. As outlined in our terms and conditions, https://www.infopro-digital.com/terms-and-conditions/subscriptions/ (clause 2.4), an Authorised User may only make one copy of the materials for their own personal use. You must also comply with the restrictions in clause 2.5.
If you would like to purchase additional rights please email info@risk.net
More on Model risk
Banks building GenAI governance in parallel to model risk
Risk Benchmarking: Majority have established AI governance committees, but ownership is fragmented
Lower-risk models face excessive reviews, banks say
Risk Benchmarking: Validation workload stretching teams, amid emerging regulatory divergence
Model risk managers see growing regulatory divergence
Risk Benchmarking study finds most banks expect easing of model risk supervisory scrutiny in the US, but tightening in Europe
Many banks do not document failure plans for Tier 1 models
Risk Benchmarking: Strong predeployment validation gives way to ad hoc escalation of breaches, even at some large lenders
Model risk managers are being asked to do more with less
Risk Benchmarking study finds function being handed expanding AI workload, on flat resources
Banks are automating GenAI testing, but scope varies widely
Risk Benchmarking: LLM-as-judge offers model testing at scale, but few lenders use it to facilitate autonomous sign-off
Model Risk Benchmarking 2026: explore the data
View interactive charts from Risk.net’s 44-bank study, covering model inventories, resourcing, GenAI governance, validation and regulation
A third of banks do not maintain logs for GenAI models
Risk Benchmarking study finds few banks review prompt logs systematically, with larger firms focusing on higher risk use cases