Underwriting
  • Research and White Papers
  • July 2026

Reimagining Model Validation: Innovative approaches in the age of GenAI

An RGA case study

By
  • Fontaine Chan
  • Karin Wu
  • Todd Seabaugh
Skip to Authors and Experts
Woman looking at a computer monitor
In Brief
RGA developed and applied a pragmatic, outcome-focused validation approach to assess the DigitalOwl underwriting tool, an RGA-implemented solution leveraging a third-party vendor GenAI model to streamline the underwriting process. The findings: Insurance carriers can rigorously validate a GenAI model using another GenAI model as part of an independent validation pipeline while maintaining strong governance, auditability and human oversight.

Key takeaways

  • RGA independently validated the DigitalOwl underwriting tool, a GenAI-enabled underwriting workflow solution, using an outcome-focused validation pipeline designed for a system with limited transparency.
  • A structured framework that combines vendor due diligence, independent testing, and ongoing monitoring to enable carriers to adopt GenAI responsibly while meeting model risk and regulatory expectations.
  • Using two complementary evaluation methods – LLM-as-a-judge and vector similarity – RGA assessed output quality at two different perspectives, providing a traceable audit trail for each evaluated case.
  • Human-on-the-loop oversight, KPI-based monitoring, and the use of LLM judges offer a practical blueprint for governing GenAI tools in critical underwriting workflows.

 

Introduction

GenAI is moving quickly from experimentation into everyday underwriting, with insurers adopting AI-assisted tools faster than traditional model risk and compliance frameworks were designed to govern them. As these tools become more embedded in underwriting, trust cannot be assumed; it must be demonstrated through clear evidence that model outputs are reliable, transparent, resilient, explainable, interpretable, and fair, and that the model remains fit for purpose in practice.

That challenge is especially acute for third-party GenAI tools. Regulators and risk leaders are increasingly focused on whether companies can independently validate models whose reasoning is difficult to explain, validate, or audit.

Business users may hesitate to trust outputs they do not fully understand and with underlying logic that is proprietary and difficult to interpret. Thus, the validation focus should be on whether such GenAI tools demonstrate key risk management attributes – such as reliability, transparency, resilience, explainability, interpretability and fairness – in alignment with the National Institute of Standards and Technology (NIST) AI Risk Management Framework. In this context, validation must rely primarily on assessing a model’s reasoning, performance consistency, and governance controls.

RGA’s Model Risk Management (MRM) team, independent of the underwriting team using the solution, validated DigitalOwl’s (a Datavant company) underwriting tool, an RGA-implemented solution leveraging a third-party GenAI model to streamline underwriting decision-making. This case study demonstrates how insurers can apply a structured, outcome-focused validation framework to a GenAI model in a way that supports effective challenges and helps establish that the model is reliable, transparent, resilient, explainable, interpretable, and fair.

A comprehensive validation approach

The DigitalOwl underwriting tool leverages large language models (LLMs) to transform complex medical records into insights and summaries that support underwriting workflows. Its risk profile extends beyond performance, reflecting inherent challenges with vendor models, including limited visibility into model logic, dependence on vendor-controlled changes, and the risk of incomplete, incorrect, or biased outputs. These factors necessitate a disciplined, comprehensive validation approach.

RGA’s validation framework for GenAI models and tools

Several considerations informed the design of RGA’s validation framework for GenAI models and tools.

Human expert review serves as the benchmark for validation accuracy, yet it presents challenges in terms of efficiency, cost, and standardization. By combining independent testing, due diligence, and ongoing monitoring, the framework was designed to identify what could go wrong and to put controls around these risks over time. As a result, trust in the model is not assumed from performance alone; it is supported by a transparent, auditable process that offers greater confidence in how the tool is governed and used.

Conventional validation methods are not sufficient for GenAI models/tools. Conventional approaches assume stable model logic, transparent model design, and model outputs that can be easily verified against predefined calculations.

In an underwriting application, validation must also account for semantic meaning, medical accuracy, and clinical relevance. And validation must recognize that important information is often distributed across lengthy medical reports and documents. At the same time, medical underwriting introduces additional constraints. Source records contain sensitive policyholder data and therefore must be de-identified in accordance with applicable requirements, including removal of personally identifiable information and other sensitive content prior to use. Vendor models like DigitalOwl are proprietary by design, limiting the feasibility of conventional white-box testing methods.

RGA therefore designed a repeatable and scalable framework for these types of models and tools. This framework is centered on observable model behavior: the level of model performance, the consistency of model outputs, and the sufficiency of governance controls to manage the residual risk. The framework starts with collecting model inventory and risk tiering, then continues with vendor due diligence review, independent testing, interpretation of results, and ongoing performance monitoring. This approach ensures that model governance is proportionate to model risk and persists through ongoing performance monitoring.

DigitalOwl logo on a patterned blue background
Step boldly, but responsibly, into the world of GenAI with RGA and DigitalOwl.

How the framework was implemented

Due diligence

RGA began with due diligence on the underlying vendor model supporting the DigitalOwl underwriting tool. This included reviewing DigitalOwl’s vendor documentation, testing results, and supporting evidence to understand the model’s intended use, model design, performance claims, controls, and limitations. The validation team also engaged the vendor to inquire about testing evidence on model performance and bias, as well as supporting controls and update practices.

This step provided critical context, but it did not replace RGA’s independent validation. Under a sound model risk framework, vendor materials are one component of the review process, not the sole basis for concluding a model is fit for purpose.

Independent validation pipeline

RGA then evaluated underwriting cases using two complementary methods – one that reviews all documents as a whole and one that reviews section by section. Each underwriting case consists of multiple medical documents (e.g., medical history, lab reports, and physician notes), which can be evaluated either as a complete record or section by section:

  1. LLM-as-a-judge (whole-case review) used an alternative large language model to compare the full set of model outputs generated by the DigitalOwl underwriting tool against the verified reference data and assess each case on factual accuracy, completeness, medical precision, and overall semantic alignment. In this approach, all documents are assessed together as a single, unified case record.
  2. Vector similarity (section-level review) evaluates each section independently by comparing sections of the generated outputs (e.g., medical history) to corresponding sections in the verified reference data. This method identifies localized gaps or inconsistencies within each section.

Together, these methods produced two different but complementary views of model performance. LLM-as-a-judge evaluates how well the documents come together as a complete and coherent narrative across the full medical record, while vector similarity examines how accurately specific sections align with the verified reference data. This combination allowed RGA to evaluate both holistic case quality and localized consistency at scale.

What the validation revealed

Across underwriting categories, LLM-as-a-judge scores tended to cluster toward the high end of the quality scale, indicating that the notes generated by the DigitalOwl underwriting tool were generally assessed as strong, coherent, and clinically meaningful when the full set of documents was reviewed together as a complete record.

Vector similarity scores are lower and more dispersed than LLM-as-a-judge scores, but that pattern is expected given the design of the method. Because it evaluates specific sections independently (e.g., medical history), it cannot fully capture cases where relevant evidence is distributed across multiple sections of the source document. Rather than undermining confidence, this difference highlights why both methods were needed: one to assess holistic narrative quality and the other to detect more localized variation.

What mattered most, however, was consistency between the two approaches. When section-level scores were rated “excellent” or “good,” whole-case evaluations nearly always landed in the same quality bands. This pattern suggests that when the model aligns well with source records at the detail level, that strength carries through to the overall underwriting narrative. For a validation audience, this consistency is more meaningful than any single score in isolation, because it shows that different evaluation approaches point to the same general conclusion.

The consistency is illustrated below for two key medical underwriting inputs.

For laboratory results, a large proportion of the validation results from the LLM-as-a-judge method are rated consistently as the results from the vector similarity method. This showed strong alignment between the two evaluation methods at higher quality tiers, reinforcing confidence in the model’s performance in this core underwriting area, based on the cases evaluated.

Medical conditions followed a similar pattern, with strong agreement between localized and holistic assessments when output quality was high.

Governance through collaboration and monitoring

Validation did not occur in isolation.

  • RGA’s underwriting team provided the business-grounded perspective needed to ensure the testing reflected how the AI tool is actually used and whether outputs were meaningful for underwriting decisions.
  • Vendor documentation and its evidence for testing were reviewed and assessed.
  • RGA’s Model Risk Management (MRM) team designed and implemented independent validation using two complementary methods to evaluate model performance.
  • To complement internal validation efforts, KPMG LLP provided an unbiased perspective and helped align the framework with emerging GenAI governance industry practices.

This approach formed one input into a broader validation process that combines business input, technical testing, and external challenges.

The framework also extends beyond initial model validation. Once the model is deployed in production, it should be subject to ongoing monitoring through key performance indicators (KPIs). When these indicators fall outside defined thresholds, human reviews are triggered so underwriters and model owners can investigate and determine whether remediation or revalidation is required. This kind of human oversight, often called a human-on-the-loop mechanism, is consistent with AI governance expectations that high-risk systems should incorporate meaningful monitoring and human intervention. It further illustrates how the GenAI model supports underwriters, yet it does not replace their decision-making or judgment.

Figure 1: GenAI model validation framework: DigitalOwl underwriting tool case study

Flow chart showing the GenAI model validation framework

Conclusion: A reference point for carriers

GenAI can deliver meaningful efficiency gains by converting hundreds of pages of medical evidence into a concise, structured view of underwriter-ready insights. But for GenAI vendor models, efficiency alone is not enough. It requires a rigorous validation approach that demonstrates reliability and transparency, supports resilience and assesses fairness, and builds trust through explainability, interpretability, and strong governance.

The broader lesson from the DigitalOwl underwriting tool case is that outcome-focused, behavior-based validation is becoming essential for GenAI models and tools in underwriting and other high-impact insurance cases. As these systems become embedded in core workflows, the standard for trust will increasingly depend on documented evidence that model outputs are reliable, consistent, well-controlled, and subject to human oversight.

Used this way, “using GenAI to validate GenAI” can be practical and responsible. With structured governance, independent testing, efficient challenge, and ongoing monitoring, even proprietary vendor models can be managed to a standard relevant for high-risk GenAI models and tools.

The article outlines RGA's validation of an RGA-implemented underwriting solution. The validation results are not a guarantee of future performance and are informational only. The article does not represent underwriting, actuarial, legal, or compliance advice.


More Like This...

Meet the Authors & Experts

Fontaine Chan
Author
Fontaine Chan
Executive Director and Actuary
Karin Wu
Author
Karin Wu
Senior Data Scientist, Enterprise Risk Analytics
Todd-Seabaugh
Author
Todd Seabaugh

Vice President, Business Initiative Lead