How the framework was implemented
Due diligence
RGA began with due diligence on the underlying vendor model supporting the DigitalOwl underwriting tool. This included reviewing DigitalOwl’s vendor documentation, testing results, and supporting evidence to understand the model’s intended use, model design, performance claims, controls, and limitations. The validation team also engaged the vendor to inquire about testing evidence on model performance and bias, as well as supporting controls and update practices.
This step provided critical context, but it did not replace RGA’s independent validation. Under a sound model risk framework, vendor materials are one component of the review process, not the sole basis for concluding a model is fit for purpose.
Independent validation pipeline
RGA then evaluated underwriting cases using two complementary methods – one that reviews all documents as a whole and one that reviews section by section. Each underwriting case consists of multiple medical documents (e.g., medical history, lab reports, and physician notes), which can be evaluated either as a complete record or section by section:
- LLM-as-a-judge (whole-case review) used an alternative large language model to compare the full set of model outputs generated by the DigitalOwl underwriting tool against the verified reference data and assess each case on factual accuracy, completeness, medical precision, and overall semantic alignment. In this approach, all documents are assessed together as a single, unified case record.
- Vector similarity (section-level review) evaluates each section independently by comparing sections of the generated outputs (e.g., medical history) to corresponding sections in the verified reference data. This method identifies localized gaps or inconsistencies within each section.
Together, these methods produced two different but complementary views of model performance. LLM-as-a-judge evaluates how well the documents come together as a complete and coherent narrative across the full medical record, while vector similarity examines how accurately specific sections align with the verified reference data. This combination allowed RGA to evaluate both holistic case quality and localized consistency at scale.
What the validation revealed
Across underwriting categories, LLM-as-a-judge scores tended to cluster toward the high end of the quality scale, indicating that the notes generated by the DigitalOwl underwriting tool were generally assessed as strong, coherent, and clinically meaningful when the full set of documents was reviewed together as a complete record.
Vector similarity scores are lower and more dispersed than LLM-as-a-judge scores, but that pattern is expected given the design of the method. Because it evaluates specific sections independently (e.g., medical history), it cannot fully capture cases where relevant evidence is distributed across multiple sections of the source document. Rather than undermining confidence, this difference highlights why both methods were needed: one to assess holistic narrative quality and the other to detect more localized variation.
What mattered most, however, was consistency between the two approaches. When section-level scores were rated “excellent” or “good,” whole-case evaluations nearly always landed in the same quality bands. This pattern suggests that when the model aligns well with source records at the detail level, that strength carries through to the overall underwriting narrative. For a validation audience, this consistency is more meaningful than any single score in isolation, because it shows that different evaluation approaches point to the same general conclusion.
The consistency is illustrated below for two key medical underwriting inputs.
For laboratory results, a large proportion of the validation results from the LLM-as-a-judge method are rated consistently as the results from the vector similarity method. This showed strong alignment between the two evaluation methods at higher quality tiers, reinforcing confidence in the model’s performance in this core underwriting area, based on the cases evaluated.
Medical conditions followed a similar pattern, with strong agreement between localized and holistic assessments when output quality was high.
Governance through collaboration and monitoring
Validation did not occur in isolation.
- RGA’s underwriting team provided the business-grounded perspective needed to ensure the testing reflected how the AI tool is actually used and whether outputs were meaningful for underwriting decisions.
- Vendor documentation and its evidence for testing were reviewed and assessed.
- RGA’s Model Risk Management (MRM) team designed and implemented independent validation using two complementary methods to evaluate model performance.
- To complement internal validation efforts, KPMG LLP provided an unbiased perspective and helped align the framework with emerging GenAI governance industry practices.
This approach formed one input into a broader validation process that combines business input, technical testing, and external challenges.
The framework also extends beyond initial model validation. Once the model is deployed in production, it should be subject to ongoing monitoring through key performance indicators (KPIs). When these indicators fall outside defined thresholds, human reviews are triggered so underwriters and model owners can investigate and determine whether remediation or revalidation is required. This kind of human oversight, often called a human-on-the-loop mechanism, is consistent with AI governance expectations that high-risk systems should incorporate meaningful monitoring and human intervention. It further illustrates how the GenAI model supports underwriters, yet it does not replace their decision-making or judgment.
Figure 1: GenAI model validation framework: DigitalOwl underwriting tool case study

Conclusion: A reference point for carriers
GenAI can deliver meaningful efficiency gains by converting hundreds of pages of medical evidence into a concise, structured view of underwriter-ready insights. But for GenAI vendor models, efficiency alone is not enough. It requires a rigorous validation approach that demonstrates reliability and transparency, supports resilience and assesses fairness, and builds trust through explainability, interpretability, and strong governance.
The broader lesson from the DigitalOwl underwriting tool case is that outcome-focused, behavior-based validation is becoming essential for GenAI models and tools in underwriting and other high-impact insurance cases. As these systems become embedded in core workflows, the standard for trust will increasingly depend on documented evidence that model outputs are reliable, consistent, well-controlled, and subject to human oversight.
Used this way, “using GenAI to validate GenAI” can be practical and responsible. With structured governance, independent testing, efficient challenge, and ongoing monitoring, even proprietary vendor models can be managed to a standard relevant for high-risk GenAI models and tools.
The article outlines RGA's validation of an RGA-implemented underwriting solution. The validation results are not a guarantee of future performance and are informational only. The article does not represent underwriting, actuarial, legal, or compliance advice.
RGA experts are eager to engage with clients to better understand and tackle the industry’s most pressing challenges together. Contact us to learn more about RGA's capabilities, resources, and solutions.