Evergreen
How to Evaluate a Property Resilience Model
Evaluate a property resilience score by peril, target, data, holdout design, calibration, errors, fairness, decision impact, governance, and production drift.
How to Evaluate a Property Resilience Model
A property resilience model is useful only for a defined peril, population, time horizon, and decision. Do not begin with its score or headline accuracy. Begin with the action someone wants to take and the consequence when the model is wrong.
"95 percent accurate" is not an evaluation. Without the label, prevalence, threshold, sampling, comparator, false positives, false negatives, holdout design, and uncertainty, the number cannot tell a buyer whether the model is fit for use.
1. Define the decision before reviewing the model
Write the exact use case without naming the vendor. The decision might be to prioritize inspections, invite policyholders into a mitigation program, route a file for additional review, estimate portfolio vulnerability, support loss control, or add information to underwriting.
These uses are not interchangeable. A model that is helpful for outreach may be unacceptable for eligibility, price, limits, deductible, nonrenewal, or an adverse consumer action.
Name the peril, geography, property type, policy segment, event intensity, forecast horizon, user, affected person, current process, allowed action, prohibited action, and owner. Record whether the model informs a human decision or executes one.
The Actuarial Standards Board's ASOP 56 makes intended purpose and intended user central to model work. The standard applies to actuaries performing professional services. It is a useful evidence frame, not a certification for every vendor.
2. Define the target in operational language
Ask what the model predicts. Survivability could mean no total loss, continued structural function, a damage-state threshold, no paid claim, or a vendor-defined composite.
Identify who created the label, when it was observed, and what evidence was used. Claims data, post-event inspection, aerial imagery, public damage assessments, expert judgment, and another model produce different targets and biases.
Review ambiguous cases. A structure may remain standing while being unsafe, uninhabitable, inaccessible, without utilities, or economically impractical to repair. If the label compresses those cases into "survived," the buyer needs to know.
3. Reconstruct the evaluation population
The buyer should receive a flow from the eligible population to the final evaluation set. Every exclusion matters.
Ask how events, locations, properties, and outcomes were selected. Identify properties with missing imagery, incomplete attributes, unmatched addresses, unavailable outcomes, duplicate records, manual overrides, or post-event information.
Separate development, tuning, and holdout data. The holdout should represent the intended deployment population and remain unavailable during model selection. Event-based models often need geographic or temporal separation so nearby properties or repeated imagery do not leak the answer.
The Actuarial Standards Board's ASOP 38 addresses selecting, using, reviewing, and evaluating catastrophe models and includes reliance on experts, validation, data, and documentation. Whether a property model falls within that standard is a scope question for the responsible actuary.
4. Inspect inputs, provenance, and timing
For every input, request the source, legal right, collection method, observation date, geographic coverage, property matching method, missingness, refresh cadence, transformation, and known bias.
Aerial imagery, street imagery, policyholder submissions, assessor data, hazard models, climate analytics, inspection records, and derived property characteristics each have different limitations.
Timing is critical. Information collected after an event can leak the outcome into a retrospective test. A roof replacement, vegetation change, renovation, or deterioration can make older information wrong. A policyholder-submitted image may be current but incomplete or attached to the wrong property.
5. Review performance beyond one metric
The model report should show the count and prevalence of each outcome before any metric. Then review discrimination, calibration, threshold performance, uncertainty, and stability.
| Evidence | What it answers |
|---|---|
| Confusion matrix | How many false positives and false negatives occur at the proposed threshold? |
| Sensitivity and specificity | Which outcome is detected and which is missed? |
| Precision and negative predictive value | How often is each decision-facing classification correct? |
| ROC and precision-recall analysis | How does performance change across thresholds? |
| Calibration | Do predicted probabilities match observed frequencies? |
| Confidence intervals | How uncertain are the estimates? |
| Slice results | Does performance hold across regions, property types, values, data coverage, and affected groups? |
| Benchmark comparison | Is the model better than the current process or a simpler rule? |
Accuracy can be dominated by the common outcome. A 95 percent accuracy result could be worthless if 95 percent of properties share one label and the model predicts that label every time.
6. Examine failures as insurance events
Translate errors into operational and consumer consequences.
A false reassurance could reduce scrutiny for a vulnerable property. An overly adverse score could increase friction, trigger unnecessary inspection, or influence a decision against a property owner. Missing data may be correlated with rural areas, older housing, low-value properties, recent construction, imagery availability, or communities with less digital access.
The NAIC's AI Model Bulletin describes regulator expectations around governance, risk management, accuracy, unfair bias, data vulnerability, third parties, testing, validation, and documentation. It is not a model law or universally adopted requirement. The buyer must check the applicable jurisdiction and use.
7. Test calibration and uncertainty
If the score is presented as a probability, compare predicted values with observed frequencies in sufficiently large groups. A score of 70 should not be called a 70 percent probability unless the output was designed and validated that way.
Review confidence intervals and sample sizes. A strong aggregate result can hide unstable estimates for a smaller peril, geography, construction class, or data source.
Ask what the system does when required data is missing, stale, contradictory, outside range, or low quality. "Unable to score" can be safer than a confident number.
8. Run a shadow decision-impact pilot
Do not move directly from a vendor demonstration to a live consumer decision.
Run the exact production version in shadow mode on the intended population. Preserve the current decision, model result, reason codes, human review, disagreement, processing time, downstream action, and later outcome.
Compare the model-supported process with the current process. Measure incremental information, inspection yield, policyholder response, expense, turnaround, overrides, complaints, and error consequences. Any loss outcome needs sufficient time and credible attribution.
flowchart LR
A["Defined decision"] --> B["Independent holdout"]
B --> C["Error and calibration review"]
C --> D["Shadow workflow"]
D --> E["Consumer and business impact"]
E --> F["Approve, restrict, or reject"]
9. Establish governance before approval
Record the approved version, data sources, intended use, prohibited uses, thresholds, overrides, monitoring, refresh triggers, incident process, documentation, vendor duties, audit rights, retention, and retirement plan.
Third-party responsibility does not replace insurer accountability. Contracts should address data rights, confidentiality, security, regulatory cooperation, model changes, validation evidence, incidents, and termination.
The NAIC's current artificial-intelligence topic page shows continued regulator work on AI system evaluation and third-party data and models. A buyer should verify current state adoption and guidance immediately before use.
10. Monitor the production population
Track input coverage, missingness, score distribution, calibration where outcomes become available, overrides, complaints, exceptions, geographic shifts, event changes, and vendor releases.
Retraining or adding a peril creates a new evidence question. A model validated on one wildfire event does not inherit approval for flood, hail, hurricane, a new region, or a different insurance decision.
Approve only the use that was evaluated. If the evidence supports outreach but not underwriting, write that restriction into the operating process.
Editorial and AI disclosure
This guide was developed from the preserved E059 transcript and current primary governance sources with AI assistance for research organization, drafting, and editing. Dalton Anderson remains the named author. Publication and use require actuarial, catastrophe-model, underwriting, data-science, model-risk, insurance-regulatory, privacy, security, consumer-protection, accessibility, and legal review appropriate to the jurisdiction and decision.
This draft is not authorized for publication. It is a general evaluation framework, not approval or criticism of Faura or any other model, and not insurance, actuarial, underwriting, pricing, coverage, engineering, legal, compliance, procurement, or regulatory advice.
Sources
Follow the evidence.
- content.naic.org: catastrophe models propertycontent.naic.org
- federalregister.gov: understanding the federal registerfederalregister.gov
- linkedin.com: valkyrieholmeslinkedin.com
- home.treasury.gov: Analyses of US Homeowners Insurance Markets 2018 2022 Climate Related Risks and Other Factors 0home.treasury.gov
- faura.us: faura raises 35m seed round to transform climate risk and property insurancefaura.us
- daltonanderson.net: valkyrie holmes building faura quantifying climate riskdaltonanderson.net
- SBA: Market Research and Competitive Analysissba.gov
- faura.us: faura redefining property riskfaura.us
- 776.org776.org
- faura.us: aboutfaura.us
- faura.usfaura.us
- daltonanderson.ghost.io: valkyrie holmes building faura quantifying climate riskdaltonanderson.ghost.io
- content.naic.org: natural catastrophe risk resilience resource centercontent.naic.org
- content.naic.org: naic adopts first national climate resilience strategy insurance close coverage gaps and improvecontent.naic.org
- youtu.be: WD3ikPPzYBMyoutu.be
- actuarialstandardsboard.org: modeling 3actuarialstandardsboard.org
- open.spotify.com: 1O8itURPFeO2OuvSpDiyKzopen.spotify.com
- content.naic.org: state insurance departmentscontent.naic.org
- actuarialstandardsboard.org: actuarial standard of practice no 38 revised editionactuarialstandardsboard.org
- ecfr.gov: understanding the ecfrecfr.gov
- content.naic.org: naic members approve model bulletin use ai insurerscontent.naic.org
- fema.gov: publicationsfema.gov