Most agencies now have AI running somewhere in the service operation. Far fewer can answer the question that follows it: when the AI did that, who was accountable, and can you show the record?
That question arrives from three directions and rarely with warning. A regulator during a market conduct exam. A carrier asking how a submission was handled. A buyer in diligence reconstructing how the book was serviced before they bought it. In each case the answer that holds is the same, and it is not a demo.
This piece covers what a control actually is in an AI service operation, the three questions that get asked, why “a person reviews everything” stops being an answer at volume, and what to have in place so the record exists before anyone asks for it.
The risk is not that AI is wrong. It is that nobody can say what it did
An AI that makes a mistake is a quality problem, and quality problems are familiar. Agencies have handled those since long before any of this.
The new problem is different. When work is done by software rather than by a named person at a desk, the ordinary trail of accountability goes quiet. There is no one to ask what they were thinking, no email thread that shows the judgment call, no note in the file. The work got done. The reasoning left no trace.
That gap is what makes an operator nervous about AI, and they are right to be. It is also, unhelpfully, the gap most AI tools do not close, because being able to reconstruct a decision is not what a demo is designed to show.
The three questions that actually get asked
Regulators, carriers, and buyers ask different things for different reasons, but the questions reduce to three.
Was the AI allowed to do that? Insurance service work is not uniformly licensable. Some of it legally requires a licensed person, and some of it does not. If AI is doing work across that boundary with no enforcement, the answer to this question is a shrug, and a shrug is the worst possible answer during an exam.
What happened when it was not sure? Every automated system encounters work it cannot handle confidently. The question is what the system does in that moment. Guessing quietly is the failure mode that produces the incident you hear about later.
Can you show me? Not a policy document describing what should have happened. The record of what did happen, on a specific request, on a specific day, including who or what acted and why.
What a control looks like in practice
A control is not a paragraph in a policy. It is something that stops an action from happening, or forces it onto a different path. Three of them do most of the work.
| Control | What it does | The question it answers |
|---|---|---|
| License gate | AI never does work it is not allowed to do. The boundary is enforced by the system, not remembered by a person. | Was the AI allowed to do that? |
| Confidence threshold | Low-confidence work escalates to a person instead of completing quietly. | What happened when it was not sure? |
| Audit trail | Every action, human or AI, is logged with actor, timestamp, and reason. | Can you show me? |
The order matters. A license gate is preventive, a confidence threshold is a routing decision made in the moment, and an audit trail is what you have afterwards. An operation missing the first two can still produce a log, but the log will be a record of things you would rather had not happened.
Two more belong alongside them. Judgment stays with people, so nothing consequential goes out without a person behind it. And every AI action carries a cost, a time, and a quality score, which means the AI is managed rather than trusted. Those are not the same posture, and the difference shows up under scrutiny.
Why “a person reviews everything” is not a control
The instinct when AI enters a service operation is to put a human review step in front of everything it produces. It feels like the safe answer, and for a pilot handling a few dozen items a day, it is.
It stops working for a specific reason. Review applied uniformly is review that cannot be applied carefully. When every item needs a signature, the signature stops meaning anything, and the reviewer becomes a rubber stamp who is nonetheless now accountable for everything they waved through. That is worse than the problem it was meant to solve, because it manufactures the appearance of oversight without the substance.
The version that holds is selective. Work the system is confident about completes and is logged. Work it is not confident about escalates to a person while it still matters. Consequential actions require approval regardless of confidence. The reviewer’s attention goes to the cases that need it, which is what keeps that attention worth having.
That is also the version you can describe to a regulator without flinching, because it is a rule rather than an intention.
What the production data shows
Across the agencies we operate:
- Roughly 44% of completed service tasks are executed by AI. This is production, not a pilot.
- AI triage filters about 68% of inbound as noise, meaning carrier notifications, duplicates, and automated confirmations that never needed a person. Classification runs at a 98% triage success rate, in under 10 seconds per item.
- About 40,000 tasks a week are classified and routed with no manual intervention.
- Escalations, meaning a task where someone is stuck and has to hand it on, are down 20%, from 435 a day to 348, and still declining.
- Median handle time on a task is 3.8 minutes, down 16% from 4.5.
VERO is the AI doing that work, and it runs inside the same task structure the people do, routed by license and skill, logged the same way. That last part is the part that matters for this piece. The AI is not a separate system operating beside the operation with its own rules. It is a supply type inside it, governed identically.
COVU runs on top of your AMS, not instead of it, with 8 AMS integrations live, so the record lives alongside the system you already use rather than in a tool nobody will think to check.
S&G Mitchell went from 17.9% to 60%+ EBITDA in 12 months. Same agency, same customers, same book of business. What changed was the operating model underneath.
Where to start
You do not need an AI governance program to make progress on this. You need one workflow you can describe completely.
Pick the workflow with the highest volume, which for most agencies is certificates or endorsement requests. Then answer three things in writing. Which steps in it legally require a licensed person. What the system does when it is not confident, stated as a rule rather than a hope. And whether, for a request handled last Tuesday, you could produce the sequence of actions with an actor and a reason against each one.
If the third answer is no, that is the place to start, and it is worth starting before something makes it urgent. The agencies that handle an exam or a diligence request calmly are not the ones with the best intentions. They are the ones who could already answer the question.
Frequently Asked Questions
Can insurance agencies use AI for client service and stay compliant?
Yes, when the boundaries are enforced by the system rather than remembered by people. The three controls that matter are a license gate so AI never does work that requires a license, a confidence threshold so uncertain work escalates to a person, and an audit trail on every action. Controls like these only hold on standardized work, which is the reason AI tools alone do not fix an agency. Compliance obligations vary by state and by carrier agreement, so confirm the specifics with your own counsel or compliance lead.
What is a license gate in AI service work?
A rule enforced in the system that prevents AI, or any unlicensed capacity, from completing a step that legally requires a licensed person. Service work decomposes into steps with different requirements, so the gate operates at the step level rather than on whole requests. That step-level split is also what relieves the licensing bottleneck.
How do you audit an AI decision?
By reading the record of the action. Every action, human or AI, is logged with actor, timestamp, and reason, which lets you reconstruct a specific request on a specific date rather than describing what your process says should have happened. It depends on the work being structured as discrete tasks in the first place, which is what an agency operating system does.
Does a person still approve the work?
For anything consequential, yes. Judgment stays with people and nothing consequential goes out without a person behind it. What changes is that review is targeted at the work that needs it, rather than applied uniformly until it becomes a formality. That distinction is the practical difference between an operating model and another AI pilot.
What should an agency have ready before an audit or diligence request?
The ability to reconstruct a single request end to end: what came in, how it was classified, who or what acted at each step, what was escalated and why, and when it closed. If that exists for your highest-volume workflow, the rest is extension rather than invention. Buyers ask the same questions in diligence for different reasons, which is part of what consolidating service operations does to EBITDA.
