Agentic AI vs STP: Why STP Rate Stalls Near 10%
8 min read
Part 1 of 2: STP measures whole touchless claims, so it understates the automation sitting inside every file.
Your straight-through processing (STP) rate climbed to 12%, then stopped. Every gain came from the same place: auto glass, low-value contents, clean PIP bills. The other 88% still lands on an adjuster’s desk, and rule-tuning does little to move it.
Part of the problem is what STP rate measures. In the agentic AI vs STP debate, STP counts how many whole claims needed zero human input from open to close. That is a punishing bar, and it says almost nothing about the file on the adjuster’s desk right now, where much of the work is routine and a few steps need a person. AI automation applies to the tasks inside a claim, not only the claim as a whole. Count the work rather than the claim and nearly every file, including the litigated ones, contains substantial automatable work that STP rate never records.
What you will learn:
- Why STP rate is a whole-claim, all-or-nothing metric, and how much automation it hides
- Why automation belongs at the task level, and why the three autonomy levels matter
- How STP and agentic AI relate, and why STP understates automation
- Why the claims your metric ignores are the ones that decide your loss ratio
- The measures to track instead of STP rate alone, and a practical sequence to start
Part 2 of this blog covers what task-level AI claims automation does to the job itself: how the adjuster’s role, the queue, and the process have to be redesigned once the AI does the work.
Key definitions
STP measures whole claims that need no humans. Agentic AI automates tasks and decisions inside claims at three autonomy levels, escalating to a person for judgment and authority.
| Term | Definition |
|---|---|
| Straight-through processing (STP) | Automated end-to-end handling of a claim from FNOL to payment with no adjuster touching the file. A whole-claim, all-or-nothing outcome. |
| STP rate | The share of total claim volume processed end to end with no manual intervention. Industry personal lines rate sits near 7%; tuned programs reach 10% to 20% of volume. A claim-level metric. |
| Task-level AI claims automation | Work within claims handled without a person, at three autonomy levels: assist (extract, summarize, draft), recommend (propose a decision with evidence), execute (take an action that changes the claim). A claim can have most of its handling effort automated and still score zero on STP rate. |
| Assist / recommend / execute | The three autonomy levels. Drafting a reserve rationale (assist) differs in risk from changing the reserve (execute). The level, not the task label, drives impact, reversibility, and required human authority. |
| Judgment-and-authority tasks | The calls a person owns: coverage disputes, liability, negotiation, and regulated actions such as payments, denials, and reserve changes above a threshold. |
| LAE (Loss Adjustment Expense) | The operational cost of settling a claim. Task-level automation can cut LAE on claims that never go fully touchless. |
| Leakage | Money paid out that a tighter process would have saved. It concentrates in complex, litigated, and disputed files, the ones STP rate scores as zero. |
How straight-through processing is measured today, and what it misses
STP rate counts the share of claims that require no human intervention from open to close. It ignores the automatable work inside the 88% of claims that still touch a desk, so it understates how much of your handling effort could be automated with AI.
Industry STP rates sit around 7% across personal lines, per Datos Insights in its 2023 update, a figure that barely moved from 2021. Well-run programs go higher on narrow claim types: a tuned program auto-processes 10% to 20% of total volume and 70% to 90% of genuinely simple claims, per the Actuary CPD Tracker.
Those numbers get read as “we automated 12% of claims.” They actually say 12% of claims needed no human at any step. Most of your book lives between fully touchless and fully manual, and that middle is where the automatable work piles up, invisible to the metric.
Here is a single moderately complex claim, broken into its tasks and the agentic AI autonomy level each one supports:
| Task in the claim | AI autonomy level | STP credit |
|---|---|---|
| FNOL intake, structuring, first triage | Execute (low stakes) | None |
| Document ingestion and data extraction (PDFs, photos, statements) | Execute | None |
| Evidence reconciliation across conflicting estimates | Recommend | None |
| Reserve rationale drafting | Assist | None |
| Correspondence, status updates, reminders | Execute (within rules) | None |
| Coverage determination on a disputed clause | Human judgment | None |
| Payment above authority limit | Human authority | None |
| Task in the claim | AI autonomy level | STP credit |
|---|---|---|
| FNOL intake, structuring, first triage | Execute (low stakes) | None |
| Document ingestion and data extraction (PDFs, photos, statements) | Execute | None |
| Evidence reconciliation across conflicting estimates | Recommend | None |
| Reserve rationale drafting | Assist | None |
| Correspondence, status updates, reminders | Execute (within rules) | None |
| Coverage determination on a disputed clause | Human judgment | None |
| Payment above authority limit | Human authority | None |
Counting tasks is the wrong measure, because tasks differ enormously in effort, cost, and risk. A single coverage determination can take longer than the other six steps combined. The honest measure is the share of handling time and cost removed, not the number of tasks. Even when the human keeps the one high-effort judgment call, automating intake, extraction, reconciliation, drafting, and correspondence can remove the majority of the handling minutes on the file. STP still records that claim as a zero.
That is the core point: STP rate can read zero on a claim where AI automation removed four hours from a five-hour handling process. The plateau is partly a measurement problem. Real automation also stalls for reasons a metric cannot fix: poor data quality, legacy integration, fragmented authority, uncertain or conflicting evidence, vendor dependencies, regulation, and weak exception handling. Naming both is more useful than claiming the automation never plateaued. The measurement gap conceals substantial potential beyond the claims that can safely run touchless; capturing it still takes real operational work.
Three levels of task agentic automation: assist, recommend, execute
Task automation is not one thing. Drafting a reserve rationale, recommending a coverage position, and issuing a payment carry very different risk, so classify each task by autonomy level before you automate it.
The word “automation” hides three very different acts. Treating them as one is what makes “task automation” sound either trivial or reckless. Separating them is what makes it a usable operating model.
| Autonomy level | What the AI does | Impact if wrong | Reversibility | Human authority |
|---|---|---|---|---|
| Assist | Extracts, summarizes, drafts (e.g., drafts the reserve rationale) | Low; a person reviews before anything changes | Full | Person edits and uses the output |
| Recommend | Proposes a decision with evidence (e.g., a coverage position) | Medium; a wrong steer if the person defers without checking | Reversible before the action | Person decides |
| Execute | Takes an action that changes the claim (e.g., issues a payment, sets a reserve, sends a denial) | High; changes the claim and the money | Varies; some actions are hard to reverse | Person pre-authorizes rules and thresholds; high-stakes actions stay manual |
| Autonomy level | What the AI does | Impact if wrong | Reversibility | Human authority |
|---|---|---|---|---|
| Assist | Extracts, summarizes, drafts (e.g., drafts the reserve rationale) | Low; a person reviews before anything changes | Full | Person edits and uses the output |
| Recommend | Proposes a decision with evidence (e.g., a coverage position) | Medium; a wrong steer if the person defers without checking | Reversible before the action | Person decides |
| Execute | Takes an action that changes the claim (e.g., issues a payment, sets a reserve, sends a denial) | High; changes the claim and the money | Varies; some actions are hard to reverse | Person pre-authorizes rules and thresholds; high-stakes actions stay manual |
Read together with confidence and data quality, this is the grid that should govern where automation goes. Route a task to a higher autonomy level only when impact-if-wrong is low or reversible, confidence is high, the data is clean, and the required authority permits it. Low-stakes, reversible, high-confidence work (intake, extraction, status updates) can execute. Consequential or hard-to-reverse actions (payments, denials, reserve changes above a threshold) stay at recommend or manual, whatever the claim type. The autonomy level, not the claim label, is what carries the risk.
How STP and agentic automation relate
STP measures an end-to-end outcome. Agentic AI systems extend automation below that outcome, into the tasks and decisions within claims that never become fully touchless. They are related, not symmetric alternatives.
STP is an operational outcome and a metric. Agentic AI is a workflow and technology approach. Rule engines and RPA already automate some tasks; agentic AI extends task automation to unstructured evidence and multi-step reasoning, which is what reaches the complex files. A fixed rule set escalates the whole file the moment a claim stops matching a template.
Agentic automation works the tasks it can: it extracts the liability indicator buried on page 11 of the police report, reconciles three conflicting estimates, drafts the reserve rationale, and routes the coverage question to the adjuster with the analysis already done. The adjuster opens a structured file instead of a stack of PDFs, and the automated work counts, even though the claim was never touchless.
| Dimension | STP | Agentic automation |
|---|---|---|
| What it is | A metric: the claim ran touchless from FNOL to payment | A workflow approach that extends automation into tasks and decisions within claims |
| Question it answers | Did this claim need zero human touch? | How much of this claim’s work can we safely automate, and at what autonomy level? |
| How it treats a complex claim | Records it as zero | Automates the routine tasks, routes judgment and authority to a person |
| Where the human sits | Reviews everything that did not clear | Owns judgment and authority; supervises the rest |
| What it captures | Only fully touchless claims (~7% industry, 10% to 20% tuned) | Automation below the touchless line, on claims that never fully clear |
| Dimension | STP | Agentic automation |
|---|---|---|
| What it is | A metric: the claim ran touchless from FNOL to payment | A workflow approach that extends automation into tasks and decisions within claims |
| Question it answers | Did this claim need zero human touch? | How much of this claim's work can we safely automate, and at what autonomy level? |
| How it treats a complex claim | Records it as zero | Automates the routine tasks, routes judgment and authority to a person |
| Where the human sits | Reviews everything that did not clear | Owns judgment and authority; supervises the rest |
| What it captures | Only fully touchless claims (~7% industry, 10% to 20% tuned) | Automation below the touchless line, on claims that never fully clear |
The familiar objections to automating this work, that claims are too complex, too costly, or too risky under regulation, are increasingly answerable. McKinsey’s April 2026 work on modernizing insurance core technology argues that agentic AI systems can operate in regulated, complex environments with auditable outputs and human-in-the-loop controls. Applied to claims, that means automation can move further up the complexity curve, with a person keeping the judgment and authority calls.
“We’ve been able to significantly increase the net promoter score for claim management, and to reduce the sales cycle from several months to only a few days.”
Quentin Colmant, CEO and Co-founder, Qover
“We've been able to significantly increase the net promoter score for claim management, and to reduce the sales cycle from several months to only a few days.”
Quentin Colmant, CEO and Co-founder, Qover
See how much of your claims work is automatable.
Request a demo →Complex claims are not just simple claims with more human tasks
Complexity comes from how the facts connect, not only from having more judgment calls. A litigated file cannot always be decomposed into independent automatable units, so task automation supports the human view rather than replacing it.
It would be too neat to say a complex claim is a simple claim with a different task ratio. What makes complex claims hard is coupling: facts depend on each other, strategy evolves as evidence arrives, parties are adversarial, and the adjuster has to hold a coherent view of the whole file.
Reconciling three estimates is a discrete task; deciding a contested liability position while a suit is pending is not, because it depends on everything else in the file and on where the negotiation is heading.
So on the hardest files the value of task automation is different. It removes the assembly and administrative load, keeps the record current, and surfaces the analysis, so the adjuster spends their time on the coupled, strategic judgment rather than on gathering. The goal on these claims is a faster, better-supported human decision, not an automated one. That is a more defensible claim than “most of a litigated file automates,” and it is the one experienced claims leaders will accept.
Why the uncounted work is the expensive work in claims processing
The complex claims STP scores as zero are the ones that decide your loss ratio. Leaving the routine work on those files manual, because they will never go fully touchless, is where money and time drain out.
The money sits in the complex claim file. Insurers spend more than $23 billion a year on defense and cost containment, per EY, and third-party bodily injury indemnity has climbed 38% since 2020. EY puts leakage at 7% to 14% of total claims spend, driven by inaccurate damage evaluations, missed settlement windows, and thin investigation on liability. Little of that lives in the auto glass queue, it lives in the files STP rate records as zeros, and because the metric calls them “not automatable,” many carriers leave all of the processing work on those files manual, including the routine assembly and administration that a person does not need to do.
That double cost is the core problem. High-frequency work stays manual because the claim did not clear an STP gate, and high-severity work stays under-supported because the AI automation was never pointed at the tasks inside it. Every week a reserve sits open on a file whose routine work could have closed days ago is LAE you did not need to spend. And a policyholder attached to a serious event, a totaled car, a flooded home, an injury, waits longest on the claims your metric told you to skip. The fix is to automate the routine work on complex claims with AI while adjusters keep coverage, liability, and payment.
How to start automating past your STP rate
Keep STP where it earns. Then measure the work rather than the claim, classify tasks by autonomy level, and automate across your whole book one claim type at a time, with people owning judgment and authority.
| Step | Why it matters | |
|---|---|---|
| Change what you measure: track active human minutes per claim, touches, and effort-weighted automation coverage, not STP rate alone | STP rate cannot see automation on the 88% of claims that still touch a desk | |
| Classify each task by autonomy level (assist, recommend, execute), impact if wrong, reversibility, and confidence | Sets where automation can act and where a human authorizes | |
| Keep STP where it earns, on auto glass, low-value contents, clean PIP | Those are the claims where task-level automation already reaches 100% | |
| Map your book by where the automatable handling effort is, per claim type | Complex types hold a large share of routine effort, the hidden opportunity | |
| Add AI automation on top of your existing CMS | Extends automation into complex claims with no migration | |
| Keep regulated actions manual; payments, denials, and reserve changes above threshold stay with a person | Protects authority and auditability on every claim | |
| Roll out one claim type at a time, proving each with outcome data before widening | Builds evidence and guardrails, and catches quality regressions early |
| Step | Why it matters | |
|---|---|---|
| ✓ | Change what you measure: track active human minutes per claim, touches, and effort-weighted automation coverage, not STP rate alone | STP rate cannot see automation on the 88% of claims that still touch a desk |
| ✓ | Classify each task by autonomy level (assist, recommend, execute), impact if wrong, reversibility, and confidence | Sets where automation can act and where a human authorizes |
| ✓ | Keep STP where it earns, on auto glass, low-value contents, clean PIP | Those are the claims where task-level automation already reaches 100% |
| ✓ | Map your book by where the automatable handling effort is, per claim type | Complex types hold a large share of routine effort, the hidden opportunity |
| ✓ | Add AI automation on top of your existing CMS | Extends automation into complex claims with no migration |
| ✓ | Keep regulated actions manual; payments, denials, and reserve changes above threshold stay with a person | Protects authority and auditability on every claim |
| ✓ | Roll out one claim type at a time, proving each with outcome data before widening | Builds evidence and guardrails, and catches quality regressions early |
Measure automation together with the results it is meant to improve. Automating more tasks does not necessarily lead to better claim outcomes—and those outcomes may still get worse. Watch active human minutes per claim, exception frequency and handling time, cycle time by severity, cost per claim, reopen and rework rates, leakage and outcome accuracy, complaints and claimant experience, and override and appeal outcomes. The goal is better claim outcomes at an appropriate autonomy level, not maximum automation coverage.
Automating the work also has a second-order effect most programs underestimate: it changes the job. When the AI does the work and the human owns the calls, the caseload, the queue, and the process itself have to be redesigned, which is the subject of covered in Part 2 of this blog.
How a global MGA cut claim cost 35% by automating the work, not the claim
Qover, an embedded-insurance MGA, handles claims across 32 countries. Cycle times ran into months, and volume could not scale without adding headcount in every market.
Rather than chase a touchless-claim rate, Qover put Clive AI to work on the tasks across each claim, on top of its existing setup: intake, document handling, evidence assembly, and the routine steps that fill a file, with people holding the judgment calls. Claim cost fell 35% and cycle time compressed from months to days across all 32 countries. (These are Five Sigma customer outcomes; task-level automation was one part of a broader change, so treat them as directional proof rather than an isolated causal measure.)
35%
lower claim cost
Months to days
shortened claims cycle
33%
in adjuster productivity
What this means for your operation
If your STP rate has stalled, the automatable work has not run out. The claims that fit an all-or-nothing definition have. The routine work is still there, on most of your claims, sitting on a desk because the metric called those files not automatable.
Clive AI, Five Sigma’s agentic AI claims expert, works the tasks across a claim’s lifecycle on top of a carrier’s existing system, from FNOL and coverage through reserving, damage assessment, payment, and closure, advancing each step under the insurer’s SOPs at the autonomy level you set, and logging every action. On a simple claim that adds up to full touchless handling. On a complex one it does the intake, extraction, reconciliation, drafting, and correspondence, and stops for the adjuster on the calls that carry stakes. McKinsey’s Insurance 2030 work projects personal lines above 90% STP and more than half of all claims activities automated by 2030. That second figure, claims activities rather than claims, is the task-level view, and it is the one that scales.
See how agentic automation handles your complex claim, task by task.
Request a demo →Key takeaways
- STP rate is a whole-claim, all-or-nothing metric that counts claims needing no human in the loop at any step, which is why industry rates sit near 7% and tuned programs cap at 10% to 20% of volume.
- The plateau is partly a measurement problem and partly real friction (data, integration, authority, evidence, regulation). STP understates automation because it ignores the routine work inside claims that still need a person.
- Measure effort, not task counts. STP can read zero on a claim where AI automation removed most of the handling minutes.
- Classify AI automation by autonomy level: assist, recommend, execute. Drafting a reserve rationale differs in risk from changing the reserve, and the level drives impact, reversibility, and required authority.
- The claims STP scores as zero decide the loss ratio: more than $23 billion a year on defense and cost containment, leakage at 7% to 14% of spend. Automate their routine work and keep humans on judgment and authority.
Frequently asked questions
What is the difference between straight-through processing and agentic AI in claims?
STP is a whole-claim outcome: a claim that ran from FNOL to payment with no human touch. Agentic AI is a workflow approach that automates tasks and decisions inside claims and escalates the ones needing judgment or authority. STP measures an outcome; agentic AI extends automation below it.
Why does STP rate stall near 10%?
Because it only counts claims where every step needed no human intervention, and most claims have at least one that does. It also stalls for real reasons: data quality, legacy integration, fragmented authority, and uncertain evidence. The metric understates automation and the friction is real.
Can agentic AI handle complex insurance claims?
It automates the routine work on them, intake, extraction, evidence assembly, drafting, correspondence, and supports the adjuster on coverage, liability, and payment. On coupled, strategic files the goal is a faster, better-supported human decision, not an automated one.
What should we measure instead of STP rate?
Active human minutes per claim, effort-weighted AI automation coverage, exception rate and handling time, cycle time by severity, cost per claim, reopen and rework rates, leakage, and complaints. Track automation against outcomes, since automating trivial tasks can mask worse results.
What are the three levels of task automation in claims processing?
Assist (extract, summarize, draft), recommend (propose a decision with evidence), and execute (take an action that changes the claim). Each carries different risk, reversibility, and required human authority, so classify tasks by level before automating.
Frequently asked questions
What is the difference between straight-through processing and agentic AI in claims?
STP is a whole-claim outcome: a claim that ran from FNOL to payment with no human touch. Agentic AI is a workflow approach that automates tasks and decisions inside claims and escalates the ones needing judgment or authority. STP measures an outcome; agentic AI extends automation below it.
Why does STP rate stall near 10%?
Because it only counts claims where every step needed no human intervention, and most claims have at least one that does. It also stalls for real reasons: data quality, legacy integration, fragmented authority, and uncertain evidence. The metric understates automation and the friction is real.
Can agentic AI handle complex insurance claims?
It automates the routine work on them, intake, extraction, evidence assembly, drafting, correspondence, and supports the adjuster on coverage, liability, and payment. On coupled, strategic files the goal is a faster, better-supported human decision, not an automated one.
What should we measure instead of STP rate?
Active human minutes per claim, effort-weighted AI automation coverage, exception rate and handling time, cycle time by severity, cost per claim, reopen and rework rates, leakage, and complaints. Track automation against outcomes, since automating trivial tasks can mask worse results.
What are the three levels of task automation in claims processing?
Assist (extract, summarize, draft), recommend (propose a decision with evidence), and execute (take an action that changes the claim). Each carries different risk, reversibility, and required human authority, so classify tasks by level before automating.
Michael Krikheli
Co-Founder & CTO, Five Sigma
Michael Krikheli co-founded Five Sigma to bring AI-native automation to P&C claims past rule-based automation into agentic AI that handles the routine work on every claim, with people engaged where judgment matters.