Banking · Banking exception handling

The Cost of Exceptions: Why High-Volume Banking Workflows Accumulate Rework

Straight-through processing gets the attention. The residual exception population often carries the real operational cost — investigation, repair, controls, hand-offs and repeated work across systems.

High-volume banking operations can look efficient from the top down.

Volumes are known. SLAs are reported. Straight-through-processing rates are tracked. Automation has removed obvious manual tasks.

Then you look at the cases that do not follow the happy path.

They move between queues. Data is checked again. Somebody investigates. A control creates another hand-off. Evidence is gathered from a second system. The case loops back. A small exception becomes its own workflow.

The exception is not one extra task. It is often a second operating model hiding underneath the first.

What’s normal

The standard management response focuses on the visible measures:

  • increase STP or automation rates;
  • attack the backlog;
  • reduce average handling time;
  • rationalise systems;
  • add workflow tooling;
  • automate the remaining manual steps.

Again, none of that is inherently wrong.

The problem is that an exception population is rarely homogeneous.

A missing field, a policy rule, a genuine risk case, poor upstream data and a legacy control can all appear in the same “manual” queue while having completely different economics and completely different solutions.

Why it fails

When the simple cases have already been automated, “manual work” becomes a weak proxy for opportunity.

The remaining cases are often manual precisely because they are different.

If you treat them as one automation backlog, you risk spending heavily to automate complexity rather than reducing it.

The other problem is visibility.

Banking workflows cross systems, teams and control points. Each team can appear locally efficient while the end-to-end case accumulates wait time and repeat handling.

A dashboard may show the final SLA breach.

It does not automatically show which upstream variant caused three extra touches two days earlier.

This is why exception cost tends to accumulate quietly.

What I do differently

I want to reconstruct the actual case journey before deciding what to automate.

That can involve process intelligence, data joins, direct process discovery and SME evidence. The goal is not a perfect enterprise process map. It is enough evidence to distinguish the economically important variants.

For the exception population, I am interested in questions such as:

  • what caused the case to leave the happy path;
  • how many additional touches it created;
  • which role handled those touches;
  • where it waited;
  • whether the control or investigation is genuinely required;
  • whether the root cause sits upstream of the team carrying the work.

That moves the conversation from “how much is manual?” to “which exception causes are worth removing?”

Diagnostic

What this looks like in practice

For one high-volume workflow, I would want the operating team to be able to answer:

  • What percentage completes straight through, and what happens to the rest?
  • Which three exception classes consume the most total handling minutes or capacity?
  • How many systems and teams does an exception touch compared with a normal case?
  • Where are the repeat checks, loops and hand-backs?
  • How much elapsed time is waiting versus active work?
  • Which controls are required by policy or regulation, and which have simply accumulated over time?
  • Has existing automation reduced total cost-to-serve, or mainly moved the difficult work into a smaller residual queue?

The purpose is not to create a giant process-mining exercise. It is to find enough evidence to rank the value pool and decide which causes deserve intervention.

Evidence from the work

In a major UK bank, I used process-intelligence approaches to follow lending journeys through to underwriting, paying particular attention to the cases that did not move straight through. I also worked on complaints analysis where the evidence suggested that some blanket handling assumptions could be challenged by looking at the actual risk behaviour of lower-risk cases.

The recurring pattern was the same: averages hide variants.

You have to see the journey of the difficult cases before you can decide what should change.

Earlier, in investment operations, I worked across back-office automation and operational excellence where reconciliation included thousands of transaction lines, including cases split across multiple lines, alongside document-based cash-transfer processing and other operational flows.

Across five back-office teams, the programme put six automations into the pipeline within four months with approximately £100,000 per year of expected savings.

The lesson was not “automate every exception”.

It was that process architecture, value-stream evidence and opportunity assessment need to sit in front of the technology decision, especially once the easy work has already gone.

The questions worth asking

Do we know the cost of an exception — or only our STP rate?

Which exception cause consumes the most capacity when handling time and recurrence are combined?

How much of the residual work is genuine risk/control activity, and how much is correcting avoidable upstream failure?

Can we reconstruct a case across systems well enough to see its loops and hand-backs?

Has automation reduced end-to-end unit cost, or simply removed the clean cases?

If we eliminated one recurring cause upstream, which downstream queue would shrink?

That is where the next wave of banking productivity is more likely to sit than in another generic use-case collection exercise.

The point

High-volume banking workflows do not become expensive only because they are manual.

They become expensive when exceptions create extra journeys through people, systems and controls.

Once you can see those journeys, you can separate necessary judgement and control from preventable repair — and choose the right intervention with much stronger evidence.

Have this problem in a live workflow? Discuss the workflow.

Discuss the workflow

← All insights