Collections · decision support

Help the agent
decide. Don't
decide for them.

A recommendation layer for NBFC recovery agents, sequenced so the model is useful without becoming the agent's starting point.

If you read nothing else
  1. The agent commits first. The suggestion is inert until their own call is written and immutable.
  2. The model can decline. On a thin file it returns nothing and names the three missing fields.
  3. Disagreement is logged, not corrected. Every divergence becomes an auditable pair with reasons attached.
Self-initiatedLive prototype Open it ↗
Shibil Mohammed · August 2026
Case CG-317082
Priya Venkataraman
Revolving Credit · Bucket 2
₹54,200Overdue
Days past due47
Hardship flagJob loss · documented
Last payment₹6,200 · 12 May
Promises2 kept · 1 broken
Active settlementNone
1 2
Contact history
21 AugRequested EMI restructure. Follow-up Friday.
18 AugNo answer
14 AugCited job loss since March
12 MayPayment received ₹6,200
3
1

47 days and a hardship flag pull opposite ways. One says escalate, the other says restructure. This is the judgement the agent is paid for.

2

A broken promise sits beside two kept ones. Any tool that reduces this to a score has thrown away the thing that decides the call.

3

Contact history on the same screen, not behind a tab. The agent has under ninety seconds per case.

01Where the pressure comes from

Forty cases,
one shift.

The console isn't opened once. An agent works a queue, and the case screen is one stop inside it. Every design decision downstream has to survive being made forty times before lunch.

That's the argument for the whole product. A tool that costs ten extra seconds per case costs seven minutes a shift. A tool that lets an agent skip thinking saves them nothing and costs the company the recovery.

Why the queue is slide one

Design an AI panel in isolation and you optimise for the demo. Design it as the fourth thing an agent looks at in ninety seconds and you get different answers.

My queue
Today · 40 assigned
14 worked
Priya Venkataraman47d₹54,200
Suresh Iyer38d₹31,400
Anand Rao22d₹18,900
Meera Shetty61d₹87,300
Vikram Joshi19d₹12,050
1 2
1

14 of 40 worked is the number the agent is actually measured on, so it sits in the header rather than a dashboard.

2

Flags are dots, not badges. At forty rows, badges become a wall of colour and stop meaning anything.

02The problem

The suggestion arrives
before the judgement.

Every decision-support tool I looked at puts its recommendation on screen the moment the case opens. The agent reads it before weighing the 47 days, the hardship flag, or the broken promise.

From there they aren't deciding what to do. They're deciding whether to disagree, which is a different and much lazier task. At forty cases a shift, nobody disagrees.

Anchoring

It doesn't feel like influence from the inside. It feels like the tool saving you time.

Suggested action
Convert to EMI
1
Your decision
30-day extension
Convert to EMI
Single settlement
Escalate to lead
2
1

Read first, so it becomes the frame. Everything below is now measured against it rather than against the case.

2

The options are still live, but the decision has effectively been made. This is the industry default, not a strawman.

03What I tried

Four ways to
sequence it.

The question was never whether to show the model's output. It's when. Three of these fail for the same reason, and the reason isn't obvious until they sit side by side.

Reading the sketches

Grey blocks are case facts and options. The solid dark bar is the model's suggestion. The black bar is commit.

A
Visible from case open
The anchoring problem, unmodified. What the tools already do.
B
Collapsed behind a toggle
Still a choice to peek, and under queue pressure everyone peeks. Now it also looks like the agent chose to be anchored.
C
Revealed on selection
Selection isn't commitment. The agent can still switch before anything is written, so there's no honest record of what they thought first.
D · chosen
Revealed after commit
The only order where the agent's judgement exists as a fact before the model's does. Everything else follows from this.
04Mechanism

Commit, then
compare.

The agent picks an action and commits. The support panel is inert until they do, and says plainly why rather than just showing a lock.

On the commit control

This is the highest-consequence button in the product, so the line under it states what commit means: the decision is written before the suggestion appears. A control that changes what you can un-know should say so.

Your decision
30-day extension
Convert to EMI
Single settlement
Escalate to lead
Recorded before the suggestion is shown.
Both calls stay in the audit history.
1 2

Decision support locked

Commit your decision to reveal. Your call is logged first so the comparison stays honest, the model can't anchor you, and you can't anchor to it.

3
1

Dark, full width, no competing action. Nothing else on the panel can be mistaken for it.

2

Two lines of consequence directly beneath. Most products put this in a modal after the click, which is after it matters.

3

The locked state explains its reasoning rather than asserting a rule. An agent who understands why won't try to route around it.

05Insufficient data

The state most
tools don't have.

Second case. No contact note, no recorded payment, no promise history. A model that answers anyway is guessing in a confident voice, and the agent can't tell that apart from a good answer.

So it returns nothing, itemises what's missing, and routes to a lead. The framing is what the agent can do, not a confidence score they can't act on.

Why this matters most

A system that is always sure is a system you eventually stop checking. The refusals are what make the answers worth reading.

Case CG-411290
Suresh Iyer
Personal Loan · 38 DPD
₹31,400Overdue
Suggested action
None
Insufficient case data. A reliable suggestion can't be produced from this file.
Last contactNot recorded
Payment historyNot recorded
Promise historyNot recorded
§3.3, §4.2 both require a documented contact event
Recovery Policy v3.4 · prototype framework
Flag for team lead
1 2 3
1

The slot stays and reads None. Hiding the panel would leave the agent wondering whether it failed or was never called.

2

Three named gaps, not a percentage. "Confidence 34%" tells an agent nothing they can act on. Three missing fields tells them exactly what to go and get.

3

Every dead end has one exit. There is no state in this product where the correct move is to stare at the screen.

06Failure

The model going
down is a state,
not an exception.

The request fails or returns something unreadable. Most tools show a toast and an empty panel, stranding the agent mid-case with nothing to act on.

This says three things: what happened, that the decision stands and is already recorded, and what to use instead.

The ordering pays off here

Because the human commits first, an outage costs the agent nothing. The work is already done. That isn't error handling, it's a consequence of the sequence.

Suggested action
Unavailable
The model couldn't be reached or returned an unreadable response. Your decision stands and is already logged. Use the policy sections below.
§4.2 Hardship restructure
Eligible when a hardship event is documented, DPD falls between 30 and 90, and no settlement is active.
Recovery Policy v3.4 · prototype framework
§3.3 Extensions
Flag for team lead
1 2
1

"Your decision stands and is already logged" is the whole message. The agent's first fear is that they've lost the work.

2

The policy clauses render in full, with the eligibility test. When the model is gone the interface has to do the model's job.

07Evidence

Suggestion, then
the facts it
rests on.

The agent isn't made to read a paragraph to reach the answer. The action comes first, then the case facts that triggered it, then the clause, then the source document.

The test I used

An agent defends this call to a lead, and eventually to an auditor. Every layer exists because someone downstream asks for it, not because more explanation is friendlier.

Suggested action
Convert to EMI
Why this was suggested
47 days past due · within 30–90 band
Hardship flag documented
Income disruption recorded 14 Aug
No active settlement plan
§4.2 Hardship restructure
Eligible when a hardship event is documented, DPD falls between 30 and 90, and no settlement is active.
Recovery Policy v3.4 · prototype framework
Logged, matched your call · pending policy check
1 2 3 4
1

Answer. Three words. On the fast cases the agent reads this and nothing else.

2

Evidence. Four case facts, each checkable against the file. Not prose, a list an auditor can tick.

3

Policy. The clause and its eligibility test, so the agent can confirm the model applied it correctly rather than trust that it did.

4

Provenance. "Pending policy check". The system doesn't claim the call is approved, only that it's recorded and matched.

08Difference

Disagreement is
the signal.

When the two calls differ, nothing is corrected and nothing is blocked. The difference is recorded, and the interface states where the two reasonings diverged rather than only flagging that they did.

Every disagreement becomes an auditable pair: a human judgement and a machine judgement on the same file, each with a stated reason.

What it's for

Agreement rate is a vanity metric. The disagreements are where you find out whether the model is wrong, or the policy is.

Decision difference
Your decision
Convert to EMI
Suggested action
30-day extension
Where they diverge
The model weighted days-past-due above the hardship flag. Your call weighted the documented income disruption above the DPD band.
Review evidenceDispute suggestion
1 2
After dispute
Under lead review. Both calls stored; your note will be requested.
3
1

Side by side, equal weight, neither styled as correct. Putting the model's call in a highlight box would answer the question this screen is meant to leave open.

2

The divergence is explained as what each side weighted. That's the sentence a lead needs; "they disagree" isn't.

3

Dispute reverses nothing. It routes, stores both, and asks for the agent's note, which is the input the model actually needs.

09Record

The log is a
by-product of
the sequence.

Collections is regulated. The useful consequence of committing first is that the audit trail falls out of the interaction. There's no separate logging step, because the order of events is the log.

The same record rolls up per session, which is what makes the word calibration mean anything: an agent can see their own divergence rate before anyone else does.

Where I'd want a lead's input

This record is the product's real asset and its real risk. It's a model-improvement loop if a team lead reads it, and performance surveillance if a manager does. That's a policy decision, not a design one, and it has to be settled before this ships.

Audit trail · CG-317082
11:42:06Case opened by agent_2841
11:42:51Agent decision Convert to EMI
11:42:51Decision committed · immutable
11:42:56Suggestion generated 30-day extension
11:42:56Difference logged to calibration record
11:43:19Suggestion disputed by agent
11:43:19Routed to lead review
1 2
Calibration log · this session
CG-317082EMI vs extensionDisputed
CG-411290No suggestion returnedLead
CG-298114SettlementMatched
CG-330627EscalateMatched
4 cases · 2 matched · 1 diverged · 1 no suggestion
3
1

Five seconds between the human call and the model's. The gap is the whole product, and it's visible in the timestamps.

2

Marked immutable at the moment of commit. If the record could be edited afterwards, none of the rest of this means anything.

3

The agent sees their own divergence rate first, in their own session. Whether anyone else sees it is the unresolved question two slides from here.

10Not resolved

Three things I haven't worked out.

Can the agent change their mind after the reveal?

Right now, no. If I allow it, the second decision is contaminated by definition and the calibration record has to store both. If I don't, the tool is telling an agent to ignore information it has just handed them. I don't know which is worse, and I don't think I can know without watching someone use it.

Who reads the calibration record?

A lead reading it to find bad policy is a feedback loop. A manager reading it to rank agents is surveillance. Same data, same screen, entirely different product, and agents will work out which one it is inside a week.

I've stamped agent_2841 on the audit trail, which quietly leans toward the second reading. I left it because collections records are regulated and have to be attributable, but I'm aware that's a decision I made without deciding it.

No recovery agent has used this.

I couldn't get access to one. Everything here rests on published collections process, RBI recovery guidance, and reasoning about how queue work behaves under time pressure. The anchoring effect is well evidenced in the literature; my specific fix for it is untested.

CASE OPEN │ facts · contact history · promise record ▼ AGENT REVIEWS EVIDENCE │ no suggestion on screen ▼ AGENT SELECTS ACTION │ ▼ COMMIT ── written · immutable │ ├──────────────┬──────────────┐ ▼ ▼ ▼ model up model down data too thin │ │ │ ▼ ▼ ▼ suggestion policy no suggestion + evidence fallback + gap list │ │ │ ▼ ▼ ▼ COMPARE flag lead flag lead │ ├────────┬─────────┐ ▼ ▼ ▼ match differ disputed │ │ │ ▼ ▼ ▼ log calibration lead record review │ │ │ └────────┴─────────┘ ▼ AUDIT TRAIL
11The whole thing

Every branch
ends somewhere.

Three failure paths, and none leave the agent stuck. That's the test I kept returning to: if the model does nothing useful, does the work still get done?

Why this is near the end

Shown first, this is an architecture diagram. Shown here, every box on it is a decision you've already watched me argue for.

12What changed
ConventionalThis
OrderSuggestion firstHuman commits first
EvidenceBehind an expanderFacts and clause on the panel
UncertaintyAlways answersDeclines and names the gaps
DisagreementTreated as agent errorLogged as a calibration signal
Model failureBlocks the workflowDecision already stands
AuditReconstructed afterA by-product of the sequence
How it was built

A browser prototype running against a local LLM, so the interaction was tested against real model output, including the runs where it was inconsistent or wrong, rather than scripted responses. The HTML scaffold was AI-assisted; the system prompt, the failure states and the decision architecture are mine.

Open the prototype ↗

Shibil Mohammed · ntshibil@gmail.com
August 2026

01 / 13