The situation
A general insurer replacing a policy administration system. 1,900 tables, column names like FLG_7, and the people who designed it left in 2016.
Home / Use cases
Use cases
Six engagements, in detail: the situation, why it was hard, the architecture we built, and what changed. Including the parts that needed a human.
Data migration
A general insurer replacing a policy administration system. 1,900 tables, column names like FLG_7, and the people who designed it left in 2016.
The vendor quoted eighteen months, most of it manual archaeology. Nobody could say which tables were still written to, and a wrong mapping is a mispriced policy.
Architecture
Profiling every table for real usage, formats and relationships, rather than trusting the documentation that did not exist.
High-confidence mappings ran unattended; anything below 85% queued for an underwriter, with the evidence attached.
Each migration run compared row counts and column checksums back to source. A mismatch stopped the run, it did not get logged.
The whole migration ran nightly against production copies for six weeks before cutover.
Infrastructure modernisation
Why
A 400-server estate across two datacentres and three cloud accounts, grown by acquisition. Four teams each held part of the picture and none held all of it.
Hard
Every migration plan stalled on the same question: if we move this, what breaks? The dependency list lived in people's heads, and two of those people had resigned.
Architecture
Two weeks of agents reading netflow, configs and traces, so the dependency graph came from traffic rather than memory.
Clustering grouped tightly coupled systems together, so each wave could move and be verified on its own.
Spiky and stateless went to public cloud; steady and regulated to private; a handful stayed put with a review date rather than a pretence.
Finance signed off on wave two knowing its monthly figure, not an estate-wide estimate.
Cybersecurity
Four analysts receiving roughly 11,000 alerts a week across endpoint, identity and network tooling. Real incidents were in there. So was everything else.
Analysts triaged the newest alerts and the backlog aged out silently. The two incidents that mattered that quarter were both found by someone outside the team.
Architecture
Every alert was enriched with asset criticality, identity and recent change history, because the same alert means different things on a test box and a payment server.
Forty related alerts became one incident with a timeline, instead of forty tickets racing each other.
Closures carried the reason and stayed auditable, so the team could challenge the logic rather than trust it.
The system ranks and escalates. It does not contain, isolate or block on its own.
Document automation
Why
A lending operation receiving bank statements, KYC packs and income proofs as phone photographs and fax-quality scans. Eleven people doing data entry.
Hard
Off-the-shelf OCR handled the clean 60% and silently guessed at the rest. A wrong income figure is a credit decision, so nothing could be trusted unreviewed.
Architecture
One uncertain balance sends one field to a human. The other fourteen fields on that page still post automatically.
The reviewer gets the cropped image region and the proposed value, not the whole document. Three seconds, not three minutes.
They moved it twice in the first month. That is the correct amount of control to hand over.
Every human correction is labelled training data, so the low-confidence pile shrinks month over month.
Predictive maintenance
Rotating equipment across remote compressor stations. Vibration and thermal data existed, but it was collected monthly on a laptop by a technician on site.
Any cloud-based approach assumed a connection the sites did not have. The previous pilot stopped producing alerts for nine days before anyone noticed the link was down.
Architecture
The edge box scores anomalies locally. The network carries reporting, not decisions.
Seven days of local retention, because the honest answer to how long links stay down was five.
If the edge box stops scoring, that is itself an alert. The previous pilot's nine quiet days were the real defect.
Technicians set thresholds for their own machines, because they knew which vibration signature was normal for a 1998 unit.
Legacy modernisation
Why
A pricing engine written in 2009, 240,000 lines, no tests and no specification. It worked. Every proposed change was deferred because the risk was unknowable.
Hard
You cannot safely refactor code you cannot describe, and you cannot test behaviour nobody has written down. The business had routed around it for four years.
Architecture
What the code actually does under real inputs, including the three edge cases that turned out to be load-bearing.
Ninety days of real inputs and outputs became 11,400 test cases that pin current behaviour exactly.
With the net in place, a rewrite became a reviewable diff rather than a leap.
Two tests were updated on purpose, with a signed note explaining why. Everything else stayed byte-identical.
Also asked
Shorter write-ups on request — or ask about one that is not here.
Grounded answers over your own policies and runbooks, with the citation attached.
Every unit checked on the line instead of one in fifty, with defect classes your team retrains.
Forecasts that carry their own confidence interval, so planners know when to ignore them.
Routing and drafting, with the hard conversations still going to a person.
Finding what you pay for twice across an estate that grew by acquisition.
Controls evidence gathered continuously rather than assembled in a panic each audit.
Where it starts
Not with a model. With two weeks working out whether the problem was worth solving, and what a wrong answer would cost.
Residency and ownership shape the architecture more than the model does.
That gap sets where the human sits.
Most projects fail after launch, not during it.
The problem everyone agrees is real and nobody has scoped. Those are the ones worth a discovery.