
Fair helps people pursue compensation for mis-sold car finance. Every claim moves through the same journey: someone submits a claim, their identity and credit history get checked, their old finance agreements get matched, and those matched loans either qualify for compensation or they don't, depending on the lender and when the loan was taken out.
The problem was not that Fair lacked data. The problem was that the data lived inside a system built for running the product, not for answering business questions. A claim's real status, how many loans were matched to it, and whether those loans actually qualified existed buried inside raw records that changed shape from one claim to the next. Answering a simple question like "how many of this month's claims are actually going to pay out" meant someone manually piecing it together by hand.
As claim volume grew, so did the cost of not having a fast answer to a few core questions: where are claimants dropping off before they finish, how many loan matches and credit checks belong to each claim, which of those matched loans actually qualify once the real lender and date rules are applied, and is marketing spend bringing in claims that convert or just clicks that don't.
These are not exotic questions. Any claims business at scale needs to answer them daily. But when the underlying claim and loan data lives in a format built for the product rather than for reporting, every one of these questions turns into a manual investigation instead of a number someone can just look up.
Every system Fair relied on, connected into one warehouse:
The data existed everywhere. It just wasn't built to answer business questions in its raw form, so every report meant rebuilding the picture by hand, and no two people rebuilding it got the same answer.
The instinctive move is to hire a data engineer to pull the raw claim and loan records into a report and call it done. That works if the data is simple. It doesn't work when the real business logic, which lenders count, which dates qualify, what a completed identity check actually means, has to be applied the exact same way every single time, or the numbers quietly stop matching each other.
Fair didn't need someone manually reconciling claim data on request. It needed one system that decided, once, what "qualified" means and applied that same definition everywhere, so operations, underwriting, and growth were always looking at the same answer.
The approach ran on one rule: raw claim and loan data gets cleaned and standardized first, and every business decision, qualification, risk, claim value, gets built on top of that same clean foundation, never recalculated from scratch each time.
Every answer to "is this claim going to pay out" now comes from one place, not from whoever rebuilt the spreadsheet that week.
The work moved in two phases
Phase 1: Getting the foundation right. Every source of claim and loan data, along with risk, claims operations, and marketing data, got connected into one warehouse and cleaned into a consistent shape. This is the unglamorous part that most builds skip, and skipping it is exactly why the numbers stop matching later.
Phase 2: Making the data answer real business questions. With clean data in place, the team built the actual decision layer: which loans qualify, what a claim is worth, which marketing channels produce claims that convert. Dashboards went live on top of that layer for three audiences: day-to-day operations, case quality and prioritization, and growth.
Before, knowing a claim's real status, how many loans matched it, and whether those loans actually qualified meant manually reconstructing it from raw records every time someone asked. Marketing spend and claim outcomes lived in separate systems with no way to connect them.
Nobody on the team is reconstructing a claim's story from raw records anymore. The system does that once, and everyone works from the same answer.


The data foundation your whole team can work from — built by Datum labs.
Get in touch


Get notified when we publish. The patterns we're actually seeing across real client stacks, not theory.