how it works

Not a data feed. A read on what happens next.

Anyone can buy a copy of the register. Scout is something else: all 5,734,779 UK companies watched in real time, their filed history held point-in-time, enriched from every tier of source we can reach — open registers, accredited feeds, and proprietary layers we built ourselves — and then run through models that put a calibrated probability on the thing you actually care about: whether a business is heading for a sale.

Real time, whole market Every UK company, not a curated slice — and live: a business files, changes hands or takes on debt, and its score is recomputed within seconds, not on next month's data drop.
Machine learning Gradient-boosted models over 51 features and seven annual vintages of history, calibrated so a stated 10% means close to 10% of that cohort really did it.
Named patterns 40 hand-built archetypes — succession windows, cash fortresses, quiet compounders — each measured against the backtest, so you can hunt a shape and not just a score.
Where does the data come from?

We don't resell anybody's database. Scout is an index we assembled ourselves, and it draws on every tier of source available on the UK market:

  • Open public records — the statutory registers every UK company is legally obliged to file into, plus insolvency notices, public contract awards and property records. This is what makes coverage a census rather than a sample: the filing obligation means no company can be invisible to us.
  • Restricted and accredited feeds — sources that are not downloadable and not public: credentialed real-time channels and per-entity interfaces we hold access to, which carry depth the open files never contain.
  • Our own proprietary layers — and this is the part that cannot be bought from anyone. Web intelligence gathered first-hand on each shortlisted company; a machine read of the narrative text inside filed accounts, not just the tagged numbers; an identity graph that resolves the same operator across every company they touch; and seven years of point-in-time history we captured and keep ourselves so the past can be replayed honestly.

How those tiers are ingested, reconciled and kept in sync is our own engineering, and we keep the recipe in-house — it is a meaningful part of what makes the index hard to copy. What we will always be explicit about is the provenance of any individual figure we show you: every number in a dossier can be traced to the filing or the observation it came from.

What data do you hold on each business?

Financials come from the tagged accounts every company must file — so we have real balance-sheet series, not estimates:

30.3M
annual account rows parsed
6.2M
companies with filed history
2010
history reaches back to
  • Financial — net assets, cash, creditors, trade debtors, stock, fixed assets split into land & buildings / plant & machinery / vehicles, turnover, gross profit, profit, and average headcount. Plus derived measures: gearing, current ratio, Altman Z, Zmijewski and Taffler scores. Every figure is stored with the date we saw it, so the past can be replayed without hindsight.
  • Non-financial — ownership from the PSC register including the owner's birth month and year (so, their age — the single most under-used field in UK company data), corporate ownership chains up to the ultimate parent, directors and their appointment history across companies, charges and who holds the security (a high-street bank versus a private lender is itself the credit tell), filing punctuality, sector, incorporation date and region.
  • The text nobody indexes — an LLM reads the actual notes and audit report inside the filed accounts and extracts what the tagged data cannot carry: material going-concern uncertainty, the audit opinion, director loans and their direction, post-balance-sheet events, customer concentration, restatements.
  • Beyond the statutory filings — insolvency notices, public contract awards, property holdings and our own web intelligence, joined onto the same company record so one profile carries all of it.

Headcount deserves a note: small companies are exempt from disclosing turnover, but nearly all must disclose average employees. That makes the staff trend the best free proxy for revenue across the long tail — and it is why our models lean on it so heavily.

Do these businesses have a weak online presence?

Yes — dramatically so, and we measure it rather than guess. Our own web intelligence layer has profiled 68,962 shortlisted companies first-hand:

57%
no findable website at all
56%
of those with a site: no LinkedIn
14%
still serve over plain HTTP

Among the 43% that do have a site: 73% run no careers page, 13% are on visibly dated stacks, 10% aren't even mobile-responsive, and roughly one in eight scores 40+ on our neglect index — copyright year years out of date, a dead blog, a homepage the Wayback Machine says hasn't changed in three years.

We treat that as signal, not as a data-quality problem. A profitable, cash-generative business whose owner stopped investing in its shop window is precisely the succession story: the owner is coasting to retirement. So a neglected web presence sitting on a healthy balance sheet raises a company's ranking rather than disqualifying it.

The practical consequence for outreach is real, though: for this cohort you should expect the phone and the letter to work better than a form fill, and expect to find the decision-maker through the register — which is why every dossier links the owner to their full portfolio.

How honest are the probabilities?

The models are gradient-boosted rankers (LightGBM) over 51 features. They are trained on the 2018–2022 annual vintages and tested strictly walk-forward on later ones — the model never sees a year it is scored against, so nothing leaks backwards from the future.

Probabilities are then calibrated on a held-out vintage of 2.96 million companies over a 24-month outcome window, and the reliability is read off rows the calibration itself never touched. When Scout says 10%, close to 10% of that cohort really did change hands.

0.54%
market base rate for a sale in 24 months
17.3%
base rate for company failure
51
features behind each score

That first number is the whole point of the product: sales are rare. A card reading "≈13%, 23× the average" is not a modest number — it is a company the model puts two orders of magnitude above the field. Every score also ships its drivers, so you can see which facts moved it and decide whether you agree.

And it stays a signal, not a verdict. These are cohort estimates about a population, never a claim about one business's intentions.

What are the patterns, and how are they different from the score?

A score ranks. A pattern explains — and it is something you can point at and hunt. Each of the 40 archetypes is an explicit set of conditions over the register: Succession window (retirement-age owner, solvent, real trading company), Cash fortress (cash over half of net assets while headcount stalls — the owner is converting the business into liquidity), Quiet compounder, Never borrowed, Hidden distress, and so on.

Every pattern carries a measured lift from the walk-forward backtest, so you can see which shapes actually precede a sale rather than merely sounding clever. They are recomputed the moment a company files, and entering a pattern is itself an alert — that transition is the origination moment.

The same idea runs over people: nine owner archetypes (serial exiter, accumulating, winding down…) so you can spot the operator who is quietly selling down a portfolio before any single company of theirs looks interesting.

How current is it, and what does "real time" actually mean?

A persistent connection to the register's filing stream, running around the clock. When a filing lands we re-fetch what changed, recompute the company's patterns and re-score it. In practice that is seconds between a filing being published and the index reflecting it.

Underneath, each layer of the index refreshes on its own schedule — some daily, some weekly, some monthly — so financial depth and ownership stay current without waiting on a quarterly rebuild. The score history is journalled, so you can see a company's probability climbing year on year. Momentum usually matters more than the level: a score rising three years running is a better lead than a high flat one.

Can I get a list built to my own thesis?

Yes — that is the main way people use it. Tell us the shape you hunt (sector, size, region, owner age, balance-sheet posture, whatever defines your mandate) and we encode it as a pattern, run it across the whole index and its history, show you the backtest, and keep it re-screened live exactly like the standard set. New matches reach you the day they qualify.

Use the request form on your unlocks page, or reach Elijah Podavalkin on LinkedIn.

Is this legal, and what about people's data?

Everything here is computed from official public records — principally Companies House, which is published under the Open Government Licence, alongside other public sources. Owner names and ages come from the PSC register, which exists precisely so that company control is public.

That said, they are still personal data, and we treat them that way: public pages are anonymised and banded with a minimum cohort size so a filter can never single out an individual business, identifying detail sits behind an account, and we keep the provenance of every figure. Nothing here is an automated adverse decision about anybody, and the framing stays "worth a closer look", never a verdict.

Want a specific number of companies, to your own brief? Tell us the shape and the volume — Elijah Podavalkin comes back to you personally, usually the same day.
Request a brief