DealSmart

Methodology

How the data is built.

Every dataset is worth exactly what its build is worth. Here is ours, stage by stage, including the parts that are hard and the parts we deliberately leave out.

1. The base layer

The foundation is UK statutory filing. Every company registered in the United Kingdom files with Companies House: incorporation details, annual accounts, the officers of the company, persons with significant control, registered charges, and a dated history of every document submitted. Those are separate datasets published in different formats and at different cadences, and none of them is joined to any of the others.

That is the raw material. It arrives as bulk files, document archives and per company endpoints — not as a database, and not in a shape anyone can query usefully. Turning it into one is the work described below.

2. Processing

Each dataset is parsed, normalised and reconciled against the others until every company resolves to a single record.

  • Officers. 29.5 million appointment records read out of a fixed-width bulk file, resolved to stable person identities, and joined to the companies they belong to.
  • Ages. Birth data is published as a month and a year with no day, one record at a time. We resolve it into a usable age for 96% of active companies and carry the uncertainty with it rather than pretending to a precision that does not exist.
  • Ownership. Persons-with-significant-control entries are decoded into who controls what, kept as the bands they genuinely are.
  • Accounts. Filed documents are parsed line by line out of their tagged source, so figures come from the filing rather than from a summary of it.
  • Charges. Secured lending read per company: outstanding, satisfied, floating, and who the lender is.
  • History. Every filing dated and classified, so a company's last four years reads as a sequence of events rather than a list of PDFs.

3. Enrichment

Processing produces something complete. Enrichment produces something useful. This is where AI and our own analysis do the work that cannot be done by parsing alone: classifying companies into sectors that reflect what a business actually does rather than the code it registered under, cleaning the inconsistencies that thirty years of filings accumulate, resolving entities that appear under several names, and transforming statutory language into fields a buyer can read.

Every enriched value is checked against the filed record it came from. Where the two disagree, the filing wins.

4. The signals layer

This is the part that does not exist anywhere else, and the reason the dataset is worth having. Signals are computed across the entire market at once, not looked up per company, so the question a searcher asks is which of these companies rather than does this one qualify.

Succession signals

  • A sole director past sixty, with no successor on the board
  • A board unchanged since incorporation
  • Two or more directors sharing a surname
  • An owner thirty years into the same company who has started nothing since
  • Concentrated ownership in a single pair of hands

Transaction signals

  • Secured debt newly taken, or newly satisfied
  • Capital events and share allotments
  • Board appointments and resignations out of pattern
  • Group restructuring, and where a company sits in a group
  • Dormancy, insolvency history, and overdue filings

5. Provenance, on every value

Each field in a company record carries where it came from, because a dataset that cannot tell you that is asking to be trusted rather than earning it. Four states, and they are never interchangeable:

  • Filed — read from a filing, shown exactly as filed.
  • Derived — computed by us, carrying the basis it was computed from.
  • Not filed — the company did not file it. A fact about the business, not a gap in our data.
  • Unreadable — we could not read it. It may well exist.

A zero standing in for “not filed” is a lie about a business, so the record never prints one.

6. What we deliberately do not hold

Some of the source material contains personal data that serves no purpose in a deal-origination product. It is dropped at the parse step, so it never lands in the dataset at all:

  • Officer service addresses. For a company this size that is very often somebody's home.
  • Full dates of birth. A year is enough to answer a succession question; a full date is identity data.
  • Street addresses. Company records carry a postcode, which keeps every geography filter working at a fraction of the exposure.

A field that never lands cannot leak, cannot be joined to by a query nobody reviewed, and does not have to be defended later.

7. One caveat we would rather state than bury

A registered office is where the post goes, not necessarily where the business is — roughly a third of small companies register at their accountant's address. Geography filters are labelled accordingly, and where a trading address can be established it is shown alongside rather than instead.

Attribution

Contains public sector information licensed under the Open Government Licence v3.0. Processing, enrichment, the signals layer and the resulting dataset are DealSmart's own work.

← Back to DealSmart