The Size and Shape of Public Entity Records
The structured records that identify organisations run to billions of facts, but organisational coverage in them is thin next to the number of companies that exist.
The records that machines use to identify an organisation are countable. Their headline totals are very large and their coverage of organisations specifically is not. A Wikipedia-derived knowledge graph holds roughly 346,000 organisations worldwide. The United Kingdom alone had 5,479,045 companies on its register in March 2026. The two numbers describe the same class of thing, and they are more than an order of magnitude apart. The gap is where most organisations sit.
The graph with no current figure
Google's Knowledge Graph is the record that matters most commercially and is the one with the least published detail. Writing in May 2020, Danny Sullivan, Public Liaison for Search at Google, stated that the graph had amassed over 500 billion facts about five billion entities. No method was disclosed and Google has published nothing newer. Any use of that figure has to carry its date; it describes 2020 and there is no 2026 equivalent.
The comparison with launch shows the trajectory. In May 2012 Amit Singhal announced the graph as containing more than 500 million objects and more than 3.5 billion facts. Between the two announcements entities grew tenfold and facts by roughly 143 times. Facts per entity rose from about seven to about a hundred. The graph did not mainly widen; it deepened around entities it already held.
The open records
The public equivalents publish live counters. Wikidata reported 123,079,455 items when retrieved on 2 September 2026, alongside 2,538,331,313 edits since launch and 41,077 active users. English Wikipedia reported 7,234,394 articles on the same date, across 66,185,860 total pages and 1,368,026,899 edits. Both pages are generated by live site counters, are cached, and change daily; either figure is a reading taken at a moment and should be stamped with the date it was taken.
The editor count is the figure that constrains everything else. Wikidata's 123 million items sit against 41,077 active users and 2,538,331,313 edits made across the project's life. That is roughly 3,000 items and 62,000 edits per active user. The identifier layer that a large part of the machine-readable web resolves against is maintained by a five-figure population, and its size is a published figure rather than an inference.
The ratio between the item and article counts is worth noting on its own. Wikidata holds roughly seventeen items for every English Wikipedia article. An item is not an article, and the two counts measure different things: identifiers on one side, written description on the other. The identifier corpus is much the larger of the two.
Where organisational coverage actually sits
Item counts do not say how many organisations are covered, because the classes are unevenly populated. DBpedia's published release statistics do say. Its Snapshot 2022-12 release, announced in March 2023, holds more than 850 million facts and 345,523 entities of the Organization class, against 1,792,308 Persons, 1,933,436 Species, 748,372 Places and 610,589 Works. The release carries 55,000 properties, of which 1,377 belong to the DBpedia ontology. These are DBpedia's own per-class counts, extracted from Wikipedia dumps, and Snapshot 2022-12 is the most recent release with a reachable announcement.
Organisations are outnumbered in that graph by species. A structured record derived from an encyclopaedia inherits the encyclopaedia's inclusion rules, and those rules were not written to enumerate commercial entities.
The registry, which is exact
The contrast is with an administrative census, where the count is complete by construction. Companies House reported that the total UK register size at the end of March 2026 was 5,479,045, an increase of 28,681 companies or 0.53%, with an effective register of 4,930,634 once companies in dissolution or liquidation are excluded, and 204,612 incorporations in the quarter alone. That is one jurisdiction. The quarter's new incorporations are more than half the organisation count of the entire DBpedia release.
A registry and a knowledge graph are answering different questions, which is why the counts diverge so far. The register records that an entity was incorporated, and inclusion follows automatically from the act. The graph records that an entity was written about, and inclusion follows from an editorial judgement made by someone else. Nothing in the first process triggers the second.
Aggregated registry data exists but is not quantified publicly. OpenCorporates states that it draws on over 140 government registries and other official sources around the world. It publishes no total entity count, and figures circulating for one should not be attached to it.
Implication
Two facts sit awkwardly together. The graphs that describe organisations to machines are large, and their organisational layer is small relative to the number of organisations that exist. Most companies are not absent from the machine-readable record by an editorial decision about them. They were never candidates for it. Registry presence is automatic and knowledge-graph presence is not, and nothing connects the two.
Control over the entry, where one exists, is weaker than it looks. Google describes the process precisely: many knowledge panels can be claimed by the subject they are about, and claiming allows a verified subject to provide feedback about potential changes or to suggest things such as a preferred photo. The operative verbs are provide feedback and suggest. Verification confers standing to request an amendment, not the ability to make one.
Two figures that would settle the practical questions do not exist. No published measurement states what share of registered companies hold a Wikidata item, and none states what share of Google knowledge panels derive from Wikipedia. Both are frequently asserted and neither is sourced. A firm that needs either number has to run the query itself against the public endpoints and publish it as its own measurement, with the query and the date attached, which is a different kind of claim from a citation and should be labelled as one.