PEPPOL Statistics update

2026-09-18

We have another update to the way the Peppol Statistics numbers are generated. In this article I will describe what changed, why it was changed, and why the numbers are now slightly different from what they were.

The data source for our statistics

The statistics on our Peppol Statistics page are created from the data as provided by the Peppol Directory.

It should be noted that it only uses that data; we do not use a scanner or scraper to verify whether each record still exists, and whether they are in fact published with the document types as presented on the Peppol Directory. One should also keep in mind that publication in the directory is optional for Peppol end-users, so not every participant on the network is published.

Thirdly, the data is not what you might call completely clean and consistent; while most data in the directory is, there are also a number of cases where the name of the organization is simply ‘Unknown’, or ‘[nd]’ (the most occurring case).

Finally; and this is where actual changes in numbers come from: we try to aggregate on actual organizations, not peppol identifiers. Many participants use multiple publications, for instance one or more GLN numbers, or both their legal identifier from the national business register and their VAT number. However, on the network, and in the directory, these are published as separate records. Therefore we try to aggregate the records that appear to be representing the same organization, which is not an exact science; some organizations have different names for their publications, some are in countries where the VAT number is the same as their legal identifier. Some publications have information from their SP in the organization name, etc.

Why change the counting algorithm

With the growth of the network, the way we used to count was not suitable for the limited resources our servers had available for this. We’ve changed the backend, and with it came another evaluation of what exactly we’re counting.

Changes from the previous version

We are now using a different input file format, which may have some minor effects on names with special characters. We are also using a different statistics backend, which, while it should theoretically not result in any changes, may have some effects on the margins as well.

But the most important change is probably that we have stopped trying to compensate for ‘unknown’-type organization names. All organizations with the same name are now counted as a single participant on the network. We used to compensate for the names ‘unknown’, ‘—’ and ‘deleted company’, but another look in the data showed that these are not even near the top of the list of ‘shared’ names. This will cause a slight dip in numbers.

The full current algorithm

All numbers are based on the number of organizations, for which we aggregate the data using the following steps:

  1. VAT identifiers that are the same as local legal identifiers are counted as a single organization, for the following schemes:
  • NO:VAT (9909) and NO:ORG (0192)
  • BE:VAT (9925) and BE:EN (0208)
  • DE:VAT (9930) and DE:LWID (0204)
  • EE:VAT (9931) and EE:CC (0191)
  • LV:VAT (9939) and LV:URN (0218)
  1. If the name contains a section in brackets, for instance ‘Example Inc. (provided by FooBar)’, that section is removed.

  2. The name is lowercased, and leading whitespace is removed.

Some numbers are now different

We have not retroactively applied these changes to past data, so there may be a minor dip in some numbers; from a quick scan the dip is less than 1% but for the larger countries, it may be noticable.