> ## Documentation Index
> Fetch the complete documentation index at: https://docs.thedatacity.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Website Matching in Germany

> How we assign websites to companies in Germany, why it works differently to the UK for now, and why reporting a match helps.

<Warning>
  **The German product is in alpha.** Coverage and documentation are early and incomplete, and both will change.
</Warning>

Assigning websites to companies is a foundation of what we do at The Data City. The more accurate and precise our website matching is, the better our Real-Time Industrial Classifications ([RTICs](/our-data/rtics/what-are-rtics)), Real-Time Klassifikation der Wirtschaftszweige (RWZs), and your Smart Lists are.

<h2 id="methodology">
  Methodology
</h2>

For Germany, website matching currently works differently to our [UK approach](/our-data/key-data-and-definitions/website-matching). Rather than a dedicated machine learning model, we get each company's website from Creditsafe-sourced website data, supplemented by manually reviewed and reported matches.

The reason is training data. A reliable machine learning model needs a large volume of manually checked matches to learn from, and in the UK that volume took roughly a decade to build up. Germany doesn't have that history yet, so a model trained on what we hold today would not be stable enough for production. Creditsafe-sourced data and manual verification cover it in the meantime.

We monitor match quality in Germany. Feedback so far suggests the current approach is good enough for what users need, so nearer-term development has gone to other markets. We'll revisit a dedicated German model as the training data grows, or if quality needs change.

<Note>
  That makes reporting a match particularly useful at the moment. A report fixes the record you're looking at, and it adds to the training data a future model would need.
</Note>

<h2 id="current-coverage">
  Current coverage
</h2>

We have a website listed against roughly 900,000 German companies, and scraped website text for around 700,000 of those. The gap is expected: some listed websites are no longer live, and others haven't been scraped yet. A further 1.87 million companies have no website assigned at all.

Closing those two gaps, the unscraped websites and the unmatched companies, is the main way we improve German coverage in the near term. A dedicated model is the longer-term answer.

<h2 id="why-reporting-matters">
  Why reporting matters
</h2>

Every report goes back into our matching process: we check it, correct the record for that company, and add the verified match to the data set a German model would eventually train on. Wrong matches, missing ones, and out-of-date websites are all worth flagging. The more reports we get, the faster the coverage gaps above close and the sooner a model becomes viable.

<h2 id="enhancing-coverage">
  Enhancing coverage
</h2>

Our first enhancement of website matching coverage will be to run a machine learning model on those companies which don't have a website reported by Creditsafe. This will be trained on the website matches provided by Creditsafe, as well as on manual reports made by yourselves to close the coverage gap. This is an experimental methodology as we work towards running a stable model over the full universe of companies. This will affect only companies that have no website at present, so any introduced instability of website matching will be restricted to those companies. Currently matched companies in your lists, and those that you report, will not change except for if manually reported.

Spotted a website that looks wrong? Email [support@thedatacity.com](mailto:support@thedatacity.com) and let us know.
