# What is The Data City?
Source: https://docs.thedatacity.com/about/what-is-the-data-city
Learn more about The Data City platform and what we do.
### About The Data City
The Data City helps organisations understand what companies actually do. We built our platform to clear the fog around how the modern economy gets classified, so investors, government teams, and analysts can see it as it actually is.
Founded in 2017 and headquartered in Leeds, England, The Data City is a global data provider and SaaS platform built to fix a problem baked into economic data for decades: most business data still relies on classification systems and manual research that don't reflect how modern companies operate.
### The problem with SIC codes
Standard Industrial Classification (SIC) codes were never built for how the economy works today. They're assigned once, at incorporation, and rarely updated. A company working in AI, Net Zero, or FinTech can end up filed under a decades-old code that has nothing to do with what it actually does. Multiply that across the economy, and you lose the ability to find, measure, or invest in the sectors that matter most.
### **How Industry Engine solves it**
Industry Engine, The Data City's platform, classifies companies in real time using website text and machine learning, not by what SIC code was assigned when a company was set up. That's the basis for our Real-Time Industrial Classifications (RTICs): Over 500 sector classifications covering emerging and fast-moving industries, from AI and FinTech to AgriTech and Net Zero, updated continuously as sectors evolve.
Platform users build their own classifications and company lists with Smart Lists, explore individual companies in EXPLORE, spot trends in ANALYSE, and compare sectors side by side in COMPARE. Data comes from a mix of our own classification work and third-party sources, including Companies House, Lightcast, Creditsafe, Specter, Innovate UK, and 360Giving.
### Our story
In 2017, the founders of The Data City ran into the limits of SIC codes first-hand. Our CEO, Alex Craven, noticed that his own digital agency, and others doing the same work, weren't showing up in digital agency lists. The reason: there was no SIC code for what they did.
That gap became the starting point for The Data City. Alex and the founding team built RTICs and Industry Engine to solve the problem properly, not just for digital agencies, but for every emerging sector traditional classifications miss.
### **Who we work with**
Government departments, local authorities, financial institutions, investors, policy teams, and academic institutions use The Data City to get accurate, up-to-date insight for decisions that matter: where to invest, where to focus policy, and where the next wave of growth is happening.
### **Where we operate today**
The Data City started with deep coverage of the UK economy, and that's still where our data runs deepest. But the mission has always been bigger than one country. Today we deliver products and datasets across the US, France, Ireland, and Germany too, with more markets on the way.
### **Our mission**
The Data City's mission is to build the new global industrial classification system: replacing outdated codes with data that reflects how companies actually work, wherever they are.
# What method do you use for your machine learning classification?
Source: https://docs.thedatacity.com/faqs/machine-learning-classification-technique
The Data City's classification technique relies on the website text of a training set of companies.
The Data City's classification technique relies on the website text of a training set of companies. Companies like those we want to identify more of are selected, added to the training set, and labelled as includes. Companies unlike those we want to identify more of are selected, added to the training set, and labelled as excludes. The classification process takes all words and word pairs (tokens) from all the websites of the companies in this training set and then represents each company's website as a normalized vector (length 1) of the frequency of these tokens. The vectors for each company are multiplied by a vector (the classifier vector) made up of all the tokens, with each token having a variable score. These word scores are shown in the product UI.
The token weights in the classifier vector are varied until the two sets (includes and excludes) of companies in the training set are most highly separated, with as many of the includes as possible having scores above 0 and as many of the excludes as possible having scores below 0. This classifier vector is then used to score all website matched companies for which we have website text. The optimisation step described as "the token weights in the classifier vector are varied until the two sets (includes and excludes) of companies in the training set are most highly separated" can never be performed exactly due to the huge search space. But it can be performed efficiently and with reproducible results in almost all cases using a number of well-developed and tested algorithms.
In order to achieve the speed and explainability of our AI system, which is essential to allow sector experts to iteratively improve lists, we have developed our own custom algorithms for this purpose. Our on-going R\&D work focuses on creating: more detailed automated quality assessments of our lists, broadening the results so we miss fewer companies in every sector, and reporting a confidence score for every company's inclusion in a given sector.
# Getting started with The Industry Engine
Source: https://docs.thedatacity.com/getting-started-with-the-data-city
Get to know the core Data City platform and learn the basics of our company data.
Once you've created your account and received your login details, you're ready to start using the platform. Here's how to get going.
### Welcome to The Data City
Once logged in, the first thing you will see is your dashboard.
When you log in, you'll land on your dashboard. From here you can jump into saved searches, browse Real-Time Industrial Classifications (RTICs), and see what's new on the platform.
Any Smart Lists or EXPLORE lists you've saved or shared are accessible from your dashboard too.
### RTICs
Real-Time Industrial Classifications, or RTICs, are The Data City's proprietary industry classifications. The platform holds over \[confirm count] emerging economy sector classifications, from AgriTech and Net Zero to AI and FinTech, and it's the best place to start your research.
You can see key details for each RTIC from the RTICs tab, including RTIC code, creation date, revised date, companies, description, and verticals.
Click the Summary tab to see how each RTIC ranks across the platform. Filter and order sectors by turnover, employees, investment, and growth, or filter by region to spot breakout sectors near you. From there, click through to ANALYSE for a deeper look at sector trends.
**Tip**: You can also access RTICs directly from the Filters in EXPLORE, to see the full list of companies in a sector.
### **Smart Search**
Smart Search lets you find companies using phrases, concepts, or even full paragraphs, not just keywords. If you can describe the kind of company you're looking for in a sentence, Smart Search can find it. It's the fastest way to explore the database when you don't know the exact RTIC or SIC code you need.
**Building an Smart list**: Want to find out more? [View our full in-depth, step-by-step Smart List guide today.](/using-industry-engine/tools/building-an-ml-list)
### Explore
EXPLORE is where you see the company database, RTICs, and your own lists in full. It shows companies with key data highlighted for each, useful for training a Smart List or vetting prospects at a glance.
Click “Full Info” on any company for the complete picture: contact information, financial and funding data, growth measures, employee stats, sector classifications, locations, and more.
**Note**: Read more about our data, and how to make the most of it, [on our data page](https://help.thedatacity.com/knowledge/our-data).
Tailor your search with a wide range of filters, whether you're starting from scratch or refining an existing list, by RTIC, SIC code, location, keyword, financials, and growth measures.
**Using EXPLORE**: Find out more about our EXPLORE tool and how it works in our [Using Explore guide](/using-industry-engine/tools/explore).
### Analyse
ANALYSE takes your research further with a suite of dashboards and statistics, built for getting a macro view of a sector or list so you can spot trends and opportunities across larger datasets.
Every graph and table in ANALYSE can be filtered or edited: change the field, the data format, or the chart type to suit your analysis.
**Tip**: In the Locations dashboard, customise tables and graphs by local authority, OECD functional urban area, constituency, ITL1 region, and more.
Load Smart Lists directly into ANALYSE, jump into RTICs, or build your own analysis using the same filters as EXPLORE. Download your data for offline use, or view the full company list back in EXPLORE.
## **Building a Smart List**
Alongside EXPLORE's filters, you can build your own Smart Lists directly in the platform. The Data City's AI helps you build bespoke company lists in minutes, using the same technology behind our expert-backed RTICs.
Use Smart Lists to find companies in a niche sector, or upload a list of prospects to find lookalikes. To start, head to My Lists and select “Create a New List.”
You'll train the AI with example data: a minimum of five companies similar to the list you're trying to build.
Example: Building a FinTech list? Add five FinTech companies you already know, such as Revolut, Monzo, Starling Bank, Klarna, and Experian.
Once you've added at least five companies, the platform starts populating your list. The more inclusions and exclusions you make, the sharper your results get.
From here, use your Smart List in EXPLORE: view individual companies, refine with filters, download data, or explore insights in ANALYSE.
Note: Want the full walkthrough? Read our [step-by-step Smart List guide](https://docs.thedatacity.com/using-industry-engine/tools/building-an-ml-list).
### My Lists
Head to My lists for a full view of your lists and searches.
Create new lists, revisit recent searches, jump into Explore lists, or edit list details directly from here. You can also create folders to organise bigger projects across multiple lists.
### **Working with the data programmatically?**
Want to pull data directly into your own systems? The Data City's API gives you programmatic access to Industry Engine, including RTICs, company data, and our Global Company Data API covering the US, France, Germany, and Ireland. Head to the [API quickstart to get set up.](https://docs.thedatacity.com/using-industry-engine/features/api#api-access)
### What's next?
Now you know the basics, you're free to explore. Our Knowledge Base has step-by-step guides, use cases, and FAQs for everything else.
### Quick links
* [Glossary / Data Dictionary](https://docs.thedatacity.com/data-dictionary) - every data point we hold, in one place
* [Using our data safely](https://docs.thedatacity.com/our-data/using-our-data/using-our-data-safely#using-our-data-safely) - things to bear in mind, including a few limitations
Need some support? Please don't hesitate to get in touch with your account manager or a member of The Data City team.
# The Data City Knowledge Base
Source: https://docs.thedatacity.com/index
Help, methodology and how-tos for The Data City platform - RTICs, Industry Engine, our data sources and more.
**Looking for `help.thedatacity.com`?** You're in the right place. The knowledge base has moved here.
## Start here
A guided walkthrough of Industry Engine - list builder, EXPLORE, ANALYSE and RTICs.
A short primer on what the platform is and the problems it solves.
## Browse the knowledge base
RTICs, key data definitions, proprietary metrics, and third-party data sources.
EXPLORE, ANALYSE, COMPARE, Smart Search, Smart lists, filters, downloads and API access.
Common questions about classifications, locations, similar companies and more.
Information to support tenders that use The Data City data and platform.
## Popular topics
## Need more help?
Reach the team for product questions, demos, or to give feedback.
# How do you get greenhouse gas emissions data per company?
Source: https://docs.thedatacity.com/our-data/faqs/how-do-you-get-green-house-emissions-data
We source our greenhouse gas emissions data from official statistics published by the ONS and the Business Register Employment Survey.
The methodology behind greenhouse gas emissions data at the company level is very similar to [GVA's](/our-data/key-data-and-definitions/gva-data). Therefore, the limitations of greenhouse gas emissions data relate to the limitations of GVA estimates per company.
The original data we use to calculate the greenhouse gas emissions data [Atmospheric emissions: greenhouse gases by industry and gas](https://www.ons.gov.uk/economy/environmentalaccounts/datasets/ukenvironmentalaccountsatmosphericemissionsgreenhousegasemissionsbyeconomicsectorandgasunitedkingdom) and the [Business Register Employment Survey](https://www.ons.gov.uk/surveys/informationforbusinesses/businesssurveys/businessregisterandemploymentsurvey). These two tables allow us to develop a standard figure of greenhouse gas emissions per employee per SIC.
Then, we use the SIC and employee data we have available at the company level to calculate an estimate for each company. Hence, you will need to consider:
* **Companies may include international employees in their accounts**, so the total greenhouse gas emissions data may represent emissions produced by employees abroad
* **The SIC groupings developed by the ONS and BRES are very broad and we match them down to 5-digit SICs**. This may mean that two companies that do something different may be given the same standard greenhouse gas emissions measure.
* **Companies choose more than one SIC**. At this stage, we divide the number of employees of a company equally across all SICs selected by a company. Then, we multiply the split value for each SIC’s greenhouse gas emissions measure.
You will want to pay special attention to the SIC some companies select.
For example, *SHELL PLC* provides energy products, but they selected *SIC 70100: Activities of head offices*. This means we are estimating Shell's greenhouse gas emissions using the standard figure for *SIC 70100: Activities of Head Offices*, which has a much lower GHG value than one for an energy generation SIC.
Hence, GHG emissions estimates may not be representative of a company if a company selects a SIC not related to its actual economic activity.
**Limitations**: For more information about our data and its limitations, please make sure you read our [Using Our Data Safely guide](/our-data/using-our-data/using-our-data-safely).
# How are 'Similar Companies' identified?
Source: https://docs.thedatacity.com/our-data/faqs/similar-companies
You can find the most similar companies in the 'Full Info' of a selected company in either EXPLORE or ANALYSE. What does this mean, and how should it be used?
We offer two methods for finding similar companies: **Semantic Similarity** and **Composite Similarity**.
**Semantic Similarity** leverages cutting-edge LLM-based methods to identifying similarity between companies based on the text on their websites.
More detail can be found [here](https://thedatacity.com/blog/building-the-industry-engine-similarity-score-update/).
**Composite Similarity** combines the approach of the **Semantic Similarity** with a measure of similarity also based on structured business characteristics, such as sector, location, and employee count.
**Quick Guide**
* **Semantic** - broader exploration, more cross-sector matches. Useful for finding companies when building a Smart List, etc.
* **Composite** - sector-aware recommendations, structural alignment. Useful when you need matches that share fundamental business attributes, i.e. when company size and location are important filters or signals, or you're looking for true industry peers or competitors.
Try both and compare if you're unsure. Both return results in the same format.
**Please note**: This methodology is not the same as our classification engine, which we use to build our RTICs. It does not generate a comprehensive list of a sector, instead it only identifies the companies which are most similar to the given company
# Why are there duplicate companies?
Source: https://docs.thedatacity.com/our-data/faqs/why-are-there-duplicate-companies
No companies are duplicated in our database. What may appear as duplicates are actually multiple distinct entities.
The single source of our input data is [Companies House](/our-data/third-party-data/companies-house). Companies will often register multiple entities of what seems like the same company.
We hold group structure data to identify where companies are linked to one another in a parent-child relationship.
**Tesco as an example**
These are two of [Tesco](https://www.tesco.com/) entities:
They are registered as separate entities on Companies House even though you could refer to them as the same company.
As you can see in the screenshot below, TESCO STORES LIMITED is a child (or sub-child) of the ultimate parent, TESCO PLC:
We are currently researching a method to effectively collapse all child companies under the ultimate parent company within the platform.
You can remove duplicate companies using "Remove subsidiary companies":
You can read more about removing subsidiary companies [here](/our-data/using-our-data/what-are-and-how-to-remove-subsidiary-companies).
# Why do all companies not have RTICs?
Source: https://docs.thedatacity.com/our-data/faqs/why-do-all-companies-not-have-rtics
Find out why some companies in our platform are not classified in an RTIC.
We use website text data to classify companies into [RTICs](/our-data/rtics/what-are-rtics). Not all companies on the platform have an RTIC associated with them. These are the most common reasons as to why:
* We have not yet matched a company in [Companies House](/our-data/third-party-data/companies-house) with a website. We can only obtain text data after successfully linking a company from Companies House to its website. If we do not have website text data for a company, we won't be able to classify it into an RTIC. We have >4.5 million individual companies on the platform, of which >1.2m have a URL match.
* We built an RTIC before a company was founded or changed its website text. If we build an RTIC before the foundation of a company or a company's change of processes/technologies, it won't be in the RTIC until the next iteration.
* Some companies do not work in the emergent sectors we target. This especially applies to companies working in foundational industries.
* Some companies may use relevant processes or technologies but the way they describe them is not similar enough to the companies selected as training data.
* They are false negatives and we missed them. We acknowledge some companies' classification falls through the cracks. When we use the term "false negative" we mean the algorithm classified a company as not being relevant for the sector when it should be. Our [quality assurance process](/our-data/rtics/how-do-we-build-rtics) is designed to recover genuine companies that might otherwise be missed, and this is regularly fixed in later iterations of RTIC building and with the help of industry experts and our users.
If you find a company that has not been included in an RTIC you can report it on the platform. Reports are what trigger a [Hotfix](/our-data/rtics/updating-rtics#types-of-rtic-update) — a correction released outside the regular update cycle. You can follow the steps shown in the image below:
**RTICs**: Find out more about RTICs and how they are built in our [What are RTICs? guide](/our-data/rtics/what-are-rtics).
# Company locations
Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/company-locations
We use Companies House, company websites, and third-party data to provide company locations — including how we identify verified operating addresses and filter out non-genuine sites.
Our operating address definition has changed. In our June 2026 release, operating addresses are evidence-backed locations where we have a signal that a company trades from that site. A location can now be both a registered office and an operating address — companies that operate from their registered office are no longer excluded. We've also introduced a data-driven blacklist to filter virtual offices and formation agents, with exemptions for companies with website evidence of genuine presence.
## Where location data comes from
Our location data comes from three sources:
* Companies House (registered address)
* Company websites
* Creditsafe (external provider)
### Companies House
For the overwhelming majority of companies, a registered address is provided via Companies House. We use this postcode as the **registered address**. Companies are required to keep their own addresses up-to-date.
### Company websites
We extract additional postcodes from the text on company websites. We only look for postcodes on specific pages that are likely to contain correct location information, such as `www.example.com/locations`. The postcodes extracted from websites must match a standard UK postcode format. We then validate these against the ONS Postcode Directory.
### Creditsafe
We use additional postcodes provided by Creditsafe. Creditsafe use an external provider for trading addresses.
## Verified operating addresses
A location is classified as a **verified operating address** when there is evidence that a company conducts business from that site. This evidence can come from the company's own website or from Creditsafe trading-address data, and helps distinguish genuine trading locations from addresses used only for registration or correspondence.
A registered office is not treated as an operating address simply because it is registered with Companies House. When we find evidence that the company trades from that site, the same location can be both the registered office and a verified operating address. Many companies, particularly SMEs, operate out of their registered office.
### How we filter non-genuine operating addresses
Some postcodes, particularly those associated with virtual offices, serviced address providers, and professional formation agents, contain an unusually high number of registered companies. To prevent these from appearing as genuine operating locations, we apply a blacklisting rule.
A postcode is a candidate for blacklisting if it is an operating address and meets one of the following criteria:
* **High density**: The postcode falls within the top 0.5% by registered company frequency.
* **Professional services**: The postcode is flagged by Creditsafe as associated with solicitors or accountants, and the company's SIC code matches a relevant professional services classification.
### Exemptions
Even if a postcode is blacklisted, an operating address is retained if:
* **Website evidence**: The company explicitly lists the address on its own website.
* **SIC self-evidence**: The company is itself a solicitor or accountant operating from its registered office.
For the most extreme outlier postcodes, website evidence alone is not sufficient to override the blacklist — the address must also be the company's registered office.
## Filtering companies by location
When you filter a company list by a location (a local authority, a postcode and radius, or a custom area), we return any company with at least one address in that area. By default this considers **both registered and operating addresses**, so a company can appear on the strength of a single branch in the area even when its registered office is elsewhere.
You can narrow this:
* **Registered address only**: Match companies whose Companies House registered office is in the area.
* **Operating address only**: Match companies with a verified operating presence in the area (an address found on the company's website or via Creditsafe, and not blacklisted), wherever they are registered.
A *genuine* operating location is one that is an operating address and is **not** blacklisted. Blacklisting does not remove the operating-address flag; the two are separate, so a virtual-office or formation-agent postcode can be both an operating address and blacklisted.
### Telling addresses apart in a download
A download lists **every** address we hold for each company that matched, not only the address that fell inside your filter. So a download can include a company's addresses outside your area. These columns distinguish them:
* **Is Registered Postcode**: The Companies House registered office.
* **Is Operating Address**: There is evidence the company trades from this site.
* **Is Blacklisted**: A likely virtual office or formation agent. Combine with **Is Operating Address** to find genuine operating locations (operating and not blacklisted).
* **Is Subsidiary Address**: The address belongs to a subsidiary of the company rather than the company itself.
## Important notes
If multiple companies are matched to the same website, they will have the same postcodes extracted from that website.
A registered address is *not necessarily* the location of a company's head office — but it can be, and our system now correctly identifies these cases.
# Company sizes
Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/company-sizes
Our company sizes are based on the UK's legislative definition on company sizes.
Here are our company size definitions:
| **Company Size** | **Annual Turnover (£)** | **Number of Employees** | **Total Assets** |
| ---------------- | ----------------------- | ----------------------- | ------------------ |
| Micro | Less than 1m | Less than 10 | Less than £850,000 |
| Small | Less than 15m | Less than 50 | Less than £7.5m |
| Medium | Less than 54m | Less than 250 | Less than £27m |
| Large | More than 54m | More than 250 | More than £27m |
Our company sizes are based on the UK’s legislation definition on [company sizes](https://www.legislation.gov.uk/ukpga/2006/46/part/15). We also have a category for no employees or turnover.
A given company must meet **at least two** of the above conditions to be categorised accordingly. For example, a medium company may have 20 employees (small), but has £20m in turnover (medium) and £10m in assets (medium).
A company must have a registered postcode to be categorised into a company size bracket.
# Estimated turnover, employees and growth
Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth
We estimate turnover and employees for companies where we have enough data. We calculate growth rates to provide these estimates.
### Intro
Our data on company employee counts and company turnovers is provided by CreditSafe and is based on financial reportings to Companies House. We get data, per company, per year.
There is often no data or missing data for all or some years. Employee count data is more common than turnover data. Since there is a lag in financial reporting, we always use estimated employees and estimated turnover for the current year's values (this also helps to address missing data). Where we cannot estimate these values, we do not report them.
We have developed our own methods for estimating company growth rates even when only limited data is available, for example, where employee counts are reported infrequently or have not been reported recently. We only estimate values where we have enough data to do so reliably.
We use our company growth rates to estimate employees and to estimate turnover.
You will see our best estimates for current turnover and employee count in the company summaries for roughly half of all businesses. Those businesses without an estimate are very likely to have zero employees. You can expect these estimates to change monthly as these estimates update as we receive more data.
### Growth in our UI
The Growth tab shows our estimates of company employees and turnover by year: best estimate employee growth percentage per year and best estimate turnover growth percentage per year.
The best way for you to get a feel for our estimation algorithm is to look at the graphs in this growth tab for a few companies that you know well.
If a company has reported employee count for three years or more we fit an exponential curve to those years and use this to calculate an annual employee growth rate. We do the same for turnover.
Growth rate does not refer to year on year growth, but rather the average growth rate from the curve we fit. A year on year growth rate would not be able to handle missing data as well.
Across the product we default to using the employee based growth rate, for example, in filtering. This is to account for inflation.
Where a company has never reported employee count we assume zero employees.
Where a company has never reported turnover we assume zero turnover.
Where a company has reported employee count for only one or two years we estimate an employee count that starts in the first year an employee count is reported and continues at the level of the most recent year up to 2024 (the final year of our projections).
**Note**: Companies much more frequently report employee count than turnover. In the case of a company reporting only employee counts we estimate turnover based on the average turnover per employee for that company’s SIC codes.
### Caveats
* We constrain projections for sensible results.
* The method here works very well for the vast majority of companies. But some companies do very strange things with their annual accounts and these edge case can affect aggregate results especially when those companies are very large. You can read more about how we're addressing this [here](/our-data/proprietary-data/what-are-companies-with-potential-anomalies).
# Gross Value Added (GVA)
Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/gva-data
What GVA data does The Data City have? How should I use this data? What caveats are there?
**Please use this metric with caution.** We are happy to talk through your particular use case should you wish. [Get in touch to find out more](mailto:support@thedatacity.com).
The Data City estimates Gross Value Added (GVA) at company and RTIC level. This is available in ANALYSE and [EXPLORE](/using-industry-engine/tools/explore).
In EXPLORE, you will find GVA information in the financials tab of a company profile.
In ANALYSE, you will find GVA data in the analysis summary box.
GVA data in ANALYSE
### How is GVA estimated?
We have produced an estimated GVA measure at the company level using official GVA and employment data.This is an estimate of a company's UK GVA, if they have overseas operations.
1. The ONS produces GVA for 104 categories that can be matched to SIC codes. Similarly, the Business Register and Employment Survey (BRES) publishes sectoral employee data that can also be related to SIC codes. With this information, we have produced a measure of average GVA contribution per employee per SIC code.
2. The Data City have employee and SIC information for each company. Our employment data comes from Companies House, but we use profiles data from Lightcast to better estimate how many employees are UK based. It is therefore possible to estimate a GVA value per company by multiplying the company's UK employee count (sourced from Data City) by the average GVA employee contribution associated with that same company's SIC.
3. A company can have multiple SIC codes. Where this is the case, we split employees equally across the SIC codes and multiply each share of employees by the GVA per employee for each SIC code.
Expect GVA per employee values to range from £50,000 to £200,000 in magnitude.
### How to use this data?
When using GVA data for an RTIC or a list of companies, we strongly recommend reviewing the RTIC or list. You should check that the largest companies, by employee count, are appropriate for your analysis.
* For large companies, is this employee count in line with what you would expect for this company - is the employee count accurate?
* You may also want to consider whether large companies primarily operate in this sector, or whether their UK operations are not focused on this sector. You may want to remove companies that do not have significant UK operations, or if only a small part of their business falls within the selected sector or RTIC.
We try and estimate the UK GVA of companies using data on where employees are located from Lightcast. We have a large coverage for this data, but it's important to note not all companies have this data point therefore it should be used with caution still. We cannot estimate GVA for a company if we don't know which sector (SIC) they're in. There is a small proportion of companies that do not have SICs.
Not all companies have GVA data therefore estimated GVA per employee is calculated by dividing the total UK-adjusted GVA (sum of each company's EstimatedGVA multiplied by their UK employee proportion) by the total UK employees from companies that have GVA data available. These values are presented in the pop-up box.
> Σ(EstimatedGVA × UKProportion) /Σ(BestEstimateUKEmployees\*)
>
> \*where companies have EstimatedGVA > 0
You might want to consider the impact that selecting the "exclude all ultimately foreign companies" filter has on the GVA value. Note: this will not remove UK based multinational firms that have large global employment, such as BP.
**Note**: You should also be aware that our estimate is of GVA per full-time employee. GVA might be overestimated in sectors which are more likely to have part-time employees, for example, in recruiting agencies.
# Industrial Strategy Classifications
Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/industrial-strategy-classifications
The UK government's Industrial Strategy identifies eight sectors — the IS-8 — as the drivers of UK economic growth over the next decade. The Data City has built precise, data-driven definitions for each, designed to be tracked and evaluated over time.
This page explains how those definitions work, why they differ from traditional classification approaches, and what evidence underpins them.
***
## The problem with SIC codes
Standard Industrial Classification (SIC) codes are the conventional way of categorising UK businesses. They have three well-documented limitations for Industrial Strategy analysis:
**They are backwards-looking.** SIC codes were last updated in 2007. Sectors like AI, quantum computing, or clean energy cannot be reliably identified using them — companies in these sectors typically register under broad codes like "IT consultancy activities".
**They only capture primary activity.** A company's SIC code reflects what it mainly does. Dual-use businesses — common in defence and fintech — are systematically miscategorised. A financial firm using AI has a different SIC code from a tech firm providing financial services, even though both are FinTech companies.
**Self-selection introduces noise.** Businesses choose their own SIC codes, and many choose incorrectly or not at all. Around 26,000 active companies declare "activities of head offices" — a vague catch-all that obscures what they actually do.
The government acknowledges these limitations explicitly. For three of the IS-8 sectors (Clean Energy, Defence, and Digital and Technologies), it states that no reliable SIC-based definition exists.
***
## How The Data City defines the IS-8
The Data City uses two proprietary classification systems to overcome the limitations of SIC.
### Real-Time Industrial Classifications (RTICs)
RTICs identify companies operating in sectors that SIC codes cannot capture. They are built using machine learning trained on company website text, combined with expert input from government departments, industry bodies and academics.
RTICs are live — they update continuously as companies' activities evolve — and explainable, meaning analysts can inspect why a company has been included or excluded.
Frontier sectors within each IS-8 sector are mapped to dedicated RTICs. Many of these have been co-designed directly with government, including DSIT, Innovate UK, and DCMS.
### Real-Time SIC Codes (RSICs)
RSICs address the self-selection problem. For sectors where SIC codes exist but are unreliable, RSICs use website text and machine learning to assign companies up to four more accurate SIC codes, correcting for vague declarations and outdated classifications.
RSICs are used in sectors like Creative Industries, Financial Services and Professional and Business Services, where the underlying SIC framework is sound but its application by businesses is not.
***
## Confidence ratings
The Data City infers a confidence rating for each IS-8 sector based on how well existing government definitions capture the sector's activities.
| Rating | Sectors | What it means |
| ---------- | ------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------- |
| **Low** | Clean Energy, Defence, Digital and Technologies | Government acknowledges no reliable SIC definition exists. Definition relies entirely on RTICs. |
| **Medium** | Advanced Manufacturing | A SIC-based proxy exists but misses significant portions of the sector. RTICs provide a more precise view. |
| **High** | Creative Industries, Financial Services, Life Sciences, Professional and Business Services | SIC codes broadly capture the sector. RSICs correct misclassifications; RTICs extend coverage to frontier areas. |
***
## The IS-8 sectors
### IS01 — Advanced Manufacturing
Covers medium-high technology firms across advanced materials, aerospace, agritech, automotive manufacturing, batteries and space.
SIC codes are used by government as a proxy, but analysis shows they identify only 4% of companies in the sector as innovative, compared to 12% under The Data City's RTIC-based definition — a three-fold difference. Both approaches produce a broadly similar overall business count (\~25,000 companies), indicating that the RTIC definition is capturing the right population with greater precision.
Frontier RTICs were developed in partnership with DSIT and Innovate UK. Three were co-created directly with government teams.
***
### IS02 — Clean Energy Industries
Covers carbon capture, heat pumps, hydrogen, nuclear fission, nuclear fusion, and offshore and onshore wind.
The government acknowledges that SIC codes are too restrictive for this sector and currently relies on the Low Carbon and Renewable Energy Economy (LCREE) survey. Survey-based estimates carry an inherent margin of error, making it difficult to attribute changes in sector size to specific policies with confidence.
The Data City's RTIC-based definition provides a company-level view that can be consistently monitored over time without survey uncertainty. The Net Zero RTIC underpins this approach and has been adopted in published academic research, including work that informed the Skidmore Review.
***
### IS03 — Creative Industries
Covers advertising and marketing, film and TV, music, performing and visual arts, and video games. Also includes **CreaTech** — the 4,879 companies at the intersection of creative industries and digital technology.
The government's Creative Industries definition dates to 1998 and is well-established. The Data City aligns with it but uses RSICs to correct misclassifications. Our business count is broadly consistent with official data.
CreaTech is defined as companies operating in both the creative industries (excluding software development) and the Digital and Technologies sector. The government is committed to better capturing this subsector through future SIC revisions.
***
### IS04 — Defence
Covers companies in land, sea, air, cyber and dual-use defence technologies.
The government acknowledges SIC codes are insufficient for this sector, primarily because defence businesses commonly serve both civil and defence markets. A company's SIC code records its primary activity, making dual-use firms invisible in SIC-based analysis.
The Data City's methodology combines RSICs and a dedicated Defence RTIC. Company website text captures dual-use activity that a single SIC code cannot. This definition is being expanded in partnership with ADS Group.
***
### IS05 — Digital and Technologies
Covers artificial intelligence, cybersecurity, engineering biology, quantum technologies, semiconductors and advanced connectivity.
DSIT has acknowledged that SIC codes cannot define this sector with the granularity modern policy requires. Around half of companies in this sector carry SIC codes that fall outside the digital economy as DSIT defines it, meaning a SIC-based approach would miss them entirely.
The Data City's definition relies entirely on RTICs. The RTIC-based approach is well-established: it underpins the government's own revised methodology for measuring the UK digital economy, developed with DSIT, Cambridge Econometrics and The Innovation and Research Caucus and published in July 2025. The specific RTICs used here differ from that publication but share the same methodological foundation.
RTICs covering AI, Cyber, Quantum and Advanced Connectivity were co-designed directly with government departments.
***
### IS06 — Financial Services
Covers asset management, capital markets, FinTech, insurance and reinsurance, and sustainable finance.
Financial Services is one of the highest-confidence IS-8 sectors. SIC codes capture most of the sector well; RSICs correct systematic misclassifications in complex group structures where companies declare head office SIC codes, ensuring the sector is accurately sized.
Where SIC falls short is at the frontier. FinTech is captured through a dedicated RTIC developed with Innovate Finance, and Sustainable Finance through the Green Finance RTIC, developed with WPI Economics.
***
### IS07 — Life Sciences
Covers biopharma and medtech, alongside broader life sciences activity in medical devices, diagnostics and research.
The government advocates for the definition created by the Office for Life Sciences (OLS), which publishes an open dataset of companies in Life Sciences frontier sectors. The Data City's definition aligns with this. RSICs correct common misclassifications, and RTICs extend coverage to frontier areas. This makes Life Sciences one of the best-evidenced IS-8 sectors, with an open, shared evidence base already in place.
***
### IS08 — Professional and Business Services
Covers accounting and audit, legal services and management consultancy, alongside a wide range of B2B professional activity.
Professional and Business Services is a high-confidence sector. RSICs correct the misclassifications that affect SIC-based analysis — particularly the \~26,000 active companies that declare "activities of head offices" but are actually in management consultancy, IT services or other professional activities — ensuring the sector is accurately represented.
***
## Further reading
Full methodology, sector comparisons and business count data are available in The Data City's discussion paper:
\[Open Sourcing the Industrial Strategy]\([https://thedatacity.com/reports/open-sourcing-the-industrial-strategy/](https://thedatacity.com/reports/open-sourcing-the-industrial-strategy/))
# Location Quotients
Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/location-quotients
What are location quotients? How do I interpret them? What do I need to know about the data?
### What are location quotients?
Location quotients are a measure of relative concentration of an industry in an area.They are useful for comparing areas of different sizes.
Traditional location analysis sometimes overlooks the size of the area. Location quotients account for this.
### How are they calculated?
Location quotients compare the local presence of an industry with the national presence of the industry.
I.e. they measure the presence or size of the industry against what would be expected for an area of this size, based on the national average for the industry.
The formula for calculating employee based location quotients is included at the bottom of the page, as an example, to illustrate how location quotients are calculated.
### Example
They help us to answer questions such as "Which industries are overrepresented in Leeds?" and “Where is industry X overrepresented?”.
Leeds is likely to have a smaller absolute count of businesses than London, but after accounting for the sizes of the geographical economies, location quotients could identify Leeds as having the greater relative concentration of businesses in a specific sector.
### How do I interpret location quotients?
Location quotients are a unit-less measure.
A value of 1 means that the proportion of companies in industry X in an area is the same as the proportion found nationally. This means you are equally as likely to find industry X in an area as you would across the nation.
A value of 2 means you are twice as likely to find industry X in an area as you would across the nation.
Conversely, a value of 0.5 means you are half as likely to find industry X in an areas as you would across the nation.
Location quotients are found on ANALYSE, under the locations tab.
Above: The top 5 local authorities where Agency Market companies are overrepresented.
### What should I know about the data?
The Data City provide location quotients for business count, employees and turnover. Location quotients using employees and turnover are subject to [the wider caveats of our employee and turnover data](/our-data/using-our-data/using-our-data-safely#creditsafe-employee-data-incorrect).
1. Similarly, local authorities that have a very small business base are more prone to very large location quotient values. For example the Isles of Scilly has 82 businesses.
Imagine that across the UK that 5% of UK businesses are restaurants. If the Isles of Scilly had 30 restaurants, that would not be a particularly high number of restaurants in absolute terms. In relative terms this would be a large proportion of the overall business base (37%). This would indicate a strong overrepresentation of restaurants in Isles of Scilly, with a location quotient of 7.4. **Some care is required in using location quotients for particularly small local authorities.**
2. Location data can be subject to biases, such as the registered office effect. UK companies of all sizes often register in London despite having minimal *real* activity in London.
Large global companies are more likely to be registered in London than other regions. This results in oddities, such as London having large location quotients for tobacco and mining companies.
3. Additionally, multinational companies' employee counts, as reported in their annual accounts, often report [all global employees](/our-data/using-our-data/using-our-data-safely#creditsafe-employee-data-global). This will have an impact on location quotients.
#### Location quotient formula example - employees
# Scale-up definition
Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/scale-ups
Our scale-up definition is the OECD's definition of a scale-up, applied to filed employment data. This page sets out the exact calculation, step by step.
**Definition**
Our scale-up flag looks at a single growth window: it ends at the company's most recent filed accounts and starts at the most recent filed year at least three years earlier. Within that window the company needs:
* 10 or more employees at the start of the window
* at least 4 filed years of employment data with 10 or more employees
* no rise of more than tenfold between consecutive filed figures
* annualised employment growth of 20% or more per year from start to end
This follows the Eurostat-OECD definition of a high-growth enterprise, the basis of the term "scale-up", measured on employment over the most recent three years of filed data. The rest of this page sets out exactly how we apply it, including the edge cases, so you can reproduce any company's flag from its filed accounts.
## The data behind the flag
* **Filed accounts only.** Employee figures come from companies' filed annual accounts. Each figure is attached to the year the accounts were made up to, giving at most one employment figure per filed year. Our estimated and projected employment series are never used for the scale-up flag.
* **Employment, not turnover, by decision.** The OECD definition allows growth to be measured by employees or by turnover. We apply the employment measure only: employee counts are disclosed far more consistently in UK filings than turnover, which smaller companies often do not file. A company scaling revenue on a flat headcount will not be flagged.
* **Anomaly filtering.** A figure only counts if it is greater than zero and has not been flagged by our anomaly detection (for example an implausible headcount for the size of the business). Anomalous figures are skipped entirely; they never influence the flag.
* **Missing figures are skipped, not zero.** Roughly 4 in 10 filed accounts carry no employee figure at all, most commonly because the filing does not include a captured headcount. A year without a figure is not treated as zero employees; the window simply starts or ends at the nearest year that has one.
* **Growth is growth in the filed headcount, however it arose.** We cannot fully distinguish organic hiring from acquisitions, intra-group staff transfers, or changes in how a group allocates employees between its entities; the OECD guidance acknowledges the same limitation. The tenfold cap below removes the worst of it: mis-transcribed figures and switches to consolidated group accounts arrive as huge single-year steps, while 99% of genuine qualifiers never step more than about eightfold between filings. A large step under the cap is still worth checking against the company's accounts before reading it as organic scaling.
* **Each release uses the six most recent years of accounts.** The flag is recalculated with every data release; we hold the six most recent years of filed accounts (currently accounts made up to 2020 onwards). A company's flag can therefore change between releases as new accounts arrive or old years leave the held history.
## The calculation, step by step
1. Take the company's reliable filed employment figures (positive, non-anomalous), ordered by year.
2. The window **ends at the most recent** of those figures.
3. The window **starts at the most recent filed year at least three years earlier**. If no filed year is that old, the company cannot qualify yet.
4. The company is a scale-up if all four of these hold:
* the start-year figure is 10 or more employees
* at least 4 filed years inside the window (start and end inclusive) have 10 or more employees
* no figure inside the window is more than ten times the previous filed figure
* annualised growth across the window is at least 20% per year
Annualised growth is the compound rate between the two window endpoints:
$$
\text{annualised growth} = (\text{end employees} \, / \, \text{start employees})^{1/n} - 1
$$
where $n$ is the window length in years (end year minus start year). A company qualifies when this is at least 0.20. No rounding is applied before the comparison, and the years between the endpoints do not enter the growth formula; they matter only for the four-filed-years requirement.
The window is fixed by data availability alone: it is always the shortest one the filings allow. The start moves to an older year only when nearer years carry no figure, a filed start is never skipped in search of a better growth rate, and a start below 10 employees fails rather than reaching further back. There is exactly one window to check per company, which keeps the flag focused on recent growth and makes it straightforward to reproduce.
## Worked examples
These are real companies, with employee counts exactly as filed in their annual accounts and held in our July 2026 release. You can look any of them up on the platform by company number. Their figures, and in some cases their flags, will change as new accounts arrive.
**Steady growth qualifies.** Principle Estate Services Limited (11056986):
| Year | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 |
| --------- | ---- | ---- | ---- | ---- | ---- | ---- |
| Employees | 19 | 30 | 44 | 54 | 68 | 99 |
The window runs from 2022, the most recent filed year at least three years before the 2025 accounts. It starts at 44 (10+), has 4 filed 10+ years, and grows $(99/44)^{1/3} - 1 = 31.0\%$ per year. Scale-up.
**Growth must be recent.** Olfasense UK Ltd (02900894):
| Year | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 |
| --------- | ---- | ---- | ---- | ---- | ---- | ---- |
| Employees | 22 | 22 | 46 | 44 | 47 | 48 |
The window runs from 2022 to 2025 and annualises at $(48/46)^{1/3} - 1 = 1.4\%$ per year: not a scale-up. The company more than doubled between 2021 and 2022, and a window drawn from 2021 would average $(48/22)^{1/4} - 1 = 21.5\%$ per year, but 2022 is a filed year, so the window starts there. An early jump cannot carry a recent plateau.
**The flag follows the window as new accounts arrive.** UD Restaurants Ltd (10515301):
| Year | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 |
| --------- | ---- | ---- | ---- | ---- | ---- | ---- |
| Employees | 19 | 26 | 48 | 40 | 41 | 45 |
When the 2023 accounts were the latest, the window ran from 2020 and annualised at $(40/19)^{1/3} - 1 = 28.2\%$ per year: a scale-up. With the 2024 accounts the window moved to 2021 and fell to $(41/26)^{1/3} - 1 = 16.4\%$. With the 2025 accounts it moved to 2022 and fell to $(45/48)^{1/3} - 1 = -2.1\%$. The flag dropped as soon as the growth stopped being recent.
**Years without figures are bridged.** TSC Kent Ltd (10853210):
| Year | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 |
| --------- | --------- | ---- | --------- | ---- | ---- | ---- |
| Employees | no figure | 16 | no figure | 37 | 36 | 38 |
Three years before the 2025 accounts is 2022, which carries no figure, so the window starts at the next older filed year, 2021. It starts at 16 (10+), has 4 filed 10+ years (2021, 2023, 2024, 2025), and grows $(38/16)^{1/4} - 1 = 24.1\%$ per year. Scale-up. The missing year still counts towards elapsed time (we divide over 4 years, not 3 observations), so bridging never inflates a growth rate.
**Bridging stops at the nearest filed year.** Royal London Asset Management Limited (02244297):
| Year | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 |
| --------- | ---- | ---- | --------- | --------- | ---- | ---- |
| Employees | 221 | 384 | no figure | no figure | 534 | 591 |
The 2022 and 2023 accounts carry no employee figure, so the window starts at 2021 and annualises at $(591/384)^{1/4} - 1 = 11.4\%$ per year: not a scale-up. A window from 2020 would average $(591/221)^{1/5} - 1 = 21.7\%$, but 2021 is a filed year and is never skipped in search of a better rate.
**Implausible steps are rejected.** A real series, anonymised: a national charity that employs around 5,000 people, whose early years were transcribed wrongly at source:
| Year | 2020 | 2021 | 2022 | 2023 | 2024 |
| --------- | ---- | ---- | ---- | ---- | ---- |
| Employees | 50 | 56 | 5193 | 5165 | 5163 |
The window from 2021 to 2024 would annualise at $(5163/56)^{1/3} - 1 = 350\%$ per year, but the 2021 to 2022 step is a 93-fold rise: far beyond anything hiring can do, and the signature of a data error or a switch to consolidated group accounts. Any rise of more than tenfold between consecutive filed figures inside the window disqualifies the company. Steps under the cap pass: 99% of genuine scale-ups never exceed about eightfold.
**Reaching 10 employees only recently is not enough.** Alexa Capital Limited (10759666):
| Year | 2020 | 2021 | 2022 | 2023 | 2024 |
| --------- | ---- | ---- | ---- | ---- | ---- |
| Employees | 3 | 3 | 6 | 8 | 10 |
The window starts at 2021, which has 3 employees: below the 10-employee floor, so the company is not a scale-up, however fast it is growing. This is the OECD's own threshold, which stops very small bases producing inflated growth rates. The floor applies at the actual window start; a start below 10 is never bridged past.
## Reproducing the flag from an export
Platform exports that include the year-by-year financial history (`CompanyFinancialsCreditSafe`) contain everything the flag is computed from:
1. Take the `Reported_Numberofemployees` figures by year, dropping years where the figure is missing or zero, or where `DeclaredEmployeesAnomalous` is true.
2. The window ends at the latest remaining year and starts at the most recent remaining year at least three years earlier.
3. Check the four criteria: 10 or more employees at the start, four 10-or-more years inside the window, no figure more than ten times the previous filed figure inside the window, and annualised growth of at least 0.20 between the two endpoints, using the formula above.
If a recomputation disagrees, the usual causes are skipping the anomaly filter in step 1 or comparing across releases: the platform recalculates with every data release, and both figures and flags move as new accounts arrive and old years leave the six-year history.
## Background
The definition of an OECD scale-up company is what we have implemented on the platform. This is the OECD definition ([source](https://committees.parliament.uk/writtenevidence/109251/pdf/)):
"*All enterprises with average annualised growth greater than 20% per annum, over a three-year period should be considered as high-growth enterprises. Growth can be measured by the number of employees or by turnover.*"
Our general growth-rate metric and company size definitions are different: those do use estimated and projected figures. You can read more about those estimates and average annual growth rates [here](https://thedatacity.com/blog/focusing-on-company-growth/).
The scale-up filter is located within the growth tab:
Using the OECD's definition to determine a company's growth stage enhances the ease of making global comparisons. The OECD does not provide a definition for a start-up.
We appreciate there are many different definitions of company size and company growth stages, and which one is right for you will depend on the purpose of your analysis.
Our platform still allows for custom definitions using the filters bar, particularly within the financial tab:
# Website Matching
Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/website-matching
Assigning websites to companies is a foundation of what we do at The Data City. The more accurate and precise our website matching is, the better our Real-Time Industrial Classifications (RTICs), Real-Time Standard Industrial Classifications (RSICs), and your Smart Lists are.
### The short version
From a range of sources, we assemble a set of potential websites for each company. We call these *candidates*.
We then scrape up to 25 pages of each candidate website.
Finally, we apply a machine learning model to select the best candidate for each company.
This model is trained on manually checked website matches for thousands of companies collected over the last decade.
We measure the success of our model continuously. The two most important metrics are:
* *Accuracy*: how often we pick the correct website, or no website if the company does not have one.
* *Precision*: how often we pick the correct website, or no website (regardless of whether the company does actually have a website).
**Accuracy** tells us how likely a website match that you see in our product is to be correct.
**Precision** is a broader measure. It considers that sometimes we won’t match small or newly incorporated companies with basic websites to that website.
In our latest releases our model regularly exceeds 90% accuracy, and 97% precision.
### The long version
V6 of our industry engine represented our biggest ever step forward in website matching.
We switched from a logical scoring method to a machine learning model trained on top of year's worth of manually collected data.
You can read about this in more detail in our blog post *[Industry Engine V6: What's new?](https://thedatacity.com/blog/industry-engine-v6-whats-new/)*
# Environmental, Social and corporate Governance (ESG)
Source: https://docs.thedatacity.com/our-data/proprietary-data/esg-statements
We identify sub-pages of company websites which likely relate to ESG
## Why we provide ESG data
ESG refers to a set of standards used to measure an organisation's environmental and social impact.
The increasing importance of ESG made us curious about what we could contribute.
Since we scrape company websites, it is possible for us to indentify pages of a company's website that are likely to contain ESG statements.
## How we indentify ESG statements
When we scrape a company's website, we collect information from up to 25 pages of the website.
We compare each URL that we scrape to a list of ESG keywords. If the URL contains any of these words, we assume that the web content for that sub-page relates to ESG.
A few examples of the words we're looking for are: *sustainable*, *environment* and *gender*.
In addition, we check the internal links on each scraped page for any further matches to our list of ESG keywords.
## Where to find ESG data
On a company page, click the `ESG (beta)` tab. If we have identified any ESG statement pages for that company, they will appear here.
## Next steps
This data is currently released under beta version status. That means it could contain some errors.
We like to ship early and often, and we will continue to improve this data.
# Gender Data
Source: https://docs.thedatacity.com/our-data/proprietary-data/gender-data
What gender data is available? How do The Data City estimate the founders of companies? What do I need to know to use this data?
**New in the June 2026 update** — gender leadership is now a single, mutually-exclusive category per company (six in total), applied separately to directors and founders, with a symmetric two-thirds threshold for a "majority". Every company is placed in exactly one category, so the counts aggregate cleanly.
### What is available and how are founders identified?
The Data City's platform can analyse the leaders and founders of a company.
Leaders are the active directors (officers) listed on Companies House. We place every company in one of six mutually-exclusive gender categories based on its active directors — and, separately, on its founders (see below).
We identify founders by looking at directors appointed **within 2 years and 28 days** of a company's incorporation. We add a 28-day buffer to account for standard administrative reporting delays permitted under UK corporate law. Specifically, the [Companies House "14+14" rule](https://www.gov.uk/guidance/people-with-significant-control-pscs) grants a company up to 14 days to update its internal register of People with Significant Control (PSC) following an appointment, and an additional 14 days to formally submit this information to the public register. Founders are specifically a **person with significant control** and they are *still* an **active director** at the company.
A [person of significant control](https://www.gov.uk/guidance/people-with-significant-control-pscs) is defined by Government. In short it means a person likely has voting rights, shares, or a controlling influence in a company.
Companies are required to declare persons of significant control. It is possible to have information protected, however.
### How we categorise leadership and founders
Gender is assigned from each officer's Companies House title (e.g. Mr, Mrs). Titles without a gender (e.g. Dr, Prof) count as **unknown** and are still included in the total. The Data City **does not** use machine learning, and never uses names, to estimate gender.
Each company is placed in exactly one of six categories, based on the share of active directors with a female-gendered title (`X`) or a male-gendered title (`Y`):
| Category | Definition |
| ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **All women led** | `X = 100%` — every active director has a female-gendered title |
| **Majority women led** | `66.7% < X < 100%` — more than two-thirds, but not all |
| **Mixed led** | `33.3% ≤ X ≤ 66.7%` — between a third and two-thirds have a female-gendered title |
| **Majority men led** | `66.7% < Y < 100%` — more than two-thirds, but not all, have a male-gendered title |
| **All men led** | `Y = 100%` — every active director has a male-gendered title |
| **Uncertain** | none of the above — including companies with no active directors, and boards where unknown-gender titles mean the category can't be determined without guessing |
The percentages are rounded for display; the boundaries are applied as exact thirds, so there is no rounding drift. A clear majority (more than two-thirds, but not all) is only possible on a team of four or more, so most companies fall into *All*, *Mixed* or *Uncertain* rather than a *Majority* category.
The same six categories are applied separately to a company's **founders** (using the share of founders with each title).
Because the categories are mutually exclusive, every company sits in exactly one — so you can aggregate them safely (for example, "women-led businesses are X% of companies"). A company is **women led** when it is *All women led* or *Majority women led* — i.e. more than two-thirds of its active directors have a female-gendered title.
**Why "Uncertain"?** If a company has, say, two women, four men, and one director with an unknown-gender title, that final title could tip the company into either *Mixed led* or *Majority men led*. Because we never guess gender from a name, the honest answer is *Uncertain*.
In some instances we are not able to identify founders, or identify the genders of founders.
Firstly, we rely on persons of significant control from Companies House. Persons of significant control is legislation that was introduced in 2015 and our ability to identify founders before then is reduced. An example is included in [appendix one](#appendix-one). In addition, sometimes the true founders of a company are not listed as persons of significant control, because of their initial level of equity in a company not meeting the threshold for significant control. Lastly, we are not able to identify the gender of a founder (or leader) where they have used a unisex title — these companies fall into the *Uncertain* category.
In both leader and founder data, gender is based on the declared titles of officers on Companies House. Data City **does not** use machine learning to estimate the gender.
You should be careful when using this data to compare summary statistics between men and women led (or founded) businesses. We recommend you [remove outliers](/our-data/proprietary-data/what-are-companies-with-potential-anomalies) from your analysis.
### Where is the data?
#### ANALYSE
In ANALYSE, you can find gender data in the analysis summary box and in the company details panel.
In the company details panel, the **Gender** section shows two 100% stacked bars — the founder gender mix and the director gender mix — each split into the six categories above. This communicates the proportion of companies in your list that are women led, men led, mixed, or uncertain, for both founders and directors.
#### EXPLORE
In EXPLORE, you can find gender data in the people tab of each company page, under the **Woman led statistics** header. This includes the women founder and women officer counts, whether the company is women led, and the company's director and founder gender categories.
#### Filtering
In ANALYSE and EXPLORE you can filter companies by any of the six categories — separately for directors (leaders) and for founders. Ticking more than one category in a group returns companies in any of them. These options are available under the company filter.
In [November 2023](https://thedatacity.com/blog/new-founder-gender-data-in-platform/#:~:text=The%20women%20led%20statistics%20here,a%20majority%20women%20led%20business.) we wrote an article detailing definitions, coverage and drawbacks of this approach.
### Appendix one: ANN SUMMERS LTD. (01034349)
As per Ann Summers' website:
> 10th December 1971, the first Ann Summers shop opened in Marble Arch. Their Executive Chair, Jacqueline Gold, joins her dad's business as an intern with a brilliant idea - Tupperware parties, but for Ann Summers Product. The Ann Summers party is born.
On Companies House, their incorporation date is [10 December 1971](https://find-and-update.company-information.service.gov.uk/company/01034349). However, on Companies House, their earliest officer (director) is David Gold who was appointed "before 1991":\
There are two issues here for founder analysis using Companies House data:
1. We do not know the exact year David was appointed as a director.
2. The incorporation date is 20 years before the director appointment date.
So although it's quite clear from the website, it's unclear using the data hence why we are unable to identify any founders regardless of gender:
# Innovation Score
Source: https://docs.thedatacity.com/our-data/proprietary-data/innovation-score
What does the Innovation score represent, and how is it calculated?
### The short version:
Our Innovation indicator uses a proprietary Machine Learning model to estimate how innovative every company in our database is based on their websites, where available. In the absence of any better method, we use R\&D intensity as a proxy for innovation.
The model is trained on 980 companies with known R\&D intensities (R\&D £ expenditure per employee) and applied to 1.6 million companies in the UK to estimate whether they are innovative or not. A 3-star rating system is used to indicate our confidence in the estimation.
### The long version:
Defining and measuring "innovation" is difficult.
Data which indicates whether a company is innovative or not does not exist for all UK companies, and predicting unknown innovation is also tricky.
The Central Bureau of Statistics of the Netherlands (CBS) have shown that [the website text of a company can accurately predict its innovativeness score.](https://www.cbs.nl/-/media/innovatie/using-website-texts-todetect-innovative-companies.pdf) But obtaining the training data to replicate this in the UK is not straightforward. Where proxy data which could be used to estimate unknown company innovation does exist, it is either private (see [the ONS UK Innovation Survey](https://www.ons.gov.uk/surveys/informationforbusinesses/businesssurveys/ukinnovationsurvey)), or difficult to capture.
R\&D spending, a proxy for innovation, is not required in annually reported accounts, and is rarely voluntarily reported. Where it is reported, it is regularly marked improperly, rendering the field in the machine-readable XBRL format accounts unreadable.
After parsing 1.5TB of machine-readable accounts in XBRL format and experimenting with OCR at scale, we have managed to capture R\&D spending data for 980 UK registered companies. These companies operate across all regions of the UK and all industrial sectors, and cover a wide range of R\&D spending and business sizes. From this we have calculated R\&D intensity (R\&D spending £ per employee).
Combined with company website text, this provides a solid source of training data from which we have developed a Machine Learning method to estimate if a company is significantly more likely to be highly innovative, based only on the content of its website. A measure of 0-3 stars is then applied based on our confidence in the innovation likelihood predicted (0 stars - not innovative; 1 star - innovative, low confidence; 2 star - innovative, medium confidence; 3 stars - innovative, high confidence).
Though we do not recommend users use the innovation indicators to build lists, this star rating system allows users to filter for innovative companies within the entire company database, or within specific lists/RTICs.
**Filtering companies by innovation score is included in both our EXPLORE and ANALYSE platforms.**
#### The Innovation Score appears as a number and not a category. What is this?
If you're using the API, you will see the raw innovation score. To convert the raw score into the confidence rating mentioned above, use the following logic:
3 star = score >= 3
2 star = 1.5 >= score \< 3
1 star = 0 \< score > 1.5
# Linked companies
Source: https://docs.thedatacity.com/our-data/proprietary-data/linked-companies
## What are linked companies?
Linked companies data is an extension of our group structure data.
Sets of companies can form “groups” that are connected, but are registered as separate companies for financial, legal and organisational reasons.
Group structure data explains how the sets of companies are connected, the structure of their connections, and the directors that connect them. Creditsafe uses a rules-based criteria to identify companies that form these groups.
However, sometimes companies can be “linked” together, but are not marked as being part of the same group structure by Creditsafe because they do not meet their criteria.
By eye, we can see that these companies are linked in some way; usually it is because they have similar (or identical) company names, websites, directors, shareholders, registered addresses, or SIC codes and more.
Here is an example. The company [**A Shade Greener Limited**](https://products.thedatacity.com/companypage/?company_id=06922318) is not part of a group structure. It is its own ultimate parent company.
Our linked companies tab shows that there are in fact 63 companies linked to **A Shade Greener Limited**. All of these share the same postcode, and many have similar names.
For example, [A Shade Greener Member LLP](https://products.thedatacity.com/companypage/?company_id=OC386772). Both companies are registered at the same address *STERLING HOUSE MAPLE COURT, MAPLE ROAD, S75 3DP* and have the same person of significant control *Mr Stewart James Davies* born in 1951.
We think these companies are likely to be part of the same organisation, even though our group structure data does not show this link.
When these type of companies are *not* considered to be in a group structure, we can end up “double-counting” companies that are effectively one economic unit. This can make economic analysis more challenging by inflating the figures.
## How do we identify Linked Companies?
We use a machine learning model trained on hundreds of thousands of data points to estimate the probability that any two companies are linked. We consider companies to be linked if they exceed 97% probability.
The data behind this model includes
* company names
* company addresses
* company websites
* director names
* director addresses
* director birthdates
* persons of significant control (shareholders)
and more. These data points are used to find potential links. We then train our model by providing it pairs of companies that are extremely likely to be linked.
## Where can I find Linked Companies data?
You can see a company's *linked companies* by clicking on the "Linked Companies" tab on a company page.
We currently have around 460k "clusters" or "groups" of linked companies, which contain approximately 1.5million companies in total.
As of Industry Engine v6.1, this data is released under "beta" version status, which means it could contain some errors. We are actively working to keep improving this data.
## How we're using Linked Companies data
When counting companies in a sector, multiple registered companies can represent the same economic entity, for example, if they share a director, location, or website. Linked Companies identifies these relationships.
Distinct Brands already aimed to solve this by counting the likely unique economic entities rather than raw company counts. It's now powered by Linked Companies data: where a group of linked companies exists, only the earliest-incorporated one counts as a Distinct Brand.
This makes Distinct Brands a more accurate measure of the true number of independent companies in your analysis.
# What are companies with potential anomalies?
Source: https://docs.thedatacity.com/our-data/proprietary-data/what-are-companies-with-potential-anomalies
The Data City has identified companies with accounts that are likely misrepresented. We will answer: Why have we added this feature? What are some examples of anomalies in accounts? What is the basis for the predicted anomalies?
### Why have we added this feature?
Our data covers all active companies registered at Companies House. As well as a breadth of data, we have depth of data. We have the ability to drill down to company level financials.
To offer this depth of data, Companies House data uses the submission of financial accounts by each company. This is a mandated process.
However, self-declared (and especially unaudited) company accounts can contain mistakes. Companies House are not responsible for verifying company accounts:
*"We carry out basic checks on documents received to make sure that they have been fully completed and signed, but we do not have the statutory power or capability to verify the accuracy of the information that companies send to us."*
\~[Companies House](https://resources.companieshouse.gov.uk/serviceInformation.shtml)
Outliers affect less than 0.05% of our companies but their impact, by nature, can be large. Identifying possible outliers will allow our subscribers to review and remove the very small number of companies that have bad data.
### What are some examples of anomalies in accounts?
Below is a **non-exhaustive** list of anomalies that can occur in financial accounts.
#### Companies can report another financial variable as their employees.
When a company is filling in their accounts, in some instances, they will copy values from another field.
For example, [GILLARDS FARMS LIMITED](https://find-and-update.company-information.service.gov.uk/company/00981261/filing-history) in their 2021 accounts report assets as their number of employees. You can see this in the two images below:
***
#### Companies can report the year as their number of employees.
For example, [ABBOTT & ABBOTT LIMITED](https://find-and-update.company-information.service.gov.uk/company/09150841/filing-history) revised their 2016 employees as '2016', in their 2017 accounts.
#### Companies can report their wage costs as the number of employees
For example, [AR CARS CLUB LIMITED](https://find-and-update.company-information.service.gov.uk/company/10798168/filing-history) in their 2017 accounts report director renumeration as the number of employees.
In a small number of cases, anomalies are introduced by parsing of company accounts. The number of examples of this is very low, with a general accuracy of over 99.9%.
The model has been trained to identify these outliers too. In addition, we are working with our data provider to improve the parsing process, which will reduce the frequency of the outliers further in the future.
The examples mentioned above are where financial accounts are misleading. In training the model to identify unusual financials, we have also identified companies that correctly have unusual financials.
In particular, the model will also identify companies that can have high levels of employments with low levels of resources. Specific examples of these include recruitment or healthcare agencies.
In these companies, employees are added to a company's books, but it is another company that is funding salaries through their sales. The model will identify companies likely using signficant temporary or part-time employment, if it appears the companies financials are not sufficient to support that level of employment.
Removing agencies, or companies with temporary workers, will be beneficial for any analysis using Gross Value Added or turnover. Our calculations of GVA rely explicitly on the number of employees referring to the number of full-time employees. To estimate turnover, we rely implicitly on the assumption that each employee is a full-time employee.
### What is the basis for predicted anomalies?
We trained a model to predict whether a company has anomalous financials. To do this, we started with known examples of anomalies. We used the model to predict anomalies, validating the predictions, and incorporating these into a training set. We completed this iteration over 20 times.
The validation process was manual. For thousands of companies, this involved inspecting their accounts on Companies House and understanding where the reported values had come from.
We now have a training set of over 10,000 companies and we will continue to review predictions, incorporating new kinds of outliers, if and when we find them. If you are aware of outliers that we are missing, please get [in touch](mailto:andrew.purdy@thedatacity.com).
# What are Distinct Brands?
Source: https://docs.thedatacity.com/our-data/proprietary-data/what-are-distinct-brands
Long-time users of our data will know that group structure has always been a challenge. Behind what looks like a single business, there can be layers of subsidiaries, brands and parent organisations shaping activity, employment and revenue. And there’s no consistent way in which this happens.
A group can contain many different brands. In the case of Tesco, this includes the overarching Tesco PLC brand, but also Dunhumby – Tesco’s analytics subsidiary – and Tesco Finance.
“Distinct brands” has been created to help you quickly find the meaningful companies within a group. For reliability, companies that file dormant accounts, or appear inactive, cannot be considered distinct brands.
## Where can I find Distinct Brands?
To better answer the question of ‘how many companies are there in a sector or list?’, we’ve brought distinct brands to analyse.
The image above shows the number of distinct brands in the Net Zero sector. We think there are 20,500 meaningful companies in the Net Zero sector, across 28,000 registered companies.
You can also find flags for distinct brands in downloads from explore, and on a company's group structure tab.
# What are RSICs and how do we ensure data quality?
Source: https://docs.thedatacity.com/our-data/proprietary-data/what-are-rsics
Companies can misreport their activities, by selecting the wrong SIC code. We have solved this issue.
RTICs are great for the emerging economy. For the foundational economy, where there is more likely to be an appropriate SIC, an issue remains. A company can choose the wrong SIC code, or the SIC code they've selected does not match their activities.
We fixed this issue. Real-Time Standard Industrial Classifications (RSICs) use machine learning and a company's website text to better classify company's activities.
For example, Shell PLC is a large energy company. Because they are large, the only SIC code they file at Companies House is activities of head offices. A better description of their activities is provided through RSICs (and RTICs): extraction of crude petroleum and natural gas, mineral oil refining, and the wholesale of fuels and petroleum products.
RSICs follow the same structure as SICs.
A reminder, our RSICs:
* Fill gaps where SIC codes are missing
* Correct inaccuracies in existing SIC codes
* Add granularity where SIC codes are vague
You can read more about RSICs [here](https://thedatacity.com/blog/sic-codes-fixed-introducing-real-time-sic-codes-rsics/).
### Data Quality
To ensure quality we focus on *methodological integrity*:
**Trust in the methodology**
Primarily, we’ve built trust in our RSICs *within* the RSIC methodology itself. We do this through three distinct layers:
1. **Evidence, not prediction:** We treat classification as an evidence problem, not a prediction problem. RSICs are not arbitrary predictions. Instead, we evaluate the *empirical likelihood* of a classification based on the company’s website text, and what we uniquely understand about companies in each sector. If the data doesn't support the code, we don't assign it.
2. **Coherence Filtering:** This validation layer which rejects codes that lack alignment with the company's specific niche. This allows us to distinguish between a company *mentioning* a topic and actually *doing* it. We identify this distinction, and we classify appropriately.
3. **Specificity:** We also penalise generic classifications. Broad, "catch-all" codes are rarely useful for decision-making, so we deprioritise them in favour of precise definitions. Companies spread across more of the classification instead of piling into a handful of catch-all codes.
**Trust in transparency**
Unlike black box AI models where the logic is hidden, our RSIC system is built on transparency. Every classification is traceable back to the specific evidence that supports it. The framework is auditable, and we remain in control.
**Quality in everything**
Quality RSICs rely on quality inputs. By prioritising quality in everything, beginning with high-fidelity website matching and cutting-edge website text analysis, we build trust at every step of the pipeline.
That includes knowing when an input isn't good enough. Not every page we find behind a company's website is really a website: some are bot checks, cookie walls, holding pages or errors. A language model will use those perfectly happily, and the result looks well-formed while telling you nothing true. So we check every website before we use it. Where we can't produce RSICs we trust, a company gets no RSIC rather than a misleading one.
# How do we build RTICs?
Source: https://docs.thedatacity.com/our-data/rtics/how-do-we-build-rtics
How an RTIC is built: The Data City's approach to classifying real economies in real time.
An [RTIC (Real-Time Industrial Classification)](/our-data/rtics/what-are-rtics) is The Data City's classification of UK companies into precisely defined sectors and sub-sectors. Standard [SIC codes are often too vague, or too out of date](/our-data/rtics/what-is-the-difference-between-rtics-sic-codes), to capture emerging and fast-moving industries — RTICs are built to describe what companies actually do, based on the evidence they publish about themselves.
### How it works
Every RTIC is produced through a rigorous, multi-stage methodology that combines our proprietary AI technology with expert analyst judgement at every step:
1. **Define** — the sector is given a rigorous written definition — its scope, its boundaries, and real companies that exemplify it — developed with input from industry and academic experts wherever possible.
2. **Classify** — our proprietary AI identifies the companies that genuinely belong in the sector, assessing each against the sector's definition using the evidence of what that company actually says about itself. Our technology draws on a continuously refreshed view of hundreds of thousands of [matched UK company websites](/our-data/key-data-and-definitions/website-matching).
3. **Quality-assure** — every list goes through multiple independent layers of AI-assisted quality assurance, designed to catch false positives, recover genuine companies that might otherwise be missed, and sense-check the sector's overall shape and statistics before anything is published.
4. **Sign off** — analysts review the results throughout, and a manager formally signs off every RTIC before publication. No AI output goes live on its own.
These stages run as a continuous loop, not a one-way pipeline: when real companies test a sector's boundaries, the definition is refined and the affected work re-run — for as long as it takes to get the sector right.
Because classification is grounded in website evidence, a company needs a [matched website](/our-data/key-data-and-definitions/website-matching) to be classified into an RTIC. See [why not all companies have RTICs](/our-data/faqs/why-do-all-companies-not-have-rtics).
### Kept current, not built once
Published RTICs don't stand still. Newly identified UK companies are continuously assessed against every published sector, and an analyst reviews every proposed addition before it joins a live list. Periodic full updates refresh the sector definition and company list together — and each update comes with a report explaining what changed and why, so you can always understand how your list has evolved.
Three types of update keep an RTIC current: an annual **Full Update**, a monthly **Maintenance Update**, and ad hoc **Hotfixes** raised from reports on the platform. Each is reflected in the RTIC's version number. See [RTIC update types and versioning](/our-data/rtics/updating-rtics).
### What you can rely on
* **Human oversight, always.** An analyst reviews, and can override, every AI output. Ambiguous and borderline companies are decided by expert judgement, not by an algorithm alone.
* **Evidence, not guesswork.** Classification decisions are grounded in what companies actually publish about themselves. Where evidence can't be found, a company is flagged for review — never guessed at.
* **Full traceability.** Every decision is versioned and auditable: any classification can be traced back to the sector definition and evidence that produced it.
* **Honest accuracy.** Where we quote a confidence or accuracy figure, it comes with the reasoning behind it — we show our working, not just a headline number.
A detailed description of our methodology is available on request — email [support@thedatacity.com](mailto:support@thedatacity.com).
# RTIC update types and versioning
Source: https://docs.thedatacity.com/our-data/rtics/updating-rtics
The three types of update that keep RTICs current, and how an RTIC's version number reflects its update history.
A quick-reference overview of how [RTICs](/our-data/rtics/what-are-rtics) are kept up to date, and how that update history is reflected in an RTIC's version number.
### Types of RTIC update
RTICs are kept accurate and relevant through three types of update:
| Update type | Cadence | What it covers |
| ---------------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
| **Full Update** | Annual | A full refresh of the RTIC vertical: taxonomy review, updated training set, and a rebuilt company list. |
| **Maintenance Update** | Monthly | New companies are classified into existing verticals using our proprietary AI, then reviewed before release. |
| **Hotfix** | Ad hoc | Released in response to mismatch or suggestion reports raised on the platform; corrects a small number of companies outside the regular cycle. |
Hotfixes are how your reports reach the live data. If you spot a company that's missing from an RTIC, or classified into one it doesn't belong in, [report it on the platform](/our-data/faqs/why-do-all-companies-not-have-rtics) — corrections are released without waiting for the next scheduled update.
### Versioning
Every update is reflected in the RTIC's version number, using semantic versioning: Major.Minor.Patch (for example, 1.3.0).
| Version segment | Maps to | Meaning |
| ----------------- | ------------------ | -------------------------------------------------------------------------------------------------------------- |
| **Major** (X.0.0) | Full Update | A full refresh of the RTIC vertical, typically once a year; may include reclassification and taxonomy changes. |
| **Minor** (0.X.0) | Maintenance Update | Monthly classification of new or updated companies into existing verticals; reviewed before release. |
| **Patch** (0.0.X) | Hotfix | Ad hoc correction for a small number of companies, raised via mismatch/suggestion reports. |
**Example:** V 3.0.4 — this RTIC vertical has undergone three Full Updates, and four Hotfixes.
A detailed description of our update process is available on request — email [support@thedatacity.com](mailto:support@thedatacity.com).
# What are RTICs?
Source: https://docs.thedatacity.com/our-data/rtics/what-are-rtics
Learn what RTICs are, why you can trust them and how they are used.
### What is an RTIC?
Real-Time Industrial Classifications (RTICs) are The Data City's classification of UK companies into precisely defined sectors and sub-sectors, built with our proprietary AI technology and grounded in the evidence companies publish about themselves.
RTICs are a modern classification approach, unlike the traditional SIC system that relies on predetermined and static categories. They are based on how companies describe themselves on their websites.
RTICs are exclusively frontier sector classifications — sectors and activities that sit outside SIC's coverage altogether. They're not a replacement for SIC codes; for correcting or extending classification within SIC's existing scope, see [RSICs](/our-data/proprietary-data/what-are-rsics).
The output is a dataset that gathers companies working in the same field. In this dataset, you will find all the data available for the companies in an RTIC. The data can be explored on the platform or downloaded.
**How are RTIC different from SIC codes?** Find out more about the differences between the two classification systems [here](/our-data/rtics/what-is-the-difference-between-rtics-sic-codes)
### Why can you trust RTICs?
Every RTIC combines our proprietary AI technology with expert analyst judgement: an analyst reviews, and can override, every AI output, classification decisions are grounded in what companies actually publish about themselves, and every decision is versioned and auditable. See [how we build RTICs](/our-data/rtics/how-do-we-build-rtics) for the full methodology.
RTICs are also built alongside experts in their respective fields. As an example, our various 'Space' RTICs (space economy, space energy, in-orbit space manufacturing etc.) were built in collaboration with the Satellite Applications Catapult. They helped us both build the taxonomy, as well as check which companies belonged in the training set.
Some of our RTICs were built with our own internal experts. For example, Software Development was built in collaboration with our development team.
**Kept current**: Published RTICs are continuously maintained — newly identified companies are assessed against every published sector, and periodic full updates refresh the definition and list together. See [how we build RTICs](/our-data/rtics/how-do-we-build-rtics#kept-current-not-built-once).
### How can you use RTICs?
RTICs provide unique insights into emerging economies that SIC codes cannot.
View RTICs at a top level to get a glimpse into the sector as a whole, or combine them with filters to get a more refined view.
RTICs can be found under the RTICs section in the platform.
Use location filters to get regional insights, or growth filters to find high-growth companies and much more.
This page offers an overview of all the RTICs available on the platform, including their definitions, creation date, latest update, number of companies, and details about the included verticals. These RTICs can be selected for further exploration using the [EXPLORE](/using-industry-engine/tools/explore) function.
**EXPLORE**: Find out more about our EXPLORE tool and how to search our company database [here](/using-industry-engine/tools/explore).
# What is the difference between RTICs & SIC codes?
Source: https://docs.thedatacity.com/our-data/rtics/what-is-the-difference-between-rtics-sic-codes
RTICs classify frontier sectors that sit outside SIC's scope entirely. Find out how that differs from the static, predetermined categories of SIC codes.
Understanding the difference between Real-Time Industrial Classifications (RTICs) and Standard Industrial Classification (SIC) codes is crucial, especially in the context of emerging UK economy.
### Definition and purpose
* **RTICs**: They represent a modern approach to classifying companies, developed by The Data City. RTICs focus on providing a real-time, accurate representation of the emerging economy by leveraging advanced technology like machine learning and expert input. They are designed to reflect the latest developments and innovations in various sectors.
* **SIC Codes**: These are traditional systems for classifying industries, intended to categorise companies by their primary business activities. However, these codes are now criticised for being outdated and not reflecting modern economic sectors. A major issue with SIC codes is their assignment process. Companies must select one to four SIC codes at their inception, but often choose codes out of convenience or due to limited options. This leads to misclassification and inaccuracies in economic data. Moreover, the SIC system hasn't been updated since 2007, failing to capture new sectors like AI and cybersecurity. This outdated framework limits its usefulness in economic analysis, policy-making and investment planning, highlighting the need for a more accurate and dynamic classification system.
### Technology and methodology
* **RTICs:** Utilise a combination of machine learning algorithms and expert training. This way we can continuously [update classifications](/our-data/rtics/how-do-we-build-rtics#kept-current-not-built-once) based on the most recent market trends and company activities by monitoring their online presence and other data sources.
* **SIC Codes:** Rely on a predetermined and static set of categories that were established through traditional economic studies and analyses. These categories are not dynamically updated and may not capture emerging industries effectively.
### Scope and coverage
* **RTICs:** Cover over 400 emerging economy sectors, including Net Zero, AgriTech, FinTech, Artificial Intelligence, and more. This wide range is particularly effective in identifying and classifying companies in cutting-edge and rapidly evolving sectors. Find all our latest RTICs [here.](https://thedatacity.com/rtics/)
* **SIC Codes:** Have a more limited scope in terms of emerging industries. They are less effective at categorising companies in newer sectors that have developed since the last major update of the SIC system. The latest (2007) update of the SIC codes can be found [here.](https://resources.companieshouse.gov.uk/sic/)
### Customisation and flexibility
* **RTICs**: Offer the ability to create custom classifications, allowing users to tailor the system to their specific needs and to map out niche or emerging sectors effectively.
* **SIC Codes**: Lack the flexibility for customisation. They adhere to a fixed set of categories, which might not suit all analytical or industrial needs, especially in the context of modern, dynamic economies.
### Accuracy and relevance
* **RTICs:** Aim to provide a more accurate and up-to-date picture of the economy. They are particularly useful for identifying companies that are often hidden in the 'not elsewhere classified' (n.e.c.) categories of SIC codes.
* **SIC Codes:** Can be less accurate in representing the current state of the economy, especially for newer industries, due to their infrequent updates and static nature.
| SIC Code | SIC Description | # of companies |
| -------- | -------------------------------------------------------------- | -------------- |
| 82990 | Other business support service activities n.e.c. | 249,233 |
| 96090 | Other service activities n.e.c. | 162,888 |
| 74909 | Other professional, scientific and technical activities n.e.c. | 83,693 |
| 64209 | Activities of other holding companies n.e.c. | 68,595 |
| 43999 | Other specialised construction activities n.e.c. | 64,272 |
| 85590 | Other education n.e.c. | 47,328 |
| 63990 | Other information service activities n.e.c. | 28,409 |
| 32990 | Other manufacturing n.e.c. | 22,181 |
| 42990 | Construction of other civil engineering projects n.e.c. | 20,921 |
| 93290 | Other amusement and recreation activities n.e.c. | 20,874 |
*This table shows the top 10 n.e.c. SIC codes by the amount of companies classified.*
RTICs aren't a replacement for SIC codes — they're a frontier-only classification, covering sectors and activities that sit outside SIC's scope altogether. Within SIC's existing scope, [RSICs](/our-data/proprietary-data/what-are-rsics) correct and add granularity to SIC codes themselves.
# 360Giving
Source: https://docs.thedatacity.com/our-data/third-party-data/360giving-data
360Giving uses a wide variety of sources which publish data about their grants to the 360Giving Data Standard.
As per their website, 360Giving has developed the [360Giving Data Standard](https://www.threesixtygiving.org/data-standard/) so it can be easily compared with data from other organisations. People can have a more informed understanding of the UK grantmaking picture.
360Giving uses a wide [variety of datasets](https://data.threesixtygiving.org/) to help funders publish open data about who, what and where they fund, using the 360Giving Data Standard. These datasets have been curated into a search-engine for grants data.
We process this data monthly and present it on a per company basis on the platform.
A great example to view this data to view ROLLS-ROYCE PLC's (01003142) company page. Within the growth tab:
## Negative, or zeroed, award amounts
Sometimes award amounts will have 0 or even negative values. This can signify that the grant has been refunded, or reissued. You can read more about this [here](https://www.360giving.org/explore/before-you-start/what-to-look-for/#Negative-zero-grants).
# B Corps
Source: https://docs.thedatacity.com/our-data/third-party-data/b-corps
Certified B Corp™ data is available in the Industry Engine. Filter companies by certification status, see how many companies in a list are certified, and include the flag in your **downloads**.
B Corp™ status is matched to company records using company name and postcode data provided by B Lab UK. Where no confident match is found, records are left unmatched. A small number of matches may still be inaccurate, so verify B Corp™ status directly via the [directory](https://bcorporation.eu/find-a-b-corp/) or publicly available B Corp™ [impact data](https://data.world/blab/b-corp-impact-data).
This uses the same general company-matching approach we apply across our third-party data — see [How We Match Third-Party Data](/our-data/third-party-data/matching-third-party-data).
# Companies House
Source: https://docs.thedatacity.com/our-data/third-party-data/companies-house
Find out more about our Companies House data.
Within our platform, basic company data such as; company number, name, address etc. is powered by the [Company Data Product](https://download.companieshouse.gov.uk/en_output.html) provided by [Companies House](https://www.gov.uk/government/organisations/companies-house).
This is our single source of truth and all other data stems from Companies House. It is our most important source of data.
This is then combined with our own proprietary data and a number of third party data sources e.g. [Creditsafe](https://www.creditsafe.com/gb/en.html), [our investment data provider](/our-data/third-party-data/investment-data) etc.
More information on the Company Data Product from Companies house can be found here: [https://resources.companieshouse.gov.uk/infoAndGuide/faq/publicDataProduct.shtml](https://resources.companieshouse.gov.uk/infoAndGuide/faq/publicDataProduct.shtml)
# Creditsafe
Source: https://docs.thedatacity.com/our-data/third-party-data/creditsafe
The Data City has a strong partnership with Creditsafe to source accessible financial information on each company.
[Creditsafe](https://www.creditsafe.com/gb/en.html) is the world’s most used provider of online business credit reports. They have worked hard to change the way business information is used worldwide.
Creditsafe states their data is the largest wholly owned database in the industry, providing accurate and reliable data to over 200,000 subscribers across the globe. They gather data from local, trusted partners and combine it with a scoring algorithm.
You can read more about their data/methodology on [their website](https://www.creditsafe.com/gb/en/more/about/our-data.html).
### Data
Creditsafe provide The Data City with a wide range of data assets which includes the likes of company financials, operating addresses, shareholders, persons with significant control and more.
Creditsafe are particularly valuable in digitising and simplifying manually submitted financial documents to Companies House (likely a PDF). Although the percentage of companies [submitting digital accounts is increasing](https://www.gov.uk/government/statistical-data-sets/companies-house-management-information-april-2023-to-march-2024), in 2024, \~9% of all annual accounts are still not submitted in a digitally readable format. These are often the largest companies (which have the greatest impact on analysis).
We have combined our data with Creditsafe's to build a more complete view of each sector. [Our RTICs](/our-data/rtics/what-are-rtics) provide a comprehensive and accurate picture of the economy, going beyond the limitations of traditional systems such as [SIC codes](/our-data/rtics/what-is-the-difference-between-rtics-sic-codes).
RTICs combined with digitised Companies House information is a rich dataset which can be used for economic analysis.
**Our data**: Keen to find out more about our data? Make sure you check out our [full list of data guides and knowledge base articles](https://help.thedatacity.com/knowledge/our-data).
### Matching
Our data stems from [Companies House](/our-data/third-party-data/companies-house). Creditsafe process documentation submitted to Companies House thus no matching is required between Creditsafe's and The Data City's data. We use Creditsafe for their digitisation of assets which power many data points on the platform and our downloads. This has allowed us to develop a paid, yet fully reproducible and downloadable economic analysis platform.
We incorporate financial data from Creditsafe for further analysis. Our unique platform stands out by offering company-level data access, setting it apart from the IDBR and ONS datasets, which are not publicly accessible.
### Timeseries
Financial data spans the past six years.
### Yearly data
A given company files their annual financial accounts with a “made-up-to date”. We receive this data from Creditsafe.
We take the day of the year it was filed and represent it as a number between 1 and 365. We check whether this number falls in the first half of the year or the second half.
If the “made up to date” is in the first half of the year, the code assigns the previous year to their financials.
If the “made up to date” is in the second half of the year, the code assigns the current year to their financials.
We do this to answer this question: “for which year, does financial performance largely cover this one or last?”
In our platform, we apply this method to simplify calculations within analyse based on years. We also do it to simplify multiple filings for a given company within a single year.
### Why we believe Creditsafe is the best at what they do
Although Companies House provides publicly available financial filings for all UK companies, many organisations still rely on credit reference agencies because of the way they handle and enhance that data. We believe Creditsafe is the best at this.
Creditsafe takes the raw underlying Companies House information and combines it with additional sources to create a more accurate and timely view of a company’s financial position. They bring in data from CCJs, insolvency notices, and other sources, and then cross-check and validate it to correct inconsistencies or errors in the official filings. Most of this can be automated, but there is still some human involvement required. This enrichment process gives users access to cleaner, more comprehensive data.
While Companies House provides separate filings year by year, credit agencies structure that information into consistent historical datasets, allowing for better trend analysis and comparison across companies and sectors. They also standardise formats across different types of accounts, enabling meaningful comparison between businesses that file full, abridged, or micro accounts. This is particularly important.
Based on the [Companies House management information](https://www.gov.uk/government/statistical-data-sets/companies-house-management-information-april-2024-to-march-2025): in 2024-25 **92.07%** of companies filed their accounts digitally. There is a lack of standardisation in field names in these digital accounts. Creditsafe handles this expertly. Of the remaining \~8% that do not file their accounts digitally, the significant majority are large companies which make up most of the economic outlook of any given sector.
In essence, Companies House is the original source of corporate filings, but credit reference agencies (Creditsafe) transform that raw information into a more complete and reliable dataset rather than relying solely on what has been formally submitted to the registry.
# Innovate UK
Source: https://docs.thedatacity.com/our-data/third-party-data/innovate-uk-data
We append successful Innovate UK grant data to our company data.
Innovate UK funding brings together a series of grants that support innovation and research, and sit within [UK Research and Innovation (UKRI)](https://www.ukri.org/)'s umbrella. They publish [open data](https://www.ukri.org/publications/innovate-uk-funded-projects-since-2004/) on successful grants from 2004.
We can easily append this information to our company data for organisations with a Company Registration Number. This data is then available at the company level on the platform and in our downloads. You can also see top-level insights for a group of companies on Analyse and Compare.
Some considerations about the data:
* This data includes the [Industrial Strategy Challenge Funds](https://committees.parliament.uk/work/1006/the-industrial-strategy-challenge-fund/) grants.
* More than one organisation can participate in the same funded project. In these cases, the data shows the total funding given to each organisation.
* Some of the organisations included in Innovate UK data are not a company. These cases will not appear in our database as our database stems from [Companies House](/our-data/third-party-data/companies-house).
### Why is this data relevant?
Knowing which companies have received funding from Innovate UK is a good indication of their R\&D capabilities.
This can help platform users to find companies that actively work on research and innovation. We do not use InnovateUK data in our innovation indicator.
You can filter companies on the platform considering if they received Innovate UK funding, and how much.
Our financial filter available on all the platform's functions (EXPLORE, ANALYSE, COMPARE and the list-building engine) has a search box that makes this possible.
This data can also support further research to understand funding priorities across time and collaboration networks. You can see an example of this type of research in this [blog](https://thedatacity.com/blog/collaboration-and-communities-fuel-quantums-funding/).
### Data Dictionary
UKRI are working on metadata for the exact data source we use.
In the meantime, you may find their other data dictionary useful. You can find that [here](https://gtr.ukri.org/resources/GtR-User-Guide.docx).
# Investment & Funding Data
Source: https://docs.thedatacity.com/our-data/third-party-data/investment-data
The Data City matches UK companies on Companies House to third-party investment data and presents investment rounds, IPOs and acquisitions for each company.
The Data City tracks investment activity for UK companies — investment rounds, IPOs and acquisitions — sourced from a specialist third-party data provider (currently Specter) and matched to Companies House.
**Provider change**: Before July 2026, this investment data was provided by Dealroom. We've since moved to a new provider. Two things to know:
* We've matched the methodology as closely as possible, but there may be small differences between the two datasets — for example in exact round amounts, valuations, or the timing of when a round appears — particularly when comparing historic data across the switchover.
* The change also brings more ways to access the investment data, including through [our API](/using-industry-engine/features/api).
We have combined this data with our own to build a more complete view of each sector. Our [RTICs](/our-data/rtics/what-are-rtics) provide a comprehensive and accurate picture of the economy, going beyond the limitations of traditional systems such as SIC codes. RTICs combined with investor data creates a rich dataset which can be used for economic analysis.
We've developed a paid, yet fully reproducible and downloadable economic analysis platform which stems from Companies House.
We incorporate investment data for further analysis. Our unique platform stands out by offering company-level data access, setting it apart from the IDBR and ONS datasets, which are not publicly accessible.
**Our data**: Keen to find out more about our data? Make sure you check out our [full list of data guides and knowledge base articles](https://help.thedatacity.com/knowledge/our-data).
### Method for matching
Specter's data does not include Companies House numbers, so The Data City matches each company to Companies House using our own matching. Matching runs daily, so new companies and changed records are re-matched as they arrive. Acquisition events are matched to the specific registered entity involved in the deal, while investment rounds and IPOs roll up to the group parent company.
See [How We Match Third-Party Data](/our-data/third-party-data/matching-third-party-data) for more on how this process works generally.
### Data
We extract specific fields from our investment data which are most useful for analysis.
These fields are based on per-round investment information — one row per investment round, IPO, or acquisition. The data is updated daily and we take this on a timeseries basis.
This information is available within the growth tab on the per company page view.
### Temporal coverage
Investment data is a rolling, continuously updated dataset rather than a fixed window. Coverage is comprehensive from roughly the mid-2010s onward and is densest across the last decade, with round activity peaking around 2020–2023. There is a thinner tail of earlier deals going back to the 2000s, but pre-2010 coverage is best treated as indicative rather than complete.
Keep three things in mind when reading the figures:
* **Recency lag:** the most recent weeks and months undercount, because funding rounds take time to be announced and then ingested. Recent counts rise as later refreshes land, so don't read the latest period as final.
* **Historical thinning:** pre-2010 activity is sparse and shouldn't be presented as exhaustive coverage.
* **Quote an "as of" date:** because the dataset is live, any figure should be quoted against the refresh it came from rather than as a fixed number.
### Investment round types
Each investment round carries one of the following types. "Counts towards total funding" shows whether rounds of that type are included in a company's total funding figure.
| Round type | Counts towards total funding |
| --------------------- | ---------------------------- |
| Angel | Yes |
| Convertible Note | Yes |
| Corporate Round | Yes |
| Equity Crowdfunding | Yes |
| Grant | No |
| Pre Seed | Yes |
| Private Equity | Yes |
| Seed | Yes |
| Series A – Series J | Yes |
| Series Unknown | Yes |
| Debt Financing | No |
| Initial Coin Offering | No |
| Non Equity Assistance | No |
| Post IPO Debt | No |
| Post IPO Equity | No |
| Post IPO Secondary | No |
| Product Crowdfunding | No |
| Secondary Market | No |
| Undisclosed | No |
Grants appear as rounds but do not count towards total funding. Grant funding on the platform is carried by our dedicated [Innovate UK](/our-data/third-party-data/innovate-uk-data) and [360Giving](/our-data/third-party-data/360giving-data) datasets, so counting it here as well would count it twice. Grants funded solely by Innovate UK or UK Research and Innovation are left out of the investment data entirely for the same reason.
Total funding is filterable within the filters bar, and the latest investment round is a company's most recent round.
Where the provider identifies who led a round, we also carry the lead investor name(s) alongside the full investor list. Around 6 in 10 rounds with investors have a named lead, and a round can have several co-leads.
Beyond investment rounds, we also track exit and M\&A events as their own rows. These never count towards total funding. Acquisitions, mergers and buyouts appear from both sides of each deal: once for the company that was bought (its exit) and once for the company that made the acquisition, marked with "(MADE)".
| Event | Shown as |
| ----------- | ---------------------------------- |
| IPO | IPO |
| Acquisition | ACQUISITION and ACQUISITION (MADE) |
| Merger | MERGER and MERGER (MADE) |
| Buyout | BUYOUT and BUYOUT (MADE) |
Each company also carries three flags derived from these events: whether it has made an acquisition, has been acquired, or has listed via an IPO. These power the corresponding filters on the platform.
**What you can get for M\&A and IPO events.** On the platform these events are represented by the three flags above, not as dated rows. The per-round detail on a company's growth tab covers funding rounds only, and the bulk download and Insights datasets exclude M\&A and IPO rows too.
The detail does exist in our investment data: every acquisition, merger and buyout event carries the deal date and the counterparty's name, with a deal price on roughly one in five. It is simply not surfaced on the platform today. So if someone asks for dated M\&A history, treat it as a data request rather than something to self-serve from the platform.
On analyse and explore you have the ability to filter for companies that have raised certain rounds of investment funding.
**Investment data**: For more detailed information about our investment data, make sure you [view our full data glossary and dictionary](/data-dictionary).
### Total funding metrics
The provider supplies two company-level totals:
* **Total funding raised** — pre-exit venture investment. This excludes grants, debt, acquisitions, and post-IPO equity, and is the total the platform reports and filters on.
* **Total capital raised** — lifetime external capital, i.e. the same set of rounds plus equity raised on the public market after an IPO. This exists in the source data but is not currently shown on the platform.
These totals are calculated across the provider's whole tracking period, not for one particular year.
### Currency
Funding amounts are converted to GBP at the time of the round.
### Unannounced investment
Some investment rounds are never publicly announced at the time they happen.
Specter has identified coverage of unannounced investment rounds as a strategic product priority and is actively investing in data acquisition and tooling improvements to expand portfolio coverage and capture previously undisclosed financings. Enhancements in this area are currently on the roadmap, with further updates expected from late Q3 through Q4.
# Lightcast
Source: https://docs.thedatacity.com/our-data/third-party-data/lightcast-data
The Data City matches UK companies on Companies House to Lightcast companies and presents jobs and skills data for each company.
[Lightcast](https://lightcast.io/about/company) is a global pioneer in the collection and big-data analysis of information on the labor market.
They provide the world’s most detailed information about occupations, skills in demand, and career pathways.
You can read more about their data/methodology on [their website](https://lightcast.io/about/data).
We have combined our data with Lightcast's to build a more complete view of each sector. [Our RTICs](/our-data/rtics/what-are-rtics) provide a comprehensive and accurate picture of the economy, going beyond the limitations of traditional systems such as [SIC codes](/our-data/rtics/what-is-the-difference-between-rtics-sic-codes). RTICs combined with job and skills information is a rich dataset which can be used for economic analysis. This data is not available to all customers. It is available as an add-on.
We've developed a paid, yet fully reproducible and downloadable economic analysis platform which stems from Companies House.
We incorporate jobs and skills data from Lightcast for further analysis. Our unique platform stands out by offering company-level data access, setting it apart from the IDBR and ONS datasets, which are not publicly accessible.
**Our data**: Keen to find out more about our data? Make sure you check out our [full list of data guides and knowledge base articles](https://help.thedatacity.com/knowledge/our-data).
### Method for matching
This uses the same general company-matching approach we apply across our third-party data — see [How We Match Third-Party Data](/our-data/third-party-data/matching-third-party-data).
We take all UK companies we have matched to a website on Companies House and all of Lightcast's companies then we create a shortlist of Lightcast possible matches to TDC companies using a combination of different fields.
We score the shortlist based on similarity metrics and add some sensible restraints. The highest scored matched is assigned the match subject to the match score exceeding a threshold.
Our match rate is currently around 90% accurate.
Due to issues with companies registering multiple entities of the same company, we may accidentally match one Lightcast company to [many Companies House companies](/our-data/using-our-data/using-our-data-safely#lightcast-double-counting).
### Data
We extract specific datasets from [Lightcast's API](https://docs.lightcast.dev/datasets) which are most useful for analysis.
We consider the most useful datasets to be: Jobs by Standard Occupational Classification ([SOC4](https://www.ons.gov.uk/methodology/classificationsandstandards/standardoccupationalclassificationsoc)), Jobs by Lightcast Occupation Taxonomy ([LOT](https://lightcast.io/resources/blog/new-occupation-taxonomy)), Specialised Skills, Common Skills and Certification Skills. The data is updated monthly and we take this data on a timeseries basis.
Here's how it looks on the platform:
Lightcast timeseries data spans over a historical 5-year period.
# How We Match Third-Party Data
Source: https://docs.thedatacity.com/our-data/third-party-data/matching-third-party-data
How The Data City links external datasets, like investment, jobs, and grants data, to the right company on Companies House.
Companies House is our single source of truth for UK company data. Every other dataset we bring in, from [investment data](/our-data/third-party-data/investment-data) to [Lightcast jobs data](/our-data/third-party-data/lightcast-data) to [B Corp certifications](/our-data/third-party-data/b-corps), is written and stored by someone else, in their own format, without a Companies House number attached. Before we can show that data on a company's page, we first have to work out which Companies House company each record actually belongs to.
We call this process **matching**.
## Why matching is hard
Companies don't always describe themselves the same way twice. A single business might appear as "Acme Ltd", "Acme Limited", or "Acme (UK)" across different sources, with different addresses, or under a trading name rather than its registered name. Some sources don't include a company number at all.
So instead of a simple lookup, we compare each incoming record against Companies House using several pieces of evidence together, things like company name, website domain, registered address, and social media handles, and work out the most likely match.
## How we do it
We use a well-established statistical matching technique (built on an open-source tool called [Splink](https://moj-analytical-services.github.io/splink/)) to compare records at scale. In short:
1. **Narrow the field.** For each incoming record, we first shortlist a small set of plausible Companies House candidates, rather than comparing it against all 5+ million companies.
2. **Score the evidence.** We compare each candidate against the incoming record across multiple signals (name, website, address, and so on) and combine them into an overall confidence score.
3. **Pick the best match.** If a candidate's score clears our accuracy threshold, we accept it as a match. If nothing scores highly enough, the record is left unmatched rather than guessed at.
4. **Resolve corporate groups.** Many companies belong to a wider corporate group. Where that matters, for example deciding which entity an investment round or acquisition should be attributed to, we apply additional rules to roll matches up (or keep them distinct) in a way that reflects how the business actually operates, rather than just its group structure on paper.
Each data source has its own quirks, so the exact combination of signals we use varies a little from one to the next. You can find source-specific notes on the relevant [third-party data pages](/our-data/third-party-data/companies-house).
## Keeping matches fresh
Matching isn't a one-off exercise. As new companies are incorporated, existing companies change their details, and new records arrive from our data providers, we re-run matching regularly to keep everything up to date. For some sources, this happens daily.
## What this means for accuracy
No automated matching process is perfect. We aim for a high degree of accuracy, and we'd rather leave a record unmatched than force a low-confidence guess. That said, a small number of matches can still be inaccurate, for instance where a company has registered multiple similar entities, or where the source data itself is incomplete.
# What is Trade Data?
Source: https://docs.thedatacity.com/our-data/third-party-data/trade-data
Where does the trade data come from? Use data on overseas income, or data on the goods traded, to better understand companies' connections to global markets.
There are two ways that we can understand companies' international trade. This is based on two data sources that we have for company exports.
1. Companies House. Companies can choose to declare their **overseas income** - this can be either through exports, or through foreign operations. This details the value of overseas income, but not which goods *or services* have been exported.
2. HMRC. Companies that import and export **have to declare** the goods traded **and when they were traded**. This includes information on imports. **However, detail on the value of goods traded is not provided**. Does not provide information on services that may have been traded.
Companies House
You can filter companies by the amount of overseas income.
HMRC data
We extract data from [UK Trade Info](https://www.uktradeinfo.com/find-uk-traders/) and you can interact with this data in two main ways:
**Trade data on company pages**
**Trade data filters**
Connecting our data to UK Trade Info's requires matching their data to a Companies House number.
# University spinouts
Source: https://docs.thedatacity.com/our-data/third-party-data/university-spinouts
How The Data City identifies university spinout companies using the HESA spin-out register.
Industry Engine identifies university spinout companies using the [HESA spin-out register](https://www.hesa.ac.uk/data-and-analysis/business-community/spin-out-register).
A spinout is a company created to commercialise research or technology developed at a university or research institution. Identifying these companies can help you analyse how research moves into the economy and compare spinout activity across sectors and places.
## About the source
HESA publishes the register as open data under the [Creative Commons Attribution 4.0 licence](https://creativecommons.org/licenses/by/4.0/). As the licence requires, we credit the Higher Education Statistics Agency (HESA) as the source.
The source includes:
* the spinout's name and, for UK companies, its Companies House registration number;
* the associated higher education provider, its UK Provider Reference Number, and the provider's region and country;
* incorporation and foundation years;
* company status and website;
* whether several higher education providers were involved;
* whether the company is a social enterprise; and
* broad subject areas: Medicine, Health and Life Sciences; Physical Sciences, Engineering and Mathematics; Social Sciences; and Arts and Humanities.
## How we process the register
We download HESA's published register each month.
We format Companies House numbers as eight characters, convert yes or no fields to boolean values, standardise month values and clean website URLs.
We validate the data against an expected schema. We keep one record for each company and higher education provider relationship, so a spinout associated with several providers can retain each relationship.
We use the register to populate the `Spinout` field on company records in Industry Engine.
## Find spinouts in Industry Engine
Use the **Spinouts** option in the **Growth** filters to focus your results on companies identified in the register. See [using filters](/using-industry-engine/features/using-filters#growth) for more information.
The company-level `Spinout` field is also available in company data. You can find its definition in the [data dictionary](/data-dictionary/classified-company#spinout).
The HESA register is the source of our company-level spinout status. The deprecated `Spinout` field attached to investment rounds is always `0` and should not be used to identify spinout companies.
## Limitations
* This is a source-based classification, not a prediction. Coverage depends on the records HESA publishes.
* Industry Engine is built around Companies House records. A register entry that cannot be linked to a company record may not appear on the platform.
* A company can be associated with more than one higher education provider. The company-level field shows whether it is a spinout, not the number of provider relationships.
* We process the source monthly. Changes to the register appear after the next pipeline and platform update.
# Company Status
Source: https://docs.thedatacity.com/our-data/using-our-data/company-births-and-deaths
Our active companies are those with a Companies House Company status of 'Active' and whose latest accounts filings are not for a dormant company.
Since adding company births and deaths, we need to be clear as to what we're presenting with our active companies. Our data is updated monthly via [Companies House](/our-data/third-party-data/companies-house). Each company has a 'Status' field.
### Active companies
Our active companies are those where their Company status is **only** the word 'Active'. We are an [active company](https://find-and-update.company-information.service.gov.uk/company/10958787) on Companies House:
### Active but in an unusual state
On Companies House, there are a range of other statuses. For example: "Active - Proposal To Strike Off", "Liquidation", "In Administration" etc. We consider these to be "Active but in an unusual state". If a given company has a Company status of "Active" but their latest account filings are accounts for a dormant company, they are considered to be "Active but in an unusual state". Here's an example:\\
### Struck off (Dissolved)
Companies which are struck off from the register are considered non-active (dead). On Companies House, these are known as dissolved companies. This is an [example](https://find-and-update.company-information.service.gov.uk/company/09091602) of a struck off (dissolved) company and here's how it looks on Companies House:\\
### UI Implementation
In Explore, you can filter using any or all of the three options mentioned above. The default filter is set to 'Active' companies only. Analyse will **only** analyse active companies.
### Background
Our implementation is different from the ONS' definition in their [business demography analysis](https://www.ons.gov.uk/businessindustryandtrade/business/activitysizeandlocation). Their definition is as follows:
> "Being active means that the business had either turnover or employment at any time during the reference year."
We talk more about this in [one of our blogs](https://thedatacity.com/blog/website-matches/).
We currently track companies back to 2015.
### Company Deaths in RTICs
At the moment we don't track company deaths in RTICs and the cumulative active companies by year visualisation reflects the total number of companies that have existed in this sector at *any* point.
Company deaths do occur within RTICs, and we are working on a way of better capturing these, using historical classifications of RTICs.
### Filtering
The "Active but in an unusual state" filter includes companies filing dormant accounts. Analyse will **only** analyse active companies.
The "Active but in an unusual state" filter includes these dormant companies to ensure our active company filter aligns more closely with national statistical definitions and provides a more accurate picture of the UK’s business landscape. By removing companies that file dormant accounts from our “Active” company filter and categorising them in our “Active but in an unusual state”, we exclude entities that are technically registered but not actively trading, making our data more representative of real economic activity.
We believe this approach brings us in line with methodologies used by organisations such as the Office for National Statistics (ONS) and the UK Business Data Survey. As a result, you can trust that when you filter for active companies, you are seeing businesses that are genuinely operating, leading to more precise insights.
# Data Update Frequency
Source: https://docs.thedatacity.com/our-data/using-our-data/data-update-frequency
How often is our data updated and how many years data do we hold?
### Update cycle
We aim to update the data on a monthly basis, in line with the monthly release cycle of [Companies House](/our-data/third-party-data/companies-house).
Once we have the Companies House data we need to combine this with our other data sources, process it and then package it to form the underlying dataset we use for the platform.
Therefore it is usually towards the middle, or end of the month, before the data is reflected in our platform.
**Example**: If Companies House data is available on 3rd of the month, it will be present in the platform after the 15th.
### Daily incremental updates
In addition to the monthly cycle, we run daily incremental updates for reported company-website matches. These updates may include:
* **A new match** — a company is matched to a website for the first time
* **A reassigned match** — a company's match is moved to a different website
* **A removed match** — a previously matched website is no longer associated with a company
These changes have cascading effects across the rest of the data, including operating addresses, RTIC classifications, and other derived fields that depend on web-scraped content.
### Data depth
This will vary from company to company but where possible we will have up to 7 years of financial information.
Other company data such as name, address, and contacts are updated monthly.
**Our data:** You can see a full list of our datapoints and sources [in our Data Dictionary](/data-dictionary).
# How can I access data on what companies do where?
Source: https://docs.thedatacity.com/our-data/using-our-data/how-can-i-access-data-on-what-companies-do-where
Where can I find profiles data? How does the Data City split employees/turnover across locations? How do you estimate the employee counts by occupation?
Companies can have multiple locations and companies are not required to provide information on what they do at each location, or how big these locations are.
The Data City have integrated data from Lightcast to better answer the question of what companies do where. Specifically, we use Lightcast profiles data to understand the occupations and locations of employees for companies. We sometimes refer to this as WhatWhere data.
You'll find this data in analyse and explore.
### Explore
On explore, you'll find WhatWhere data on the locations tab of a company, where we have it.
Not all companies have WhatWhere data, but a significant proportion of large companies do.
The counts listed below are the number of Lightcast profiles in each location and occupation.
### Analyse
On analyse, we've integrated profiles data in many places. Currently, we treat profiles data differently than on explore.
Where explore is the raw count of profiles, on analyse we've combined companies' declared employees with the percentage of profiles in a location or occupation.
This is done as a company's employee count can be different than the number of profiles we have for it, and the declared employee count is our preferred source.
| Page Location | Field | Detail |
| ------------------- | ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Analyse Summary Box | Total UK Employees | We use profiles data to better understand how many employees are based in the UK vs the Rest of the World. |
| Analyse Summary Box | Best Estimate Total GVA | We use profiles data to focus our GVA estimate on UK activity. This is done by using the percentage of UK employees, of total employees. |
| Locations | Employees by geography | For each of the geographies (Local Authority, Strategic Authority, etc.) we use profiles data where available. This allows us to better represent firms' activity across the country. Where we do not have profiles data, we fall back to splitting a company's employees equally across the number of locations it has. |
| Locations | Turnover by geography | We do not apply profiles data to inform turnover by geography. Turnover by geography is the result of splitting a company's total turnover across the number of sites it has. |
| Jobs and Skills | Occupation counts | We use profiles data, to estimate the number of employees by occupation. This draws upon Lightcast's extensive occupational taxonomy. Thousands of individual occupations are aggregated to form the categories presented. |
# How do you identify or analyse different kinds of companies?
Source: https://docs.thedatacity.com/our-data/using-our-data/identifying-different-company-types
How do you find community interest companies/limited by guarantee companies/publicly listed companies, etc.? Our platform has a filter that allows you to filter on company type.
There are many different types of companies. The most familiar of these types are limited companies.
For some questions you may want to analyse other kinds of companies, for example publicly listed companies, or community interest companies.
To analyse a specific kind of company, you should apply a company category filter. The screenshot below shows how to analyse publicly listed companies.
In the sectors filter tab, select category, then select Public Limited Company and click update.
Below is a full list of the company types that we have on our platform. The numbers in bracket refers to the number of companies that exist of that type, as of June 2024.
| | Charitable Incorporated Organisation (35,380) |
| - | -------------------------------------------------------------------------------------------------- |
| | Community Interest Company (32,334) |
| | Converted/Closed (1) |
| | Industrial and Provident Society (160) |
| | Investment Company with Variable Capital (617) |
| | Investment Company with Variable Capital (Securities) (9) |
| | Investment Company with Variable Capital(Umbrella) (69) |
| | Limited Liability Partnership (51,889) |
| | Limited Partnership (58,503) |
| | Old Public Company (16) |
| | Other company type (14,710) |
| | Other Company Type (3) |
| | Overseas Entity (30,740) |
| | PRI/LBG/NSC (Private, Limited by guarantee, no share capital, use of 'Limited' exemption) (38,527) |
| | PRI/LTD BY GUAR/NSC (Private, limited by guarantee, no share capital) (116,245) |
| | PRIV LTD SECT. 30 (Private limited company, section 30 of the Companies Act) (15) |
| | Private Limited Company (5,040,169) |
| | Private Unlimited (95) |
| | Private Unlimited Company (4,234) |
| | Protected Cell Company (5) |
| | Public Limited Company (4,665) |
| | Registered Society (10,563) |
| | Royal Charter Company (898) |
| | Scottish Charitable Incorporated Organisation (6,405) |
| | Scottish Partnership (267) |
| | United Kingdom Economic Interest Grouping (258) |
| | United Kingdom Societas (17) |
# Product roadmap
Source: https://docs.thedatacity.com/our-data/using-our-data/product-roadmap
Interested in seeing what's coming soon?
We are always working on improvements and adding new features to our product. The details of short and long term goals for the platform can be viewed in our [release notes](https://thedatacity.com/blog/data-explorer-release-notes/#product-roadmap).
# Using our data safely
Source: https://docs.thedatacity.com/our-data/using-our-data/using-our-data-safely
There are many things that you should bear in mind when using our data, including some limitations.
### Summary
Some of the issues here relate to incorrect or limited primary data. For example, employee counts may be reported incorrectly from source, or not provided. Where possible The Data City have tried to minimise these issues. In the case of employee counts, we've [created an automated way of identifying outliers](/our-data/proprietary-data/what-are-companies-with-potential-anomalies). Outliers are removed from our data before creating estimates of turnover or employees.
#### Issues in primary data
Issues and errors in primary data are compounded when this data is used to generate secondary metrics. For instance, the range of issues associated with employee data will cause subsequent issues with estimated turnover data and GVA data.
Some of the issues relate to other core data points, such as URL matches, or company locations. These issues and errors are compounded when these data points are used to filter, assign, or generate secondary data points.
For example, mismatched URLs can lead to attribute data being incorrectly linked to the wrong company, including location information and webtext. This can subsequently lead to the company being incorrectly selected in location filters and/or the company being incorrectly classified in RTICs and/or allocated an incorrect Innovation Score.
#### Third party data
We also inherit some other uncertainties and errors from our data providers, e.g. location information provided by CreditSafe is of unknown accuracy, our investment data provider's ability to identify investments has improved over time (introducing uncertainty when taking a timeseries view), and Lightcast data might contain duplicates.
Matching third-party data to our own data based on our core attribute of Companies House ID can also introduce some uncertainty or data duplication, e.g. matching Lightcast IDs to our Companies House ID.
***
### Creditsafe Data
#### Employee data can be incorrect
Some employee data is entered incorrectly at source. For example, GILLARDS FARMS LIMITED (00981261) reports 1,299,450 employees in its 2021 accounts filed with Companies House. This is incorrect and is the same value as it submits as its ‘net worth’.
Other data is misprocessed by CreditSafe. For example, CHOUDHURYS VENTURES LTD (08238569). In 2019, their financial report to Companies House reports 2 employees but CreditSafe have reported 61k.
The Data City have trained a model to identify data like this. Being 98% accurate, the model is extremely accurate for known examples. We remove the bad data before we create estimates of turnover and employees. We do this as sometimes outliers will just affect one year out of multiple filings and removing the company from the analysis entirely may be misleading, where a company is particularly important in a sector.
You can read more about our approach to outliers [here.](/our-data/proprietary-data/what-are-companies-with-potential-anomalies)
#### Employee data is sometimes global
International businesses often report their global headcount as their employee count to Companies House. Whilst the number may be correct, it is important to understand that many, perhaps most of these employees, may be based outside of the UK.
The Data City have done a considerable amount of work to address this issue. The Data City have introduced data that allows us to better understand what companies do where. We've integrated this into our product, including on regional analysis. You can read more about this [here](https://thedatacity.com/blog/what-companies-do-where/).
**Guidance:** Despite adding data to aid with this issue, we recommend manually checking for outliers like this and removing them prior to analysis.
A great way to do this is to view the list in explore and sort by turnover or employees highest to lowest. Larger companies are more likely to have global operations and errors in larger companies will have a larger effect on analysis. Another quick way of sense checking the data is by using the dots on a map visualisation on analyse. The dots on this map are sized by the number of employees that are estimated to be at that location, so companies reporting a global headcount will appear oversized for a particular area.
#### ANALYSE reports employees which exceeds official statistics
The base of our data is Companies House. Official Statistics are created using a separate sampling frame, called the Interdepartmental Business Register (IDBR). Not all companies registered at Companies House are also on the IDBR. Therefore, our employee counts may differ from official statistics slightly.
We are working to address these differences, through group structure, identifying outliers, and improving data on what companies do where.
#### **Some** company operating addresses are incorrect or missing
A UK company is only required to register one address with Companies House and this address does not necessarily have to be an operating address but we assume it to be. We estimate additional operating addresses for companies from their website text and other sources e.g. CreditSafe.
We can never be certain we have identified all operating addresses. We can never be certain the identified operating addresses are correct.
**Guidance:** Be aware of potentially missing/incorrect addresses when analyse the data. Include a caveat in your report to state the [accuracy of our addresses](/our-data/key-data-and-definitions/company-locations). If reported values seem unusual, consider this to be a possible cause.
***
### Lightcast Data
#### Multiple companies, one Lightcast ID
One Lightcast ID can match to multiple companies in our dataset. This is likely to happen within group structures. We remove subsidiaries where we can but there may be cases where double counting still occurs.
You can read more about the implementation of [Lightcast data here](/our-data/third-party-data/lightcast-data).
At this stage, we do not split the job postings across all companies which are allocated to the same Lightcast ID.
**Guidance:** Caveat the data appropriately to take this into account.
#### These are job postings, not positions filled
This is the advertisement of jobs. These jobs are not necessarily filled.
**Guidance:** Be clear in your analysis that these are job postings. You may want to add the caveat that they are not necessarily filled.
#### Job postings are not assigned to specific locations
For any location filter applied on the platform, the job postings returned will be all job postings assigned to the companies with at least one address (either operating or registered) in the filtered geography.
Job postings are not split across locations. This leads to duplicates.
**Guidance:** If comparing geographies, be conscious that job postings will be duplicated (assigned to both geographies) where a company has a location in both geographies.
The Data City does not have job postings by region at company level on the platform.
***
### Investment Data
**Provider change**: Before July 2026, this investment data was provided by Dealroom. We've since moved to a new provider (currently Specter); while we've matched the methodology as closely as possible, there may be small differences between the two datasets, particularly when comparing historic data across the switchover. See our [investment data guide](/our-data/third-party-data/investment-data) for more detail.
#### Timeseries limitations
Our investment data provider's ability to track deals has improved over time, and this has also been true of prior providers we've used.
**Guidance:** Use caution when doing timeseries analysis of investment data before 2015. Providers make efforts to backdate their tracking, however there will be increases in coverage as a provider's tracking has scaled.
Consider outliers carefully. The total investment figure presented includes everything from seed to IPOs. ANALYSE does not provide a breakdown.
#### Outliers
Investment data is susceptible to outliers. Often an RTIC's total investment can be dominated by one or two large companies.
**Guidance:** Errors in larger companies will have a larger effect on analysis. We recommend manually checking for outliers and removing them prior to analysis if required. A great way to do this is to download the data and sort by Total Funding highest to lowest.
#### Venture capital funding may be incorrectly assigned to a company
Our investment data provider matches some of their companies to a Companies House company. The Data City also matches some of the provider's companies to a Companies House company. The overall match is highly accurate (>98%) however sometimes the match is incorrect. Sometimes a match from the provider to a Companies House company does not meet the threshold for an accurate match.
An incorrect and/or a missing match may misrepesent the total invesment into a given sector/location etc.
**Guidance:** Caveat the data appropriately to take this into account.
***
### Proprietary Metrics
#### Company GVA estimate will be wrong if employee numbers are wrong
GVA for a company is estimated using its estimated number of employees, registered SIC codes, and published national statistics on average GVA per employee for those SIC codes. The [known issues with employee data](/our-data/using-our-data/using-our-data-safely#creditsafe-employee-data-incorrect) such as companies reporting foreign employees will cause GVA to be overestimated.
**Guidance:** Problems with GVA over-estimates can be reduced by manually reducing GVA estimates for companies known to have reported a large number of foreign employees in their UK accounts. Since large companies contribute the most to GVA it is best to start checking large companies first by using the Sort by Employees: High to Low function in Explore. Companies that do not have significant operations in the UK should be removed from analysis.
**Company GVA estimates do not measure regional productivity**
Our approach to estimating GVA for a company assumes that all companies in that sector have the same GVA per employee (productivity). In reality, productivity varies by region.
**Guidance**: This is an estimate of GVA primarily for understanding sectors where official statistics don't exist and regional productivity comparisons are not possible at the moment.
#### Company GVA estimates do not estimate split of function within a company
Where a company's estimated employee count is correct we are able to make a good estimate of the company's GVA based on the average GVA per employee in its sectors of operation. Users should be careful however that a company's GVA cannot be fully assigned to any of its RTICs. For example, many companies use artificial intelligence as part of their work but it would probably be wrong to assign a company's full GVA as being produced in the AI sector.
**Guidance:** When writing up GVA analysis phrasing such as "companies using AI contribute £Xbn to the UK economy" would be more correct than saying "AI contributed £Xbn to the UK economy".
#### Incorrect data points based on mismatched URLs
There will be instances of [mismatched URLs in our data](https://help.thedatacity.com/knowledge/url-matching). We have recently improved our match rate to 95%, but mismatches remain.
If a URL mismatch has occurred, several datapoints which are either scraped from the URL directly, or derived from the webtext, including (but not limited to) operating locations, company description, RTICs, and Innovation Score, will be incorrectly assigned to the company with the mismatched URL.
**Guidance:** Caveat the data appropriately. Report URLs that are incorrect/missing. You may also want to exclude/include companies from your analysis which you know are mismatched/missing.
#### Turnover and Employee counts are estimates, and there are known limitations of these estimates
Not all companies report turnover/employee count and/or there is a lag in report date. The Data City [attempt to fill turnover/employee data where missing](https://help.thedatacity.com/knowledge/estimated-turnover-employees).
**Guidance:** For reports you will want to do a QA of the list. Errors in larger companies will have a larger effect on analysis. View the list and sort by turnover and employees high to low to identify outliers and remove where appropriate.
#### There might be no turnover or employee data for a company
The Data City [attempt to fill turnover/employee data where missing](https://help.thedatacity.com/knowledge/estimated-turnover-employees). We cannot provide an estimate without appropriate data to do so. This could lead to underestimating the size of the sector/economy in a location. This is usually only affects small companies so the overall impact is likely to be somewhat negligible.
**Guidance:** In your analysis make sure to refer to these datapoints as ‘estimated turnover’ or ‘estimated employment’. For clarity, you may want to report the number of companies which we do provide estimates for. The largest companies have the highest influence on summary statistics and these less likely to be impacted by missing data.
#### Estimated turnover and employees are evenly split by location
Companies are not required to declare their financial accounts for each operating address that they have. Previously, we did not know how many employees work in each operating address.
Through our work on what companies do where, we now have an ability to estimate how many employees are in a particular location. This is based on profiles data from Lightcast and where we have the data for a company, we use this to understand the size of their operations in a particular location.
Where we do not have Lightcast profiles data, we estimate by equally splitting the number of employees or amount of turnover across the number of operating addresses.
Lightcast profiles data does not inform us of the relative turnover contribution of a company's sites. In this case, we split turnover equally across the number of locations.
**Guidance:** Caveat the data appropriately. If reported values seem unusual, consider this to be a possible cause.
#### Assumed split of employees across company locations can lead to potentially incorrect local growth.
Employees incorrectly linked to either an address, or a specific activity in a region, will be included in the calculation for the growth of that sector in the region. The Data City works hard to [source location data](/our-data/key-data-and-definitions/company-locations) for each company and has to make subjective decisions as to whether an address should be included as a company location.
Example 1: Should all of BP's petrol stations be included in the location data or just the office locations? We know revenue is generated at each petrol station but we may choose to attribute this revenue to only the offices.
**This varies across companies and this may have an impact on summary statistics.**
Example 2: If you were choosing to analyse AdTech, you would likely include Amazon in your list. Amazon has warehouse locations across the UK. We track these warehouse locations but we do not know which are offices and which are warehouses without further research. We would therefore assign employees working in AdTech to warehouses falsely.
**Guidance:** Be mindful of this issue and explore the location data of the largest companies in your list. Make a sensible decision on how you might overcome assigning employees to the correct locations.
#### Location Quotients could be distorted
We generate location quotients based on all addresses [associated with companies](/our-data/key-data-and-definitions/company-locations) (both registered and operating addresses).
This could introduce known location issues, such as including locations of properties used as default registered addresses, e.g. accountants, where activity relating to the company might not actually take place.
Location quotients are also subject to the registered office effect. London is the registered office location for many multinational companies (and perhaps head office). This can create distortions. For example, when using location quotients, some of the top specialisations for London are tobacco or coal mining. This is due to the registered office of tobacco and mining firms in London.
Therefore there are several reasons why the resulting location quotient value may misrepresent the location.
**Guidance:** Take care when applying location filters and interpreting the location quotients. Be aware of these caveats.
#### Company size definitions should be well understood
The company size definitions are based on the Companies Act 2006 definition. Other definitions are available. You can read about the [implementation of this definition here](/our-data/key-data-and-definitions/company-sizes).
A misunderstanding of these definitions could lead to incorrect conclusions made from the data.
**Guidance:** Clearly state the definitions used in your analysis.
Alternatively, you can generate custom definitions yourself using the filter options.
#### A company might be assigned an incorrect company size
This would arise due to incorrect turnover or employment data, which could occur as a result of [known issues with employee and turnover data](/our-data/using-our-data/using-our-data-safely#creditsafe-employee-data-incorrect).
This could be a declaration error, e.g. an error in a company’s financial accounts. Alternatively, it could be an error in estimated values for a company’s turnover or employment.
**Guidance:** Company sizes are based on The Data City’s estimated values of turnover and employment. The Data City have robustly estimated turnover and employment, but a company can be miscategorised if errors do occur. Users should be aware of this.
#### RTIC time series analysis
The RTIC analysis within ANALYSE provides insight into the historical financial performance of the companies which are **currently** within the RTIC.
It is important to note that this is not a time-series analysis of the industry as a whole, but specifically of these companies. Currently, our approach does not include historical classifications or track the inception and closure of companies (births and deaths).
**Guidance:** Caveat the data appropriately to take this into account. Be mindful that this is the case when interpreting the results.
#### The innovation indicator should be properly understood and used appropriately
Our innovation indicator uses a proprietary machine learning model to estimate how innovative every company in our database is based on their website text, where available. R\&D intensity is used as a proxy for innovation.
You can read more about our [innovation indicator here](/our-data/proprietary-data/innovation-score).
Misunderstanding or misinterpreting our innovation indicator can lead to misleading results.
**Guidance:** We urge users to exercise caution whilst using our innovation indicator, viewing it as a valuable tool when analysing lists, rather than a definitive measure.
It should also not be used to build company lists, instead, it is more appropriate to use it as a filter in analysis.
#### We are not able to identify founders for all companies
We are able to [identify founders](/our-data/proprietary-data/gender-data) for approximately 80% of companies.
It is worth noting that our ability to identify founders is very dependent on the age of a company.
The Data City tracks the active companies on Companies House and these tend to be younger companies.
Misrepresenting founder information can lead to error in your analysis.
**Guidance:** Be careful when comparing companies with founder information against the baseline.
For example, comparing women founded businesses to the UK business base will be tricky.
Since we are not able to identify founders for older companies, women founded businesses will be younger than the business base, so a national comparison might not be suitable.
Instead, we advise comparing women founded businesses to non-women founded businesses.
# What are subsidiary companies and how can they be removed?
Source: https://docs.thedatacity.com/our-data/using-our-data/what-are-and-how-to-remove-subsidiary-companies
Removing subsidiaries can avoid double counting of employees and make your lists easier to view.
We remove subsidiaries from analyse and explore by default. You can choose to also look at all companies using the companies filter.
Group structure is an area that we’ve been working on for a long time and with this release, we have included a feature that will simplify any lists you have created and improve your analysis.
To show how this will be useful, let’s consider Specsavers. Specsavers has over 1000 companies in its group structure. The largest of these (by employee count) is shown below.
The subsidiary companies of Specsavers often respond to a store's location - many stores are incorporated as their own companies. Many of the subsidiaries of Specsavers Optical Superstores Limited have employees attributed to them. Importantly, these employees are also counted in the ultimate parent company, **so there is the potential to double count employees**.
### Why remove subsidiaries?
A typical question that someone might answer in The Data City platform would be ‘how many people work in a given sector?’.
It’s normal to take the list of companies and look at the number of employees associated with these companies. It annoys us a little to admit it is more complex than that. At the Data City we aim to simplify complexity wherever we can.
One of the challenges that we’ve had for a long time is that companies can spread their operations over multiple corporate entities. This introduces the potential for double counting of financial variables, like employees or turnover.
#### How subsidiary accounts can include double counting
For the sake of simplicity, let’s answer the question of ‘how many people does a company employ?’. Let’s use BP.
Is BP’s total employment the sum of its employees across all its operating companies? In short, it’s not.
BP employed 79,400 people in 2023 according to the financial statements of the parent company. This is what is known as a consolidated statement. [\[1\]](#_ftn1)
BP has 194 UK-based corporate entities. If we sum across these entities, there are 87,600 employees. The extent of double counting by including subsidiaries will vary, depending on the number of entities, and whether they file subsidiary accounts. In this case, the double counting is around 10%.
BP, as a large multinational, is a good example of the complexity of group structure. In total, there are 6 levels of corporate entities underneath the parent company. For example, the parent company has 23 children organisations. The 23 children organisations have 54 children of their own.
As mentioned before, the parent company of BP filed a consolidated statement. A consolidated statement means that BP have gone through the work of removing double counting from across its group.
“Where a parent company prepares Companies Act group accounts, all the subsidiary undertakings of the company must be included in the consolidation, subject to \[the following] exceptions.”[\[2\]](#_ftn2)
#### Matched Websites vs Group structure
The purpose of this example is to provide a bit more information on what we mean by group structure. Specifically, this example will explain the difference between matched websites and group structure, as well as explaining the source of group structure data.
A core part of our business is matching companies to websites. Matching to websites and scraping those sites allows us to classify companies into what they do.
For some companies, we’ll match a large number of entities to one domain. Peel Group is a good example. There are 203 companies matched to the peel.co.uk domain.
Does that mean these 203 companies belong to the same group? Unfortunately, no. There can be multiple groups within one domain. In this case there are 9 groups assigned to the peel.co.uk domain.
**The implication of this is that we won’t be able to ‘deduplicate’ down to one company per domain**. There are further reasons why we won’t always be able to consolidate a list of grouped companies down to one company. Our priority is to simplify your analysis without introducing errors. We will, however, be able to consolidate a large number of companies.
If matched websites and group structure are separate, what is the source of the group structure? Our group structure data comes from our data supplier CreditSafe.
CreditSafe parse financial statements to gather group structure data. Below is an example of where you can find the raw group structure data. For example, Peel Media Management Limited although it is a group in itself, has an ultimate parent called Tokenhouse Limited. [\[3\]](#_ftn3)
### FAQ
1. Why do you require a list to contain the parent AND the subsidiary to consolidate?
Let’s imagine that you are analysing a list of personal finance providers in the UK. You have built a machine learning list using known examples.
When you review the machine learning list, you see that Tesco Bank has accurately been included in your list. Tesco Bank is a subsidiary of Tesco PLC. [Tesco Bank](https://products.thedatacity.com/v2/companypage/?company_id=SC173199) reported 3,593 employees in 2023. [Tesco PLC](https://products.thedatacity.com/v2/companypage/?company_id=00445790) reported 225,659 in 2023.
Tesco PLC has not been included in your personal finance list. They have a different domain (tesco.com vs tescobank.com). We should not consolidate Tesco Bank into Tesco PLC in this case because:
1. Tesco PLC was not identified using your filters
2. It would be inaccurate to include all of Tesco PLC’s employees.
So, we will only consolidate companies from your list, and we will not introduce companies from outside of your list.
**2) I’ve used the Remove Subsidiary Companies Filter and there are still duplicate companies**
Group structure is complicated. Companies under a certain size do not have the same reporting requirements as large companies. Small groups are not required to file consolidated accounts.
A small group is one with fewer than 50 employees, less than £10.2m in net turnover or a net balance sheet less than £5.1m.[\[4\]](#_ftn4)
[\[1\]](#_ftnref1) BP Financial Statements 2023. Page 245. Available at: [https://www.bp.com/content/dam/bp/business-sites/en/global/corporate/pdfs/investors/bp-annual-report-and-form-20f-financial-statements-2023.pdf](https://www.bp.com/content/dam/bp/business-sites/en/global/corporate/pdfs/investors/bp-annual-report-and-form-20f-financial-statements-2023.pdf)
[\[2\]](#_ftnref2) Companies Act 2006. Section 405.
[\[3\]](#_ftnref3) Group of Companies Accounts Made up to March 2023. Page 26. [https://find-and-update.company-information.service.gov.uk/company/07861087/filing-history](https://find-and-update.company-information.service.gov.uk/company/07861087/filing-history)
[\[4\]](#_ftnref4) Companies Act 2006. Section 382.
# The Data City Tender Materials
Source: https://docs.thedatacity.com/tender/the-data-city-tender-materials
This is the basic information to complete tenders using The Data City data and platform
### What is The Data City?
The Data City is a UK-based Global data provider and SaaS platform focusing on emerging sectors, high-growth companies, innovation clusters and investment.
We work with many of the UK's leading government departments, local authorities, investment firms and policymakers, to provide real-time insights that help drive better investment decisions.
The Data City was founded in 2016 and incorporated as a business in 2017. It was founded to do accomplish two main objectives:
1. Fix the fundamental limitations of SIC codes [limitations of SIC codes](/our-data/rtics/what-is-the-difference-between-rtics-sic-codes) using open data, the web and data science.
2. Build data infrastructure to allow measurement and analysis of the economy at individual company level in real-time.
As a result, The Data City has created the best available Data Assets and developed the best AI-powered Data as a Service product. The Data City's platform, the Industry Engine, comprises all the data sourced by The Data City, innovative solutions like RTICs, The Industry (Search) Engine supported, or machine learning list building.
### The Data City's data
The Data City holds company-level data for all companies classed as active in Companies House. The Data City collates company data from different providers and develops proprietary metrics that describe the performance of individual companies. Below there is a summary of the main data sources used by The Data City and the proprietary metrics it develops. You can find more information about these (including limitations and advice on how to use the data safely) in [this article](/our-data/using-our-data/using-our-data-safely):
* CreditSafe: we source observed data from company accounts from CreditSafe. This includes time series data for a large range of variables, like turnover, number of employees or net worth.
* Investment data: The Data City developed an integration with third-party investment data that allows users to see how much private funding a company has raised along funding rounds
* Lightcast: The Data City developed an integration with Lightcas that allows users to understand job postings for a group of companies. This is an add-on to the regular license fee and is not directly available to all platform users.
* Proprietary metrics: we develop extra metrics for each company. We advise reading the articles linked to each metric to understand whether these proprietary metrics can be used safely in the context of your project/research:
* [Turnover, employee and company growth estimates](/our-data/key-data-and-definitions/estimated-turnover-employees-growth)
* [Location quotients](/our-data/key-data-and-definitions/location-quotients)
* [Innovation Score](/our-data/proprietary-data/innovation-score)
* [Founders Gender data](/our-data/proprietary-data/gender-data)
* [GVA Estimates](/our-data/key-data-and-definitions/gva-data)
* [Greenhouse Gas Emissions Estimates](/our-data/faqs/how-do-you-get-green-house-emissions-data)
The Data City also provides access to its proprietary data product, Real-Time Industrial Classifications or RTICs. These are built-in sector datasets that group companies working in emergent economy sectors.
### Our Platform and Technology
The Data City has developed a platform (the Industry Engine) powered by machine learning technology that allows users to browse and query available company data, and build machine learning lists of companies from scratch. These functionalities are divided into different sections of the platform:
* [EXPLORE](/using-industry-engine/tools/explore): this function of the platform allows users to query all company data available. It includes existing RTICs, SICs and other key filters like [keyword search](/using-industry-engine/features/using-keyword-filters) (instant keyword search for companies' website text), geographies, financials and growth metrics.
* [ANALYSE](/using-industry-engine/tools/using-analyse): this in-built analytical engine makes it possible to extract a wide range of summary statistics, proprietary metrics and insights for a group of companies. It includes the same filters mentioned above.
* [COMPARE](/using-industry-engine/tools/compare): this in-built analytical engine allows users to compare side-to-side similar metrics to those produced by the ANALYSE function for two groups of companies.
* [MY LISTS](/using-industry-engine/tools/my-lists): this function allows users to create their own machine-learning lists or sector classifications.
### Real-Time Industrial Classifications (RTICs)
Real-Time Industrial Classifications (RTICs) are The Data City's proprietary form of industrial classification, based upon web-scraped data and built using machine learning.
RTICs are a modern classification approach, unlike the traditional SIC system that relies on predetermined and static categories. They are based on how companies describe themselves on their websites.
The output is a dataset that gathers companies working in the same field. In this dataset, you will find all the data available for the companies in an RTIC. The data can be explored on the platform or downloaded.
**How are RTIC different from SIC codes?** Find out more about the differences between the two classification systems [here](/our-data/rtics/what-is-the-difference-between-rtics-sic-codes)
#### Why can you trust RTICs?
Our RTICs are built not only by people who understand our data but alongside experts in their respective fields.
As an example, our various 'Space' RTICs (space economy, space energy, in-orbit space manufacturing etc.) were built in collaboration with the Satellite Applications Catapult. They helped us both build the taxonomy, as well as check which companies belonged in the training set.
Some of our RTICs were built with our own internal experts. For example, Software Development was built in collaboration with our development team.
**Updating RTICs**: Find out more about how we keep RTICs current in our [build methodology article](/our-data/rtics/how-do-we-build-rtics#kept-current-not-built-once).
#### Quality Assurance
The Data City’s quality assurance processes are applied to ensure data integrity and robustness. TDC implements a manual quality assurance process before the publication of RTICs.
The process consists in manually checking 30 random companies from each subcategory identified in the taxonomy. The manual check includes a review of:
* **Accuracy of the URL-company match**: The Data City manually checks whether the URL assigned to a particular company is correct. This is done by comparing the company name, company postcodes and director details.
* **Accuracy of the machine learning classification method at the RTIC level**: The Data City manually checks whether the company that has been classified in an RTIC provides products and/or services for the sector. In short, it implies checking whether the captured companies work in the sector. This is corroborated by the companies’ website information.
* **Accuracy of the machine learning classification method at the vertical level**: manual check of the activity of the companies captured in each vertical.
Through this process, The Data City can estimate the accuracy of its datasets — see our [quality assurance process](/our-data/rtics/how-do-we-build-rtics) for how. RTICs are updated through an annual Full Update, a monthly Maintenance Update, and ad hoc Hotfixes raised from reports on the platform. See [RTIC update types and versioning](/our-data/rtics/updating-rtics).
**RTIC updates**: Learn more about how we keep RTICs current [here](/our-data/rtics/how-do-we-build-rtics#kept-current-not-built-once).
#### How can you use RTICs?
RTICs provide unique insights into emerging economies that SIC codes cannot.
View RTICs at a top level to get a glimpse into the sector as a whole, or combine them with filters to get a more refined view.
RTICs can be found under the RTICs section in the platform.
Use location filters to get regional insights, or growth filters to find high-growth companies and much, much more.
This page offers an overview of all the RTICs available on the platform, including their definitions, creation date, latest update, number of companies, and details about the included verticals. These RTICs can be selected for further exploration using the [EXPLORE](/using-industry-engine/tools/explore) function.
**EXPLORE**: Find out more about our EXPLORE tool and how to search our company database [here](/using-industry-engine/tools/explore).
You can find more information on [RTIC methodology in this article](/our-data/rtics/how-do-we-build-rtics).
### Real-Time Standard Industrial Classifications (RSICs)
RTICs are great for the emerging economy. For the foundational economy, where there is more likely to be an appropriate SIC, an issue remains. A company can choose the wrong SIC code, or the SIC code they've selected does not match their activities.
We fixed this issue. RSICs use machine learning and a company's website text to better classify company's activities.
Many of our users want to link companies to official statistics, or compare a sector internationally. SICs are widely used for official statistics, and RSICs will allow you to better understand the activities of a company, and the most appropriate sector to compare to.
### Data Services
The Data City has a team of data analysts and data scientists that can be contracted to work on a project. The daily rate for data analysts and data scientists is 1000/day. Below are described the main services provided by our team:
* Data and Platform Support: one of our analysts will work with you to make sense of the data to make sure that you understand the potential and capabilities of the platform and answer your research questions. Likewise, you can contract one of our data scientists to discuss advanced methodologies and approaches to analyse our data in relationship with other data sources.
* Data Analysis: one or a team of data scientists will perform an agreed analysis and provide outputs that are not currently available on the platform.
* RTIC building: a team of data analysts will work with you to create your own custom RTIC and publish it on the platform. This includes taxonomy development, machine learning list production and review, and publishing.
* Bespoke work: our team of data analysts, scientists and developers can work together with partners to develop new products or tools.
### Commercials
If you are already a subscriber to the Data City platform – please get in touch to find out how we can help you use our platform and Industry Engine with our Data Services Team to successfully win your bid.
If you are **not** already a subscriber to the Data City platform – please get in touch to find out how we can help you use our platform and Industry Engine (under our pitch license, in simple terms you get all the functionality but without permission to use it commercially) with our Data Services Team to successfully win your bid.
# Why do I get different locations to the ones I've selected?
Source: https://docs.thedatacity.com/using-industry-engine/faq/additional-locations-included
Companies have two kinds of address; these are registered and operating addresses.
Companies have two kinds of address; these are registered and operating addresses.
Every company has just one registered address. This is their legal address. Greggs' registered address is Greggs house in Newcastle.
A company can have one or more operating address. Greggs has more than 1 operating address. We find 1300 locations for Greggs.
### Different locations in ANALYSE
When you apply a location filter, you will be shown companies that have at least one address in that location.
1. When you select "only include companies with registered address within filter locations" this returns companies that are registered in that area. On ANALYSE you will see the operating addresses that these companies have as well. These may be outside of the location that you have selected, unless you select "Perform analysis based on registered addresses only".
2. If you select 'Only filter by company operating addresses' this returns companies that have an operating address in the selected area. On ANALYSE you will see the operating addresses that these companies have as well. These may be outside of the location that you have selected, unless you select "Perform analysis based on registered addresses only".
3. If you select "Perform analysis based on registered addresses only", the analysis will be just on registered addresses. You will be shown companies that only have registered addresses in the selected location.
Above: Analysing Greggs when I have "Perform analysis based on registered addresses only" selected.
Below: Analysing Greggs when I **do not** have "Perform analysis based on registered addresses only" selected.
# How do I check my Smart List is done?
Source: https://docs.thedatacity.com/using-industry-engine/faq/how-do-i-check-my-ml-list-is-done
Find out more about Smart List building and how to know when your list is complete.
To know if your [Smart List](/using-industry-engine/tools/building-an-ml-list) is done, you will need to answer the following two questions:
### 1. Is my Smart List representative of the sector?
This refers to the content of the list itself and requires checking the type of companies captured.
It requires making sure that you do not have false positives on the list (companies outside of the sector of interest) and that the size and value of the sector are cohesive with previous research.
#### A) Checking for false positives:
Keyword searches are a quick and easy way to do this. Add all the keywords you have used to the 'Does not contain' keyword filter as in the image below.
The search will output all companies that do not contain the relevant language. If you get large numbers, this means that you still need to continue training the algorithm to get rid of those
#### B) Desk research and using ANALYSE:
You can use the ANALYSE function to verify the results. Hit 'Analyse Results' and the platform will give you an insight into how the companies perform as a group.
Compare your results with existing information: is your sector much larger or smaller in terms of number of companies, employees, or turnover? Are you including big companies whose main activity is not representative of the sector?
Remember that, if you are mapping an emergent sector, you may find more companies on our platform. That is OK and the purpose of our technology. This is more of a sense check that reaffirms that your results are cohesive with economic trends.
#### C) Random sampling:
Go to different pages of your list randomly and check the type of companies captured.
### 2. Have I missed part of the sector while building my list?
#### Check for false negatives:
Companies part of the sector of interest that have been left out of the classification. There are two easy checks that you can do to find it out.
#### A) Change the score cutoff value:
As a default, the list will show all companies with a score above 0 and exclude the rest. Change the score value as in the video below to -0.5/-1.
#### Check the companies in that range:
If you see relevant companies with a negative score that are relevant to the sector, you have trained the algorithm too strictly.
Add some of those to your positive training set and repeat until you find a very few of them (which could be added manually to the list at the end of the process).
**Video to re-embed.** This guide originally embedded a HubSpot video (id 152568649375). Add an updated embed link in this article.
#### B) Check keywords in EXPLORE
Make some keyword searches on [EXPLORE](/using-industry-engine/tools/explore) with your keywords and check if any companies are not on your list. You can check if a company is on your list by using the search box at the top right corner.
**Building an Smart List**: You can find out more about our Machine Learning List Builder and how to use it effectively in our [Building an Smart List step-by-step guide](/using-industry-engine/tools/building-an-ml-list).
# Reverting downloads to old locations format
Source: https://docs.thedatacity.com/using-industry-engine/faq/reverting-downloads-to-old-locations-format
This guide shows you how to use Excel's Power Query to convert your company data from the new two-tab format back to the old single-table format with concatenated location data.
### Converting Company Data from New Format to Old Format
What you'll get:
One row per company with all location names combined in a single column, separated by commas.
Time required:
5-10 minutes
#### Overview
Your download now comes in two separate tabs:
* Companies tab: Contains company information (Companynumber, Companyname, etc.)
* Locations tab: Contains location data (Companynumber, Postcode, LAname, LAcode, etc.)
This guide uses Excel's Power Query to combine these into the old single-row format where each company has all its location data in one concatenated column.
#### Steps
1. Prepare your data
* Download the data from the platform and ensure both Companies and Locations tabs exist
* Open a blank Excel file
2. Load data into Power Query
* Go to Data > Get Data > From Other Sources > From Workbook
* Select your Excel file you've just downloaded and choose both tables
* Load them into Power Query Editor by selecting Transform Data (bottom right)
**2a. Ensure Excel has promoted the headers for you**
* Sometimes Excel will do this for you but you need to promote the headers so we can used named columns.
**3. Join the tables**
* In Power Query Editor, select the Companies table
* Go to Home > Combine > Merge Queries as New
* Select Locations table as the second table
* Join on Companynumber column from both tables
* Choose Left Outer join type
4. Create concatenated columns
* In the ribbon, click Add Column tab
* Click Custom Column
* New column name: Type "LAname\_Combined"
* Custom column formula: Copy and paste this exactly:
```text theme={null}
Text.Combine([Locations][LAname], ", ")
```
* Click OK
5. Repeat Step 4 as necessary
* For LAcode data, click Add Column > Custom Column
* New column name: Type "LAcode\_Combined"
* Custom column formula:
```text theme={null}
Text.Combine([Locations][LAcode], ", ")
```
* Click OK
* For Postcode data, click Add Column > Custom Column
* New column name: Type "Postcode\_Combined"
* Custom column formula:
```text theme={null}
Text.Combine([Locations][Postcode_withspaces], ", ")
```
* Click OK
* Repeat for any other location columns you need (Constituencies, ITL regions, etc.)
6. Clean up and load
* Remove intermediate columns
* Close & Load to get your final table
# API access
Source: https://docs.thedatacity.com/using-industry-engine/features/api
Find out more about our API and how to access it.
We have a JSON based REST API that allows you to download the data behind the lists you have created (ML and EXPLORE), search for company details and more.
Documentation for our API lives at [**/api-reference**](/api-reference), with a [quickstart](/api-reference/guides/quickstart), [authentication](/api-reference/guides/authentication) guide, and a live playground for every endpoint.
Please note: an additional license is required to access data via the API.
**Interested in accessing our API?** Speak to one of the team or [contact us today](https://thedatacity.com/contact-us/).
# Getting started with our API
Source: https://docs.thedatacity.com/using-industry-engine/features/api-getting-started
This guide has moved to the new API reference.
The getting-started guide has moved.
Make your first request to The Data City API in under five minutes — Bearer token, base URL, end-to-end example.
For everything else, see the [API reference](/api-reference).
# Downloading data
Source: https://docs.thedatacity.com/using-industry-engine/features/downloading-data
Find out how to export data for offline analysis
If you want to export data from our platform for further analysis, you can do so in CSV or XLSX formats.
You have the option to select which parts of the data you wish to export, which you will receive as a row-based view\* of company data. All downloads will be restricted to the first 20,000 results. If you select any field within ‘yearly financials’, the data will contain multiple rows for each company (one row for each financial year).
\*we have removed our old column-based view of company data (Jul 2024).
To download an export just press the three dots next to your list name and select **Download List** from the menu.
You can download data directly from [EXPLORE](/using-industry-engine/tools/explore), [ANALYSE](/using-industry-engine/tools/using-analyse) and [Smart ListS](/using-industry-engine/tools/building-an-ml-list).
You will then be presented with a choice of formats and data to include. If you choose CSV, you will receive a zip file with a number of separate CSV files. The XLSX version will be one file with multiple sheets.
Each checkbox represents a column in the spreadsheet. There are over 80 different checkboxes, grouped into 8 categories. Financial and employee data is broken down yearly, and only given if it is available for at least one company.
### BASIC
This covers [essential data](/our-data/third-party-data/companies-house) about a company e.g. company number, incorporation date and RTICs.
### COMPANY
Contains further company information that may be of interest e.g. [innovation score](/our-data/proprietary-data/innovation-score).
### CONTACTS
Email, phone number and socials of each company if known.
### YEARLY FINANCIALS
In-depth financial data, including [estimated turnover](/our-data/key-data-and-definitions/estimated-turnover-employees-growth), wages/salaries and profit after tax. This data is broken down yearly where available.
### GROWTH
Metrics relating or contributing to the growth of a company, e.g. best estimate growth percentage per year, and [investment](/our-data/third-party-data/investment-data) & [Innovate UK](/our-data/third-party-data/innovate-uk-data) funding data.
### LOCATION
Geographical information about a company, for example, its registered address, constituency and Strategic Authority. [Make sure you understand the difference between the available geographies](/using-industry-engine/features/using-filters#locations).
### PEOPLE
Employee data and [founder/director gender data](/our-data/proprietary-data/gender-data).
### SECTORS
View which SIC and/or [CIC](/using-industry-engine/tools/my-lists#cics) code a company is classified within.
### In summary, with the download, you’ll receive:
* A row-based view of the company data you chose to include (unless you choose financial fields).
* Directors’ details for each company within the list
If downloading from an [Smart List](/using-industry-engine/tools/building-an-ml-list), you will also receive:
* A breakdown of the scoring system used by the classifier
* Details of the training set you provided when you created the list e.g. companies to include
**API**: If you need more than 20,000 records, you may find our API a better solution, information of that can be [found on our API page](/using-industry-engine/features/api).
# Using Filters
Source: https://docs.thedatacity.com/using-industry-engine/features/using-filters
Use filters across The Data City platform to perform in-depth searches.
LEPs: These are being removed in October's update.
[Strategic Authorities:](#locations) These have been added in September's update.
Our platform offers a powerful filtering system designed to enhance your user experience by helping you pinpoint precisely what you're looking for.
These filters are located in the [EXPLORE](/using-industry-engine/tools/explore), COMPARE and ANALYSE section of our platform.
Filters are divided into eight categories, each serving a distinct purpose. Users can select a number of filters across a range of categories to bring about a specific set of results.
### RTICs
This filter helps users navigate our database and discover results for a specific Real-Time Industrial Classification (RTICs). You can scroll or use the search filter to select as many as you want.
Each RTIC has verticals you can see by clicking its dropdown menu beside it. For example, Adtech has 5.
### CICs
CICs are Custom Industrial Classifications. Essentially they are a custom RTIC that's only available to selected group of users.
From here you can view and select any CICs available to you.
**Note**: The filter is only available if the user has access to a CIC.
### Sectors
Filter companies by specific sectors within the Standard Industrial Classification (SIC) Codes. The sectors are divided into 3 sections:
* [**SICs**](https://resources.companieshouse.gov.uk/sic/): Target companies by their Standard Industrial Classification (SIC) codes. Search by specific codes or browse through categories. (Data source: [Company house](https://www.gov.uk/government/organisations/companies-house))
* **SIC** **sections**: Focus on broad industry groups using SIC Sections. Choose from 21 high-level categories like 'Manufacturing' or 'Construction'.
* **Categories**: Identify company types and structures. Filter by categories like 'Limited Liability Company' or 'Non-Profit Organisation'.
### Locations
Refine your results by a specific location, region or city.
We group our locations into:
* **Cities**: The cities filter consists of UK functional urban areas (FUA) as defined by the [OECD](https://www.oecd.org/regional/regional-statistics/functional-urban-areas.htm). An FUA is composed of a core city and its commuting zone. The cities filter reflects the core city and its commuting zone. You can read more about this [here](https://thedatacity.com/blog/product-update-our-cities-filter-just-got-even-smarter/).
* **LAs**: The local authorities (LA) filter is composed of local authority districts Feb 2024 [#1](https://www.data.gov.uk/dataset/e3a0a920-d0e0-49c0-b030-9db882eca78f/local-authority-district-april-2023-to-lau1-to-itl3-to-itl2-to-itl1-january-2021-lookup) [#2](https://geoportal.statistics.gov.uk/datasets/e832e833fe5f45e19096800af4ac800c/about) of the United Kingdom.
* **ITL1 Regions (International Territorial Level 1):** The regions filter is comprised of [UK ITL1](https://www.ons.gov.uk/methodology/geography/ukgeographies/eurostat) regions.
* **ITL2 Regions (International Territorial Level 2):** The regions filter is comprised of [UK ITL2](https://www.ons.gov.uk/methodology/geography/ukgeographies/eurostat) regions.
* **Constituencies**: The constituencies filter is comprised of [UK parliamentary constituencies.](https://www.parliament.uk/about/how/elections-and-voting/constituencies/)
* ****Strategic Authorities****: Strategic Authorities reflect the UK’s updated regional policy landscape, following the phase-out of Local Enterprise Partnerships (LEPs). Strategic Authorities are [replacing](https://www.gov.uk/government/publications/english-devolution-and-community-empowerment-bill-guidance/english-devolution-and-community-empowerment-bill-guidance) Combined Authorities and they are now responsible for driving local economic development.
* **Parent Nation:** The parent nation filter enables filtering by the nationality of a company's ultimate parent. This is available where we have data on company group structure.
In the locations filter, there is a checkbox: "Exclude companies known to be ultimately foreign owned". Selecting this checkbox will select all companies for where **we know** the ultimate parent company **and** we know that company is in another country (**not the UK)**. In some cases, the ultimate parent company is the same as the parent company.
### Keywords
Refine your search using keywords that match your interests.
You can find companies that use or do not use specific keywords on their website.
'**Contains' -** Filter by keywords found on a company website.
'**Does not contain' -** Filter by keywords not found on a company website.
These can be used in combination to refine your results. You can read more about the use of the [keyword filter here](/using-industry-engine/features/using-keyword-filters).
### Financials
Allows users filter by Turnover, Employee Count, Total Funding, Pre Tax Profit, Shareholder Funds, Export, EBITDA, and Total Innovate UK Funding – set minimum and maximum values for precise searches.
* **Turnover**: The total revenue generated by a company in a year. The turnover filter includes both reported and estimated turnover and refers to the turnover for the latest complete year. (Data source: [Creditsafe](http://www.creditsafe.com) )
* **Employee Count**: The number of people employed by a company. The Employee Count filter includes both reported and estimated Employee Count and refers to the Employee Count for the latest complete year.
* **Total Funding**: The total amount of money a company has received from investors which includes both private and grant funding. (Data source: our [investment data](/our-data/third-party-data/investment-data))
* **Pre-Tax Profit**: The company's earnings before taxes are deducted.
* **Shareholder Funds**: The net worth of a company, representing the value of its assets minus its liabilities. (Data source: Companies House ([CreditSafe](http://www.creditsafe.com)))
* **Export**: The total value of goods and services a company sells to other countries.
* **EBITDA**: Earnings Before Interest, Taxes, Depreciation, and Amortization, a measure of a company's operational profitability.
* **Total Innovate UK Funding:** The amount of grants or investments a company has received from Innovate UK, a UK government agency supporting innovation. (Data source: [Innovate UK grant funding](https://www.ukri.org/publications/innovate-uk-funded-projects-since-2004/))
### Companies
Allows users streamline their search by targeting specific companies. This includes:
* Only include companies with matched URLs
* Exclude foreign companies
* Filter by [gender category](/our-data/proprietary-data/gender-data) — for directors (leaders) and founders separately. Each has six mutually-exclusive options: all women led, majority women led, mixed led, majority men led, all men led, and uncertain. Ticking more than one in a group returns companies in any of them.
Users can also add companies they want to be on their list their company number or comapany name.
### Growth
Use the 'Growth' filter to explore high-growth companies in a specific sector.
* **Growth Rate:** Allows users Select from 'Shrinking fast', 'Shrinking', 'Stable', 'Growing' or 'Growing fast' to focus on specific trends. ([Link](https://thedatacity.com/blog/focusing-on-company-growth/))
* **Company Growth Percentage Per Year**: This filter allows users to pinpoint companies by their annual growth rate by entering a specific percentage range.
* **Incorporation Date**: Narrows results by company age, settable from 1900 to 2024.(Data source: [Company house](https://www.gov.uk/government/organisations/companies-house))
* **Company Size:** Filters by employee count (no employees, micro, small, medium, large). (Data source: [Company size](https://www.legislation.gov.uk/ukpga/2006/46/part/15)).
* **Spinouts and Scaleups**: Separate options to focus on spinouts or OECD-defined scaleups. ([Link](https://thedatacity.com/blog/focusing-on-company-growth/))
**Clear Filter:** Clear your filters and start anew with a simple tap of the Clear button.
**
**
# Using keyword Filters
Source: https://docs.thedatacity.com/using-industry-engine/features/using-keyword-filters
Find out how to use the keyword filter and structure keywords for improved company searches.
#### What is the keyword filter and where can I find it?
Our users commonly use keywords to find companies for training sets in [Smart Lists.](/using-industry-engine/tools/building-an-ml-list) They are also used to fine-tune searches in [EXPLORE](/using-industry-engine/tools/explore) and to retrieve more relevant and targeted companies for analysis.
The keyword searches generally follow boolean logic.
**Using filters**: To find out more about our range of search filters, make sure you [read our 'Using Filters' guide](/using-industry-engine/features/using-filters#keywords).
#### Common search queries
Use the ***‘ALL of these words’*** dropdown when you have a set of keywords and want to find companies that use all of the keywords you have mentioned in their website text.
If you want to find companies using any one of the listed keywords, then select ‘***ANY of these words***\*’.\*
*
*
1. **Using inverted commas.** surrounding phrases helps in identifying companies with the same keyword referenced in your search. For example: “data analytics”.
2. **Using NEAR in searches**. for example: NEAR (“Data” “insights” , 5) you are letting the platform output companies that have ‘data’ and ‘insights’ within 5 words of each other. Be careful to get the syntax right.
3. **Utilising an asterisk *(\*)*** ***for broader search terms.*** For instance, entering “data analy”\* allows the platform to query for companies which include phrases starting with 'data analy', such as data analysis, data analytics, etc. This method broadens the search to capture all relevant variations that stem from the root 'data analy'.
4. **Combine multiple searches in keywords.** For example, the following search returns all the companies that may work with machine learning or data analytics within the artificial intelligence sector. *(“Artificial Intelligence” OR “AI”) AND (“Data analy”\* OR “machine learning”)*
5. **Complex keyword searches**. In the most advanced applications of this technology we develop complex keyword searches such as: *(“artificial intelligence” OR “ai” OR “neural networks” OR “deep learning” OR “machine learning”) AND (“data analytics” OR “data analysis” OR “insight”\* OR “big data” OR “data process”\* OR “data governance”)*
6. **Combining keywords with AND and OR queries**. For example: “Data Analytics” AND “Machine learning” will return companies that use both data analytics and machine learning in their website text. This approach is equivalent to utilising the provided dropdown menu, but it offers a more sophisticated method for those who prefer to use keyword filtering in a more advanced manner.
**False positives**: Read more about how to use keywords to identify false positives in [our Smart Lists article](/using-industry-engine/faq/how-do-i-check-my-ml-list-is-done).
#### Caveats
Keywords are powerful tools for targeting specific areas, but must be used with caution.
1. **Keyword Bias**: When using the keyword filter in EXPLORE, it brings up any company with the specified keyword, regardless of its context. This may increase the chance of having false positives. The specific keyword has to appear just **once, anywhere** on its whole website! To avoid this, it is recommended only to layer keywords on top of RTIC data.
2. **Complexity**: When creating advanced keyword searches, it is recommended to avoid constructing excessively long keyword strings, as they can potentially slow down the search and they are prone to bugs (due to input formatting).
3. **Limited Accuracy**: Depending solely on keyword searches in the EXPLORE section to map a sector is not recommended as it can impact the overall accuracy of the list. For such cases, using Smart Lists is preferable to achieve a more precise and accurate sectoral mapping.
# Building an Smart List
Source: https://docs.thedatacity.com/using-industry-engine/tools/building-an-ml-list
Our platform enables you to generate your own Machine Learning (ML) list, underpinned by our proprietary classification system. Find out how to build a list.
Table of contents
1. [Overview](#overview)
2. [Getting started](#getting-started)
3. [Training your list](#training-your-list)
4. [Finalising your list](#finalising-your-list)
### Overview
Smart Lists allow you to build and classify company in a matter of minutes, using our AI tech.
To build a machine learning list you will need to train the algorithm to capture companies you are interested in and exclude the rest.
This is an iterative process: you will build your lists in stages until you get the desired output.
To start building your list you will need:
* A set of companies that you know are good representatives of the industry
AND/OR
* A set of keywords that define the activity of the companies in the industry.
The algorithm will analyse the companies that you provide as training data and create a model to classify the rest of the companies available in our database.
This will assign each company with a score value based on how similar the company is to the ones that you provided as good examples of the sector. The output list will be all companies that have a score value above 0.
Smart Lists are available to all Data City users and can be set up directly from MY LISTS in the platform.
### Step 1 - Getting started
To get started, head to MY LISTS and click 'Create a New List'.
Add the first example companies using one or both of the methods shown in the video. You can either insert company numbers or use the keyword search engine to find good examples.
**Video to re-embed.** This guide originally embedded a HubSpot video (id 152442612400). Add an updated embed link in this article.
The output list ranks all companies against the ones you selected (positive training set). The score value indicates how similar the website text is to the companies in the training set.
Hit "CLASSIFIER TERMS" to understand the language that is leading the classification. You should be checking this at different stages to make sure it remains relevant.
### Step 2 - Training your list
Start training your list to exclude companies that you don't want in the list by adding these to the negative training set by clicking the "x" button.
You can use the keyword filter or the score value to identify these companies - companies at the end of the list will have lower score values.
**Video to re-embed.** This guide originally embedded a HubSpot video (id 152447307109). Add an updated embed link in this article.
### Step 3 - Finalising your list
Repeat this process until you get a representative list. You can add more companies to your positive training set by clicking the "+" button.
Important considerations:
* The companies that you add to your training sets lead the classification, not the keywords that you use to filter the list.
* The keywords that you use to filter the list are helpful to find companies for the training sets.
* Once the algorithm has generated a list, you can manually add or remove companies. This is useful when you want to add a company without impacting the training sets or the classifier terms, or when a company does not have an available website.
**Video to re-embed.** This guide originally embedded a HubSpot video (id 152473222933). Add an updated embed link in this article.
* Check the big companies and manually exclude them if they are not relevant. Big multinational companies may use relevant language on their website, but might not be a good representative of the sector in question. This is important because they will skew the results for the overall performance of the sector in subsequent analysis if left in.
# Using COMPARE
Source: https://docs.thedatacity.com/using-industry-engine/tools/compare
The COMPARE tool allows users to easily visualise two different sectors or Explore lists side-by-side.
### Introducing COMPARE
Our COMPARE page tool allows you to quickly compare and contrast two datasets through visualisations of overview statistics.
It is based around our separate ANALYSE tool. Essentially, it allows you to combine two separate ANALYSE page views side-by-side.
### Getting started
The COMPARE tool is a core feature of our platform and can be accessed from the link in the header bar.
When you first navigate to the page you will be presented with a blank view and two sets of controls, one for each dataset.
The control bars allow you to define your comparison and can be configured in two ways:
* Using filters to define a sector from scratch.
* Loading a pre-existing EXPLORE list.
### Filters
The COMPARE page uses the same set of filters as other Data City tools. Company searches can be refined based on properties such as [RTICs (Real-Time Industry Classifications)](/our-data/rtics/what-are-rtics), location and financials.
**Using filters:** for more information about our filters and how to use them please visit our [Using Filters page.](/using-industry-engine/features/using-filters)
### EXPLORE Lists
COMPARE allows you to quickly load in previously saved lists created on our [EXPLORE tool](/using-industry-engine/tools/explore).
To do this click the three dots in the top left of either of the two control bars and pick the "Load list" option.
Then pick the relevant list from the dropdown options. The control bar filters will automatically be updated.
**Creating lists in EXPLORE:** for more information about our EXPLORE tool please visit our [Using EXPLORE page](/using-industry-engine/tools/explore).
### Naming your datasets
Each COMPARE dataset can be given a custom name.
To do this click on the textbox in either control bar and enter a new value. If you load in an EXPLORE list the dataset will automatically be renamed to match the list.
The two datasets are named "Group 1" and "Group 2" by default. Their names can be updated at any time without needing to recalculate the results.
### Results
Once you are happy with your input hit the "Calculate" button at the top of the page to start processing the results.
The results for each dataset are displayed alongside each other. The COMPARE results contain all the same data fields available on our ANALYSE tool.
Total turnover, business counts and RTIC vertical membership are just some examples of the data available in COMPARE.
These are grouped into sections, and results are ordered by the combined count of both values for a datapoint.
Pressing the three dots in the top-right corner of each data chart will allow you download the results for that particular field in CSV format.
Through this menu you also have the option to see a company list view of either dataset in the EXPLORE page.
***
**Overall the COMPARE page is a powerful tool for quickly visualising the differences between two lists. It offers a new way of easily seeing the strengths or weaknesses of industrial sectors.**
# Using EXPLORE
Source: https://docs.thedatacity.com/using-industry-engine/tools/explore
The EXPLORE function is a powerful tool that allows users to search, sort, and filter companies, providing detailed company-level information.
### Introducing EXPLORE
The **EXPLORE** function is the core of the Industry Engine application. Here users can create, view and refine company lists as they see fit.
You can use EXPLORE in a number of ways:
* **From scratch** - Use filters to define a list direct from EXPLORE.
* **From an RTIC** - Jump straight into an RTIC and view companies in EXPLORE.
* **From ANALYSE** - Take your analysis from ANALYSE and view as a list in EXPLORE.
EXPLORE by default shows all the companies listed on the platform (approximately 5.4 million). The filter function helps to narrow down this huge list of companies based on the user needs.
Get to know the key features of the tool below:
### Key features
#### 1. EXPLORE header
The header of EXPLORE indicates what dataset you are viewing.
By default EXPLORE showcases 'All UK Companies' when clicked through the main navigation, however you can open other lists from [RTICs (Real-Time Industry Classifications)](/our-data/rtics/what-are-rtics) and other parts of the platform.
The three dots beside the header provides option to save a list created, download the list, open in ANALYSE and more.
The user can built a custom list using filters or the search bar to get specific set of companies and then save the list on the platform or download it in various formats including json, xls, csv.
The list can be downloaded either with basic information or detailed based on the complexity of the analysis.
#### 2. Filters
The filter section can be used to refine the search based on several criteria.
This includes RTICs (Real-Time Industry Classifications), CICs (Custom Industrial Classification), Sectors, Locations, Keywords, Financial data, Companies, and Growth metrics.
**Using filters**: you can find out more about our filters and how to use them [on our Using Filters page](/using-industry-engine/features/using-filters).
#### 3. Sort by
This dropdown helps to sort the dataset based on different parameters, with the current setting being 'Sector keyword counts: High to low'.
The following screenshot shows the different types of sorting available. This is useful for finding specific companies in larger datasets and lists.
#### 4. Search Bar
The search functionality on top right helps the user to look up for companies by name, number, or URL.
### Building a list in EXPLORE
#### Using the Search Bar
Users can create a list in various ways, either by applying filters or by searching for a group of companies using the search bar.
For instance, if a user enters "03457" in the search bar, as illustrated in the snapshot,
the platform will return a list of companies that include "03457" in their company number, as shown below.
The platform returns a list of 290 companies that meet the entered condition. This method can also be used to search for companies based on URLs or names.
#### Using Filters
Using filters is a great way to build user specific list.
**Example**: Say a user wants to know all the companies that provide Legal Services and are based either in Manchester or London.
This can be achieved in few simple steps. First we select the Legal Services RTIC from the RTIC option in the filter menu. More details about what is an RTIC can be found [here.](https://thedatacity.com/rtics/)
The Legal Services RTIC is further divided into many verticals to get more specific companies in the Legal Services domain. For this example we consider all the companies falling in the Legal Services category.
Next, we need to select Manchester and London from the Location option in the filter menu. The location filter has various options like filtering out only companies with registered address in selected location or filter by [ITL regions.](https://en.wikipedia.org/wiki/International_Territorial_Level)
The filters used in this example will generate a list of companies that offer Legal Services and are either registered or operating in Manchester or London.
Once the final list is compiled, it can be saved on the platform to revisit again or downloaded using the options provided in the header.
You can also filter geographically by postcode. In the Location filter's **Postcodes** tab you can select one or more postcode outcodes (for example `LS1`), search a single full postcode within a chosen radius, or paste a list of up to 2,500 specific postcodes to match exactly — handy when you already have a set of postcodes and want only the companies operating there.
**Using filters**: For more on Filters, [view our full guide here](/using-industry-engine/features/using-filters).
### Company overview
The following screenshot shows the information shown for each company in EXPLORE.
1. **Record Count:** The '1 / 5,467,419' indicates the current company's position in the overall dataset and the total number of companies available.
2. **Company Record:** This section provides the name of the company with the associated company number, incorporation date and registered postcode.
3. **Company Website:** This section provides the link to the company's website. Users can also report discrepancies between the listed company website and the actual company. The callout also provides the reason why a particular URL was selected for that company.
4. **Description:** A brief description of the company's activities or offerings.
5. **RTICs:** Listed industries or sectors the company is associated with based on real-time industry classification. Read more about RTICs [here](https://thedatacity.com/rtics/).
6. **CICs:** Custom Industrial Classifications are identical to RTICs but are only provided to a select group of platform users. They are developed using the same process used to build RTICs and are tailored to meet the needs of a specific user base.
7. **Sector Keyword Count:** Represents the frequency of specific keywords related to a business sector found on a company's website, providing a quantitative measure of the company's association with industry trends.
8. **Innovation Score & more**: This section provides the [Innovation Score](https://thedatacity.com/blog/innovation-measure-updated/), Estimated turnover and Estimated Employee Count for the company.
**Our data**: You can find out more about our data and how to use it [on our data page](https://help.thedatacity.com/knowledge/our-data).
### Company pages
To find out more about a specific company to see more in-depth data and insights, simply click on the 'Full Info' button from EXPLORE. This button leads to the individual company page.
On this page you can access key contact information, financial and funding data, growth measures and estimates, employee stats, sector classifications, locations and much more.
This is the perfect place for conducting due diligence and analysing individual companies.
***
Overall the EXPLORE tool offers a straightforward way for users to delve into detailed company information and view lists in a number of ways.
The company-level data allows users to easily conduct research, find new prospects, refine Smart Lists and more.
# MY LISTS
Source: https://docs.thedatacity.com/using-industry-engine/tools/my-lists
Access and simplify your list management process through the My Lists section.
### Introducing My LISTS
The MY LISTS page allows users to save and view different lists ranging from [Machine Learning (ML) lists](/using-industry-engine/tools/building-an-ml-list), EXPLORE Lists to CICs.
The page simplifies the management and organisation of your lists.
### Smart Lists
These are lists of companies built using Machine Learning (ML).
They are built by first defining a training set of companies that are examples of the industry vertical you wish to investigate.
Companies are ranked by how representative they are of the companies selected in the training set. User can easily create their own Smart List by selecting the **CREATE A NEW LIST** tab.
**Building an Smart List**: Find out how to build a list in our detailed [Building an Smart List guide](/using-industry-engine/tools/building-an-ml-list).
Users can view the Smart Lists created by them or any lists shared in the most recent list section.
This section shows only the top few lists, you can access all the Smart Lists in **SEARCH ALL LISTS** section. Lists can be added to folders, shared with other users, edited, and duplicated within this area.
### EXPLORE lists
It is a list of companies created by [filtering](/using-industry-engine/features/using-filters) our data in EXPLORE.
Users can save, access, share, and duplicate these lists, similar to Smart Lists.
**Using EXPLORE**: Find out how to make the most out of this feature in our [Using Explore guide](/using-industry-engine/tools/explore).
### CIC's
CIC or Custom Industrial Classifications are identical to [RTICs](/our-data/rtics/what-are-rtics) but are only provided to a select group of platform users.
They are developed using the same process used to build RTICs and are tailored to meet the needs of a specific user base. The CIC's that are available to the user can be accessed in this area.
# Smart Search
Source: https://docs.thedatacity.com/using-industry-engine/tools/smart-search
The Smart Search feature allows users to find companies based on phrases, concepts, sentences and even complete paragraphs. This is a move away from keyword search and a drastic improvement to the discoverability of companies.
### Introducing Smart Search
The **Smart Search** feature uses a Large Language Model (LLM) based approach of semantic search, matching the meaning *and* context of the search query to the meanings and context of entire web pages. It then selects the top (up to) 500 companies (where available) with web text that is most similar to the input search phrase.
### Functionality
Once a list has been created using a Smart Search, it can be filtered down using the same filtering options as are available on the [Explore Page](/using-industry-engine/tools/explore). Or the list can be taken through to the Explore page itself to be further filtered.
When a list is filtered in-place (within the Smart Search page), a fresh “top 500” companies will be selected to the original list with the filters applied. For example, if a location filter is applied, such as “Leeds”, the list will now rank all companies based on their similarity to the search phrase, it will filter for “located in Leeds”, and then select the top 500 companies from this list. As such, applying a filter will not, necessarily, reduce the number of companies in the resulting list, as might be expected.
### How is it useful?
This is especially useful when looking for companies in emerging sectors where the hot topics are constantly developing, and buzz words come and go. Or even in identifying sectors which use phrases synonymously to convey the same thing (think “recruitment”, “head-hunting”, “candidate resourcing”, “talent acquisition”, etc.).
It also could be used as another resource in the initial phase of Smart List building, offering as an alternative to a keyword taxonomy when seeking to identify relevant companies in the development of training sets.
### Best use cases
The tool is very new, and we are still exploring best use cases.
From our experience so far, longer search phrases generate “better” results.
The more descriptive text relating to the kind of companies that the user is looking to identify generally serves a better chance of identifying more of the target companies.
It is not, at this stage, recommended to include terms in the search which would be better applied as filters later, such as location identifiers “AI companies *in Leeds*”. Nor would we recommend including growth indicators in the search “Fastest growing AI companies”, since these terms are unlikely to be found in the company’s web text. These markers are better-applied post-search using the filters.
Good examples can be found underneath the search box as a guide.
See [this supporting blog](https://thedatacity.com/blog/building-the-industry-engine-smart-search/) post for more suggestions.
# Using ANALYSE
Source: https://docs.thedatacity.com/using-industry-engine/tools/using-analyse
ANALYSE is the quickest way to understand a list of companies on our platform.
### Introducing ANALYSE
ANALYSE is the home of summary statistics on our platform. ANALYSE offers the ability to conveniently understand the total employment, amongst many other metrics, of a list of companies.
The list of companies could be in the form of companies in an RTIC, a machine learning list that has been built, or simply a list of company numbers.
The creation of summary statistics requires careful interpretation and careful consideration of the companies that are included in a list.
This article will provide a summary of the considerations to be made for each data point, with more detail available in the relevant Knowledge Base article for that data point.
### Starting points for analysis
#### Analysing RTICs
The filter bar, shown above, contains all of the filters you need to hone and then analyse your list. From here you can choose to analyse RTICs.
[Read more about our filters.](/using-industry-engine/features/using-filters)
#### Analysing Personal Lists
If you have created your own Machine Learning (ML) list using [our tool](https://products.thedatacity.com/v2/define), then you can also analyse these here. Click lists in the top left, and from the dropdown, select the Smart List you have created.
[Find out more about Smart Lists and how to build them.](/using-industry-engine/tools/building-an-ml-list)
#### Analysing a List of Company Numbers
If you want to quickly get a sense of the size of a set of companies you can paste in a list of Company House numbers.
In the Companies filter, select the COMPANY NUMBERS tab, and paste your list of company numbers into the box at the bottom. Click update and you will now analyse these companies.
How to save a list? If you want to access this view again, you can save these companies as a list. At the top of the page, select VIEW LIST.
Then from the dropdown, select SAVE LIST. Give your list a name and click save. From here you can share your list with other platform users. Sharing is caring, especially with lists. How do I share a list?
Click SHARE LIST, and enter the email address of the user you would like to share the list with. The list will appear under the user's MY LISTS page. Users are not automatically notified that a list has been shared with them.
###
### The Analysis Summary Box
When you analyse the list or RTIC that you've selected, the first thing you'll see is the Analysis summary box. This is shown below.
Within ANALYSE, the summary box is the quickest way to understand a list. There is one thing to bear in mind when interpreting the **Total Employees** figure: it's a best estimate of UK employees and, while highly accurate, may in some cases include overseas employees where these can't be cleanly separated from a company's group-wide figures. [Read more about how we estimate employees.](/our-data/using-our-data/using-our-data-safely#creditsafe-employee-data-global) **Total Turnover**, and turnover figures throughout ANALYSE, always reflect a company's total global turnover, since we don't have reliable data on how much of a company's turnover is UK-based versus overseas.
**Companies Considered**: The unique number of companies in the list.
A company can be in multiple RTIC verticals or multiple RTICs. You can use companies considered to report the total number of companies without double counting.
**Total Employees**: This is a best estimate of the total number of UK employees across the companies in the list. This estimate is highly accurate, but in some cases it may include overseas employees where these cannot be cleanly separated from a company's group-wide reporting. Not all companies have declared employees for every year; The Data City estimates employees where it is accurate to do so. The number of companies for which employees are declared and estimated are also provided.
**Total Turnover**: The total turnover of companies in the list. This will include global turnover, for UK registered multinational companies. This turnover is across all of a company's operations and is not specific to a selected RTIC.
**Total Investment Funding:** This is the total amount of funding received by the companies, according to our [investment data](/our-data/third-party-data/investment-data).
**Total Innovate UK Grant Funding:** This is the total value of Innovate UK grants won by companies in this list.
**Estimated Growth Per Year**: The Data City provide an estimate of growth for companies in the list. This is based on a combination of turnover and employment.
[Read more about employee and turnover data, as well as how we estimate growth.](https://help.thedatacity.com/knowledge/estimated-turnover-employees)
**Best Estimate Total GVA**: The Data City provide an estimate of Gross Value Added for the list of companies or selected RTIC. [Please read this note before using this metric in your analysis.](/our-data/key-data-and-definitions/gva-data)
**Estimated GVA Per Employee**: This is the estimated Gross Value Added per employee. [Please read this note before using this metric in your analysis.](/our-data/key-data-and-definitions/gva-data)
**Women Founded Companies**: This contains information on the number of women founded companies, the number of women led companies, and the number of women directors. [Read more on this information.](/our-data/proprietary-data/gender-data) [https://help.thedatacity.com/knowledge/gender-data](/our-data/proprietary-data/gender-data)The asterisked number refers to the total number of directors.
### Locations
You are able to analyse the following geographies on our platform: local authorities, OECD functional urban areas, constituencies, Strategic Authorities and ITL1 and ITL2 regions. [Make sure you understand the differences between these geographies.](/using-industry-engine/features/using-filters#locations)
You can analyse the geographical distribution of businesses, employees and turnovers.
Employees and turnovers are [split equally](/our-data/using-our-data/using-our-data-safely#proprietary-turnover-employees-evenly-split) across locations for multi-site companies.
If you apply a filter to your analysis, for example, looking at a specific sector, or apply financial criteria, you'll be able to analyse using [location quotients](/our-data/key-data-and-definitions/location-quotients).
NOTE: It is possible to download the data behind many of the visualisations, by clicking the download button in the top right corner of the widget.
**Business counts by local authority**\
The number of businesses in each local authority. Where a business has operating addresses in more than one local authority, it is counted more than once. If you are filtering to specific geographies you may see results outside of those geographies where a company has additional operating addresses there.
Only businesses with at least one UK employee are included. A business is counted only when its **best estimate UK employees** value is greater than zero—that is, when The Data City has an estimated UK headline employee count for that company. Businesses with zero or unknown UK employees are excluded from these counts.
This is different from **Employees by local authority** and **Turnovers by local authority**, which use per-location employee and turnover values split across a business’s operating addresses. A company can therefore appear in the employee or turnover charts (with a zero attributed share at a given location) but not in the business count for that area if its overall best estimate UK employees is zero or unknown.
[Read more on how we estimate turnover and employment.](/our-data/key-data-and-definitions/estimated-turnover-employees-growth)
**Employees by local authority**\
The number of employees in each local authority. Where a business has operating addresses in more than one local authority, it is counted more than once. If you are filtering to specific geographies you may see results outside of those geographies when a company has additional operating addresses there.
**Turnovers by local authority**\
The total turnover of businesses in each local authority. Where a business has operating addresses in more than one local authority, it is counted more than once. If you are filtering to specific geographies you may see results outside of those geographies when a company has additional operating addresses there.
You can find out about the [types of geography](/using-industry-engine/features/using-filters#locations) we have on our platform and you can find out more about multiple company locations [here.](/using-industry-engine/faq/additional-locations-included)
### Website Analysis
As you scroll down the page, the next tab is Website keywords.
From the field dropdown you are able to choose either sector keywords or innovation keywords.
#### What are sector and innovation keywords?
Sector keywords are from a dictionary of keywords that we created alongside partners. Sector keywords identified the keywords for 18 broad sectors and themes.
Innovation keywords are keywords associated with innovative processes. The keywords range from "new markets", "continuous professional development" and "lean".
The Data City have scraped up to 90 pages of website text for companies. We then count how many times the sector or innovation keywords are mentioned; **this is the keyword count**.
#### Keyword enrichment
Keyword enrichment is a measure of how overrepresented a key word is compared to the average website. In this case "small molecule" is found 1209% (or 13x) more than the typical website.
Sector keywords are not keywords from the machine learning output. If you are interested in machine learning based keywords for a list that you have made, you should read about [the classifier terms](/using-industry-engine/tools/building-an-ml-list).
### Company Details
The following tab contains demographic information on the companies.
From the field dropdown, you can choose **Gender**, **Employees**, **Company growth**, **Company births and deaths**, or **Company trade data**.
#### Gender
The Gender section shows two 100% stacked bars — the **founder gender mix** and the **director gender mix** — each split into the six mutually-exclusive categories: all women led, majority women led, mixed led, majority men led, all men led, and uncertain. [Get more information on the gender data](/our-data/proprietary-data/gender-data), including how the categories are defined and why some companies are uncertain.
#### Employees
**Employees by year** is the total employment of the list or RTIC. The dashed line shows our estimated value of employees. The solid line is the measured level of employees. [Read more on why we estimate turnover and employment.](https://help.thedatacity.com/knowledge/estimated-turnover-employees)
**Company size by employees** shows the distribution of companies by their best estimate UK employees. In this case there are 1400 companies with zero (or unknown) UK employees.
You can change the line chart to a bar chart using the controls at the top of the widget.
**
**
#### Company growth
**Company size**
The distribution of company sizes. This takes into account the number of employees as well as turnover to define company size. Specifically, [it uses the Companies Act 2006 definition.](/our-data/key-data-and-definitions/company-sizes)
**Company growth rate**
This shows the distribution of growth rates for companies. For example, there are 508 companies in this list that are shrinking fast.
**Founding dates**
The last field of company details looks at the founding dates of companies.
At the time of writing The Data City only track active companies. Younger companies are more likely to be active than older companies. This can suggest that the sector is growing quickly in business count, but older companies have dissolved.
The Data City are working to add company births and deaths, so the rate of incorporation over time can be more accurately tracked.
**
**
### Jobs and Skills
We partner with [Lightcast](/our-data/third-party-data/lightcast-data) to be able to offer granular labour market insights on our platform. If you have subscribed for Lightcast, you will see widgets focussed on jobs and skills.
**SOC4 code**
Our widgets here show the job postings and average advertised salary by Standardised Occupational Code (SOC). Think SIC, but for occupations.
**Skills**
We show the number of advertised job postings for:
* [Common skills](https://thedatacity.github.io/documentation/docs/data-dictionary.html#company-job-postings-by-common-skill)
* [Certification skills](https://thedatacity.github.io/documentation/docs/data-dictionary.html#company-job-postings-by-certifications-skill)
* [Software skills](https://thedatacity.github.io/documentation/docs/data-dictionary.html#company-job-postings-by-software-skill)
* [Specialised skills](https://thedatacity.github.io/documentation/docs/data-dictionary.html#company-job-postings-by-specialized-skill)
These are advertised job postings and may not necessarily be filled positions.
### RTICs and Sectors
We provide a breakdown for all types of sectors on our platform: RTICs, [CICs](/using-industry-engine/tools/my-lists#cics) and [SICs](/our-data/rtics/what-is-the-difference-between-rtics-sic-codes).
It's worth noting that a company can be in multiple RTICs, CICs, or SIC codes. The visualisations here show the number of companies in a sector. This does **not** remove double counting.
At the time of writing, there are 212,000 companies that are in an at least one RTIC. Because a company can be in multiple RTICs, or multiple RTIC verticals within the same RTIC, if you sum the data from the widgets below, you will be double counting. Summing RTIC sector counts results in a number of 280,000. **Be aware of double counting here.**
**RTIC sector counts**\
The RTIC sectors of the companies in your list that match your filters. Where a company has more than one RTIC sector it is counted more than once.
**Employees by RTIC sector**\
The number of UK employees in each RTIC sector, based on each company's best estimate UK employees. Where a company has more than one RTIC sector it is counted more than once.
**Job postings by RTIC sector**\
Job postings data is sourced from [Lightcast](/our-data/third-party-data/lightcast-data). Currently this data is not available to all customers. We are looking at offering this as an upgrade.
**Turnovers by RTIC sector**\
The total turnover of businesses in each RTIC sector. Where a company has more than one RTIC sector it is counted more than once.
### Financials
Our final tab focusses on financial details and funding information.
#### Financial details
**Net worth:** Total current assets - Total liabilities for companies in this list.
**Turnover by year:** Total revenue for companies in this list
**Profit after tax:** Profit after tax for companies in this list. Not all companies declare this. A more simplistic method of estimating profit after tax is used.
**Total current assets:** Assets that can be converted into cash in the next year.
**Total liabilities:** Total amount due.
#### Funding
**Total investment funding by year:** The value of funding received in each year, from our investment data.
**Total investment funding rounds by year:** The number of rounds of funding in each year. Some rounds do not have data on the amount raised.
**Total Innovate UK grants by year** The value of Innovate UK grants received by this list.
# Use cases
Source: https://docs.thedatacity.com/using-industry-engine/use-cases/use-cases
You can find our use cases [here.](https://thedatacity.com/blog/?category=guide)
# Classify companies
Source: https://docs.thedatacity.com/api-reference/classification/classify-companies
/api-reference/openapi.json post /v2/classification/classify
Classify companies.
# Check companies exist
Source: https://docs.thedatacity.com/api-reference/companies/check-companies-exist
/api-reference/openapi.json post /v2/companies/check
# Download an export file
Source: https://docs.thedatacity.com/api-reference/companies/download-an-export-file
/api-reference/openapi.json get /v2/companies/export/download/{file}
Stream an export file through the public API.
# Export list
Source: https://docs.thedatacity.com/api-reference/companies/export-list
/api-reference/openapi.json post /v2/companies/export
# Get company details
Source: https://docs.thedatacity.com/api-reference/companies/get-company-details
/api-reference/openapi.json get /v2/companies/{companyNumber}
# Get company group structure
Source: https://docs.thedatacity.com/api-reference/companies/get-company-group-structure
/api-reference/openapi.json get /v2/companies/{companyNumber}/group-structure
Group structure (parent / subsidiary rows) for a company.
# Get multiple company details
Source: https://docs.thedatacity.com/api-reference/companies/get-multiple-company-details
/api-reference/openapi.json post /v2/companies/batch
Get the details of multiple companies in a single request.
# List companies by filter
Source: https://docs.thedatacity.com/api-reference/companies/list-companies-by-filter
/api-reference/openapi.json post /v2/companies
# List companies by multiple website URLs
Source: https://docs.thedatacity.com/api-reference/companies/list-companies-by-multiple-website-urls
/api-reference/openapi.json post /v2/companies/by-website/batch
Union of active companies matching any of the website URLs (paged).
# List companies by website URL (GET)
Source: https://docs.thedatacity.com/api-reference/companies/list-companies-by-website-url-get
/api-reference/openapi.json get /v2/companies/by-website
Active companies for one normalised website URL (paged; query string).
# List companies by website URL (POST body)
Source: https://docs.thedatacity.com/api-reference/companies/list-companies-by-website-url-post-body
/api-reference/openapi.json post /v2/companies/by-website
Same as GET by-website but accepts the URL in the JSON body (for long URLs).
# Get director details
Source: https://docs.thedatacity.com/api-reference/directors/get-director-details
/api-reference/openapi.json get /v2/directors/{personNumber}
# Get explore list details
Source: https://docs.thedatacity.com/api-reference/explore/get-explore-list-details
/api-reference/openapi.json get /v2/explore/{listId}
Get explore list details.
# Get explore lists by domain
Source: https://docs.thedatacity.com/api-reference/explore/get-explore-lists-by-domain
/api-reference/openapi.json get /v2/explore/by-domain/{domain}
Get explore lists by domain.
# Get explore lists by email
Source: https://docs.thedatacity.com/api-reference/explore/get-explore-lists-by-email
/api-reference/openapi.json get /v2/explore/by-email/{email}
Get explore lists by email.
# Get filters
Source: https://docs.thedatacity.com/api-reference/filters/get-filters
/api-reference/openapi.json get /v2/filters
# Get filters with criteria
Source: https://docs.thedatacity.com/api-reference/filters/get-filters-with-criteria
/api-reference/openapi.json post /v2/filters
# Authentication
Source: https://docs.thedatacity.com/api-reference/guides/authentication
Use a Bearer token in the Authorization header. Keys are issued manually by The Data City.
The Data City v2 API uses Bearer-token authentication. Every request must send an `Authorization` header.
## Request format
```http theme={null}
Authorization: Bearer YOUR_API_KEY
```
Tokens are sent as plain HTTP Bearer credentials — no signing, no rotation header, no expiry parameter on the request itself.
Treat your API key like a password. Never check it into source control, ship it in a frontend bundle, or paste it into a shared chat. Store it as an environment variable or in a secrets manager.
## Getting a key
API keys are issued manually by The Data City. To request one, email [support@thedatacity.com](mailto:support@thedatacity.com) with:
* The name of your organisation.
* The email address of the person who will hold the key.
* A short description of what you intend to use the API for.
A member of the team will provision the key and reply with the value.
## Replacing a key
If a key is leaked, lost, or you want to rotate it, email [support@thedatacity.com](mailto:support@thedatacity.com) and ask for a replacement. The old key will be disabled.
## Failure modes
| Status | Body | Meaning |
| ------ | ----------------------------------------- | --------------------------------------------------------------------------- |
| `401` | `{"message":"Unauthenticated."}` | Header missing or token wrong. |
| `403` | `{"message":"Invalid ability provided."}` | Token is valid but doesn't have the ability for that endpoint. |
| `429` | `{"message":"Too Many Attempts."}` | Per-minute limit hit. See [Rate limits](/api-reference/guides/rate-limits). |
See [Errors](/api-reference/guides/errors) for the full envelope and other status codes.
## Abilities
Each token is scoped to one or more abilities (for example `companies:read`, `rtic:read`, `companies:companyDetails`). The endpoints you can call are the intersection of what your token allows and what the route requires. If a request returns `403` despite a valid token, the missing piece is an ability — contact support to widen the scope.
## Admin routes
A small number of `/admin/...` routes use a different `X-API-Key` header for internal use only. These are filtered out of the public reference and are not callable with a customer Bearer token.
# Batching & company-number sets
Source: https://docs.thedatacity.com/api-reference/guides/batching
Send many companies in one request. For very large lists, upload a set once and reference it by ID.
The API supports two patterns for working with many companies at once.
## Batch endpoints
Several endpoints take an array of company identifiers in a single request:
| Endpoint | Body shape | When to use |
| ------------------------------------- | ------------- | ----------------------------------------------------------- |
| `POST /v2/companies/batch` | Array of CRNs | Up to a few thousand companies — fetch details in one call. |
| `POST /v2/companies/check` | Array of CRNs | Check which exist; lightweight existence check. |
| `POST /v2/companies/by-website/batch` | Array of URLs | Resolve many websites at once. |
These are the preferred path for "I have a list of N items, give me data for all of them" workflows.
## Company-number sets — for very large lists
Above roughly **10,000 CRNs** (and certainly above 50,000), packing them inline into every `POST /v2/companies` request becomes wasteful: each call re-sends the same big JSON array. Instead, upload the list **once** as a set, get a `setId`, and reference it from filter requests.
### Upload a set
```bash theme={null}
# Plaintext body, one CRN per line, UTF-8.
curl -X POST 'https://product-api.thedatacity.com/api/v2/company-number-sets' \
-H 'Authorization: Bearer YOUR_API_KEY' \
-H 'Content-Type: text/plain' \
--data-binary $'12345678\n87654321\n11223344\n'
```
Response:
```json theme={null}
{
"setId": "set_abc123",
"count": 3,
"expiresAt": "2026-05-19T10:23:00Z"
}
```
Sets expire after **roughly 24 hours**. Call `DELETE /v2/company-number-sets/{setId}` to remove one early.
### Use a set in filter requests
```bash theme={null}
curl -X POST 'https://product-api.thedatacity.com/api/v2/companies' \
-H 'Authorization: Bearer YOUR_API_KEY' \
-H 'Content-Type: application/json' \
-d '{
"returnCount": 1000,
"CompanyNumberSetId": "set_abc123",
"AllowLargeAnalysis": true
}'
```
When `CompanyNumberSetId` is set, omit `CompanyNumbers` — the API ignores it and the JSON stays small.
### Gzip the upload
For very large bodies, gzip the request body:
```bash theme={null}
gzip -c crns.txt | curl -X POST 'https://product-api.thedatacity.com/api/v2/company-number-sets' \
-H 'Authorization: Bearer YOUR_API_KEY' \
-H 'Content-Type: text/plain' \
-H 'Content-Encoding: gzip' \
--data-binary @-
```
### Error shapes
| HTTP | Meaning |
| ----- | --------------------------------------------------- |
| `400` | Malformed body (empty, non-UTF-8, bad CRNs). |
| `429` | `TooManyConcurrentSets` or `TooManyCompanyNumbers`. |
See [Errors](/api-reference/guides/errors) for the wider envelope and [Rate limits](/api-reference/guides/rate-limits) for the throttle behaviour.
## Picking the right path
| Workload | Best path |
| ---------------------------------- | ------------------------------------------------------ |
| \< 1000 companies, one-off | Inline `CompanyNumbers` array |
| 1000–10k companies, repeated calls | Inline `CompanyNumbers` array |
| 10k–50k companies | Either inline or set (set is cheaper if you re-use it) |
| > 50k companies | Upload a set, reference by `setId` |
# Changelog
Source: https://docs.thedatacity.com/api-reference/guides/changelog
Material changes to The Data City API, newest first.
This page records material changes to the v2 API — new endpoints, new fields, breaking changes, and behaviour shifts worth knowing about.
The Data City does not maintain a separate API changelog today. This page is the new canonical home; entries will accrue here as the v2 API evolves. For a precise record of past changes between two points in time, diff the OpenAPI spec at [`/api/documentation/spec/v2`](https://product-api.thedatacity.com/api/documentation/spec/v2).
## Conventions
Entries are tagged so you can scan for the kind of change that affects you:
* **Added** — new endpoint, field, or capability. Safe to ignore if you don't use it.
* **Changed** — non-breaking change to existing behaviour (new optional field, tightened error message, performance improvement).
* **Deprecated** — still works but scheduled for removal. Start migrating.
* **Removed** — no longer available. Will only appear in a major version bump.
* **Fixed** — bug fix bringing actual behaviour in line with documented behaviour.
## History
*No entries yet. Check back after the next deploy, or subscribe to release notifications by emailing [support@thedatacity.com](mailto:support@thedatacity.com).*
# Errors
Source: https://docs.thedatacity.com/api-reference/guides/errors
All error responses share the same JSON envelope. HTTP status codes follow standard conventions.
The Data City API uses standard HTTP status codes. Every error response shares a single JSON envelope.
## Error envelope
```json theme={null}
{
"message": "Human-readable description of what went wrong."
}
```
The `message` field is intended for log lines and developer diagnostics. Don't surface it verbatim to end users — phrasing may change between releases.
## Status code matrix
| Status | When to expect it | What to do |
| ------ | --------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- |
| `200` | Successful request. | Process the response body. |
| `400` | Malformed input (bad JSON, invalid query parameter, validation failed). | Inspect `message` and fix the request. |
| `401` | Missing or invalid Bearer token. | Re-check the `Authorization` header. See [Authentication](/api-reference/guides/authentication). |
| `403` | Token is valid but lacks the required ability for the endpoint. | Email support to widen the token's scope. |
| `404` | The resource (company, list, set) does not exist. | Confirm the identifier; some IDs expire after 24 hours — see [Batching](/api-reference/guides/batching). |
| `422` | Request well-formed but semantically invalid (for example, mutually exclusive filter fields). | Inspect `message`, adjust the body. |
| `429` | Too many requests. | Back off and retry. See [Rate limits](/api-reference/guides/rate-limits). |
| `500` | Unexpected server error. | Retry with exponential backoff; if persistent, contact [support@thedatacity.com](mailto:support@thedatacity.com). |
| `503` | Service unavailable, usually during a deploy. | Retry after a short delay. |
## Retrying
Retry `429`, `500`, `502`, `503`, and `504` with exponential backoff. Do **not** auto-retry `400`, `401`, `403`, `404`, or `422` — the request will fail again the same way.
# Filtering
Source: https://docs.thedatacity.com/api-reference/guides/filtering
Build a list of companies that match your criteria — RTICs, CICs, or your own portfolio.
The `POST /v2/companies` endpoint takes a JSON body describing which companies you want and returns the matching list. It's the workhorse of most integrations.
## Minimal request
```bash theme={null}
curl -X POST 'https://product-api.thedatacity.com/api/v2/companies' \
-H 'Authorization: Bearer YOUR_API_KEY' \
-H 'Content-Type: application/json' \
-d '{
"returnCount": 20,
"rtics": ["RTIC-12345"]
}'
```
This returns the first 20 companies classified under RTIC `RTIC-12345`.
## Core filter fields
| Field | Type | Purpose |
| -------------------- | ---------- | ------------------------------------------------------------------------------------------------------------------------------- |
| `rtics` | `string[]` | Selected RTIC codes. Companies matching **any** of the listed RTICs are returned. |
| `cics` | `string[]` | User-defined Custom Industrial Classification codes. |
| `CompanyNumbers` | `string[]` | Inline UK Companies House numbers. Use for portfolios up to roughly 50k entries. |
| `CompanyNumberSetId` | `string` | An uploaded set ID — use instead of `CompanyNumbers` for very large portfolios. See [Batching](/api-reference/guides/batching). |
| `AllowLargeAnalysis` | `boolean` | Required when filtering against a large uploaded set. |
The full filter shape is documented on the [`POST /v2/companies` reference page](/api-reference) — these are the most common starting points.
## Pairing with `insights`
Add `?insights=true` to get dashboard-style aggregations alongside the company rows:
```bash theme={null}
curl -X POST 'https://product-api.thedatacity.com/api/v2/companies?insights=true' \
-H 'Authorization: Bearer YOUR_API_KEY' \
-H 'Content-Type: application/json' \
-d '{"returnCount": 20, "rtics": ["RTIC-12345"]}'
```
Aggregations are computed over the **full matched set**, not just the page. That means you can keep `returnCount` small on the first call to grab the insights cheaply.
## Patterns
### Filter a portfolio you already have
```json theme={null}
{
"returnCount": 1000,
"CompanyNumbers": ["12345678", "87654321", "11223344"],
"rtics": ["RTIC-12345"]
}
```
Restricts results to the intersection of your portfolio and the RTIC.
### Use a large uploaded set
```json theme={null}
{
"returnCount": 1000,
"CompanyNumberSetId": "set_abc123",
"AllowLargeAnalysis": true
}
```
Upload the set first (see [Batching](/api-reference/guides/batching)), then filter against the `setId`. Don't combine `CompanyNumberSetId` with `CompanyNumbers` — the set wins.
## Available filters
`GET /v2/filters` returns the catalogue of filter codes (RTICs, sectors, etc.) currently available. Call it once at startup, cache the result, and use the codes in your `POST /v2/companies` body.
## Pagination
Filter responses paginate with `returnCount` and `skip`. See [Pagination](/api-reference/guides/pagination).
# Pagination
Source: https://docs.thedatacity.com/api-reference/guides/pagination
Page through list responses with returnCount and skip. Default page size is 40; maximum is 1000.
List endpoints page their results with two parameters: `returnCount` (page size) and `skip` (offset). Where you put them depends on the endpoint:
* **GET list endpoints** (for example `GET /v2/companies/by-website`, `GET /v2/explore/{listId}`) take both as **query parameters**.
* **POST filter endpoints** (for example `POST /v2/companies`) take both as **fields in the JSON request body**, alongside your filter criteria.
When in doubt, check the [API reference](/api-reference) — each operation page lists exactly where its parameters go.
## Parameters
| Parameter | Type | Default | Range | Purpose |
| ------------- | ------- | ------- | ------ | ------------------------------------------- |
| `returnCount` | integer | `40` | 1–1000 | Maximum number of records to return. |
| `skip` | integer | `0` | ≥ 0 | Number of records to skip before returning. |
## Walking a list
### GET endpoints — query parameters
```bash theme={null}
# Page 1 — first 100 records
curl 'https://product-api.thedatacity.com/api/v2/companies/by-website?url=acme.com&returnCount=100' \
-H 'Authorization: Bearer YOUR_API_KEY'
# Page 2 — next 100
curl 'https://product-api.thedatacity.com/api/v2/companies/by-website?url=acme.com&returnCount=100&skip=100' \
-H 'Authorization: Bearer YOUR_API_KEY'
```
### POST endpoints — body fields
```bash theme={null}
# Page 1
curl -X POST 'https://product-api.thedatacity.com/api/v2/companies' \
-H 'Authorization: Bearer YOUR_API_KEY' \
-H 'Content-Type: application/json' \
-d '{"rtics": ["RTIC-12345"], "returnCount": 100}'
# Page 2
curl -X POST 'https://product-api.thedatacity.com/api/v2/companies' \
-H 'Authorization: Bearer YOUR_API_KEY' \
-H 'Content-Type: application/json' \
-d '{"rtics": ["RTIC-12345"], "returnCount": 100, "skip": 100}'
```
Stop when a page returns fewer than `returnCount` records.
## Picking a page size
* **Small portfolios (\< 1000 companies).** Set `returnCount` to the total — one request, no pagination.
* **Large filtered lists.** Set `returnCount=1000` and walk with `skip`. This is the most efficient mode.
* **Insights enabled.** When `?insights=true`, aggregations cover the *full matched set*, not just the current page. Keep `returnCount` low (40–100) on the first call to get the aggregations cheaply; raise it on subsequent paging requests if you also want the rows.
## Limits
`returnCount > 1000` is rejected with `400`. There's no hard ceiling on `skip`, but very deep pagination (`skip > 100,000`) is slow — use [filters](/api-reference/guides/filtering) or [company-number sets](/api-reference/guides/batching) to narrow the set instead.
## Per-endpoint variations
A few endpoints document their own page-size parameters with different defaults (for example `/v2/companies/by-website` has its own `returnCount`). The [API reference](/api-reference) is authoritative — check each operation's parameters before assuming the default.
# Quickstart
Source: https://docs.thedatacity.com/api-reference/guides/quickstart
Make your first request to The Data City API in under five minutes.
This guide walks you from zero to a working API call against The Data City v2 API.
## 1. Get an API key
API keys are issued manually by The Data City. Email [support@thedatacity.com](mailto:support@thedatacity.com) to request one. See [Authentication](/api-reference/guides/authentication) for the full flow.
## 2. Make your first request
Every request needs an `Authorization: Bearer ` header. Try fetching details for a UK company by its Companies House number:
```bash theme={null}
curl -X GET 'https://product-api.thedatacity.com/api/v2/companies/12345678' \
-H 'Authorization: Bearer YOUR_API_KEY'
```
A successful response is a JSON object with the company's classification, financials, and metadata. If you get `{"message":"Unauthenticated."}`, the token is missing or wrong.
## 3. Explore from there
Every endpoint, with a live playground.
How tokens work and what to do if yours stops.
Build a list of companies that match RTICs, CICs, or your own criteria.
Bulk endpoints and company-number sets for large lists.
# Rate limits
Source: https://docs.thedatacity.com/api-reference/guides/rate-limits
Customer endpoints are limited to 200 requests per minute. Hitting the limit returns 429.
The Data City API throttles per-minute request volume to protect the service for every customer.
## Limits
| Endpoint group | Limit |
| ----------------------------- | --------------------------- |
| Customer v2 API (`/v2/...`) | **200 requests per minute** |
| Internal admin (`/admin/...`) | 60 requests per minute |
The limit is applied per authenticated token. Each request counts as one, regardless of the response status or payload size.
## Hitting the limit
When you exceed the per-minute quota, the API responds with:
```http theme={null}
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
{"message":"Too Many Attempts."}
```
See [Errors](/api-reference/guides/errors) for the full envelope and other status codes.
## Backing off
Retry `429` responses with exponential backoff. A pragmatic starting point:
```text theme={null}
attempt 1 → wait 2 s
attempt 2 → wait 4 s
attempt 3 → wait 8 s
attempt 4 → give up, surface the error
```
Add a small random jitter (0–500 ms) so concurrent clients don't synchronise their retries.
## Designing around it
If you have a workload that comes near 200/min:
* **Batch where the API supports it.** Endpoints like `/v2/companies/batch`, `/v2/companies/by-website/batch`, and `/v2/companies/check` take many inputs in a single request — one HTTP call instead of many.
* **Use [company-number sets](/api-reference/guides/batching) for large portfolios.** Filtering against a 200k-company set is one request, not 200,000.
* **Cache stable data.** Company details don't change second-to-second; cache results for as long as your domain tolerates staleness.
If your use case genuinely needs a higher steady-state rate, email [support@thedatacity.com](mailto:support@thedatacity.com).
# Versioning
Source: https://docs.thedatacity.com/api-reference/guides/versioning
All current endpoints live under /v2. Older v1 routes still exist for historical clients but are not recommended for new work.
The Data City versions its API by URL prefix.
## Current version: v2
Every endpoint documented in the [API reference](/api-reference) lives under `/v2/...`. This is the version to build against — `info.description` in the spec reads:
> API v2. Use these endpoints for new integrations.
v2 is stable. Backward-compatible changes (new optional fields on responses, new endpoints, new optional query parameters) ship without a version bump. Breaking changes would ship as `/v3/...` alongside `/v2/...`.
## What counts as a breaking change
Breaking, requires a version bump:
* Removing or renaming an endpoint, query parameter, or response field.
* Changing the type of an existing response field.
* Tightening validation on a parameter that previously accepted broader input.
* Changing the JSON error envelope.
Non-breaking, ships without notice:
* New endpoints.
* New optional response fields.
* New optional query parameters or body fields.
* Bug fixes that align actual behaviour with documented behaviour.
## Legacy endpoints
A small set of older routes is still served for historical integrations and is exposed via a separate spec at `/api/documentation/spec/legacy`. These are **not** documented here, are not maintained, and may be retired without notice. Migrate to v2 if you're still using them.
## Discovering changes
* Watch the [Changelog](/api-reference/guides/changelog) for material changes.
* The full OpenAPI spec is at [`/api/documentation/spec/v2`](https://product-api.thedatacity.com/api/documentation/spec/v2) — diff it between deploys if you want a precise record.
* For breaking-change notice timelines, see the changelog or email [support@thedatacity.com](mailto:support@thedatacity.com).
# Industry Engine API
Source: https://docs.thedatacity.com/api-reference/index
REST API for The Data City Industry Engine. Every v2 endpoint, with a live playground.
Access classifications, financials, growth signals, and group structure for UK companies. The REST API exposes the same data that powers Industry Engine — same coverage, same definitions, fresh on the same cadence.
Authenticate with a Bearer token, send JSON, get JSON back.
Make your first request in under five minutes.
How Bearer tokens work and where to get one.
The error envelope and status-code matrix.
200 requests per minute, with backoff guidance.
## Working with lists
`returnCount` and `skip` — defaults, maximums, patterns.
Build a list with RTICs, CICs, or your own portfolio.
Bulk endpoints and company-number sets for large lists.
## Reference
Every operation is documented under the **Endpoints** group in the sidebar. Each page includes a live playground — paste your token, fill the parameters, send the request.
Base URL:
```text theme={null}
https://product-api.thedatacity.com/api
```
All current endpoints are under `/v2/...`. See [Versioning](/api-reference/guides/versioning) for the policy and [Changelog](/api-reference/guides/changelog) for what's changed.
## For LLMs and agents
The full v2 surface is also available as a single curated text file optimised for LLM context. Paste it into ChatGPT/Claude or feed it to an agent for end-to-end API context in one request.
Endpoint decision tree, request/response shapes, idiomatic patterns. Regenerated on every API deploy.
# RTIC details
Source: https://docs.thedatacity.com/api-reference/rtics/rtic-details
/api-reference/openapi.json get /v2/rtics
# Full-text search
Source: https://docs.thedatacity.com/api-reference/search/full-text-search
/api-reference/openapi.json get /v2/search/fullText/{keyword}
# Search by company number
Source: https://docs.thedatacity.com/api-reference/search/search-by-company-number
/api-reference/openapi.json get /v2/search/aheadFull/{companyNumber}
# Get list details
Source: https://docs.thedatacity.com/api-reference/smart-lists/get-list-details
/api-reference/openapi.json get /v2/smart-lists/{listId}
Get list details.
# Get lists by domain
Source: https://docs.thedatacity.com/api-reference/smart-lists/get-lists-by-domain
/api-reference/openapi.json get /v2/smart-lists/by-domain/{domain}
Get lists by domain.
# Get lists by email
Source: https://docs.thedatacity.com/api-reference/smart-lists/get-lists-by-email
/api-reference/openapi.json get /v2/smart-lists/by-email/{email}
Get lists by email.
# BasicGrant360
Source: https://docs.thedatacity.com/data-dictionary/basic-grant360
11 fields on the BasicGrant360 schema — covers Funding.
**11 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Funding
| Field name | Description | Data type | Source | Updated |
| ----------------------------- | ------------------------------------------------------------------------ | --------- | --------- | ------- |
| `amountAwarded` | The amount awarded. · Download name: `AmountAwarded` | `number` | 360Giving | Monthly |
| `awardDate` | The date the grant was awarded. · Download name: `AwardDate` | `string` | 360Giving | Monthly |
| `currency` | The currency of the grant. · Download name: `Currency` | `string` | 360Giving | Monthly |
| `description` | A description of the grant. · Download name: `Description` | `string` | 360Giving | Monthly |
| `funderName` | The name of the funder. · Download name: `FunderName` | `string` | 360Giving | Monthly |
| `grantProgrammeTitle` | The title of the grant programme. · Download name: `GrantProgrammeTitle` | `string` | 360Giving | Monthly |
| `ImpactCategory` | The impact category of the grant. | `string` | 360Giving | Monthly |
| `Primaryissue` | The primary issue the grant is for. | `string` | 360Giving | Monthly |
| `recipientName` | The name of the recipient. · Download name: `RecipientName` | `string` | 360Giving | Monthly |
| `recipientURL` | The URL of the recipient. · Download name: `RecipientURL` | `string` | 360Giving | Monthly |
| `title` | The title of the grant. · Download name: `Title` | `string` | 360Giving | Monthly |
# ClassifiedCompany
Source: https://docs.thedatacity.com/data-dictionary/classified-company
85 fields on the ClassifiedCompany schema — covers Funding, Financials, Employees, Social Media, Classification, Company, Website, Trade, Growth, Job Postings, Group Structure, Investment, Contact, ESG, Gender, Acquisitions, Keywords, Inno…
**85 fields** · across **24 categories** · **24** sourced from [CreditSafe](/our-data/third-party-data/creditsafe).
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Classification
| Field name | Description | Data type | Source | Updated |
| ------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------- | --------------- | ------- |
| `CICs` | Custom Industrial Classifications (CICs) the company has been classified into. API includes the industry verticals within the same field. On the download, verticals are a separate column. These are custom RTICs. · [Related: The Data City guide](https://docs.thedatacity.com/using-industry-engine/features/using-filters#cics) | `array[companynumberrtic]` | The Data City | Ad hoc |
| `ConsolidatedAccountsCompanyName` | A list of company names within the same group structure that have consolidated accounts. This is a strong signal of the companies that are more likely to be the parent company. This is because it is common for the parent company to file consolidated accounts that include all the subsidiaries in the group structure. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `array[string]` | The Data City | Monthly |
| `ConsolidatedAccountsCompanyNumber` | A list of company numbers within the same group structure that have consolidated accounts. This is a strong signal of the companies that are more likely to be the parent company. This is because it is common for the parent company to file consolidated accounts that include all the subsidiaries in the group structure. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `array[string]` | The Data City | Monthly |
| `DistinctBrandsInGroup` | A list of distinct brands that can be found across the company's group structure. | `array[distinctbrandinfo]` | The Data City | Monthly |
| `ESG_RTIC` | A boolean value indicating whether the company is in an ESG RTIC. · Download name: `ESGRTIC` | `boolean` | The Data City | Monthly |
| `IS8Codes` | Industrial Strategy Sector codes associated with the company, derived from RTICs and RSICs. · Nested object — see [IndustrialStrategySector](/data-dictionary/industrial-strategy-sector) | [`array[industrialstrategysector]`](/data-dictionary/industrial-strategy-sector) | The Data City | Ad hoc |
| `IsLikelyAgency` | A boolean value indicating whether the company is likely to be an agency. This is based on assessment of the company's financials. | `boolean` | The Data City | Monthly |
| `LinkedCompanyNumbers` | A list of company numbers that are linked to the company, even if they are not connected through group structure. This is based on a combination of signals including shared officers, shared addresses, shared shareholders, and other signals. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `array[string]` | The Data City | Monthly |
| `RSICs` | Real-Time SIC (RSIC) codes with associated confidence score · [Related: The Data City guide](https://docs.thedatacity.com/our-data/proprietary-data/what-are-rsics) · Nested object — see [RSIC](/data-dictionary/rsic) | [`array[rsic]`](/data-dictionary/rsic) | The Data City | Monthly |
| `RTICs` | RTICs the company has been classified into. API includes the industry verticals within the same field. On the download, verticals are a separate column. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/rtics/what-is-the-difference-between-rtics-sic-codes) | `array[companynumberrtic]` | The Data City | Ad hoc |
| `SICHLUs` | Companies House data. SIC Sections. Nature of business: Standard Industrial Classification (SIC) codes. · [Related: Companies House guide](https://docs.thedatacity.com/our-data/rtics/what-is-the-difference-between-rtics-sic-codes) | `array[string]` | Companies House | Monthly |
| `SICs` | Companies House data. Nature of business: Standard Industrial Classification (SIC) codes. · [Related: Companies House guide](https://docs.thedatacity.com/our-data/rtics/what-is-the-difference-between-rtics-sic-codes) | `array[string]` | Companies House | Monthly |
| `UltimateLinkedCompanyNumber` | The ultimate linked company number in the group structure. This is the company number that was first incorporated within the collection of linked companies. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | The Data City | Monthly |
| `UltimateUKParentCompanyNames` | A list of company names belonging to companies in the group that are the highest level UK registered companies. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `array[string]` | The Data City | Monthly |
| `UltimateUKParentCompanyNumbers` | A list of company numbers belonging to companies in the group that are the highest level UK registered companies. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `array[string]` | The Data City | Monthly |
## Company
| Field name | Description | Data type | Source | Updated |
| ---------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | --------------- | ------- |
| `CompanyCategory` | Companies House data. The type of company, e.g. Publically Listed (PLC) or Limited (LTD). · Download name: `Companycategory` | `string` | Companies House | Monthly |
| `CompanyName` | Companies House data. Companies have the ability to change their company name through submission to Companies House. · Download name: `Companyname` | `string` | Companies House | Monthly |
| `CompanyNumber` | Companies House data. Unique key, doesn't change. Most important part of our data. · [Related: Companies House guide](https://docs.thedatacity.com/our-data/third-party-data/companies-house) · Download name: `Companynumber` | `string` | Companies House | Monthly |
| `CompanyStatus` | The company status as per Companies House. | `string` | The Data City | Monthly |
| `CountryOfOrigin` | Companies House data. Not to be confused with group structure (ParentCompanyNation, UltimateParentCompanyNation). · Download name: `Countryoforigin` | `string` | Companies House | Monthly |
| `IncorporationDate` | Companies House data. · Download name: `IncorporationDate/SortableIncorporationDate` | `string` | Companies House | Monthly |
| `RegisteredAddress` | Companies House data. The companies registered address as registered on Companies House. This is not to be confused with their head office. Not sourced from CreditSafe. · [Related: Companies House guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/company-locations) · Download name: `Registeredaddress` | `string` | Companies House | Monthly |
| `RegisteredPostcode` | Companies House data. Not to be confused with head office (we do not know this at this stage). Not sourced from CreditSafe. · [Related: Companies House guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/company-locations) · Download name: `Registeredpostcode` | `string` | Companies House | Monthly |
| `Spinout` | True if the company is a spinout, sourced from the HESA spinout register. A spinout is a company created to commercialise research or technology developed at a university or research institution. | `boolean` | HESA | Ad hoc |
## Investment
| Field name | Description | Data type | Source | Updated |
| ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------- | ---------------------- | ------- |
| `DealroomFunding` | Legacy field, retained for backward compatibility only; use Investment instead. · [Related: Investment guide](https://docs.thedatacity.com/our-data/third-party-data/dealroom-data) · Nested object — see [PerRoundInvestment](/data-dictionary/per-round-investment) | [`array[perroundinvestment]`](/data-dictionary/per-round-investment) | Investment | N/A |
| `HasIPO` | True if the company has listed on a public market, based on an IPO event in the investment feed. · [Related: Third-party Investment guide](https://docs.thedatacity.com/our-data/third-party-data/dealroom-data) | `boolean` | Third-party Investment | Daily |
| `Investment` | Per-round investment data matched to a CompanyNumber with high confidence. TotalInvestmentRaised\_GBP\_Million is the sum of all rounds, *except* Debt and Acquisition rounds. This is the total across the provider's tracking period and is not the total in one particular year. · [Related: Investment guide](https://docs.thedatacity.com/our-data/third-party-data/dealroom-data) · Nested object — see [PerRoundInvestment](/data-dictionary/per-round-investment) | [`array[perroundinvestment]`](/data-dictionary/per-round-investment) | Investment | Daily |
| `LatestRoundAmount` | The amount raised in the company's most recent funding round. This is the individual round amount, not the company's total funding raised across all rounds. · [Related: Third-party Investment guide](https://docs.thedatacity.com/our-data/third-party-data/dealroom-data) | `number` | Third-party Investment | Daily |
| `LatestRoundDate` | The date of the company's most recent funding round. · [Related: Third-party Investment guide](https://docs.thedatacity.com/our-data/third-party-data/dealroom-data) | `string` | Third-party Investment | Daily |
| `LatestRoundType` | The round type of the company's most recent funding round (e.g. "Seed", "Series A"). · [Related: Third-party Investment guide](https://docs.thedatacity.com/our-data/third-party-data/dealroom-data) | `string` | Third-party Investment | Daily |
| `LatestRoundValuation` | The company's valuation as reported for its most recent funding round, in GBP. Not all rounds have a disclosed or estimated valuation. · [Related: Third-party Investment guide](https://docs.thedatacity.com/our-data/third-party-data/dealroom-data) | `number` | Third-party Investment | Daily |
## Financials
| Field name | Description | Data type | Source | Updated |
| ---------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------- | --------------- | ----------------- |
| `AccountsOverdueByMoreThanAYear` | A flag indicating the company's accounts are overdue by more than a year. | `boolean` | Companies House | Monthly |
| `BestEstimateEBITDA` | Where OPERATING\_PROFIT, DEPRECIATION\_OF\_TANGIBLES and AMORTISATION\_OF\_INTANGIBLES are available, these three values are summed. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `EBITDA` | `number` | The Data City | Monthly |
| `BestEstimateTurnover` | Available from Companies House for large companies, if absent then The Data City provide an estimate. The BestEstimate will report actual (reported) where available or it will be an estimate. BestEstimateTurnover refers to the current year: where accounts covering the current year have not been reported, the value is projected forward from the most recently reported year. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) · Download name: `BestEstimateCurrentTurnover` | `integer` | The Data City | Monthly |
| `CompanyFinancials_CreditSafe` | Annual financial reporting from filings processed by Creditsafe. Creditsafe use a semi-automated process to convert scanned financial PDF documents into numerical/digital values. This process may lead to errors, e.g. reporting profit as turnover. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: CreditSafe guide](https://docs.thedatacity.com/our-data/third-party-data/creditsafe) · Download name: `CompanyFinancialsCreditSafe` · Nested object — see [CleanCreditSafeCompany](/data-dictionary/clean-credit-safe-company) | [`array[cleancreditsafecompany]`](/data-dictionary/clean-credit-safe-company) | CreditSafe | Monthly |
| `EstimatedGVA` | We have produced an estimated GVA measure at the company level. We have employee and SIC information for each company. This makes it possible to estimate a GVA value per company by multiplying a company's number of employees by the standard GVA employee contribution associated with that same company's SIC. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: ONS; BRES guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/gva-data) · Download name: `BestEstimate_CurrentGVA` | `number` | ONS; BRES | Monthly/Quarterly |
| `EstimatedUKGVA` | We have produced an estimated GVA measure at the company level. We have employee and SIC information for each company. This makes it possible to estimate a GVA value per company by multiplying a company's number of UK employees by the standard GVA employee contribution associated with that same company's SIC. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: ONS; BRES guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/gva-data) · Download name: `BestEstimate_CurrentUKGVA` | `number` | ONS; BRES | Monthly/Quarterly |
## Job Postings
| Field name | Description | Data type | Source | Updated |
| ---------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- | --------- | ------- |
| `CompanyJobPostings` | This company's job postings grouped by SOC4 occupation, one entry per occupation group, covering the most recent five years. See the CompanyJobPosting schema for the fields on each entry. Job postings data is a paid upgrade and is not included for all customers. · Nested object — see [CompanyJobPosting](/data-dictionary/company-job-posting) | [`array[companyjobposting]`](/data-dictionary/company-job-posting) | Lightcast | Monthly |
| `CompanyJobPostingsByLot` | This company's job postings grouped by Lightcast Occupation Taxonomy (LOT) occupation, one entry per occupation, covering the most recent five years. See the CompanyJobPostingByLot schema for the fields on each entry. Job postings data is a paid upgrade and is not included for all customers. · Nested object — see [CompanyJobPostingByLot](/data-dictionary/company-job-posting-by-lot) | [`array[companyjobpostingbylot]`](/data-dictionary/company-job-posting-by-lot) | Lightcast | Monthly |
| `CompanyJobPostingsByLOTCodeOverTime` | This company's monthly job posting counts and median advertised salary per LOT occupation, one entry per occupation per month, covering the most recent five years. Job postings data is a paid upgrade and is not included for all customers. · Nested object — see [CompanyJobPostingByLOTCodeOverTime](/data-dictionary/company-job-posting-by-lot-code-over-time) | [`array[companyjobpostingbylotcodeovertime]`](/data-dictionary/company-job-posting-by-lot-code-over-time) | Lightcast | Monthly |
| `CompanyJobPostingsBySkill` | This company's job postings grouped by common skill requested, covering the most recent five years. Despite the name this holds common skills only; specialised skills, software skills and certifications are returned separately and are not included here. See the CompanyJobPostingByCommonSkill schema for the fields on each entry. Job postings data is a paid upgrade and is not included for all customers. · Nested object — see [CompanyJobPostingByCommonSkill](/data-dictionary/company-job-posting-by-common-skill) | [`array[companyjobpostingbycommonskill]`](/data-dictionary/company-job-posting-by-common-skill) | Lightcast | Monthly |
| `CompanyJobPostingsBySOC4CodeOverTime` | This company's monthly job posting counts and median advertised salary per SOC4 occupation, one entry per occupation per month, covering the most recent five years. Job postings data is a paid upgrade and is not included for all customers. · Nested object — see [CompanyJobPostingBySOC4CodeOverTime](/data-dictionary/company-job-posting-by-soc4-code-over-time) | [`array[companyjobpostingbysoc4codeovertime]`](/data-dictionary/company-job-posting-by-soc4-code-over-time) | Lightcast | Monthly |
| `CompanyJobPostingsOverTime` | This company's monthly job posting counts as two time series, one totalled by SOC4 occupation and one by LOT occupation, covering the most recent five years. Use this for overall hiring trend over time; use CompanyJobPostingsBySOC4CodeOverTime or CompanyJobPostingsByLOTCodeOverTime for a trend per individual occupation. Job postings data is a paid upgrade and is not included for all customers. | `companyjobpostingsovertime` | Lightcast | Monthly |
## Social Media
| Field name | Description | Data type | Source | Updated |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | --------- | ------------ | ----------------- |
| `Bluesky` | Takes the most common anchor link which contains the social media domain. We prioritise social media links that are found on the homepage. | `string` | Web Scraping | Monthly/Quarterly |
| `Facebook` | Takes the most common anchor link which contains the social media domain. We prioritise social media links that are found on the homepage. | `string` | Web Scraping | Monthly/Quarterly |
| `Instagram` | Takes the most common anchor link which contains the social media domain. We prioritise social media links that are found on the homepage. | `string` | Web Scraping | Monthly/Quarterly |
| `Linkedin` | Takes the most common anchor link which contains the social media domain. We prioritise social media links that are found on the homepage. | `string` | Web Scraping | Monthly/Quarterly |
| `Twitter` | Takes the most common anchor link which contains the social media domain. We prioritise social media links that are found on the homepage. | `string` | Web Scraping | Monthly/Quarterly |
| `Youtube` | Takes the most common anchor link which contains the social media domain. We prioritise social media links that are found on the homepage. | `string` | Web Scraping | Monthly/Quarterly |
## Gender
| Field name | Description | Data type | Source | Updated |
| ------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------- | ------------- | ------- |
| `GenderFoundedCategory` | The company's single gender-leadership category based on its active founders. One of: All women led, Majority women led, Mixed led, Majority men led, All men led, Uncertain. Mutually exclusive — every company falls in exactly one. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/proprietary-data/gender-data) · Download name: `gender_founded_category` | `string` | The Data City | Monthly |
| `GenderLedCategory` | The company's single gender-leadership category based on its active directors. One of: All women led, Majority women led, Mixed led, Majority men led, All men led, Uncertain. Mutually exclusive — every company falls in exactly one. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/proprietary-data/gender-data) · Download name: `gender_led_category` | `string` | The Data City | Monthly |
| `WomenLedStats` | For each company we have identified the likely founders. We have used their declared title to assign them a gender. We distinguish between founders that have founded a company, but are no longer active at the company and founders that are still active at the company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/proprietary-data/gender-data) | `womanledstats` | The Data City | Monthly |
## Group Structure
| Field name | Description | Data type | Source | Updated |
| ------------------------------ | -------------------------------------------------------------------------------------------------------- | --------- | --------------- | ------- |
| `CompanyNation` | The company nation. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `ParentNation` | The nation of the parent company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `UltimateParentNation` | The nation of the ultimate parent company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
## Growth
| Field name | Description | Data type | Source | Updated |
| ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------- | ------------- | ------- |
| `CompanyGrowthTrends` | Based on the company growth trends data. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) | `companygrowthtrend` | The Data City | Monthly |
| `HighGrowth` | A boolean value indicating whether the company is a high growth company. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/scale-ups) | `boolean` | The Data City | Monthly |
| `OECDScaleup` | A boolean value indicating whether the company is an OECD defined scale-up. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/scale-ups) | `boolean` | The Data City | Monthly |
## Keywords
| Field name | Description | Data type | Source | Updated |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------- | ------------ | --------- |
| `InnovationKeywords` | We have a large list of emerging economy keywords. We loop through the web text to see if the web text contains the keyword. If they do, the keyword is assigned to the company. · [Related: Web Scraping guide](https://docs.thedatacity.com/using-industry-engine/tools/using-analyse#website-analysis) | `array[string]` | Web Scraping | Monthly |
| `ManufacturingKeywords` | We have a large list of manufacturing keywords. We loop through the web text to see if the web text contains the keyword. If they do, the keyword is assigned to the company. | `array[string]` | Web Scraping | Monthly |
| `SectorKeywords` | We have a large list of emerging economy keywords. We loop through the web text to see if the web text contains the keyword. If they do, the keyword is assigned to the company. · [Related: Web Scraping guide](https://docs.thedatacity.com/using-industry-engine/tools/using-analyse#website-analysis) · Download name: `Sectorkeywords` | `array[string]` | Web Scraping | Quarterly |
## Officers
| Field name | Description | Data type | Source | Updated |
| ---------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------- | --------------- | ------- |
| `NumberOfCurrentDirectors` | The number of directors currently appointed to the company, i.e. those with no resignation date recorded whose role is Director. Excludes other officer roles such as company secretaries, LLP members and managing officers, so a limited liability partnership will typically report zero current directors. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | Companies House | Monthly |
| `NumberOfCurrentOfficers` | The number of officers currently appointed to the company, i.e. those with no resignation date recorded. Counts all officer roles, including company secretaries and LLP members, not only directors. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | Companies House | Monthly |
| `Officers` | These are officers as per Companies House. E.g. [https://find-and-update.company-information.service.gov.uk/company/10958787/officers](https://find-and-update.company-information.service.gov.uk/company/10958787/officers) · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Nested object — see [Officer](/data-dictionary/officer) | [`array[officer]`](/data-dictionary/officer) | Companies House | Monthly |
## Other
| Field name | Description | Data type | Source | Updated |
| ------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------- | ----------------------- | ----------------- |
| `HasBeenAcquired` | True if the company has been acquired, based on funding rounds with an "Acquisition" round type where this company is the target. · [Related: Third-party Investment guide](https://docs.thedatacity.com/our-data/third-party-data/dealroom-data) | `boolean` | Third-party Investment | Daily |
| `HasMadeAcquisitions` | True if the company has acquired another company, based on funding rounds with an "Acquisition" round type where this company is the acquirer. · [Related: Third-party Investment guide](https://docs.thedatacity.com/our-data/third-party-data/dealroom-data) | `boolean` | Third-party Investment | Daily |
| `EmailAddresses` | Emails are extracted from the website text. Director emails are prioritised. For matching directors to emails: any combination of the complete surname or first name with at least the first initial of the opposing field. No AI is used. · Download name: `Email` | `array[string]` | Web Scraping | Monthly/Quarterly |
| `PhoneNumbers` | CreditSafe provide The Data City with CTPS approved phone numbers. Phone numbers are extracted from the website text. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: CreditSafe guide](https://corporate.ctpsonline.org.uk/) · Download name: `Telephone` | `array[string]` | CreditSafe | Monthly/Quarterly |
| `BestEstimateEmployees` | Employees data is available for most companies. Occasionally there are gaps in reporting. Where there are gaps in reporting, we estimate the number of employees; this includes a projection of employees forwards and backwards. BestEstimateEmployees refers to the current year. BestEstimateEmployees may include overseas employees. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) · Download name: `BestEstimateCurrentEmployees` | `integer` | The Data City | Monthly |
| `BestEstimateUKEmployees` | Similar to BestEstimateEmployees, but estimates only UK-based employees. BestEstimateUKEmployees currently refers to the latest complete year. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) | `integer` | The Data City | Monthly |
| `_360GivingFunding` | Data provided by a charity that helps organisations to publish open, standardised grants data, and supports people to use it to improve charitable giving. Source: [https://grantnav.threesixtygiving.org/](https://grantnav.threesixtygiving.org/) · [Related: 360Giving guide](https://docs.thedatacity.com/our-data/third-party-data/360giving-data) · Download name: `360GivingFunding` · Nested object — see [BasicGrant360](/data-dictionary/basic-grant360) | [`array[basicgrant360]`](/data-dictionary/basic-grant360) | 360Giving | Monthly |
| `InnovateUKFunding` | Information about the projects funded by Innovate UK from 2004. Source: [https://www.ukri.org/publications/innovate-uk-funded-projects-since-2004/](https://www.ukri.org/publications/innovate-uk-funded-projects-since-2004/) · [Related: InnovateUK guide](https://docs.thedatacity.com/our-data/third-party-data/innovate-uk-data) · Nested object — see [InnovateUKFundingClean](/data-dictionary/innovate-uk-funding-clean) | [`array[innovateukfundingclean]`](/data-dictionary/innovate-uk-funding-clean) | InnovateUK | Monthly |
| `IsBCorp` | A flag of whether a company is a B Corp. This is based on a list of B Corps that we have obtained from B Corp and matched to our companies. · [Related: BCorp and The Data City guide](https://docs.thedatacity.com/our-data/third-party-data/b-corps) | `boolean` | BCorp and The Data City | Quarterly |
| `IsLikelySPV` | A flag indicating the company is likely a special purpose vehicle (SPV), based on an ML model applied to Companies House data. | `boolean` | The Data City | Quarterly |
| `SimilarCompanies` | Our similarity score adopts measures of mathematical similarity to find similar companies using their website text. Companies with website text with the same words are more likely to report a higher level of similarity. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/faqs/similar-companies) · Nested object — see [CompanySimilarity](/data-dictionary/company-similarity) | [`array[companysimilarity]`](/data-dictionary/company-similarity) | The Data City | Monthly/Quarterly |
| `SimilarCompositeCompanies` | Our composite similarity score. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/faqs/similar-companies) · Nested object — see [CompanySimilarity](/data-dictionary/company-similarity) | [`array[companysimilarity]`](/data-dictionary/company-similarity) | The Data City | Monthly/Quarterly |
| `CompanyExports` | HMRC export trade data for the company. · [Related: HMRC guide](https://docs.thedatacity.com/our-data/third-party-data/trade-data) | `array[companytrade]` | HMRC | Monthly |
| `CompanyImports` | HMRC import trade data for the company. · [Related: HMRC guide](https://docs.thedatacity.com/our-data/third-party-data/trade-data) | `array[companytrade]` | HMRC | Monthly |
| `CompanyDescription` | Takes the description from the description meta tag from the homepage and the about page. We check this is in English. · Download name: `Description` | `string` | Web Scraping | Monthly/Quarterly |
| `Homepage_domain` | Goes through the large web-scraping/matching process. · [Related: Web Scraping guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/website-matching) · Download name: `URLs` | `string` | Web Scraping | Monthly/Quarterly |
| `ESGStatementAnchors` | This is an array of ESG statement anchors. Each anchor is a string that is found in the company's website text. · Nested object — see [ESGStatementAnchor](/data-dictionary/esg-statement-anchor) | [`array[esgstatementanchor]`](/data-dictionary/esg-statement-anchor) | The Data City | Monthly |
| `TonnesOfC02equivGHGPerYear` | We estimate greenhouse gas emissions based on a published table of total carbon emissions by SIC sector. We assume that carbon emissions are spread equally across all companies within the SIC sector in proportion to their number of employees. Therefore, we only hold an estimate for a limited number of companies. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/faqs/how-do-you-get-green-house-emissions-data) | `number` | The Data City | Monthly |
| `InnovationScore` | We present this on the platform as stars but the API is actually a numerical value which gets categorised. The result is a binary classifier. The company is either innovative or it is not. The confidence increases the higher the score. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/proprietary-data/innovation-score) | `number` | The Data City | Quarterly |
| `LocationDetails` | An array of location objects with postcode-related fields. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: Companies House guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/company-locations) · Nested object — see [PostcodeDetail\_grouped](/data-dictionary/postcode-detail-grouped) | [`array[postcodedetail_grouped]`](/data-dictionary/postcode-detail-grouped) | Companies House | Monthly |
| `URLMatchStats` | Based on the URL matching process. | `urlmatchstat` | The Data City | Monthly/Quarterly |
# CleanCreditSafeCompany
Source: https://docs.thedatacity.com/data-dictionary/clean-credit-safe-company
40 fields on the CleanCreditSafeCompany schema — covers Financials.
**40 fields** · across **1 category** · **34** sourced from [CreditSafe](/our-data/third-party-data/creditsafe).
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Financials
| Field name | Description | Data type | Source | Updated |
| ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------- | --------------- | ------- |
| `ACCOUNTANT_NAME` | Accountant name. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Accountantname` | `string` | Companies House | Monthly |
| `ACCOUNTS_FORMAT` | Format of the submitted accounts. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `AccountsFormat` | `string` | Companies House | Monthly |
| `AMORTISATION_OF_INTANGIBLES` | Amortisation of intangibles. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Amortisationofintangibles` | `integer` | Companies House | Monthly |
| `AUDITORS` | Auditors. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Auditors` | `string` | Companies House | Monthly |
| `BANK_OVERDRAFT` | Bank overdraft. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Bankoverdraft` | `integer` | Companies House | Monthly |
| `CASH` | Cash at bank in hand. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Cash` | `integer` | Companies House | Monthly |
| `CONSOLIDATED_ACS` | Boolean. True if the financials are consolidated. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `ConsolidatedAcs` | `boolean` | Companies House | Monthly |
| `CREDITORS` | Creditors. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Creditors` | `integer` | Companies House | Monthly |
| `CURRENCY` | Currency the financials are reported in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Currency` | `string` | Companies House | Monthly |
| `DEBTORS` | Debtors. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Debtors` | `integer` | Companies House | Monthly |
| `DEBTORS_DUE_AFTER_ONE_YEAR` | Debtors due after one year. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Debtorsdueafteroneyear` | `integer` | Companies House | Monthly |
| `DECREASE_OR_INCREASE_IN_CASH` | Decrease or increase in cash. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `DecreaseOrIncreaseInCash` | `integer` | Companies House | Monthly |
| `DEPRECIATION_OF_TANGIBLES` | Depreciation of tangibles. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Depreciationoftangibles` | `integer` | Companies House | Monthly |
| `EBITDA` | A calculation. OPERATING\_PROFIT + DEPRECIATION\_OF\_TANGIBLES + AMORTISATION\_OF\_INTANGIBLES · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | The Data City | Monthly |
| `EMPLOYEES_ANOMALOUS` | Boolean. True if the number of employees is declared as anomalous. · Download name: `DeclaredEmployeesAnomalous` | `boolean` | The Data City | Monthly |
| `EMPLOYEES_ANOMALOUS_PROBABILITY` | Float. The probability that the number of employees is declared as anomalous. 0.5 is the threshold. · Download name: `EmployeesAnomalousProbability` | `number` | The Data City | Monthly |
| `EXPORT` | Export. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Export` | `integer` | Companies House | Monthly |
| `GROUP_DEBTORS` | Group debtors. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Groupdebtors` | `integer` | Companies House | Monthly |
| `LIKELY_AGENCY` | Boolean. True if the company is likely an agency. · Download name: `LikelyAgency` | `boolean` | The Data City | Monthly |
| `LIKELY_AGENCY_PROBABILITY` | Float. The probability that the company is likely an agency. 0.5 is the threshold. · Download name: `LikelyAgencyProbability` | `number` | The Data City | Monthly |
| `MADE_FROM_DATE` | Accounting From Date (ISO 8601 string). · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `MadeFromDate` | `string` | Companies House | Monthly |
| `MADE_UPTO_DATE` | Accounting To Date (ISO 8601 string). · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `MadeUptoDate` | `string` | Companies House | Monthly |
| `NET_ASSETS` | Net assets. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `NetAssets` | `integer` | Companies House | Monthly |
| `NET_WORTH` | Net worth of the company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Networth` | `integer` | Companies House | Monthly |
| `NUM_OF_MONTHS` | Number of months the financials cover. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `NumOfMonths` | `integer` | Companies House | Monthly |
| `NUMBER_OF_EMPLOYEES` | Employees as submitted to Companies House. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Reported_Numberofemployees` | `integer` | Companies House | Monthly |
| `OPERATING_PROFIT` | Operating profit. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Operatingprofit` | `integer` | Companies House | Monthly |
| `PRE_TAX_PROFIT` | Pre-tax profit. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Pretaxprofit` | `integer` | Companies House | Monthly |
| `PROFIT_AFTER_TAX` | Profit after tax. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Profitaftertax` | `integer` | Companies House | Monthly |
| `SHAREHOLDER_FUNDS` | Shareholder funds. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Shareholderfunds` | `integer` | Companies House | Monthly |
| `TOTAL_ASSETS` | Total assets. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Totalassets` | `integer` | Companies House | Monthly |
| `TOTAL_CURRENT_ASSETS` | Total current assets. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Totalcurrentassets` | `integer` | Companies House | Monthly |
| `TOTAL_CURRENT_LIABILITIES` | Total current liabilities. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Totalcurrentliabilities` | `integer` | Companies House | Monthly |
| `TOTAL_LIABILITIES` | Total liabilities. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Totalliabilities` | `integer` | Companies House | Monthly |
| `TRADE_DEBTORS` | Trade debtors. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Tradedebtors` | `integer` | Companies House | Monthly |
| `TURNOVER` | Turnover as submitted to Companies House. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Reported_Turnover` | `integer` | Companies House | Monthly |
| `TURNOVER_ANOMALOUS` | Boolean. True if the declared turnover is anomalous. · Download name: `DeclaredTurnoverAnomalous` | `boolean` | The Data City | Monthly |
| `TURNOVER_ANOMALOUS_PROBABILITY` | Float. The probability that the declared turnover is anomalous. 0.5 is the threshold. · Download name: `TurnoverAnomalousProbability` | `number` | The Data City | Monthly |
| `WAGES_AND_SALARIES` | Wages and salaries. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `WagesAndSalaries` | `integer` | Companies House | Monthly |
| `YEAR` | Year ending. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Yearending` | `integer` | Companies House | Monthly |
# CompanyGrowthTrend
Source: https://docs.thedatacity.com/data-dictionary/company-growth-trend
12 fields on the CompanyGrowthTrend schema — covers Growth.
**12 fields** · across **1 category** · **3** sourced from [CreditSafe](/our-data/third-party-data/creditsafe).
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Growth
| Field name | Description | Data type | Source | Updated |
| ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------- | --------------- | ------- |
| `BestEstimateCurrentEmployees` | The best estimate of the current number of employees. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) | `integer` | The Data City | Monthly |
| `BestEstimateCurrentTurnover` | The best estimate of the current turnover. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) | `integer` | The Data City | Monthly |
| `BestEstimatedTurnoverToEmployeeRatio` | A ratio of the best estimated turnover to the best estimated employees. This is calculated across the SIC(s) a company belongs to. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) | `number` | The Data City | Monthly |
| `BestEstimateEmployeeGrowthPercentagePerYear` | Employment growth percentage per year. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) | `number` | The Data City | Monthly |
| `BestEstimateGrowthPercentagePerYear` | The company's headline annual growth rate. Always equal to BestEstimateEmployeeGrowthPercentagePerYear: the annual employment growth rate from a log-linear fit across the company's reliable filed employee years (positive, non-anomalous values; at least three required, otherwise this is null). Also drives our employment projections. Never an average with BestEstimateTurnoverGrowthPercentagePerYear, which is reported separately and plays no part here, and plays no part in the OECDScaleup flag, which is computed independently from filed accounts over a fixed window. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) | `number` | The Data City | Monthly |
| `BestEstimateTurnoverGrowthPercentagePerYear` | Turnover growth percentage per year. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) | `number` | The Data City | Monthly |
| `CompanyAgeInDays` | The age of the company in days. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) | `number` | The Data City | Monthly |
| `CompanySizeEstimate` | Company size estimate. This is a categorical value and is based on the Companies Act 2006 definition. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/company-sizes) · Download name: `CompanyGrowthStage` | `string` | The Data City | Monthly |
| `EmploymentEstimates` | This is an array with the following fields. It is based on the employment estimates data. Financials are used to estimate the number of employees. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) · Nested object — see [YearMeasure](/data-dictionary/year-measure) | [`array[yearmeasure]`](/data-dictionary/year-measure) | The Data City | Monthly |
| `OECDScaleup` | Eurostat-OECD scale-up, computed from filed accounts only. The growth window ends at the latest filed accounts and starts at the most recent filed year at least three years earlier; the company needs 10 or more employees at the window start, at least four filed years with 10 or more employees inside the window, and annualised employment growth of at least 20% per year across it. BestEstimateGrowthPercentagePerYear plays no part in this flag. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/scale-ups) | `boolean` | The Data City | Monthly |
| `SICs` | An array of SIC codes the company belongs to. | `array[string]` | Companies House | Monthly |
| `TurnoverEstimates` | This is an array with the following fields. It is based on the turnover estimates data. Financials are used to estimate their turnover over time. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) · Nested object — see [YearMeasure](/data-dictionary/year-measure) | [`array[yearmeasure]`](/data-dictionary/year-measure) | The Data City | Monthly |
# CompanyJobPosting
Source: https://docs.thedatacity.com/data-dictionary/company-job-posting
4 fields on the CompanyJobPosting schema — covers Job Postings.
**4 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Job Postings
| Field name | Description | Data type | Source | Updated |
| ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | --------- | ------- |
| `MedianSalary` | The average of the monthly median advertised salaries for this company's postings in this SOC4 occupation group, in GBP. Despite the name this is a mean of monthly medians rather than a median across individual postings. Empty where no salary was advertised. Job postings data is a paid upgrade and is not included for all customers. | `number` | Lightcast | Monthly |
| `Postings` | The number of unique job postings this company advertised in this SOC4 occupation group. Covers the most recent five years of postings: the upstream extract is limited to that window, so this is a five-year total rather than an all-time one. Job postings data is a paid upgrade and is not included for all customers. | `integer` | Lightcast | Monthly |
| `SOC4Code` | The four-digit Standard Occupational Classification (SOC) code for the occupation group these postings fall into. SOC is the UK standard classification of occupations. Job postings data is a paid upgrade and is not included for all customers. | `string` | Lightcast | Monthly |
| `SOC4Name` | The name of the occupation group identified by SOC4Code, for example 'Programmers and software development professionals'. Job postings data is a paid upgrade and is not included for all customers. | `string` | Lightcast | Monthly |
# CompanyJobPostingByCertificationsSkill
Source: https://docs.thedatacity.com/data-dictionary/company-job-posting-by-certifications-skill
3 fields on the CompanyJobPostingByCertificationsSkill schema — covers Job Postings.
**3 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Job Postings
| Field name | Description | Data type | Source | Updated |
| ------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | --------- | ------- |
| `MedianPostingDuration` | Not currently populated and always empty. Reserved for the median number of days that postings requesting this certification stayed open, but the pipeline does not load a value for it. Job postings data is a paid upgrade and is not included for all customers. | `number` | Lightcast | Monthly |
| `SkillName` | The name of a certification requested in this company's job postings, as identified by Lightcast, for example 'Certified Public Accountant'. Job postings data is a paid upgrade and is not included for all customers. | `string` | Lightcast | Monthly |
| `UniquePostings` | The number of unique job postings from this company that requested this certification. Covers the most recent five years of postings: the upstream extract is limited to that window, so this is a five-year total rather than an all-time one. Job postings data is a paid upgrade and is not included for all customers. | `integer` | Lightcast | Monthly |
# CompanyJobPostingByCommonSkill
Source: https://docs.thedatacity.com/data-dictionary/company-job-posting-by-common-skill
3 fields on the CompanyJobPostingByCommonSkill schema — covers Job Postings.
**3 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Job Postings
| Field name | Description | Data type | Source | Updated |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | --------- | ------- |
| `MedianPostingDuration` | Not currently populated and always empty. Reserved for the median number of days that postings requesting this common skill stayed open, but the pipeline does not load a value for it. Job postings data is a paid upgrade and is not included for all customers. | `number` | Lightcast | Monthly |
| `SkillName` | The name of a common skill requested in this company's job postings, as identified by Lightcast. Common skills are general workplace skills such as communication or teamwork. Job postings data is a paid upgrade and is not included for all customers. | `string` | Lightcast | Monthly |
| `UniquePostings` | The number of unique job postings from this company that requested this common skill. Covers the most recent five years of postings: the upstream extract is limited to that window, so this is a five-year total rather than an all-time one. Job postings data is a paid upgrade and is not included for all customers. | `integer` | Lightcast | Monthly |
# CompanyJobPostingByLot
Source: https://docs.thedatacity.com/data-dictionary/company-job-posting-by-lot
4 fields on the CompanyJobPostingByLot schema — covers Job Postings.
**4 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Job Postings
| Field name | Description | Data type | Source | Updated |
| ---------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | --------- | ------- |
| `LotCode` | The Lightcast Occupation Taxonomy (LOT) code for the occupation these postings fall into. LOT is Lightcast's own occupation taxonomy, separate from the UK SOC classification used by the CompanyJobPosting schema. Job postings data is a paid upgrade and is not included for all customers. | `string` | Lightcast | Monthly |
| `LotName` | The name of the occupation identified by LotCode. Job postings data is a paid upgrade and is not included for all customers. | `string` | Lightcast | Monthly |
| `MedianSalary` | The average of the monthly median advertised salaries for this company's postings in this LOT occupation, in GBP. Despite the name this is a mean of monthly medians rather than a median across individual postings. Empty where no salary was advertised. Job postings data is a paid upgrade and is not included for all customers. | `number` | Lightcast | Monthly |
| `Postings` | The number of unique job postings this company advertised in this LOT occupation. Covers the most recent five years of postings: the upstream extract is limited to that window, so this is a five-year total rather than an all-time one. Job postings data is a paid upgrade and is not included for all customers. | `integer` | Lightcast | Monthly |
# CompanyJobPostingBySoftwareSkill
Source: https://docs.thedatacity.com/data-dictionary/company-job-posting-by-software-skill
3 fields on the CompanyJobPostingBySoftwareSkill schema — covers Job Postings.
**3 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Job Postings
| Field name | Description | Data type | Source | Updated |
| ------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | --------- | ------- |
| `MedianPostingDuration` | Not currently populated and always empty. Reserved for the median number of days that postings requesting this software skill stayed open, but the pipeline does not load a value for it. Job postings data is a paid upgrade and is not included for all customers. | `number` | Lightcast | Monthly |
| `SkillName` | The name of a software skill requested in this company's job postings, as identified by Lightcast, for example 'Python' or 'Microsoft Excel'. Job postings data is a paid upgrade and is not included for all customers. | `string` | Lightcast | Monthly |
| `UniquePostings` | The number of unique job postings from this company that requested this software skill. Covers the most recent five years of postings: the upstream extract is limited to that window, so this is a five-year total rather than an all-time one. Job postings data is a paid upgrade and is not included for all customers. | `integer` | Lightcast | Monthly |
# CompanyJobPostingBySpecializedSkill
Source: https://docs.thedatacity.com/data-dictionary/company-job-posting-by-specialized-skill
3 fields on the CompanyJobPostingBySpecializedSkill schema — covers Job Postings.
**3 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Job Postings
| Field name | Description | Data type | Source | Updated |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------- | --------- | ------- |
| `MedianPostingDuration` | Not currently populated and always empty. Reserved for the median number of days that postings requesting this specialised skill stayed open, but the pipeline does not load a value for it. Job postings data is a paid upgrade and is not included for all customers. | `number` | Lightcast | Monthly |
| `SkillName` | The name of a specialised skill requested in this company's job postings, as identified by Lightcast. Specialised skills are specific to an occupation, as opposed to general workplace skills. Job postings data is a paid upgrade and is not included for all customers. | `string` | Lightcast | Monthly |
| `UniquePostings` | The number of unique job postings from this company that requested this specialised skill. Covers the most recent five years of postings: the upstream extract is limited to that window, so this is a five-year total rather than an all-time one. Job postings data is a paid upgrade and is not included for all customers. | `integer` | Lightcast | Monthly |
# CompanySimilarity
Source: https://docs.thedatacity.com/data-dictionary/company-similarity
3 fields on the CompanySimilarity schema — covers Similarity.
**3 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Similarity
| Field name | Description | Data type | Source | Updated |
| -------------------- | ------------------------------------------------------------------------- | --------- | ------------- | ----------------- |
| `CompanyOne` | The company number of the first company (the company you are looking at). | `string` | The Data City | Monthly/Quarterly |
| `CompanyTwo` | The company number of the second company (the company that is similar). | `string` | The Data City | Monthly/Quarterly |
| `Similarity` | A numerical value representing the similarity between the two companies. | `number` | The Data City | Monthly/Quarterly |
# DealroomPerRoundFunding
Source: https://docs.thedatacity.com/data-dictionary/dealroom-per-round-funding
13 fields on the DealroomPerRoundFunding schema — covers Investment.
**13 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Investment
| Field name | Description | Data type | Source | Updated |
| ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------- | ---------- | ------- |
| `CompanyId` | The investment provider's identifier for the company. | `string` | Investment | Daily |
| `LastInvestmentDate` | The last investment date. · Download name: `LastKnownInvestment` | `string` | Investment | Daily |
| `ParticipatingInvestorTypes` | The distinct Investor types of the round's participating investors (for example "Angel", "Venture Capital"), ' \| '-joined and alphabetical. Empty on acquisition and IPO rows, and where no participating investor has a type. | `string` | Investment | Daily |
| `ProviderCompanyName` | The investment provider's name for the company. | `string` | Investment | Daily |
| `RoundAmount_GBP_Million` | The round amount, in GBP millions. | `number` | Investment | Daily |
| `RoundDate` | The round date. | `string` | Investment | Daily |
| `RoundInvestorName` | The round investor name. | `string` | Investment | Daily |
| `RoundInvestorType` | Deprecated and always empty. The funding source no longer supplies an investor type per round; use ParticipatingInvestorTypes for the types of investor that took part in a round. | `string` | Investment | Daily |
| `RoundLeadInvestorName` | The lead investor name(s) for the round, ' \| '-joined and always a subset of RoundInvestorName. Empty when the provider flags no lead. | `string` | Investment | Daily |
| `RoundType` | The round type. | `string` | Investment | Daily |
| `RoundVerified` | A boolean. 1 if the round is verified by the investment provider. | `integer` | Investment | Daily |
| `Spinout` | Deprecated and always 0. The funding source no longer flags spinouts per round; spinout status is now a company-level field populated from the spinout register. | `integer` | Investment | Daily |
| `TotalInvestmentRaised_GBP_Million` | The total investment the company has raised. DEBT and ACQUISITION rounds are not included in this total. · [Related: Investment guide](https://docs.thedatacity.com/our-data/third-party-data/dealroom-data) · Download name: `TotalInvestmentMillionPounds` | `number` | Investment | Daily |
# ESGStatementAnchor
Source: https://docs.thedatacity.com/data-dictionary/esg-statement-anchor
3 fields on the ESGStatementAnchor schema — covers ESG.
**3 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## ESG
| Field name | Description | Data type | Source | Updated |
| ---------------- | ------------------------------------------------------------------------------- | --------- | ------------- | ------- |
| `Anchor` | The ESG statement anchor. | `string` | The Data City | Monthly |
| `Domain` | The domain of the company's website. | `string` | The Data City | Monthly |
| `Word` | The word in the ESG statement anchor which indicates the type of ESG statement. | `string` | The Data City | Monthly |
# GroupStructureDetails
Source: https://docs.thedatacity.com/data-dictionary/group-structure-details
14 fields on the GroupStructureDetails schema — covers Classification.
**14 fields** · across **1 category** · **14** sourced from [CreditSafe](/our-data/third-party-data/creditsafe).
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Classification
| Field name | Description | Data type | Source | Updated |
| ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------- | ------------- | ------- |
| `CanBeConsolidated` | Indicates whether the subsidiary company can be consolidated into the parent company's financial statements. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `boolean` | The Data City | Monthly |
| `ConsolidatedAccountsCompanyName` | The name of the company whose accounts are consolidated with the subsidiary. There can be multiple consolidated accounts company names if the subsidiary's accounts are consolidated with more than one other company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | The Data City | Monthly |
| `ConsolidatedAccountsCompanyNumber` | The Company Number of the company whose accounts are consolidated with the subsidiary. There can be multiple consolidated accounts company numbers if the subsidiary's accounts are consolidated with more than one other company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | The Data City | Monthly |
| `IsLikelyDistinctBrand` | Indicates whether the subsidiary is likely to be a distinct brand within the parent company's group structure. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `boolean` | The Data City | Monthly |
| `IsSubGroup` | Indicates whether the subsidiary is part of a sub-group within the parent company's group structure. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `boolean` | The Data City | Monthly |
| `ParentCompanyName` | The name of the parent company directly above the subsidiary in the group structure. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | The Data City | Monthly |
| `ParentCompanyNumber` | The Company Number of the parent company directly above the subsidiary in the group structure. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | The Data City | Monthly |
| `SubsidiaryCompanyLevel` | The level at which the subsidiary company operates within the group structure. Level 0 is the ultimate parent company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | The Data City | Monthly |
| `SubsidiaryCompanyName` | The name of the subsidiary company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | The Data City | Monthly |
| `SubsidiaryCompanyNumber` | The Company Number of the subsidiary. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | The Data City | Monthly |
| `SubsidiaryCompanyWebsite` | The website of the subsidiary company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | The Data City | Monthly |
| `TypeOfAccounts` | The type of accounts filed by the subsidiary company, which can indicate its financial reporting requirements and potential consolidation status. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | The Data City | Monthly |
| `UltimateParentCompanyName` | The name of the ultimate parent company in the group structure. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | The Data City | Monthly |
| `UltimateParentCompanyNumber` | The Company Number of the ultimate parent company in the group structure. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | The Data City | Monthly |
# Data dictionary
Source: https://docs.thedatacity.com/data-dictionary/index
Browse every field on Industry Engine — 345 fields across 31 schemas.
**345 fields across 31 schemas.** 146 of them are [CreditSafe-sourced](/our-data/third-party-data/creditsafe).
Search and filter every field below, or download the full dictionary. Downloads are public and do not require a documentation login.
## Browse by schema
**85 fields** · Classification · Company · Investment\
*24 sourced from [CreditSafe](/our-data/third-party-data/creditsafe)*
**40 fields** · Financials\
*34 sourced from [CreditSafe](/our-data/third-party-data/creditsafe)*
**26 fields** · Location\
*26 sourced from [CreditSafe](/our-data/third-party-data/creditsafe)*
**22 fields** · Funding
**18 fields** · Officers\
*18 sourced from [CreditSafe](/our-data/third-party-data/creditsafe)*
**14 fields** · Classification\
*14 sourced from [CreditSafe](/our-data/third-party-data/creditsafe)*
**13 fields** · Investment
**13 fields** · Investment
**13 fields** · Investment
**12 fields** · Growth\
*3 sourced from [CreditSafe](/our-data/third-party-data/creditsafe)*
**11 fields** · Funding
**9 fields** · Gender\
*9 sourced from [CreditSafe](/our-data/third-party-data/creditsafe)*
**7 fields** · Shareholders\
*7 sourced from [CreditSafe](/our-data/third-party-data/creditsafe)*
**6 fields** · PSC\
*6 sourced from [CreditSafe](/our-data/third-party-data/creditsafe)*
**5 fields** · Insights · Investment\
*3 sourced from [CreditSafe](/our-data/third-party-data/creditsafe)*
**4 fields** · Job Postings
**4 fields** · Job Postings
**4 fields** · Classification
**4 fields** · Job Postings
**4 fields** · URL Matching
**4 fields** · Job Postings
**3 fields** · ESG
**3 fields** · Similarity
**3 fields** · Growth\
*2 sourced from [CreditSafe](/our-data/third-party-data/creditsafe)*
**3 fields** · Job Postings
**3 fields** · Job Postings
**3 fields** · Job Postings
**3 fields** · Job Postings
**2 fields** · Job Postings
**2 fields** · Job Postings
**2 fields** · Classification
# IndustrialStrategySector
Source: https://docs.thedatacity.com/data-dictionary/industrial-strategy-sector
4 fields on the IndustrialStrategySector schema — covers Classification.
**4 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Classification
| Field name | Description | Data type | Source | Updated |
| ------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | ------------- | ------- |
| `ISCode` | Industrial Strategy Sector codes assigned to the company. These codes are based on the UK Government's Industrial Strategy framework, but created by The Data City. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/industrial-strategy-classifications) · Download name: `IS8Codes` | `string` | The Data City | Ad hoc |
| `ISCodeName` | Industrial Strategy Sectors assigned to the company. These names correspond to the Industrial Strategy Sectors based on the UK Government's Industrial Strategy framework. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/industrial-strategy-classifications) · Download name: `IS8Names` | `string` | The Data City | Ad hoc |
| `ISVerticalCode` | Industrial Strategy Vertical codes assigned to the company. These codes are sub-codes within the main Industrial Strategy Sector codes, providing a more granular classification of the company's business activities. The sub-codes correspond to the frontier industries identified by the government. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/industrial-strategy-classifications) · Download name: `IS8VerticalCodes` | `string` | The Data City | Ad hoc |
| `ISVerticalName` | Industrial Strategy Verticals assigned to the company. These names correspond to the vertical sub-categories within the Industrial Strategy Sectors, providing detailed insight into the company's specific area of operation. The sub-codes correspond to the frontier industries identified by the government. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/industrial-strategy-classifications) · Download name: `IS8VerticalNames` | `string` | The Data City | Ad hoc |
# InnovateUKFundingClean
Source: https://docs.thedatacity.com/data-dictionary/innovate-uk-funding-clean
22 fields on the InnovateUKFundingClean schema — covers Funding.
**22 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Funding
| Field name | Description | Data type | Source | Updated |
| --------------------------------------------- | ---------------------------------------------------------------------------------- | --------- | ---------- | ------- |
| `ActualSpendtoDate` | The actual spend to date. | `number` | InnovateUK | Monthly |
| `ApplicationNumber` | The application number. | `string` | InnovateUK | Monthly |
| `AwardOffered` | The value of the award. | `number` | InnovateUK | Monthly |
| `CompetitionReference` | The competition reference. | `string` | InnovateUK | Monthly |
| `CompetitionTitle` | The competition title. | `string` | InnovateUK | Monthly |
| `CompetitionYear` | The competition year. | `integer` | InnovateUK | Monthly |
| `EnterpriseSize` | The size of the company. | `string` | InnovateUK | Monthly |
| `IndustrialStrategyChallengeFundISCF` | A boolean. "Yes" if the project is part of the Industrial Strategy Challenge Fund. | `string` | InnovateUK | Monthly |
| `InnovateUKProductType` | The product type of the project. | `string` | InnovateUK | Monthly |
| `IsLeadParticipant` | A boolean. "Yes" if the participant is the lead participant. | `string` | InnovateUK | Monthly |
| `ParticipantName` | The participant name (usually the company name). | `string` | InnovateUK | Monthly |
| `ParticipantWithdrawnFromProject` | A boolean. "Active" if the participant is still active in the project. | `string` | InnovateUK | Monthly |
| `Postcode` | The postcode of the project. | `string` | InnovateUK | Monthly |
| `ProgrammeTitle` | The programme title. | `string` | InnovateUK | Monthly |
| `ProjectEndDate` | The project end date. | `string` | InnovateUK | Monthly |
| `ProjectNumber` | The project number. | `string` | InnovateUK | Monthly |
| `ProjectStartDate` | The project start date. | `string` | InnovateUK | Monthly |
| `ProjectStatus` | The status of the project. | `string` | InnovateUK | Monthly |
| `ProjectTitle` | The project title. | `string` | InnovateUK | Monthly |
| `PublicDescription` | The public description of the project. | `string` | InnovateUK | Monthly |
| `Sector` | The sector of the project. | `string` | InnovateUK | Monthly |
| `TotalCosts` | The total costs of the project. | `number` | InnovateUK | Monthly |
# ListInsights
Source: https://docs.thedatacity.com/data-dictionary/list-insights
5 fields on the ListInsights schema — covers Insights, Investment.
**5 fields** · across **2 categories** · **3** sourced from [CreditSafe](/our-data/third-party-data/creditsafe).
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Insights
| Field name | Description | Data type | Source | Updated |
| -------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | ------------- | ------- |
| `BestEstimateTotalEmployeesAttributedToLocation` | Sum of best-estimate employees across operating locations that match the active geographic filter (e.g. selected areas, postcode search and radius, or custom geography WKT). Null when no geographic filter applies or when no locations match the filter. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) | `integer` | The Data City | Monthly |
| `BestEstimateTotalGVAAttributedToLocation` | Sum of best-estimate GVA across operating locations that match the active geographic filter (e.g. selected areas, postcode search and radius, or custom geography WKT). Null when no geographic filter applies or when no locations match the filter. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/gva-data) | `number` | The Data City | Monthly |
| `BestEstimateTotalTurnoverAttributedToLocation` | Sum of best-estimate turnover across operating locations that match the active geographic filter (e.g. selected areas, postcode search and radius, or custom geography WKT). Null when no geographic filter applies or when no locations match the filter. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) | `number` | The Data City | Monthly |
## Investment
| Field name | Description | Data type | Source | Updated |
| ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------- | ---------------------- | ------- |
| `TopInvestorsByRoundCount` | The investors that have participated in the most funding rounds across companies in this list, ranked by number of rounds. · [Related: Third-party Investment guide](https://docs.thedatacity.com/our-data/third-party-data/dealroom-data) | `array[topinvestor]` | Third-party Investment | Daily |
| `TopInvestorsByValue` | The investors that have contributed the most total funding value across companies in this list, ranked by total investment amount in GBP. · [Related: Third-party Investment guide](https://docs.thedatacity.com/our-data/third-party-data/dealroom-data) | `array[topinvestor]` | Third-party Investment | Daily |
# Officer
Source: https://docs.thedatacity.com/data-dictionary/officer
18 fields on the Officer schema — covers Officers.
**18 fields** · across **1 category** · **18** sourced from [CreditSafe](/our-data/third-party-data/creditsafe).
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Officers
| Field name | Description | Data type | Source | Updated |
| --------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | --------------- | ------- |
| `AddressLine1` | The first line of the address of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `AddressLine2` | The second line of the address of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Appointedon` | The date the officer was appointed. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Careof` | The care of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Country` | The country of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `County` | The county of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `DOB` | The date of birth of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Forenames` | The forenames of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Honours` | The honours of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Nationality` | The nationality of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Occupation` | The occupation of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `PersonNumber` | The person number of the officer. Person number is a unique identifier for the officer provided by Companies House. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `PostTown` | The post town of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `ResignationDate` | The date the officer resigned (if applicable). · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Role` | The role of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Surname` | The surname of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Title` | The title of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `UsualResidentialCountry` | The usual residential country of the officer. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
# PersonWithSignificantControl
Source: https://docs.thedatacity.com/data-dictionary/person-with-significant-control
6 fields on the PersonWithSignificantControl schema — covers PSC.
**6 fields** · across **1 category** · **6** sourced from [CreditSafe](/our-data/third-party-data/creditsafe).
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## PSC
| Field name | Description | Data type | Source | Updated |
| ---------------------------- | ----------------------------------------------------------------------------------------------------- | --------- | --------------- | ------- |
| `CountryOfResidence` | The country of residence of the person. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `DoB` | The date of birth of the person. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Name` | The name of the person. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `NatureOfControl` | The nature of the person's control. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Title` | The title of the person. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
| `Type` | The type of the person's control. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | Companies House | Monthly |
# PostcodeDetail_grouped
Source: https://docs.thedatacity.com/data-dictionary/postcode-detail-grouped
26 fields on the PostcodeDetail_grouped schema — covers Location.
**26 fields** · across **1 category** · **26** sourced from [CreditSafe](/our-data/third-party-data/creditsafe).
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Location
| Field name | Description | Data type | Source | Updated |
| --------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | --------------- | ------- |
| `AssociatedAddress` | The associated address found with the postcode. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | The Data City | Monthly |
| `BestEstimateLocationEmployees` | The best estimate of employees for this location. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | The Data City | Monthly |
| `BestEstimateLocationGVA` | The best estimate of GVA attributed to this operating location. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/gva-data) | `number` | The Data City | Monthly |
| `BestEstimateLocationTurnover` | The best estimate of turnover for this location. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `number` | The Data City | Monthly |
| `FoundOnWebpage` | A boolean value to indicate if the postcode was found on the company's website. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `boolean` | The Data City | Monthly |
| `IsBlacklisted` | A boolean value to indicate if the postcode is blacklisted as an operating address. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `boolean` | The Data City | Monthly |
| `IsOperatingAddress` | A boolean value to indicate if this postcode is an operating address for the company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `boolean` | The Data City | Monthly |
| `IsSubsidiaryAddress` | A boolean value to indicate if this postcode is a subsidiary's address rather than the company's own. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `boolean` | The Data City | Monthly |
| `ITL1` | The ITL1 region code the postcode resides in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | ONS | Monthly |
| `ITL1NM` | The ITL1 region name the postcode resides in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `ITL1` | `string` | ONS | Monthly |
| `ITL2` | The ITL2 (International Territorial Level 2) region code the postcode resides in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `ITL2Code` | `string` | ONS | Monthly |
| `ITL2NM` | The ITL2 (International Territorial Level 2) region name the postcode resides in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `ITL2` | `string` | ONS | Monthly |
| `ITL3` | The ITL3 (International Territorial Level 3) region code the postcode resides in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `ITL3Code` | `string` | ONS | Monthly |
| `LAcode` | The local authority code the postcode resides in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | ONS | Monthly |
| `LAname` | The local authority name the postcode resides in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | ONS | Monthly |
| `Latitude` | The latitude of the location. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `number` | ONS | Monthly |
| `Longitude` | The longitude of the location. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `number` | ONS | Monthly |
| `LSOA` | The Lower Layer Super Output Area the postcode resides in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | ONS | Monthly |
| `MSOA` | The Middle Layer Super Output Area the postcode resides in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | ONS | Monthly |
| `OECDFUACorePlusCommuting` | The OECD Economic Classification Data - Functional Urban Area Core Plus Commuting code the postcode resides in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | OECD | Monthly |
| `Postcode` | A postcode for the company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Postcodes` | `string` | Companies House | Monthly |
| `Postcode_withspaces` | A postcode for the company with formatted spaces. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `PostcodeWithspaces` | `string` | Companies House | Monthly |
| `RegisteredPostcode` | A boolean value to indicate if the postcode is the registered postcode. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `boolean` | The Data City | Monthly |
| `StrategicAuthorityCode` | The strategic authority code the postcode resides in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | ONS | Monthly |
| `StrategicAuthorityName` | The strategic authority name the postcode resides in. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `string` | ONS | Monthly |
| `UKConstituency` | The constituency name the postcode resides in. The constituencies filter is comprised of UK parliamentary constituencies. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Constituencies` | `string` | ONS | Monthly |
# RSIC
Source: https://docs.thedatacity.com/data-dictionary/rsic
2 fields on the RSIC schema — covers Classification.
**2 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Classification
| Field name | Description | Data type | Source | Updated |
| ------------------ | --------------------------------------------------------------------------------------------------------------------------------- | --------- | ------------- | ------- |
| `Score` | The score of the RSIC. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/proprietary-data/what-are-rsics) | `number` | The Data City | Monthly |
| `SIC_Code` | The SIC code of the RSIC. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/proprietary-data/what-are-rsics) | `string` | The Data City | Monthly |
# ShareHolder
Source: https://docs.thedatacity.com/data-dictionary/share-holder
7 fields on the ShareHolder schema — covers Shareholders.
**7 fields** · across **1 category** · **7** sourced from [CreditSafe](/our-data/third-party-data/creditsafe).
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Shareholders
| Field name | Description | Data type | Source | Updated |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- | --------- | --------------- | ------- |
| `CURRENCY` | The currency of the share. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `Currency` | `string` | Companies House | Monthly |
| `SHARE_HOLDER_NAME` | The name of the shareholder. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `ShareHolderName` | `string` | Companies House | Monthly |
| `SHARE_HOLDING` | The holding of the share as a value. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `ShareHolding` | `number` | Companies House | Monthly |
| `SHARE_PERCENT` | The percentage of the share. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `SharePercent` | `number` | Companies House | Monthly |
| `SHARE_TYPE` | The type of the share. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `ShareType` | `string` | Companies House | Monthly |
| `TOTAL_SHARE_VALUE` | The total value of the shares. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `TotalShareValue` | `number` | Companies House | Monthly |
| `TOTAL_SHARES` | The total number of shares. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · Download name: `TotalShares` | `number` | Companies House | Monthly |
# UrlMatchStat
Source: https://docs.thedatacity.com/data-dictionary/url-match-stat
4 fields on the UrlMatchStat schema — covers URL Matching.
**4 fields** · across **1 category**.
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## URL Matching
| Field name | Description | Data type | Source | Updated |
| ---------------------------------- | ------------------------------------------------------------------------------------------------------ | --------------- | ------------- | ----------------- |
| `MatchedWebsite` | The matched website. | `string` | The Data City | Monthly/Quarterly |
| `MatchedWebsiteConfidence` | The confidence level of the website match. | `string` | The Data City | Monthly/Quarterly |
| `MatchedWebsiteReasoning` | An array of reasons for the match. | `array[string]` | The Data City | Monthly/Quarterly |
| `MatchedWebsiteScore` | The matched website score. The higher the score the more confident we are that the website is a match. | `number` | The Data City | Monthly/Quarterly |
# WomanLedStats
Source: https://docs.thedatacity.com/data-dictionary/woman-led-stats
9 fields on the WomanLedStats schema — covers Gender.
**9 fields** · across **1 category** · **9** sourced from [CreditSafe](/our-data/third-party-data/creditsafe).
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Gender
| Field name | Description | Data type | Source | Updated |
| ------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- | --------- | ------------- | ------- |
| `ActiveFemaleDirectors` | The number of active female directors. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | The Data City | Monthly |
| `ActiveFemaleFounders` | The number of female founders that are still active at the company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | The Data City | Monthly |
| `ActiveMaleDirectors` | The number of active male directors. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | The Data City | Monthly |
| `ActiveMaleFounders` | The number of male founders that are still active at the company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | The Data City | Monthly |
| `InitialFemaleFounders` | The number of female founders at the time of founding. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | The Data City | Monthly |
| `InitialMaleFounders` | The number of male founders at the time of founding. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | The Data City | Monthly |
| `TotalActiveDirectors` | The number of active total directors. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | The Data City | Monthly |
| `TotalActiveFounders` | The number of founders that are still active at the company. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | The Data City | Monthly |
| `TotalInitialFounders` | The number of founders at the time of founding. · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) | `integer` | The Data City | Monthly |
# YearMeasure
Source: https://docs.thedatacity.com/data-dictionary/year-measure
3 fields on the YearMeasure schema — covers Growth.
**3 fields** · across **1 category** · **2** sourced from [CreditSafe](/our-data/third-party-data/creditsafe).
This schema is part of Industry Engine. Use the [data dictionary browser](/data-dictionary) to search and filter every field, or jump straight to a field below. Array fields that reference another schema link to that schema's own page.
## Growth
| Field name | Description | Data type | Source | Updated |
| ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | ------------- | ------- |
| `Estimated` | The estimated number of employees (for EmploymentEstimates) or estimated turnover (for TurnoverEstimates). · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) · Download name: `BestEstimate_Numberofemployees or BestEstimate_Turnover` | `integer` | The Data City | Monthly |
| `Measure` | The measured number of employees (for EmploymentEstimates) or measured turnover (for TurnoverEstimates). · [CreditSafe-sourced](/our-data/third-party-data/creditsafe) · [Related: CreditSafe guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) · Download name: `Reported_Numberofemployees or Reported_Turnover` | `integer` | CreditSafe | Monthly |
| `Year` | The financial year. · [Related: The Data City guide](https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth) · Download name: `Yearending` | `integer` | The Data City | Monthly |
# Instant Classification API
Source: https://docs.thedatacity.com/instant-classification/index
Classify any company by description or website. Get RTICs, RSICs, RNAICs, SICs, and similar companies in a single call.
Send a company description or a website URL and get back real-time industrial classifications across multiple taxonomies, plus a list of similar companies.
## What you get back
A single `POST` request returns:
| Field | What it is |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |
| **RTICs** | Real-Time Industrial Classifications — The Data City's proprietary codes that classify companies by what they actually do, updated continuously. |
| **RSICs** | Real-time SIC codes — SIC codes refined with live data rather than the static codes on Companies House. |
| **RNAICs** | Real-time NAICS codes — the North American industry classification equivalent, built the same way as RSICs. |
| **SICs** | Standard Industrial Classification codes — the statutory UK codes registered at Companies House. |
| **Similar companies** | Companies with comparable business activities, ranked by similarity score. |
Most classifications carry a match score, so you can decide how much to trust each result. See the [response reference](/instant-classification/guides/response-reference) for what every field contains.
## Who it's for
Use this API when you need to classify companies programmatically — for lead enrichment, portfolio tagging, market sizing, or compliance screening.
## How to get started
Sign up and classify your first company in under five minutes.
OAuth2 password flow — get a JWT, use a Bearer header.
Every field in the response, and when it's empty.
Every error code, what it means, and what to do about it.
Free-tier usage quota and per-minute rate limits.
## Base URL
```text theme={null}
https://instant-classification-api.thedatacity.com/api/v1
```
All endpoints sit under `/api/v1`. The interactive playground on each endpoint page lets you paste your token and send live requests.
# Instant Classification
Source: https://docs.thedatacity.com/api-reference/classification/instant-classification
https://instant-classification-api.thedatacity.com/api/v1/openapi.json post /api/v1/instantClassification
Classify a company using the instant classification service.
This endpoint requires authentication. The user must be logged in with a valid JWT token.
# Login Access Token
Source: https://docs.thedatacity.com/api-reference/login/login-access-token
https://instant-classification-api.thedatacity.com/api/v1/openapi.json post /api/v1/login/access-token
OAuth2 compatible token login, get an access token for future requests
# Test Token
Source: https://docs.thedatacity.com/api-reference/login/test-token
https://instant-classification-api.thedatacity.com/api/v1/openapi.json post /api/v1/login/test-token
Test access token
# Get Batch Results
Source: https://docs.thedatacity.com/api-reference/match/get-batch-results
https://datamatcher-api.thedatacity.com/api/v1/openapi.json get /api/v1/match/batch/{job_id}/results
# Get Batch Status
Source: https://docs.thedatacity.com/api-reference/match/get-batch-status
https://datamatcher-api.thedatacity.com/api/v1/openapi.json get /api/v1/match/batch/{job_id}
# Match One
Source: https://docs.thedatacity.com/api-reference/match/match-one
https://datamatcher-api.thedatacity.com/api/v1/openapi.json post /api/v1/match
Match a single company to a Companies House number.
# Submit Batch
Source: https://docs.thedatacity.com/api-reference/match/submit-batch
https://datamatcher-api.thedatacity.com/api/v1/openapi.json post /api/v1/match/batch
Enqueue an async batch (max 1,000 companies). Poll GET /match/batch/{job_id}.
# Register User
Source: https://docs.thedatacity.com/api-reference/users/register-user
https://instant-classification-api.thedatacity.com/api/v1/openapi.json post /api/v1/users/signup
Create new user without the need to be logged in.
# Read Public Config
Source: https://docs.thedatacity.com/api-reference/utils/read-public-config
https://instant-classification-api.thedatacity.com/api/v1/openapi.json get /api/v1/utils/public-config
Return non-sensitive settings needed by the public frontend.
# Changelog
Source: https://docs.thedatacity.com/company-matching/changelog
Material changes to the Company Matching API, newest first.
This page records changes to the API contract and to behaviour worth knowing about: new fields, renames, breaking changes. Each entry names the release it shipped in. The release is `info.version` in the OpenAPI spec — see [Versioning](/company-matching/guides/versioning).
Entries start at release 1.0.0, the first numbered release. Changes made during the earlier beta were not tracked.
## Conventions
Entries are tagged so you can scan for the kind of change that affects you:
* **Breaking** — you must change your integration. Always listed first.
* **Added** — new endpoint, field, or capability. Safe to ignore if you don't use it.
* **Changed** — non-breaking change to existing behaviour.
* **Deprecated** — still works but scheduled for removal. Start migrating.
* **Removed** — no longer available.
* **Fixed** — bug fix bringing actual behaviour in line with documented behaviour.
## History
### 2026-09-09 — release 1.1.0
* **Changed** — the match endpoints now declare `404`, `409` and `500` in the OpenAPI spec, so the Endpoints pages and any client generated from the spec list every status a route can return. What the endpoints do is unchanged.
* **Changed** — submitted request bodies are cleared from usage records 30 days after the call. The usage metadata for that call is kept: the endpoint, the number of companies, the response status and the processing time.
* **Changed** — cached responses from upstream company-data sources are deleted after 180 days. Before this release they stopped being used at that age but were not removed.
* **Changed** — submitted company fields and match results are no longer written to application log records. A log entry identifies a request by the `id` you supply and by our own job and row identifiers.
The [response reference](/company-matching/guides/response-reference) now sets out how `route` and `rigorous_match` derive the `confidence` band. That behaviour has not changed; it was undocumented.
### 2026-09-04 — release 1.0.1
* **Fixed** — deployment fix only. No change to the API contract.
### 2026-09-04 — release 1.0.0
First numbered release. The URL prefix stays `/api/v1`.
* **Breaking** — request field `postcode` renamed to `input_postcode`.
* **Breaking** — response field `matched_postcode` renamed to `matched_registered_postcode`. It holds the registered office postcode of the matched company.
Requests that still send `postcode` are not rejected. The API returns `200` and ignores the unknown field, so the postcode stops contributing to matching and you get no error. Rename the field.
# Authentication
Source: https://docs.thedatacity.com/company-matching/guides/authentication
Customer API keys (dm_live_…) as a Bearer token on match endpoints.
## Getting a key
API keys are issued by The Data City. A key:
* Starts with `dm_live_`
* Is shown in full **once** at creation — store it immediately
* Can be revoked; revoked keys return `401`
Contact [support@thedatacity.com](mailto:support@thedatacity.com) if you need a key or a rotation.
## Using the key
Add an `Authorization` header to every match request:
```http theme={null}
Authorization: Bearer dm_live_…
```
Example:
```bash theme={null}
curl -sS -X POST "https://datamatcher-api.thedatacity.com/api/v1/match" \
-H "Authorization: Bearer dm_live_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"input_name":"John Lewis plc"}'
```
Treat your API key like a password. Store it in an environment variable or secrets manager. Never check it into source control or include it in client-side code.
## What the key is not
Dashboard login uses a separate OAuth2 password flow (`POST /api/v1/login/access-token`) and returns a JWT. That JWT is for the admin UI and administrative APIs. **Match calls must use a `dm_live_…` API key.**
If you send a JWT to `/match`, expect `401`.
## Failure modes
| Status | Meaning | What to do |
| ------ | ---------------------------------------------------------- | ------------------------------------------------------------------ |
| `401` | Missing header, malformed key, unknown key, or revoked key | Check the key value; request a new one if revoked |
| `429` | Too many match requests | Back off — see [Rate limits](/company-matching/guides/rate-limits) |
See [Errors](/company-matching/guides/errors) for the full status-code matrix.
# Batch matching
Source: https://docs.thedatacity.com/company-matching/guides/batch
Submit up to 1,000 companies, poll job status or receive a webhook, then page through results.
Use the batch API when you have a list. One submit request queues the work; matching runs asynchronously.
## 1. Submit
```bash theme={null}
curl -sS -X POST "https://datamatcher-api.thedatacity.com/api/v1/match/batch" \
-H "Authorization: Bearer dm_live_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"companies": [
{"input_name": "Sainsbury'\''s Supermarkets Ltd", "id": "demo-001"},
{"input_name": "John Lewis plc", "id": "demo-002"},
{"input_name": "Holmes Joinery Partners LLP", "id": "demo-003"}
]
}'
```
Response (`202 Accepted`):
```json theme={null}
{
"job_id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
"status": "queued",
"total": 3
}
```
Maximum batch size: **1,000** companies. Larger payloads return `422`.
Jobs are retained for **30 days** from creation (`expires_at` on the status payload).
## 2. Poll status
```bash theme={null}
curl -sS "https://datamatcher-api.thedatacity.com/api/v1/match/batch/{job_id}" \
-H "Authorization: Bearer dm_live_YOUR_KEY"
```
```json theme={null}
{
"job_id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
"status": "running",
"total": 3,
"completed_count": 1,
"error_count": 0,
"error_message": null,
"created_at": "2026-08-17T12:00:00Z",
"started_at": "2026-08-17T12:00:01Z",
"finished_at": null,
"expires_at": "2026-09-16T12:00:00Z"
}
```
| Status | Meaning |
| ----------- | ------------------------------------------------------------------------------------ |
| `queued` | Accepted, not started |
| `running` | Matching in progress (partial results may already be available) |
| `completed` | Finished |
| `failed` | Job-level failure — see `error_message` |
| `cancelled` | Stopped by The Data City before it finished. Rows matched so far are still available |
Poll every few seconds. Avoid tight polling loops.
## 3. Fetch results
Results are available once status is `running`, `completed`, `failed`, or `cancelled`. Requesting them while still `queued` returns `409`.
```bash theme={null}
curl -sS "https://datamatcher-api.thedatacity.com/api/v1/match/batch/{job_id}/results?skip=0&limit=100" \
-H "Authorization: Bearer dm_live_YOUR_KEY"
```
| Query | Default | Notes |
| ------- | ------- | --------------------------- |
| `skip` | `0` | Offset into the result list |
| `limit` | `100` | Page size, max **1000** |
Each item in `data` uses the same fields as `POST /match`. See [Response reference](/company-matching/guides/response-reference).
## 4. Completion webhook
Optional. Pass `callback_url` on submit to receive one `POST` when the job reaches `completed`, `failed`, or `cancelled`. Polling still works.
```bash theme={null}
curl -sS -X POST "https://datamatcher-api.thedatacity.com/api/v1/match/batch" \
-H "Authorization: Bearer dm_live_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"callback_url": "https://example.com/hooks/datamatcher",
"webhook_secret": "at-least-16-chars",
"companies": [
{"input_name": "Sainsbury'\''s Supermarkets Ltd", "id": "demo-001"}
]
}'
```
`callback_url` must be `https` (http is allowed only in local development). Private, loopback, and link-local hosts are rejected outside local development. Cloud metadata addresses such as `169.254.169.254` are always rejected.
The webhook body is the same shape as `GET /match/batch/{job_id}`, plus `event`:
```json theme={null}
{
"event": "job.completed",
"job_id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
"status": "completed",
"total": 3,
"completed_count": 3,
"error_count": 0,
"error_message": null,
"created_at": "2026-08-17T12:00:00Z",
"started_at": "2026-08-17T12:00:01Z",
"finished_at": "2026-08-17T12:00:10Z",
"expires_at": "2026-09-16T12:00:00Z"
}
```
`event` is `job.completed`, `job.failed`, or `job.cancelled`.
If you set `webhook_secret` (16–128 characters), the request includes:
```text theme={null}
X-DataMatcher-Signature: sha256=
```
The hex digest is HMAC-SHA256 of the **raw JSON body** using your secret. Verify against the bytes you received; do not re-serialise the JSON first.
Respond with `2xx`. Delivery is attempted once. A failed webhook does not fail the job — poll status if you miss it.
Retrying misses on a finished job queues it again; a second webhook fires when that run finishes.
## Ownership
Jobs are scoped to the user who owns the API key. Another account's `job_id` returns `404`.
# Errors
Source: https://docs.thedatacity.com/company-matching/guides/errors
Every error code the Company Matching API returns, what it means, and what to do.
The Company Matching API uses standard HTTP status codes. Most error responses include a JSON body with a `detail` field.
## Error envelope
```json theme={null}
{
"detail": "Human-readable description of the error."
}
```
Match on status codes in your error handling, not on the exact `detail` string — wording can change between releases.
**`429` is the exception.** Rate-limit responses from SlowAPI typically use an `error` key instead of `detail`. Read both keys when logging.
## Status code matrix
| Status | When to expect it | What to do |
| ------ | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- |
| `200` | Single match completed (whether or not a CRN was found) | Read `matched_company_number` |
| `202` | Batch accepted and queued | Poll status (or wait for `callback_url`), then fetch results |
| `401` | Missing or invalid API key | Fix or rotate the key — see [Authentication](/company-matching/guides/authentication) |
| `404` | Batch `job_id` unknown or not owned by this key's user | Check the id from the submit response |
| `409` | Batch results requested before the job has started | Wait and poll status until `running`, `completed`, or `failed` |
| `422` | Schema validation failed, batch larger than 1,000, or invalid `callback_url` | Fix the body; shrink the batch; use https for webhooks |
| `429` | Match rate limit exceeded | Back off — see [Rate limits](/company-matching/guides/rate-limits) |
| `500` | Unexpected server error during matching | Retry with exponential backoff |
| `503` | Matching engine not ready | Retry after a short delay |
## Retrying
Retry `429`, `500`, and `503` with exponential backoff:
```text theme={null}
attempt 1 → wait 2 s
attempt 2 → wait 4 s
attempt 3 → wait 8 s
attempt 4 → give up, surface the error
```
Add random jitter (0–500 ms) so concurrent clients do not synchronise retries.
Do **not** auto-retry `401`, `404`, `409`, or `422` with the same payload. Fix the key, id, timing, or body first.
## Common scenarios
### `401` on every call
Almost always a missing `Authorization` header, a JWT instead of a `dm_live_…` key, or a revoked key.
### `422` on batch submit
Either a company object failed validation (for example empty `input_name`), the list exceeds **1,000** companies, or `callback_url` / `webhook_secret` failed validation.
### `409` on results
You called `GET /match/batch/{job_id}/results` while the job was still `queued`. Poll `GET /match/batch/{job_id}` until status is `running`, `completed`, or `failed`, then fetch results.
### Client timeout on `POST /match`
Matching can take tens of seconds on a cold path. Set the client timeout to at least **60–120 seconds** before treating the call as failed.
# Quickstart
Source: https://docs.thedatacity.com/company-matching/guides/quickstart
Get an API key and match your first company in a few minutes.
## 1. Get an API key
Company Matching uses **customer API keys**, not login JWTs, on the match endpoints.
Contact [support@thedatacity.com](mailto:support@thedatacity.com) for a key. Keys look like:
```text theme={null}
dm_live_…
```
Treat the key like a password. Store it in an environment variable or secrets manager. Never commit it to source control or ship it in a browser bundle.
See [Authentication](/company-matching/guides/authentication) for details.
## 2. Match one company
Send a `POST` to `/api/v1/match` with at least `input_name`. Include your key in the `Authorization` header.
```bash cURL theme={null}
curl -sS -X POST "https://datamatcher-api.thedatacity.com/api/v1/match" \
-H "Authorization: Bearer dm_live_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"input_name":"Sainsbury'\''s Supermarkets Ltd","id":"demo-sainsburys"}'
```
```python Python theme={null}
import httpx
resp = httpx.post(
"https://datamatcher-api.thedatacity.com/api/v1/match",
headers={"Authorization": "Bearer dm_live_YOUR_KEY"},
json={
"input_name": "Sainsbury's Supermarkets Ltd",
"id": "demo-sainsburys",
},
timeout=120.0,
)
data = resp.json()
```
```javascript JavaScript theme={null}
const resp = await fetch(
"https://datamatcher-api.thedatacity.com/api/v1/match",
{
method: "POST",
headers: {
Authorization: "Bearer dm_live_YOUR_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
input_name: "Sainsbury's Supermarkets Ltd",
id: "demo-sainsburys",
}),
}
);
const data = await resp.json();
```
Set your client timeout to at least **60–120 seconds**. Matching can call several external sources; cold runs are slower than cached ones.
Optional fields that improve accuracy when you have them:
| Field | Purpose |
| -------------------- | --------------------------------------- |
| `input_url` | Company website |
| `input_postcode` | Improves match scoring |
| `address` | Company address |
| `sic_code` | Improves match scoring |
| `start_trading_date` | Improves match scoring (year is enough) |
| `id` | Your own row id, echoed in the response |
## 3. Read the result
A successful response looks like this:
```json theme={null}
{
"input_name": "Sainsbury's Supermarkets Ltd",
"id": "demo-sainsburys",
"input_url": null,
"input_postcode": null,
"address": null,
"sic_code": null,
"start_trading_date": null,
"matched_company_number": "3261722",
"matched_company_name": "SAINSBURY'S SUPERMARKETS LTD",
"matched_registered_postcode": "EC1N 2HT",
"active": "Active",
"confidence": "high",
"rigorous_match": true,
"route": "fast"
}
```
If no confident match is found, `matched_company_number` (and related fields) are `null`, and `confidence` is `none`. That is still a `200` — check for a CRN rather than assuming every call resolves.
See [Response reference](/company-matching/guides/response-reference) for every field, including how to read `confidence`.
## Matching a list
For more than one company, use the [batch API](/company-matching/guides/batch): submit the list, then poll status or receive a completion webhook.
## What's next
How keys work and how to send them.
Every field in the response, and when it is empty.
Status codes and retry guidance.
Full request/response schema with the interactive playground.
# Rate limits
Source: https://docs.thedatacity.com/company-matching/guides/rate-limits
Per-minute rate limits on Company Matching endpoints.
## Match endpoints
`POST /api/v1/match` and `POST /api/v1/match/batch` are limited to **30 requests per minute** (per authenticated client, via the API key).
When you exceed the limit, the API returns `429`:
```json theme={null}
{
"error": "Rate limit exceeded: 30 per 1 minute"
}
```
Every counted request counts toward the limit regardless of whether matching succeeded.
Status and results polling (`GET /match/batch/{job_id}` and `…/results`) are not subject to the same match limiter as submit. Still avoid busy-looping — poll every few seconds.
## Backing off
Retry `429` with exponential backoff:
```text theme={null}
attempt 1 → wait 2 s
attempt 2 → wait 4 s
attempt 3 → wait 8 s
attempt 4 → give up, surface the error
```
Add random jitter (0–500 ms).
## Batch size vs rate limit
A single batch submit is **one** HTTP request, even if it contains up to 1,000 companies. Prefer batching over many `POST /match` calls so you stay under the per-minute limit.
# Response reference
Source: https://docs.thedatacity.com/company-matching/guides/response-reference
Every field the Company Matching API returns, what it contains, and when it is empty.
A successful match returns a JSON object with the fields below.
## Response fields
These fields are returned on `POST /match` and on each row in batch results.
| Field | Type | When present |
| ----------------------------- | -------------- | -------------------------------------------------------------- |
| `input_name` | string | Always — echoed from the request |
| `id` | string \| null | Your optional row id, if you sent one |
| `input_url` | string \| null | Echoed if provided |
| `input_postcode` | string \| null | Echoed if provided |
| `address` | string \| null | Echoed if provided |
| `sic_code` | string \| null | Echoed if provided |
| `start_trading_date` | string \| null | Echoed if provided |
| `matched_company_number` | string \| null | Set when a match is found; otherwise `null` |
| `matched_company_name` | string \| null | Registered name for the matched CRN, when available |
| `matched_registered_postcode` | string \| null | Registered office postcode for the matched CRN, when available |
| `active` | string \| null | Active / inactive (or equivalent) when known |
| `confidence` | string | Always — `high`, `medium`, `low`, or `none` |
| `rigorous_match` | boolean | Always |
| `route` | string \| null | `"fast"`, `"slow"`, or `null` when unmatched |
A `200` with `matched_company_number: null` means the engine finished without a confident match. That is not an error. In that case `confidence` is `none` and `route` is `null`.
### Confidence
`confidence` tells you how strong the match evidence is. It is **not** a percentage probability.
| Value | Typical meaning |
| -------- | ----------------------------------------------------------------------------------------------------------------------- |
| `high` | Exact name match on the fast path, or a slow-path match that passed the stricter scoring check (`rigorous_match: true`) |
| `medium` | Slow-path match that did not pass the stricter check (`rigorous_match: false`) |
| `low` | Weak evidence for the chosen company, or a band reduced as described below |
| `none` | No CRN returned, or the match failed for that row |
If part of the pipeline was unavailable for that row (for example a website lookup), the band may drop one notch (e.g. `high` → `medium`).
Use `confidence` to triage results — for example auto-accept `high`, review `medium` or `low`. Always check `matched_company_number` to see whether a CRN was returned.
### Route
`route` records which stage produced the match:
| Value | Meaning |
| ------ | ----------------------------------------------------------------------------- |
| `fast` | Exact (normalised) name match against company data — usually quick and strong |
| `slow` | Full multi-source pipeline (search, scoring, optional website signals) |
| `null` | No match |
### Rigorous match
`rigorous_match` is `true` when the slow path cleared the stricter scoring check for the chosen company.
* A fast-path match does not need that check, so it can be `confidence: high` with `rigorous_match: false`.
* On the slow path, `rigorous_match: true` is what lifts a result to `confidence: high`, subject to the band reduction noted above.
Treat `rigorous_match` as a technical signal; prefer `confidence` for product decisions unless you are debugging match quality.
### Example (matched)
```json theme={null}
{
"input_name": "Sainsbury's Supermarkets Ltd",
"id": "demo-sainsburys",
"input_url": null,
"input_postcode": null,
"address": null,
"sic_code": null,
"start_trading_date": null,
"matched_company_number": "3261722",
"matched_company_name": "SAINSBURY'S SUPERMARKETS LTD",
"matched_registered_postcode": "EC1N 2HT",
"active": "Active",
"confidence": "high",
"rigorous_match": true,
"route": "fast"
}
```
### Example (unmatched)
```json theme={null}
{
"input_name": "Completely Unknown Trading Style XYZ",
"id": "row-42",
"matched_company_number": null,
"matched_company_name": null,
"matched_registered_postcode": null,
"active": null,
"confidence": "none",
"rigorous_match": false,
"route": null
}
```
### Example (with address)
```json theme={null}
{
"input_name": "Holmes Joinery Partners LLP",
"id": "demo-holmes",
"input_postcode": null,
"address": "Unit 4, Industrial Estate, Leeds LS1 4DY",
"matched_company_number": "OC123456",
"matched_company_name": "HOLMES JOINERY PARTNERS LLP",
"active": "Active",
"confidence": "high",
"rigorous_match": true,
"route": "slow"
}
```
## Request body fields
| Field | Required | Notes |
| -------------------- | -------- | --------------------------------------------------------------------------------------------------------------------------- |
| `input_name` | Yes | 1–512 characters |
| `input_url` | No | Website URL (up to 2048 characters) |
| `input_postcode` | No | Improves match scoring. Takes precedence over a postcode found in `address` |
| `address` | No | Full address (up to 512 characters). If `input_postcode` is omitted, a UK postcode found in the address is used for scoring |
| `sic_code` | No | Improves match scoring |
| `start_trading_date` | No | Improves match scoring; a year string is enough |
| `id` | No | Your correlation id (up to 128 characters) |
Null optional fields may be omitted from the request; the API does not require them.
# Versioning
Source: https://docs.thedatacity.com/company-matching/guides/versioning
How the Company Matching API is versioned, what counts as a breaking change, and where to find out what changed.
The API carries two version numbers. They answer different questions.
| | What it tells you | Where you see it |
| --------------------------------- | ---------------------------------------------- | ----------------------------------------------------------------- |
| **API version** (`/api/v1`) | The wire contract: paths, fields, status codes | The URL prefix on every endpoint |
| **Release** (for example `1.0.1`) | The software build serving that contract | `info.version` in the OpenAPI spec, and the Swagger UI at `/docs` |
## API version
Every endpoint lives under `/api/v1`. The prefix changes only for a breaking change to the wire contract, which would move to a new prefix such as `/api/v2`.
While the API is in beta, treat the contract as settling rather than frozen. Release 1.0.0 renamed two fields under `/api/v1` without a prefix change — see the [changelog](/company-matching/changelog). Breaking changes are always listed there.
## Release
The release follows [semantic versioning](https://semver.org/) and is published as `info.version` in the OpenAPI spec:
```text theme={null}
https://datamatcher-api.thedatacity.com/api/v1/openapi.json
```
* **Patch** (`1.0.0` → `1.0.1`): fixes and operational changes. No change to the contract.
* **Minor** (`1.0.x` → `1.1.0`): backwards-compatible additions, such as a new optional field.
* **Major** (`1.x` → `2.0.0`): a breaking change to the contract.
A release can ship with no visible change to the API. The changelog says so when that happens.
## What counts as a breaking change
Breaking:
* Removing or renaming an endpoint, request field, or response field.
* Changing the type of an existing response field.
* Tightening validation on a field that previously accepted broader input.
* Changing the JSON error envelope.
Not breaking:
* New endpoints.
* New optional response fields.
* New optional request fields.
* Fixes that bring actual behaviour in line with documented behaviour.
Write your client so that unknown response fields are ignored. That is what makes non-breaking additions safe.
## Discovering changes
* Read the [changelog](/company-matching/changelog) for material changes and the release each shipped in.
* Diff the OpenAPI spec between deploys for a precise record of contract changes.
* Email [support@thedatacity.com](mailto:support@thedatacity.com) with questions about a specific change.
# Company Matching API
Source: https://docs.thedatacity.com/company-matching/index
Match a company name to a UK Companies House registration number.
Send a company name (optionally with a website, postcode, address, SIC code, or trading date) and get back the best matching UK company registration number (CRN), plus the registered name, active status, and a confidence level.
**This API is in beta.** It is stable enough to build against and serves production data. Keys are
issued by hand — email [support@thedatacity.com](mailto:support@thedatacity.com) to get one.
Feedback now shapes what ships next.
## What you get back
| Field | What it is |
| --------------------------------- | ------------------------------------------------------------------------------- |
| **matched\_company\_number** | Companies House registration number (CRN), when a match is found |
| **matched\_company\_name** | Registered name for that CRN |
| **matched\_registered\_postcode** | Registered office postcode for that CRN |
| **active** | Active / inactive (or equivalent) status from our company data |
| **confidence** | How strong the match is: `high`, `medium`, `low`, or `none` |
| **rigorous\_match** | Whether the match passed the stricter scoring check |
| **route** | Which matching path ran: `fast`, `slow`, or `null` when unmatched |
| **Request fields** | Your `input_name`, optional `id`, URL, postcode, address, SIC, and trading date |
Each result includes a confidence level, so you can decide how much to trust it. See the [response reference](/company-matching/guides/response-reference) for what every field contains.
A single `POST /match` returns one result. For lists, submit a batch and poll until it completes, or pass `callback_url` to be notified once.
## Who it's for
Use this API when you need to attach a reliable Companies House identity to customer lists, CRM records, or enrichment pipelines.
## How to get started
Get an API key and match your first company in a few minutes.
Customer API keys (`dm_live_…`) as a Bearer token.
Every field in the response, and when it is empty.
Submit up to 1,000 companies and poll for results.
Status codes, what they mean, and what to do.
Per-minute limits on match endpoints.
The `/api/v1` contract, the release number, and what counts as breaking.
Material changes to the API, newest first.
## Base URL
```text theme={null}
https://datamatcher-api.thedatacity.com/api/v1
```
All match endpoints sit under `/api/v1/match`. The interactive playground on each endpoint page lets you paste your API key and send live requests.
# Classify companies
Source: https://docs.thedatacity.com/global-api/de/classification/classify-companies
/global-api/de.openapi.json post /v1/de/classification
Runs the classification pipeline for the posted filter and training payload.
# Filter companies
Source: https://docs.thedatacity.com/global-api/de/companies/filter-companies
/global-api/de.openapi.json post /v1/de/companies
Returns filtered companies and aggregations as host-defined JSON.
# Get a company
Source: https://docs.thedatacity.com/global-api/de/companies/get-a-company
/global-api/de.openapi.json get /v1/de/companies/{companyNumber}
Returns a single company as host-defined JSON.
# Get a company group
Source: https://docs.thedatacity.com/global-api/de/companies/get-a-company-group
/global-api/de.openapi.json get /v1/de/companies/{companyNumber}/group
Returns group or location rows for a company as host-defined JSON, or 404 when absent.
# Get companies in batch
Source: https://docs.thedatacity.com/global-api/de/companies/get-companies-in-batch
/global-api/de.openapi.json post /v1/de/companies/batch
Returns up to 1000 companies by company number as host-defined JSON elements.
# Get company locations
Source: https://docs.thedatacity.com/global-api/de/companies/get-company-locations
/global-api/de.openapi.json get /v1/de/companies/location/{companyNumber}
Returns one operating-location record as JSON, or 404.
# Get locations in batch
Source: https://docs.thedatacity.com/global-api/de/companies/get-locations-in-batch
/global-api/de.openapi.json post /v1/de/companies/location/batch
Batch lookup by company numbers (plain JSON array body).
# List filters
Source: https://docs.thedatacity.com/global-api/de/filters/list-filters
/global-api/de.openapi.json get /v1/de/filters
Returns filter catalog metadata (dimensions, code trees) as host-defined JSON.
# List dimension values
Source: https://docs.thedatacity.com/global-api/de/lookups/list-dimension-values
/global-api/de.openapi.json get /v1/de/dimensions/{dimension}
Returns a distinct string list for the named dimension when this country exposes it (e.g. `postcode`, `commune`, `state`, `town`).
# Data release
Source: https://docs.thedatacity.com/global-api/de/metadata/data-release
/global-api/de.openapi.json get /v1/de/version
Returns the deployed data release version and API build identity.
# Search companies
Source: https://docs.thedatacity.com/global-api/de/search/search-companies
/global-api/de.openapi.json get /v1/de/search
Performs a semantic search and returns hits with company details.
# Classify companies
Source: https://docs.thedatacity.com/global-api/fr/classification/classify-companies
/global-api/fr.openapi.json post /v1/fr/classification
Runs the classification pipeline for the posted filter and training payload.
# Filter companies
Source: https://docs.thedatacity.com/global-api/fr/companies/filter-companies
/global-api/fr.openapi.json post /v1/fr/companies
Returns filtered companies and aggregations as host-defined JSON.
# Get a company
Source: https://docs.thedatacity.com/global-api/fr/companies/get-a-company
/global-api/fr.openapi.json get /v1/fr/companies/{companyNumber}
Returns a single company as host-defined JSON.
# Get a company group
Source: https://docs.thedatacity.com/global-api/fr/companies/get-a-company-group
/global-api/fr.openapi.json get /v1/fr/companies/{companyNumber}/group
Returns group or location rows for a company as host-defined JSON, or 404 when absent.
# Get companies in batch
Source: https://docs.thedatacity.com/global-api/fr/companies/get-companies-in-batch
/global-api/fr.openapi.json post /v1/fr/companies/batch
Returns up to 1000 companies by company number as host-defined JSON elements.
# Get company locations
Source: https://docs.thedatacity.com/global-api/fr/companies/get-company-locations
/global-api/fr.openapi.json get /v1/fr/companies/location/{companyNumber}
Returns one operating-location record as JSON, or 404.
# Get locations in batch
Source: https://docs.thedatacity.com/global-api/fr/companies/get-locations-in-batch
/global-api/fr.openapi.json post /v1/fr/companies/location/batch
Batch lookup by company numbers (plain JSON array body).
# List filters
Source: https://docs.thedatacity.com/global-api/fr/filters/list-filters
/global-api/fr.openapi.json get /v1/fr/filters
Returns filter catalog metadata (dimensions, code trees) as host-defined JSON.
# List dimension values
Source: https://docs.thedatacity.com/global-api/fr/lookups/list-dimension-values
/global-api/fr.openapi.json get /v1/fr/dimensions/{dimension}
Returns a distinct string list for the named dimension when this country exposes it (e.g. `postcode`, `commune`, `state`, `town`).
# Data release
Source: https://docs.thedatacity.com/global-api/fr/metadata/data-release
/global-api/fr.openapi.json get /v1/fr/version
Returns the deployed data release version and API build identity.
# Search companies
Source: https://docs.thedatacity.com/global-api/fr/search/search-companies
/global-api/fr.openapi.json get /v1/fr/search
Performs a semantic search and returns hits with company details.
# Authentication
Source: https://docs.thedatacity.com/global-api/guides/authentication
Send your API key as a Bearer token. One key works across all four markets.
Every request to the Global Company Data API needs an `Authorization` header carrying your API key
as a Bearer token. There are no unauthenticated endpoints.
## Request format
```http theme={null}
Authorization: Bearer YOUR_API_KEY
```
A complete request:
```bash theme={null}
curl --request GET \
--url "https://global-api.thedatacity.com/v1/us/filters" \
--header "Authorization: Bearer YOUR_API_KEY"
```
## One key, four markets
The same key works for the United States, France, Germany and Ireland. You do not need a separate
key per market, and you do not need to tell us which markets you intend to call.
The market is chosen by the path segment after `/v1/`, not by the key.
Keys are not currently scoped to individual markets. Any valid key can reach all four. If you need
a key restricted to a subset, tell us — it is not something you can configure yourself today.
## Getting a key
API keys are issued by hand. Email [support@thedatacity.com](mailto:support@thedatacity.com) with:
* The name of your organisation.
* The email address of the person who will hold the key.
* Which markets you expect to use.
We provision the key and reply with the value.
Treat your API key like a password. Do not commit it to source control, ship it in a browser
bundle, or paste it into a shared chat. Store it in an environment variable or a secrets manager.
## Replacing a key
Email [support@thedatacity.com](mailto:support@thedatacity.com) to have a key revoked and reissued.
Revocation takes effect immediately, so arrange the swap before you ask us to revoke the old key.
## Authentication failures
A missing, malformed or revoked key returns `401`:
```json theme={null}
{
"type": "https://httpproblems.com/http-status/401",
"title": "Unauthorized",
"status": 401,
"detail": "No Authorization Header",
"instance": "/v1/us/companies"
}
```
The `detail` field distinguishes a missing header from a rejected key. See
[Errors](/global-api/guides/errors) for the full response shape.
## AI assistants
The Model Context Protocol server for each market uses the same key and the same header. See
[Connect an AI assistant](/global-api/index) on the overview page.
# Errors
Source: https://docs.thedatacity.com/global-api/guides/errors
Status codes, the problem-details response shape, and the failure that does not return an error at all.
The API uses standard HTTP status codes. Error bodies follow
[RFC 9457 problem details](https://www.rfc-editor.org/rfc/rfc9457), so they share the same field
names wherever they come from.
## The failure that is not an error
Read this before the status code table, because it is the problem most integrations hit first.
A filter value the market does not recognise does **not** return an error. It returns `200` with an
empty result, exactly as a valid filter matching no companies would.
```json theme={null}
{ "Companies": [], "TotalCount": 0 }
```
A typo in a classification code, a US state name sent to France, or a filter key that market does
not have all produce that same response. Nothing in it tells you the request was wrong.
Call `GET /filters` for the market and use the values it returns. Do not guess codes, and do not
copy them between markets — the classification schemes and location filters differ per country.
See [Compare markets](/global-api/index) for what changes.
If a query returns nothing and you expected results, check your filter values against `/filters`
before assuming the data is missing.
## Status codes
| Status | When to expect it | What to do |
| ------ | ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------ |
| `200` | Request succeeded. An empty result set is still a `200`. | Process the body. If it is empty and you expected records, check your filter values. |
| `400` | Malformed request — bad JSON, or a parameter of the wrong type. | Read `detail` and fix the request. |
| `401` | No API key, or a key that is not valid. | Check the `Authorization` header. See [Authentication](/global-api/guides/authentication). |
| `404` | Unknown path, unknown market code, a wrong HTTP method, or a company number that does not exist. | Check the path, the market code and the method. Only `us`, `fr`, `de` and `ie` are valid. |
| `422` | Well-formed request the service cannot process, such as mutually exclusive filters. | Read `detail` and adjust the body. |
| `429` | Over the rate limit. | Wait for `retry-after`. See [Rate limits](/global-api/guides/rate-limits). |
| `500` | Unexpected server error. | Retry with backoff. If it persists, contact support with the request ID. |
## Response shape
Errors come from one of two places, and the difference shows up in the `content-type`.
### Rejected at the gateway
Requests that never reach the data service are answered by the gateway: `401`, `429`, and the `404`
you get from an unknown path, an unknown market code or a wrong HTTP method. These carry
`content-type: application/problem+json` and include a `trace` object:
```json theme={null}
{
"type": "https://httpproblems.com/http-status/401",
"title": "Unauthorized",
"status": 401,
"detail": "No Authorization Header",
"instance": "/v1/us/companies",
"trace": {
"timestamp": "2026-07-30T10:18:35.786Z",
"requestId": "c73bf67a-8f05-42e7-bee7-bf84094b4b9a",
"buildId": "dbc65d42-8f18-4149-bf68-b85466262259",
"rayId": "a2339ec59d67a62f-MAN"
}
}
```
### Returned by the data service
A request that routes correctly but the service cannot fulfil returns `400`, `422`, `500`, or a
`404` for a company number that does not exist. These carry `content-type: application/json` and the
same problem-details fields, without the `trace` object:
| Field | Type | Meaning |
| ---------- | ------- | ------------------------------------------- |
| `type` | string | URI identifying the kind of problem. |
| `title` | string | Short summary of the problem type. |
| `status` | integer | The HTTP status code, repeated in the body. |
| `detail` | string | What went wrong with this specific request. |
| `instance` | string | The path that produced the error. |
Individual fields can be absent. Read `status` from the HTTP response rather than the body.
`500` responses may have no body at all. Do not assume an error response is parseable JSON —
check the status code first, then the content type.
## Reporting a problem
When contacting [support@thedatacity.com](mailto:support@thedatacity.com), include the
`trace.requestId` if the response had one, or the timestamp and full request path if it did not.
That identifier lets us find the exact request in our logs.
Do not send us your API key.
## Handling errors in code
Treat `429` and `500` as retryable, and everything else in the `4xx` range as a request you need to
fix:
```python theme={null}
response = requests.post(url, headers=headers, json=body, timeout=60)
if response.status_code == 429:
... # wait for retry-after, then retry
elif response.status_code >= 500:
... # retry with exponential backoff
elif response.status_code >= 400:
problem = response.json()
raise ValueError("%s: %s" % (problem.get("title"), problem.get("detail")))
companies = response.json()["Companies"]
```
The `detail` field is written for developers, not end users. Its wording can change between
releases, so log it rather than displaying it in your own interface.
# MCP for AI assistants
Source: https://docs.thedatacity.com/global-api/guides/mcp
Each market runs a Model Context Protocol server, so an assistant can query company data in plain language.
Every market runs a [Model Context Protocol](https://modelcontextprotocol.io) server. Point an MCP
client at it and an assistant can query company data directly, without you writing request code.
## Endpoints
| Market | URL |
| ------------- | ---------------------------------------------- |
| United States | `https://global-api.thedatacity.com/v1/us/mcp` |
| France | `https://global-api.thedatacity.com/v1/fr/mcp` |
| Germany | `https://global-api.thedatacity.com/v1/de/mcp` |
| Ireland | `https://global-api.thedatacity.com/v1/ie/mcp` |
One server per market, each authenticated with the same API key you use for the REST endpoints. See
[Authentication](/global-api/guides/authentication).
Connect more than one server when you want an assistant to compare markets. Each server only knows
about its own country, so comparing the US with Germany means connecting both.
## Connecting a client
Most clients take a URL and a header. For Claude Code:
```bash theme={null}
claude mcp add --transport http tdc-us \
https://global-api.thedatacity.com/v1/us/mcp \
--header "Authorization: Bearer YOUR_API_KEY"
```
Clients that use a JSON configuration file usually take this shape:
```json theme={null}
{
"mcpServers": {
"tdc-us": {
"type": "http",
"url": "https://global-api.thedatacity.com/v1/us/mcp",
"headers": { "Authorization": "Bearer YOUR_API_KEY" }
}
}
}
```
Name each server after its market. An assistant with `tdc-us` and `tdc-de` connected can tell them
apart; two servers both called `tdc` cannot be reasoned about.
## Tools
Each server exposes six tools. They are a deliberate subset of the REST API, not a mechanical
translation of all eleven endpoints — batch and classification endpoints are left out because they
serve bulk pipelines rather than conversational use.
| Tool | What it does |
| ----------------------- | --------------------------------------------------------------------------- |
| `list_filters` | Lists every filter the market supports, with valid values and match counts. |
| `list_dimension_values` | Lists valid values for one location dimension. |
| `search_companies` | Finds companies by name or free text, using semantic search. |
| `filter_companies` | Finds companies matching structured criteria. |
| `get_company` | Retrieves the full record for one company by registration number. |
| `get_company_group` | Retrieves the corporate group around one company. |
`list_filters` is described to the assistant as the tool to call first. The same rule that governs
the REST API applies here: valid values cannot be guessed, and an unrecognised one returns an empty
result rather than an error. See [Errors](/global-api/guides/errors).
## What you can ask
The tools are designed to be composed, so useful questions are the ones that need several steps:
* "How many industrial biotech companies are there in the US compared with Germany?"
* "Find companies in Bavaria with a website and turnover above €10m, then show me the group
structure of the largest three."
* "What filters can I use for France, and which of them do not exist for the US?"
The assistant calls `list_filters` to discover the right codes, then `filter_companies` to run the
query. You do not need to know the classification scheme before asking.
## Rate limits
MCP traffic is limited to **300 requests per minute per key**, higher than the 60 per minute that
applies to the REST endpoints, because a single question can trigger several tool calls.
A throttled tool call comes back as a *successful* MCP result whose content is a `429` document.
The assistant may read that as data and answer from it rather than reporting a failure. If answers
suddenly become vague or contradict earlier ones, check whether you are hitting the limit.
## Limits worth knowing
* The server is read-only. No tool writes, updates or deletes anything.
* A key reaches all four markets, so connecting a market server is a choice about what the assistant
should see, not a permission boundary.
* Tool results count against your quota the same way REST calls do. An assistant exploring a broad
question can spend a lot of calls quickly.
Tell us which questions your assistant handles badly at
[support@thedatacity.com](mailto:support@thedatacity.com). The tool set is curated by hand, so
what agents actually struggle with is what we change next.
# Pagination
Source: https://docs.thedatacity.com/global-api/guides/pagination
Page through company results with Limit and Offset in the request body. Maximum page size is 1000.
`POST /companies` pages its results with two fields in the **JSON request body**, alongside your
filter criteria. They are not query parameters.
## Fields
| Field | Type | Range | Default | Purpose |
| ------------------- | ------- | ------------ | --------------- | --------------------------------------------------- |
| `Limit` | integer | 1–1000 | Service default | Maximum number of companies to return. |
| `Offset` | integer | 0 or greater | `0` | Number of companies to skip before returning. |
| `IncludeTotalCount` | boolean | — | `true` | When `false`, omits `TotalCount` from the response. |
Coming from the [Industry Engine API](/api-reference/guides/pagination)? The fields there are
named `returnCount` and `skip`. This API uses `Limit` and `Offset`. The behaviour is the same.
## Walking a result set
Request the first page:
```bash theme={null}
curl --request POST \
--url "https://global-api.thedatacity.com/v1/us/companies" \
--header "Authorization: Bearer YOUR_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"NAICS": ["541715"],
"Limit": 500,
"Offset": 0
}'
```
The response carries the records and the size of the full result set:
```json theme={null}
{
"Companies": [ "..." ],
"TotalCount": 1284,
"Insights": { "...": "..." }
}
```
Request the next page by advancing `Offset` by the page size:
```json theme={null}
{ "NAICS": ["541715"], "Limit": 500, "Offset": 500 }
```
Stop when you have collected `TotalCount` records, or when a page returns fewer than `Limit`.
## Paging in code
```python theme={null}
def all_companies(session, url, headers, filters, page_size=1000):
offset, collected = 0, []
while True:
body = dict(filters, Limit=page_size, Offset=offset)
page = session.post(url, headers=headers, json=body, timeout=120).json()
companies = page["Companies"]
collected.extend(companies)
if len(companies) < page_size or len(collected) >= page.get("TotalCount", 0):
return collected
offset += page_size
```
## Page size and the rate limit
Your key is limited to 60 requests per minute, so page size is what decides whether a large export
finishes comfortably or spends most of its time throttled. Pulling 50,000 companies takes 50
requests at `Limit: 1000`, and 5,000 requests at `Limit: 10`.
Ask for the largest page you can process. See [Rate limits](/global-api/guides/rate-limits).
## Making large exports cheaper
Two fields turn off work you may not need:
* `IncludeInsights: false` skips aggregation of the insight buckets.
* `IncludeTotalCount: false` skips the total-count query and omits `TotalCount` from the response.
If you set `IncludeTotalCount: false`, you cannot use the total to decide when to stop. Page until
a response returns fewer records than you asked for.
Do not change the filters partway through paging a result set. `Offset` is a position in the
result of the query you send, so editing the filters between pages can skip or repeat records.
## Other endpoints
`GET /search` takes a `limit` query parameter and has no offset — it returns the best matches for a
name, not a walkable list. The batch endpoints take an explicit list of company numbers and return
one record per entry, so they need no paging.
# Rate limits
Source: https://docs.thedatacity.com/global-api/guides/rate-limits
Each key is limited to 60 requests per minute. Exceeding it returns 429 with a retry-after header.
Each API key is limited to **60 requests per minute**. Every request counts as one, whatever the
response status or the size of the result.
This is a different limit from the [Industry Engine API](/api-reference/guides/rate-limits), which
allows 200 requests per minute. If you use both APIs, do not reuse the same backoff settings.
## Hitting the limit
Requests over the limit are rejected at the gateway, before they reach the data service:
```http theme={null}
HTTP/1.1 429 Too Many Requests
content-type: application/problem+json
retry-after: 26
```
The body follows the same problem-details shape as every other gateway error, described in
[Errors](/global-api/guides/errors). The header you need is `retry-after`.
## Backing off
The `retry-after` header gives the number of seconds until your quota resets. Wait that long before
retrying. Do not retry immediately, and do not retry on a fixed short interval — a client that
retries every second while throttled simply stays throttled.
```python theme={null}
import time
import requests
def call(url, headers, body, attempts=5):
for attempt in range(attempts):
response = requests.post(url, headers=headers, json=body, timeout=60)
if response.status_code != 429:
return response
wait = int(response.headers.get("retry-after", 2 ** attempt))
time.sleep(wait)
raise RuntimeError("still rate limited after %d attempts" % attempts)
```
If `retry-after` is absent for any reason, fall back to exponential backoff rather than a fixed
delay.
## Staying under the limit
Most jobs that hit the limit are paginating a large result set one small page at a time. Two things
help more than backoff does:
* **Ask for larger pages.** `Limit` accepts up to 1000 records per request. Ten requests of 1000
cost ten calls; a thousand requests of 10 cost a thousand. See
[Pagination](/global-api/guides/pagination).
* **Turn off work you are not using.** Setting `IncludeInsights` and `IncludeTotalCount` to `false`
makes each request cheaper to serve when you only want the company records.
## Higher limits
If 60 requests per minute is not enough for your use case, email
[support@thedatacity.com](mailto:support@thedatacity.com) and tell us the shape of the workload —
how many records you need and how often. The limit is set per key and can be raised.
# Classify companies
Source: https://docs.thedatacity.com/global-api/ie/classification/classify-companies
/global-api/ie.openapi.json post /v1/ie/classification
Runs the classification pipeline for the posted filter and training payload.
# Filter companies
Source: https://docs.thedatacity.com/global-api/ie/companies/filter-companies
/global-api/ie.openapi.json post /v1/ie/companies
Returns filtered companies and aggregations as host-defined JSON.
# Get a company
Source: https://docs.thedatacity.com/global-api/ie/companies/get-a-company
/global-api/ie.openapi.json get /v1/ie/companies/{companyNumber}
Returns a single company as host-defined JSON.
# Get a company group
Source: https://docs.thedatacity.com/global-api/ie/companies/get-a-company-group
/global-api/ie.openapi.json get /v1/ie/companies/{companyNumber}/group
Returns group or location rows for a company as host-defined JSON, or 404 when absent.
# Get companies in batch
Source: https://docs.thedatacity.com/global-api/ie/companies/get-companies-in-batch
/global-api/ie.openapi.json post /v1/ie/companies/batch
Returns up to 1000 companies by company number as host-defined JSON elements.
# Get company locations
Source: https://docs.thedatacity.com/global-api/ie/companies/get-company-locations
/global-api/ie.openapi.json get /v1/ie/companies/location/{companyNumber}
Returns one operating-location record as JSON, or 404.
# Get locations in batch
Source: https://docs.thedatacity.com/global-api/ie/companies/get-locations-in-batch
/global-api/ie.openapi.json post /v1/ie/companies/location/batch
Batch lookup by company numbers (plain JSON array body).
# List filters
Source: https://docs.thedatacity.com/global-api/ie/filters/list-filters
/global-api/ie.openapi.json get /v1/ie/filters
Returns filter catalog metadata (dimensions, code trees) as host-defined JSON.
# List dimension values
Source: https://docs.thedatacity.com/global-api/ie/lookups/list-dimension-values
/global-api/ie.openapi.json get /v1/ie/dimensions/{dimension}
Returns a distinct string list for the named dimension when this country exposes it (e.g. `postcode`, `commune`, `state`, `town`).
# Data release
Source: https://docs.thedatacity.com/global-api/ie/metadata/data-release
/global-api/ie.openapi.json get /v1/ie/version
Returns the deployed data release version and API build identity.
# Search companies
Source: https://docs.thedatacity.com/global-api/ie/search/search-companies
/global-api/ie.openapi.json get /v1/ie/search
Performs a semantic search and returns hits with company details.
# Global Company Data API
Source: https://docs.thedatacity.com/global-api/index
Company data for the United States, France, Germany and Ireland through one authenticated API.
One API gives you company data for the United States, France, Germany and Ireland. The endpoint
shape stays the same across all four markets. The classifications, identifiers and filters do not.
United States, France, Germany and Ireland.
The same API surface in every market.
Bearer authentication across the whole API.
**This API is in beta.** It is stable enough to build against and serves production data. Keys are
issued by hand, there is no self-service signup yet, and a key currently reaches all four markets.
Tell us what you need at [support@thedatacity.com](mailto:support@thedatacity.com) — feedback now
shapes what ships next.
This API does **not** serve UK company data. For the UK, use the
[Industry Engine API](/api-reference/index).
## Choose a market
Start with the market you want to query. Each tab shows values taken from that market's OpenAPI
document. Run the request at the end of the tab first: it returns the filter values that market
accepts, which you then use to build a company query.
Base path: `/v1/us/`
Primary filter: `NAICS`
Secondary filter: `RNAICS`
Response field: `CompanyNumber`
Top-level region filter: `State`
Location filters: `State`, `Town`, `HQPostcode`
**11 endpoints** and **47 company filter fields**.
Use only classification and location values returned by this market's filters endpoint.
An unrecognised value returns an empty result with `200`, so a typo is indistinguishable
from a query that legitimately matches nothing.
```bash cURL theme={null}
curl --request GET \
--url "https://global-api.thedatacity.com/v1/us/filters" \
--header "Authorization: Bearer YOUR_KEY"
```
Base path: `/v1/fr/`
Primary filter: `NAFs`
Secondary filter: **Not available in this market**
Response field: `Siren`
Top-level region filter: **Not available in this market**
Location filters: `Commune`, `HQPostcode`
**11 endpoints** and **40 company filter fields**.
Use only classification and location values returned by this market's filters endpoint.
An unrecognised value returns an empty result with `200`, so a typo is indistinguishable
from a query that legitimately matches nothing.
```bash cURL theme={null}
curl --request GET \
--url "https://global-api.thedatacity.com/v1/fr/filters" \
--header "Authorization: Bearer YOUR_KEY"
```
Base path: `/v1/de/`
Primary filter: `WZs`
Secondary filter: `RWZs`
Response field: `CompanyNumber`
Top-level region filter: `State`
Location filters: `State`, `Town`, `HQPostcode`
**11 endpoints** and **48 company filter fields**.
Use only classification and location values returned by this market's filters endpoint.
An unrecognised value returns an empty result with `200`, so a typo is indistinguishable
from a query that legitimately matches nothing.
```bash cURL theme={null}
curl --request GET \
--url "https://global-api.thedatacity.com/v1/de/filters" \
--header "Authorization: Bearer YOUR_KEY"
```
Base path: `/v1/ie/`
Primary filter: `NACEs`
Secondary filter: `RNACEs`
Response field: `CompanyNumber`
Top-level region filter: `State`
Location filters: `State`, `Town`, `HQPostcode`
**11 endpoints** and **45 company filter fields**.
Use only classification and location values returned by this market's filters endpoint.
An unrecognised value returns an empty result with `200`, so a typo is indistinguishable
from a query that legitimately matches nothing.
```bash cURL theme={null}
curl --request GET \
--url "https://global-api.thedatacity.com/v1/ie/filters" \
--header "Authorization: Bearer YOUR_KEY"
```
## Make your first request
Use `us`, `fr`, `de` or `ie` after `/v1/` in every path. A request never searches more than one
market.
Pass your key as a bearer token. Select a market above and copy its cURL request.
Requests without a valid key return `401`.
Call `GET /filters` before you filter companies. It returns the classifications, locations and
other values accepted by that market.
Send values returned by `/filters` to `POST /companies`. Do not reuse classification or location
values from another market.
Each key can make 60 requests per minute. A `429` response includes a `retry-after` header.
There is no sandbox key. These docs provide copyable cURL with a placeholder key and do not send
requests from this page.
## Compare markets
Every market exposes the **same 11 endpoints**. What differs is the values they accept,
which is where integrations break: an unrecognised filter key returns an empty result rather
than an error, so a request written for one market can silently return nothing for another.
| | **United States** | **France** | **Germany** | **Ireland** |
| ------------------ | ----------------------------- | ----------------------- | ----------------------------- | ----------------------------- |
| Path segment | `/v1/us/` | `/v1/fr/` | `/v1/de/` | `/v1/ie/` |
| Industry codes | `NAICS` | `NAFs` | `WZs` | `NACEs` |
| Secondary codes | `RNAICS` | **Not available** | `RWZs` | `RNACEs` |
| Company identifier | `CompanyNumber` | `Siren` | `CompanyNumber` | `CompanyNumber` |
| Region filter | `State` | **Not available** | `State` | `State` |
| Location filters | `State`, `Town`, `HQPostcode` | `Commune`, `HQPostcode` | `State`, `Town`, `HQPostcode` | `State`, `Town`, `HQPostcode` |
| Filter keys | 47 | 40 | 48 | 45 |
**France has no region filter and no secondary scheme.** “Not available” is a real answer, not
a gap in this table. France filters location by `Commune` and `HQPostcode`; there is no `RNAF`
equivalent of the other markets' secondary codes. A request carrying `State` against France
matches nothing and returns `200`.
Generated from the four OpenAPI documents on 2026-09-10, so it cannot drift from what the API accepts.
## Common workflows
`GET /filters` lists the valid values for the selected market.
`GET /search` finds companies when you have a name rather than filter criteria.
`POST /companies` filters companies and returns aggregate insights with the results.
`POST /classification` adds market-specific classifications to company records.
## Endpoint reference
Use the market groups in the sidebar for parameters, request bodies and response schemas. Every
reference page comes from that market's OpenAPI document, so it matches what the API accepts.
## Connect an AI assistant
Each market also exposes a [Model Context Protocol](https://modelcontextprotocol.io) server. Use the
same API key and replace `COUNTRY_CODE` with `us`, `fr`, `de` or `ie`:
```text MCP endpoint theme={null}
https://global-api.thedatacity.com/v1/COUNTRY_CODE/mcp
```
Connect more than one market server when you want an assistant to compare countries.
# Classify companies
Source: https://docs.thedatacity.com/global-api/us/classification/classify-companies
/global-api/us.openapi.json post /v1/us/classification
Runs the classification pipeline for the posted filter and training payload.
# Filter companies
Source: https://docs.thedatacity.com/global-api/us/companies/filter-companies
/global-api/us.openapi.json post /v1/us/companies
Returns filtered companies and aggregations as host-defined JSON.
# Get a company
Source: https://docs.thedatacity.com/global-api/us/companies/get-a-company
/global-api/us.openapi.json get /v1/us/companies/{companyNumber}
Returns a single company as host-defined JSON.
# Get a company group
Source: https://docs.thedatacity.com/global-api/us/companies/get-a-company-group
/global-api/us.openapi.json get /v1/us/companies/{companyNumber}/group
Returns group or location rows for a company as host-defined JSON, or 404 when absent.
# Get companies in batch
Source: https://docs.thedatacity.com/global-api/us/companies/get-companies-in-batch
/global-api/us.openapi.json post /v1/us/companies/batch
Returns up to 1000 companies by company number as host-defined JSON elements.
# Get company locations
Source: https://docs.thedatacity.com/global-api/us/companies/get-company-locations
/global-api/us.openapi.json get /v1/us/companies/location/{companyNumber}
Returns one operating-location record as JSON, or 404.
# Get locations in batch
Source: https://docs.thedatacity.com/global-api/us/companies/get-locations-in-batch
/global-api/us.openapi.json post /v1/us/companies/location/batch
Batch lookup by company numbers (plain JSON array body).
# List filters
Source: https://docs.thedatacity.com/global-api/us/filters/list-filters
/global-api/us.openapi.json get /v1/us/filters
Returns filter catalog metadata (dimensions, code trees) as host-defined JSON.
# List dimension values
Source: https://docs.thedatacity.com/global-api/us/lookups/list-dimension-values
/global-api/us.openapi.json get /v1/us/dimensions/{dimension}
Returns a distinct string list for the named dimension when this country exposes it (e.g. `postcode`, `commune`, `state`, `town`).
# Data release
Source: https://docs.thedatacity.com/global-api/us/metadata/data-release
/global-api/us.openapi.json get /v1/us/version
Returns the deployed data release version and API build identity.
# Search companies
Source: https://docs.thedatacity.com/global-api/us/search/search-companies
/global-api/us.openapi.json get /v1/us/search
Performs a semantic search and returns hits with company details.
# Authentication
Source: https://docs.thedatacity.com/instant-classification/guides/authentication
OAuth2 password flow — exchange your email and password for a JWT Bearer token.
The Instant Classification API uses OAuth2 password-flow authentication. You send your email and password to the token endpoint and receive a JSON Web Token (JWT). Include that token as a Bearer credential in every subsequent request.
## Getting a token
Send a `POST` to `/api/v1/login/access-token` with `application/x-www-form-urlencoded` form data:
| Field | Value |
| ---------- | -------------------------- |
| `username` | Your account email address |
| `password` | Your account password |
```bash theme={null}
curl -sS -X POST "https://instant-classification-api.thedatacity.com/api/v1/login/access-token" \
-H "Content-Type: application/x-www-form-urlencoded" \
-d "username=you@yourcompany.com&password=YOUR_PASSWORD"
```
Response:
```json theme={null}
{
"access_token": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...",
"token_type": "bearer"
}
```
## Using the token
Add an `Authorization` header to every API request:
```http theme={null}
Authorization: Bearer eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...
```
Treat your access token like a password. Store it in an environment variable or secrets manager. Never check it into source control or include it in client-side code.
## Token lifetime
Tokens are valid for **8 days** from issue. After that, the API returns **`401`** and you need to request a new token.
There is no refresh-token flow — when your token expires, call `/api/v1/login/access-token` again with your credentials.
Refresh on `401`. A missing, expired or otherwise unusable token all return `401 Could not validate credentials` with a `WWW-Authenticate: Bearer` header.
This changed in July 2026. The API previously returned `403` for an expired token, which meant the usual `if 401: refresh()` branch never fired. If your client was written against `403`, switch it to `401`.
## Testing your token
To verify a token is valid without making a classification request, call:
```bash theme={null}
curl -sS -X POST "https://instant-classification-api.thedatacity.com/api/v1/login/test-token" \
-H "Authorization: Bearer YOUR_ACCESS_TOKEN"
```
A valid token returns your user profile. An invalid or expired token returns `401`.
```json theme={null}
{
"email": "you@yourcompany.com",
"full_name": "Your Name",
"role": "standard",
"is_active": true,
"is_paid": false,
"id": "3fa85f64-5717-4562-b3fc-2c963f66afa6",
"created_at": "2026-07-01T09:12:44Z"
}
```
`is_paid` tells you whether your account is still subject to the free-tier quota. See [Rate limits & quotas](/instant-classification/guides/rate-limits).
## Password recovery
If you forget your password:
1. Send a `POST` to `/api/v1/password-recovery/{email}` with your account email.
2. Check your inbox for a reset link containing a one-time token. The token expires after **48 hours**.
3. Send a `POST` to `/api/v1/reset-password/` with the token and your new password (8–128 characters).
An expired or already-used token returns `400 Invalid token`. Request a new one by repeating step 1.
The API always returns the same response regardless of whether the email exists, to prevent account enumeration.
## Failure modes
| Status | Meaning | What to do |
| ------ | -------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------ |
| `400` | Incorrect email or password. | Check your credentials and try again. |
| `400` | Inactive user. Returned on any authenticated endpoint, not just login. | Contact [support@thedatacity.com](mailto:support@thedatacity.com). |
| `401` | No `Authorization` header, or the token is expired, malformed, or no longer identifies an account. | Request a new token from `/api/v1/login/access-token`. |
| `429` | Too many token requests. See below. | Back off and retry. |
Both `/login/access-token` and `/password-recovery/{email}` are rate limited to **5 requests per minute** in production. Cache your token for its full 8 days rather than fetching a new one per request.
See [Errors](/instant-classification/guides/errors) for the full status-code matrix.
# Errors
Source: https://docs.thedatacity.com/instant-classification/guides/errors
Every error code the Instant Classification API returns, what it means, and what to do.
The Instant Classification API uses standard HTTP status codes. Error responses include a JSON body with a `detail` field describing what went wrong.
## Error envelope
```json theme={null}
{
"detail": "Human-readable description of the error."
}
```
The `detail` wording may change between releases — match on status codes in your error handling, not on the string.
**`429` is the exception.** Rate-limit responses use an `error` key instead of `detail`:
```json theme={null}
{
"error": "Rate limit exceeded: 5 per 1 minute"
}
```
Read both keys when logging errors, or you'll record an empty message for every throttled request.
## Status code matrix
| Status | When to expect it | What to do |
| ------ | --------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| `200` | Successful classification. | Process the response body. |
| `400` | Invalid request: bad JSON, empty body, rejected input, or wrong login credentials. Also returned for an inactive account. | Read `detail`, fix your request, and retry. |
| `401` | No `Authorization` header, or the token is expired, malformed, or no longer identifies an account. | Get a new token from `/api/v1/login/access-token`. See [Authentication](/instant-classification/guides/authentication). |
| `402` | Free-tier usage quota exceeded. | Contact [support@thedatacity.com](mailto:support@thedatacity.com) for a paid plan. See [Rate limits](/instant-classification/guides/rate-limits). |
| `422` | Request failed schema validation: `Text` over 1,000 characters, `Website` over 500, or Terms & Conditions not accepted on signup. | Read `detail` and adjust the request. |
| `429` | Rate limit exceeded, on classification or on the login endpoints. | Back off and retry. See [Rate limits](/instant-classification/guides/rate-limits). |
| `500` | Unexpected server error. | Retry with exponential backoff. If persistent, contact [support@thedatacity.com](mailto:support@thedatacity.com). |
| `502` | The upstream classification engine returned an error. | Retry after a short delay. |
| `503` | Classification engine unreachable or timed out after server-side retries. | Retry after a longer delay (30+ seconds). |
| `504` | Rare. A timeout that surfaced outside the normal retry path. | Treat it like `503`. |
Any credential problem is a `401`: no header, an expired token, a malformed one, or a token whose account no longer exists. Refresh on `401` and retry once.
The API returned `403` for expired tokens until July 2026. If your client still keys on `403`, update it.
## Retrying
Retry `429`, `500`, `502`, `503`, and `504` with exponential backoff:
```text theme={null}
attempt 1 → wait 2 s
attempt 2 → wait 4 s
attempt 3 → wait 8 s
attempt 4 → give up, surface the error
```
Add random jitter (0–500 ms) so concurrent clients don't synchronise retries.
Do **not** auto-retry `400`, `402`, or `422`. The same request will fail the same way. For `401`, get a new token first: retrying with the same one keeps failing.
The API already retries the upstream engine up to four times before it gives up, so a `503` means several attempts have failed. Wait longer than you would for a `500`.
## Common error scenarios
### Classification returns 400
The most likely cause is an empty request body. You must supply either `Text` or `Website` (not both):
```json theme={null}
{"Text": "Software development and consulting"}
```
```json theme={null}
{"Website": "https://example.com"}
```
`Text` is capped at 1,000 characters and `Website` at 500 characters. Exceeding either limit returns `422`.
### Classification returns 402
Your free-tier quota is exhausted. The free tier does not reset — contact [support@thedatacity.com](mailto:support@thedatacity.com) for a paid plan.
### Classification returns 502 or 503
The upstream classification engine is temporarily unavailable. This is usually transient. Wait 30 seconds and retry.
A timeout also arrives as `503`, because the API wraps timeouts in its own retry handling before returning. Don't wait for a `504`.
### Classification hangs, or your client times out
Each upstream attempt gets up to 30 seconds, and the API makes up to four attempts with 1s, 2s and 4s waits between them. A slow request can therefore run for **well over a minute** before it either succeeds or fails.
Set your client timeout to at least **120 seconds** on this endpoint. A short client timeout is the most common cause of failures that look like API errors but are actually the client giving up early.
# Quickstart
Source: https://docs.thedatacity.com/instant-classification/guides/quickstart
Sign up and classify your first company in under five minutes.
This guide takes you from zero to a working classification in three steps.
## 1. Create an account
Sign up at [instant-classification-api.thedatacity.com/signup](https://instant-classification-api.thedatacity.com/signup). You need:
* A **work email address**. Consumer providers (Gmail, Outlook, Hotmail, Yahoo, iCloud, Proton) and known disposable-email domains are rejected.
* A **password** of 8–128 characters.
* To **accept the Terms & Conditions**. Signing up without this returns `422`.
New accounts start on the free tier with a limited request quota for evaluation. See [Rate limits & quotas](/instant-classification/guides/rate-limits) for details.
## 2. Get an access token
Exchange your email and password for a JWT using the token endpoint. The token is valid for **8 days**. Cache it rather than requesting one per call, because this endpoint is rate limited to 5 requests per minute.
```bash cURL theme={null}
curl -sS -X POST "https://instant-classification-api.thedatacity.com/api/v1/login/access-token" \
-H "Content-Type: application/x-www-form-urlencoded" \
-d "username=you@yourcompany.com&password=YOUR_PASSWORD"
```
```python Python theme={null}
import httpx
resp = httpx.post(
"https://instant-classification-api.thedatacity.com/api/v1/login/access-token",
data={"username": "you@yourcompany.com", "password": "YOUR_PASSWORD"},
)
token = resp.json()["access_token"]
```
```javascript JavaScript theme={null}
const resp = await fetch(
"https://instant-classification-api.thedatacity.com/api/v1/login/access-token",
{
method: "POST",
headers: { "Content-Type": "application/x-www-form-urlencoded" },
body: "username=you@yourcompany.com&password=YOUR_PASSWORD",
}
);
const { access_token } = await resp.json();
```
A successful response looks like this:
```json theme={null}
{
"access_token": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...",
"token_type": "bearer"
}
```
Treat your access token like a password. Don't check it into source control or include it in frontend bundles.
## 3. Classify a company
Send a `POST` to `/api/v1/instantClassification` with either `Text` (a company description) **or** `Website` (a URL) — not both. Include your token in the `Authorization` header.
Send **one** of `Text` or `Website` per request. The API forwards both if you send both, which can produce unreliable results.
Set your client timeout to at least **120 seconds** on this endpoint. Classification can take over a minute when the engine is slow, because the API retries upstream before giving up. Most "API errors" reported to support turn out to be short client timeouts.
**Classify by description:**
```bash cURL theme={null}
curl -sS -X POST "https://instant-classification-api.thedatacity.com/api/v1/instantClassification" \
-H "Authorization: Bearer YOUR_ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{"Text": "Software development and consulting services"}'
```
```python Python theme={null}
resp = httpx.post(
"https://instant-classification-api.thedatacity.com/api/v1/instantClassification",
headers={"Authorization": f"Bearer {token}"},
json={"Text": "Software development and consulting services"},
)
data = resp.json()
```
```javascript JavaScript theme={null}
const resp = await fetch(
"https://instant-classification-api.thedatacity.com/api/v1/instantClassification",
{
method: "POST",
headers: {
Authorization: `Bearer ${access_token}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
Text: "Software development and consulting services",
}),
}
);
const data = await resp.json();
```
**Or classify by website:**
```bash cURL theme={null}
curl -sS -X POST "https://instant-classification-api.thedatacity.com/api/v1/instantClassification" \
-H "Authorization: Bearer YOUR_ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{"Website": "https://example.com"}'
```
```python Python theme={null}
resp = httpx.post(
"https://instant-classification-api.thedatacity.com/api/v1/instantClassification",
headers={"Authorization": f"Bearer {token}"},
json={"Website": "https://example.com"},
)
data = resp.json()
```
```javascript JavaScript theme={null}
const resp = await fetch(
"https://instant-classification-api.thedatacity.com/api/v1/instantClassification",
{
method: "POST",
headers: {
Authorization: `Bearer ${access_token}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ Website: "https://example.com" }),
}
);
const data = await resp.json();
```
A successful response includes classifications across all taxonomies:
```json theme={null}
{
"RTICs": [
{
"Code": "10.01.01",
"Description": "Software Development",
"Score": 0.92,
"WordsMatched": 3
}
],
"RSICs": [
{
"Code": "62012",
"Description": "Business and domestic software development",
"Score": 0.88
}
],
"RNAICs": [
{
"Code": "541511",
"Description": "Custom Computer Programming Services",
"Score": 0.85
}
],
"SICs": [
{
"Code": "62012",
"Description": "Business and domestic software development"
}
],
"SimilarCompanies": [
{
"CompanyNumber": "12345678",
"CompanyName": "Example Software Ltd",
"Website": "https://example-software.co.uk",
"Similarity": 0.91
}
],
"TotalWordsProcessed": 5,
"ProcessingTimeMs": 342,
"Website": "https://example.com",
"Description": "Software development and consulting services"
}
```
## What's next
Token lifetime and password recovery.
What each error code means and how to handle it.
Free-tier quota and paid rate limits.
Full request/response schema with the interactive playground.
# Rate limits & quotas
Source: https://docs.thedatacity.com/instant-classification/guides/rate-limits
Free-tier usage quota and per-minute rate limits for the Instant Classification API.
The Instant Classification API enforces two types of limit: a **usage quota** for free-tier accounts and a **per-minute rate limit** for all accounts.
## Free-tier quota
New accounts start on the free tier. Free-tier accounts are limited to **20 classification requests** in total. When your quota is exhausted, the API returns `402`:
```json theme={null}
{
"detail": "Usage quota exceeded. Contact support for elevated access."
}
```
The free tier does not reset — it is designed for evaluation. To continue using the API, contact [support@thedatacity.com](mailto:support@thedatacity.com) for a paid plan.
Only **successful** classifications count against the quota. Requests that fail with `4xx` or `5xx` are recorded but don't consume it, so a run of errors won't burn through your allowance.
The free tier gives you 20 requests to evaluate the API. If you're building an integration, contact support for a paid account before you start development.
## Per-minute rate limit
Free and paid accounts alike are limited to **5 requests per minute**, counted per authenticated user. Requests without a valid token are counted per IP address instead.
When you exceed the limit, the API returns `429` with an `error` key rather than the usual `detail`:
```json theme={null}
{
"error": "Rate limit exceeded: 5 per 1 minute"
}
```
Every classification request counts, regardless of response status or payload size.
No `Retry-After` or `X-RateLimit-*` headers are sent, so you can't read the reset time from the response. Use the backoff schedule below.
### The login endpoints are limited too
`POST /api/v1/login/access-token` and `POST /api/v1/password-recovery/{email}` are also limited to **5 requests per minute** in production. Tokens last 8 days, so fetch one and cache it rather than requesting a token per call.
## How the two limits interact
| Account type | Usage quota | Per-minute rate limit |
| ------------ | -------------------------------- | --------------------- |
| Free | 20 successful requests, lifetime | 5 requests/minute |
| Paid | Unlimited | 5 requests/minute |
If a free-tier account sends requests rapidly, the per-minute rate limit may reject a request before the quota check runs.
## Backing off
Retry `429` responses with exponential backoff:
```text theme={null}
attempt 1 → wait 2 s
attempt 2 → wait 4 s
attempt 3 → wait 8 s
attempt 4 → give up, surface the error
```
Add random jitter (0–500 ms) to prevent thundering-herd retries.
Do **not** retry `402` (quota exceeded). The request will fail the same way until your account is upgraded to a paid plan.
## Checking your limits
Call `GET /api/v1/utils/public-config` to see the current free quota:
```bash theme={null}
curl -sS "https://instant-classification-api.thedatacity.com/api/v1/utils/public-config"
```
```json theme={null}
{
"free_quota_limit": 20,
"free_quota_scope": "lifetime"
}
```
This endpoint is unauthenticated, and it reports the limit rather than your personal usage. To check whether the quota still applies to you, call [`/api/v1/login/test-token`](/instant-classification/guides/authentication) and read `is_paid`. To track consumption, count your own successful responses.
To upgrade to a paid account, contact [support@thedatacity.com](mailto:support@thedatacity.com).
# Response reference
Source: https://docs.thedatacity.com/instant-classification/guides/response-reference
Every field the Instant Classification API returns, what it contains, and when it's empty.
A successful classification returns a single JSON object. Every field is optional: the five classification arrays default to empty, and the scalar fields can be `null`. Write your client so it tolerates missing keys.
## Top-level fields
| Field | Type | What it contains |
| --------------------- | ------- | ------------------------------------------------------------------------------- |
| `RTICs` | array | Real-Time Industrial Classifications matched to the input. |
| `RSICs` | array | Real-time SIC codes. |
| `RNAICs` | array | Real-time NAICS codes. |
| `SICs` | array | Standard Industrial Classification codes. |
| `SimilarCompanies` | array | Companies with comparable activities. |
| `TotalWordsProcessed` | integer | How many words the engine processed. |
| `ProcessingTimeMs` | integer | Server-side processing time in milliseconds. |
| `Website` | string | The website you supplied, echoed back. `null` if you classified by text. |
| `Description` | string | The description you supplied, echoed back. `null` if you classified by website. |
An empty array is a normal result, not an error. If the engine finds no match in a taxonomy it returns `[]` with a `200`, so check array length before indexing.
## Classification entries
`RTICs`, `RSICs` and `RNAICs` share a shape. Only `Code` is guaranteed.
| Field | Type | Notes |
| -------------- | ------- | --------------------------------------------------------------------------------------- |
| `Code` | string | The classification code. Always present. |
| `Description` | string | Human-readable label for the code. May be `null`. |
| `Score` | number | Match strength from the classification engine. Higher is a closer match. May be `null`. |
| `WordsMatched` | integer | **`RTICs` only.** How many input words contributed to this match. May be `null`. |
`SICs` are different, because they are the statutory codes registered at Companies House rather than something the engine infers. They carry no score:
| Field | Type | Notes |
| ------------- | ------ | --------------------------------- |
| `Code` | string | The SIC code. |
| `Description` | string | The official label for that code. |
## Similar companies
Every field here is always present.
| Field | Type | Notes |
| --------------- | ------ | ------------------------------------------------------------- |
| `CompanyNumber` | string | Companies House number. Use this to join to your own records. |
| `CompanyName` | string | Registered company name. |
| `Website` | string | The company's matched website. |
| `Similarity` | number | How close the match is. Higher is more similar. |
The request takes no limit or paging parameters, so the number of companies returned is whatever the classification engine gives for that input.
## Ordering and stability
Array order is whatever the classification engine returns, and we don't guarantee it. Sort by `Score` or `Similarity` yourself if you need a ranking.
Treat scores as subject to change too, since the underlying models are retrained. If you persist results, store the `Code` values and the date you fetched them rather than assuming the same input always produces identical scores.
# Instant Classification API
Source: https://docs.thedatacity.com/instant-classification/index
Classify any company by description or website. Get RTICs, RSICs, RNAICs, SICs, and similar companies in a single call.
Send a company description or a website URL and get back real-time industrial classifications across multiple taxonomies, plus a list of similar companies.
## What you get back
A single `POST` request returns:
| Field | What it is |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |
| **RTICs** | Real-Time Industrial Classifications — The Data City's proprietary codes that classify companies by what they actually do, updated continuously. |
| **RSICs** | Real-time SIC codes — SIC codes refined with live data rather than the static codes on Companies House. |
| **RNAICs** | Real-time NAICS codes — the North American industry classification equivalent, built the same way as RSICs. |
| **SICs** | Standard Industrial Classification codes — the statutory UK codes registered at Companies House. |
| **Similar companies** | Companies with comparable business activities, ranked by similarity score. |
Most classifications carry a match score, so you can decide how much to trust each result. See the [response reference](/instant-classification/guides/response-reference) for what every field contains.
## Who it's for
Use this API when you need to classify companies programmatically — for lead enrichment, portfolio tagging, market sizing, or compliance screening.
## How to get started
Sign up and classify your first company in under five minutes.
OAuth2 password flow — get a JWT, use a Bearer header.
Every field in the response, and when it's empty.
Every error code, what it means, and what to do about it.
Free-tier usage quota and per-minute rate limits.
## Base URL
```text theme={null}
https://instant-classification-api.thedatacity.com/api/v1
```
All endpoints sit under `/api/v1`. The interactive playground on each endpoint page lets you paste your token and send live requests.
# Classify companies
Source: https://docs.thedatacity.com/global-api/de/classification/classify-companies
/global-api/de.openapi.json post /v1/de/classification
Runs the classification pipeline for the posted filter and training payload.
# Filter companies
Source: https://docs.thedatacity.com/global-api/de/companies/filter-companies
/global-api/de.openapi.json post /v1/de/companies
Returns filtered companies and aggregations as host-defined JSON.
# Get a company
Source: https://docs.thedatacity.com/global-api/de/companies/get-a-company
/global-api/de.openapi.json get /v1/de/companies/{companyNumber}
Returns a single company as host-defined JSON.
# Get a company group
Source: https://docs.thedatacity.com/global-api/de/companies/get-a-company-group
/global-api/de.openapi.json get /v1/de/companies/{companyNumber}/group
Returns group or location rows for a company as host-defined JSON, or 404 when absent.
# Get companies in batch
Source: https://docs.thedatacity.com/global-api/de/companies/get-companies-in-batch
/global-api/de.openapi.json post /v1/de/companies/batch
Returns up to 1000 companies by company number as host-defined JSON elements.
# Get company locations
Source: https://docs.thedatacity.com/global-api/de/companies/get-company-locations
/global-api/de.openapi.json get /v1/de/companies/location/{companyNumber}
Returns one operating-location record as JSON, or 404.
# Get locations in batch
Source: https://docs.thedatacity.com/global-api/de/companies/get-locations-in-batch
/global-api/de.openapi.json post /v1/de/companies/location/batch
Batch lookup by company numbers (plain JSON array body).
# List filters
Source: https://docs.thedatacity.com/global-api/de/filters/list-filters
/global-api/de.openapi.json get /v1/de/filters
Returns filter catalog metadata (dimensions, code trees) as host-defined JSON.
# List dimension values
Source: https://docs.thedatacity.com/global-api/de/lookups/list-dimension-values
/global-api/de.openapi.json get /v1/de/dimensions/{dimension}
Returns a distinct string list for the named dimension when this country exposes it (e.g. `postcode`, `commune`, `state`, `town`).
# Data release
Source: https://docs.thedatacity.com/global-api/de/metadata/data-release
/global-api/de.openapi.json get /v1/de/version
Returns the deployed data release version and API build identity.
# Search companies
Source: https://docs.thedatacity.com/global-api/de/search/search-companies
/global-api/de.openapi.json get /v1/de/search
Performs a semantic search and returns hits with company details.
# Classify companies
Source: https://docs.thedatacity.com/global-api/fr/classification/classify-companies
/global-api/fr.openapi.json post /v1/fr/classification
Runs the classification pipeline for the posted filter and training payload.
# Filter companies
Source: https://docs.thedatacity.com/global-api/fr/companies/filter-companies
/global-api/fr.openapi.json post /v1/fr/companies
Returns filtered companies and aggregations as host-defined JSON.
# Get a company
Source: https://docs.thedatacity.com/global-api/fr/companies/get-a-company
/global-api/fr.openapi.json get /v1/fr/companies/{companyNumber}
Returns a single company as host-defined JSON.
# Get a company group
Source: https://docs.thedatacity.com/global-api/fr/companies/get-a-company-group
/global-api/fr.openapi.json get /v1/fr/companies/{companyNumber}/group
Returns group or location rows for a company as host-defined JSON, or 404 when absent.
# Get companies in batch
Source: https://docs.thedatacity.com/global-api/fr/companies/get-companies-in-batch
/global-api/fr.openapi.json post /v1/fr/companies/batch
Returns up to 1000 companies by company number as host-defined JSON elements.
# Get company locations
Source: https://docs.thedatacity.com/global-api/fr/companies/get-company-locations
/global-api/fr.openapi.json get /v1/fr/companies/location/{companyNumber}
Returns one operating-location record as JSON, or 404.
# Get locations in batch
Source: https://docs.thedatacity.com/global-api/fr/companies/get-locations-in-batch
/global-api/fr.openapi.json post /v1/fr/companies/location/batch
Batch lookup by company numbers (plain JSON array body).
# List filters
Source: https://docs.thedatacity.com/global-api/fr/filters/list-filters
/global-api/fr.openapi.json get /v1/fr/filters
Returns filter catalog metadata (dimensions, code trees) as host-defined JSON.
# List dimension values
Source: https://docs.thedatacity.com/global-api/fr/lookups/list-dimension-values
/global-api/fr.openapi.json get /v1/fr/dimensions/{dimension}
Returns a distinct string list for the named dimension when this country exposes it (e.g. `postcode`, `commune`, `state`, `town`).
# Data release
Source: https://docs.thedatacity.com/global-api/fr/metadata/data-release
/global-api/fr.openapi.json get /v1/fr/version
Returns the deployed data release version and API build identity.
# Search companies
Source: https://docs.thedatacity.com/global-api/fr/search/search-companies
/global-api/fr.openapi.json get /v1/fr/search
Performs a semantic search and returns hits with company details.
# Classify companies
Source: https://docs.thedatacity.com/global-api/ie/classification/classify-companies
/global-api/ie.openapi.json post /v1/ie/classification
Runs the classification pipeline for the posted filter and training payload.
# Filter companies
Source: https://docs.thedatacity.com/global-api/ie/companies/filter-companies
/global-api/ie.openapi.json post /v1/ie/companies
Returns filtered companies and aggregations as host-defined JSON.
# Get a company
Source: https://docs.thedatacity.com/global-api/ie/companies/get-a-company
/global-api/ie.openapi.json get /v1/ie/companies/{companyNumber}
Returns a single company as host-defined JSON.
# Get a company group
Source: https://docs.thedatacity.com/global-api/ie/companies/get-a-company-group
/global-api/ie.openapi.json get /v1/ie/companies/{companyNumber}/group
Returns group or location rows for a company as host-defined JSON, or 404 when absent.
# Get companies in batch
Source: https://docs.thedatacity.com/global-api/ie/companies/get-companies-in-batch
/global-api/ie.openapi.json post /v1/ie/companies/batch
Returns up to 1000 companies by company number as host-defined JSON elements.
# Get company locations
Source: https://docs.thedatacity.com/global-api/ie/companies/get-company-locations
/global-api/ie.openapi.json get /v1/ie/companies/location/{companyNumber}
Returns one operating-location record as JSON, or 404.
# Get locations in batch
Source: https://docs.thedatacity.com/global-api/ie/companies/get-locations-in-batch
/global-api/ie.openapi.json post /v1/ie/companies/location/batch
Batch lookup by company numbers (plain JSON array body).
# List filters
Source: https://docs.thedatacity.com/global-api/ie/filters/list-filters
/global-api/ie.openapi.json get /v1/ie/filters
Returns filter catalog metadata (dimensions, code trees) as host-defined JSON.
# List dimension values
Source: https://docs.thedatacity.com/global-api/ie/lookups/list-dimension-values
/global-api/ie.openapi.json get /v1/ie/dimensions/{dimension}
Returns a distinct string list for the named dimension when this country exposes it (e.g. `postcode`, `commune`, `state`, `town`).
# Data release
Source: https://docs.thedatacity.com/global-api/ie/metadata/data-release
/global-api/ie.openapi.json get /v1/ie/version
Returns the deployed data release version and API build identity.
# Search companies
Source: https://docs.thedatacity.com/global-api/ie/search/search-companies
/global-api/ie.openapi.json get /v1/ie/search
Performs a semantic search and returns hits with company details.
# List dimension values
Source: https://docs.thedatacity.com/global-api/us/lookups/list-dimension-values
/global-api/us.openapi.json get /v1/us/dimensions/{dimension}
Returns a distinct string list for the named dimension when this country exposes it (e.g. `postcode`, `commune`, `state`, `town`).
# Data release
Source: https://docs.thedatacity.com/global-api/us/metadata/data-release
/global-api/us.openapi.json get /v1/us/version
Returns the deployed data release version and API build identity.
# Search companies
Source: https://docs.thedatacity.com/global-api/us/search/search-companies
/global-api/us.openapi.json get /v1/us/search
Performs a semantic search and returns hits with company details.