# What is The Data City? Source: https://docs.thedatacity.com/about/what-is-the-data-city Learn more about The Data City platform and what we do. ### About The Data City The Data City helps organisations understand what companies actually do. We built our platform to clear the fog around how the modern economy gets classified, so investors, government teams, and analysts can see it as it actually is. Founded in 2017 and headquartered in Leeds, England, The Data City is a global data provider and SaaS platform built to fix a problem baked into economic data for decades: most business data still relies on classification systems and manual research that don't reflect how modern companies operate. ### The problem with SIC codes Standard Industrial Classification (SIC) codes were never built for how the economy works today. They're assigned once, at incorporation, and rarely updated. A company working in AI, Net Zero, or FinTech can end up filed under a decades-old code that has nothing to do with what it actually does. Multiply that across the economy, and you lose the ability to find, measure, or invest in the sectors that matter most. ### **How Industry Engine solves it** Industry Engine, The Data City's platform, classifies companies in real time using website text and machine learning, not by what SIC code was assigned when a company was set up. That's the basis for our Real-Time Industrial Classifications (RTICs): Over 500 sector classifications covering emerging and fast-moving industries, from AI and FinTech to AgriTech and Net Zero, updated continuously as sectors evolve. Platform users build their own classifications and company lists with Smart Lists, explore individual companies in EXPLORE, spot trends in ANALYSE, and compare sectors side by side in COMPARE. Data comes from a mix of our own classification work and third-party sources, including Companies House, Lightcast, Creditsafe, Specter, Innovate UK, and 360Giving. ### Our story In 2017, the founders of The Data City ran into the limits of SIC codes first-hand. Our CEO, Alex Craven, noticed that his own digital agency, and others doing the same work, weren't showing up in digital agency lists. The reason: there was no SIC code for what they did. That gap became the starting point for The Data City. Alex and the founding team built RTICs and Industry Engine to solve the problem properly, not just for digital agencies, but for every emerging sector traditional classifications miss. The Data City board sep 2023 The Data City board sep 2023 ### **Who we work with** Government departments, local authorities, financial institutions, investors, policy teams, and academic institutions use The Data City to get accurate, up-to-date insight for decisions that matter: where to invest, where to focus policy, and where the next wave of growth is happening. ### **Where we operate today** The Data City started with deep coverage of the UK economy, and that's still where our data runs deepest. But the mission has always been bigger than one country. Today we deliver products and datasets across the US, France, Ireland, and Germany too, with more markets on the way. ### **Our mission** The Data City's mission is to build the new global industrial classification system: replacing outdated codes with data that reflects how companies actually work, wherever they are. # What method do you use for your machine learning classification? Source: https://docs.thedatacity.com/faqs/machine-learning-classification-technique The Data City's classification technique relies on the website text of a training set of companies.
The Data City's classification technique relies on the website text of a training set of companies. Companies like those we want to identify more of are selected, added to the training set, and labelled as includes. Companies unlike those we want to identify more of are selected, added to the training set, and labelled as excludes. The classification process takes all words and word pairs (tokens) from all the websites of the companies in this training set and then represents each company's website as a normalized vector (length 1) of the frequency of these tokens. The vectors for each company are multiplied by a vector (the classifier vector) made up of all the tokens, with each token having a variable score. These word scores are shown in the product UI. The token weights in the classifier vector are varied until the two sets (includes and excludes) of companies in the training set are most highly separated, with as many of the includes as possible having scores above 0 and as many of the excludes as possible having scores below 0. This classifier vector is then used to score all website matched companies for which we have website text. The optimisation step described as "the token weights in the classifier vector are varied until the two sets (includes and excludes) of companies in the training set are most highly separated" can never be performed exactly due to the huge search space. But it can be performed efficiently and with reproducible results in almost all cases using a number of well-developed and tested algorithms. In order to achieve the speed and explainability of our AI system, which is essential to allow sector experts to iteratively improve lists, we have developed our own custom algorithms for this purpose. Our on-going R\&D work focuses on creating: more detailed automated quality assessments of our lists, broadening the results so we miss fewer companies in every sector, and reporting a confidence score for every company's inclusion in a given sector. # Getting started with The Industry Engine Source: https://docs.thedatacity.com/getting-started-with-the-data-city Get to know the core Data City platform and learn the basics of our company data. Once you've created your account and received your login details, you're ready to start using the platform. Here's how to get going. ### Welcome to The Data City Once logged in, the first thing you will see is your dashboard. When you log in, you'll land on your dashboard. From here you can jump into saved searches, browse Real-Time Industrial Classifications (RTICs), and see what's new on the platform. Screenshot 2025-03-20 113716 Any Smart Lists or EXPLORE lists you've saved or shared are accessible from your dashboard too. ### RTICs Real-Time Industrial Classifications, or RTICs, are The Data City's proprietary industry classifications. The platform holds over \[confirm count] emerging economy sector classifications, from AgriTech and Net Zero to AI and FinTech, and it's the best place to start your research. You can see key details for each RTIC from the RTICs tab, including RTIC code, creation date, revised date, companies, description, and verticals. Click the Summary tab to see how each RTIC ranks across the platform. Filter and order sectors by turnover, employees, investment, and growth, or filter by region to spot breakout sectors near you. From there, click through to ANALYSE for a deeper look at sector trends. Screenshot 2025-03-20 141208 Screenshot 2025-03-20 141208 **Tip**: You can also access RTICs directly from the Filters in EXPLORE, to see the full list of companies in a sector. ### **Smart Search** Smart Search lets you find companies using phrases, concepts, or even full paragraphs, not just keywords. If you can describe the kind of company you're looking for in a sentence, Smart Search can find it. It's the fastest way to explore the database when you don't know the exact RTIC or SIC code you need. **Building an Smart list**: Want to find out more? [View our full in-depth, step-by-step Smart List guide today.](/using-industry-engine/tools/building-an-ml-list) 2 1 2 1 ### Explore EXPLORE is where you see the company database, RTICs, and your own lists in full. It shows companies with key data highlighted for each, useful for training a Smart List or vetting prospects at a glance. Screenshot 2025-03-20 114145 Screenshot 2025-03-20 114145 Click “Full Info” on any company for the complete picture: contact information, financial and funding data, growth measures, employee stats, sector classifications, locations, and more. **Note**: Read more about our data, and how to make the most of it, [on our data page](https://help.thedatacity.com/knowledge/our-data). Tailor your search with a wide range of filters, whether you're starting from scratch or refining an existing list, by RTIC, SIC code, location, keyword, financials, and growth measures. Screenshot 2025-03-20 112127-1 Screenshot 2025-03-20 112127-1 **Using EXPLORE**: Find out more about our EXPLORE tool and how it works in our [Using Explore guide](/using-industry-engine/tools/explore). ### Analyse ANALYSE takes your research further with a suite of dashboards and statistics, built for getting a macro view of a sector or list so you can spot trends and opportunities across larger datasets. Every graph and table in ANALYSE can be filtered or edited: change the field, the data format, or the chart type to suit your analysis. **Tip**: In the Locations dashboard, customise tables and graphs by local authority, OECD functional urban area, constituency, ITL1 region, and more. Load Smart Lists directly into ANALYSE, jump into RTICs, or build your own analysis using the same filters as EXPLORE. Download your data for offline use, or view the full company list back in EXPLORE. Screenshot 2025-03-20 141747 Screenshot 2025-03-20 141747 ## **Building a Smart List** Alongside EXPLORE's filters, you can build your own Smart Lists directly in the platform. The Data City's AI helps you build bespoke company lists in minutes, using the same technology behind our expert-backed RTICs. Use Smart Lists to find companies in a niche sector, or upload a list of prospects to find lookalikes. To start, head to My Lists and select “Create a New List.” 1 7 1 7 You'll train the AI with example data: a minimum of five companies similar to the list you're trying to build. Example: Building a FinTech list? Add five FinTech companies you already know, such as Revolut, Monzo, Starling Bank, Klarna, and Experian. 1 8 1 8 Once you've added at least five companies, the platform starts populating your list. The more inclusions and exclusions you make, the sharper your results get. From here, use your Smart List in EXPLORE: view individual companies, refine with filters, download data, or explore insights in ANALYSE. Note: Want the full walkthrough? Read our [step-by-step Smart List guide](https://docs.thedatacity.com/using-industry-engine/tools/building-an-ml-list). ### My Lists Head to My lists for a full view of your lists and searches. Screenshot 2025-03-20 143005 Screenshot 2025-03-20 143005 Create new lists, revisit recent searches, jump into Explore lists, or edit list details directly from here. You can also create folders to organise bigger projects across multiple lists. ### **Working with the data programmatically?** Want to pull data directly into your own systems? The Data City's API gives you programmatic access to Industry Engine, including RTICs, company data, and our Global Company Data API covering the US, France, Germany, and Ireland. Head to the [API quickstart to get set up.](https://docs.thedatacity.com/using-industry-engine/features/api#api-access) ### What's next? Now you know the basics, you're free to explore. Our Knowledge Base has step-by-step guides, use cases, and FAQs for everything else. ### Quick links * [Glossary / Data Dictionary](https://docs.thedatacity.com/data-dictionary) - every data point we hold, in one place * [Using our data safely](https://docs.thedatacity.com/our-data/using-our-data/using-our-data-safely#using-our-data-safely) - things to bear in mind, including a few limitations Need some support? Please don't hesitate to get in touch with your account manager or a member of The Data City team. # The Data City Knowledge Base Source: https://docs.thedatacity.com/index Help, methodology and how-tos for The Data City platform - RTICs, Industry Engine, our data sources and more. **Looking for `help.thedatacity.com`?** You're in the right place. The knowledge base has moved here. ## Start here A guided walkthrough of Industry Engine - list builder, EXPLORE, ANALYSE and RTICs. A short primer on what the platform is and the problems it solves. ## Browse the knowledge base RTICs, key data definitions, proprietary metrics, and third-party data sources. EXPLORE, ANALYSE, COMPARE, Smart Search, Smart lists, filters, downloads and API access. Common questions about classifications, locations, similar companies and more. Information to support tenders that use The Data City data and platform. ## Popular topics ## Need more help? Reach the team for product questions, demos, or to give feedback. # How do you get greenhouse gas emissions data per company? Source: https://docs.thedatacity.com/our-data/faqs/how-do-you-get-green-house-emissions-data We source our greenhouse gas emissions data from official statistics published by the ONS and the Business Register Employment Survey.
The methodology behind greenhouse gas emissions data at the company level is very similar to [GVA's](/our-data/key-data-and-definitions/gva-data). Therefore, the limitations of greenhouse gas emissions data relate to the limitations of GVA estimates per company. The original data we use to calculate the greenhouse gas emissions data [Atmospheric emissions: greenhouse gases by industry and gas](https://www.ons.gov.uk/economy/environmentalaccounts/datasets/ukenvironmentalaccountsatmosphericemissionsgreenhousegasemissionsbyeconomicsectorandgasunitedkingdom) and the [Business Register Employment Survey](https://www.ons.gov.uk/surveys/informationforbusinesses/businesssurveys/businessregisterandemploymentsurvey). These two tables allow us to develop a standard figure of greenhouse gas emissions per employee per SIC. Then, we use the SIC and employee data we have available at the company level to calculate an estimate for each company. Hence, you will need to consider: * **Companies may include international employees in their accounts**, so the total greenhouse gas emissions data may represent emissions produced by employees abroad * **The SIC groupings developed by the ONS and BRES are very broad and we match them down to 5-digit SICs**. This may mean that two companies that do something different may be given the same standard greenhouse gas emissions measure. * **Companies choose more than one SIC**. At this stage, we divide the number of employees of a company equally across all SICs selected by a company. Then, we multiply the split value for each SIC’s greenhouse gas emissions measure. You will want to pay special attention to the SIC some companies select. For example, *SHELL PLC* provides energy products, but they selected *SIC 70100: Activities of head offices*. This means we are estimating Shell's greenhouse gas emissions using the standard figure for *SIC 70100: Activities of Head Offices*, which has a much lower GHG value than one for an energy generation SIC. Hence, GHG emissions estimates may not be representative of a company if a company selects a SIC not related to its actual economic activity. **Limitations**: For more information about our data and its limitations, please make sure you read our [Using Our Data Safely guide](/our-data/using-our-data/using-our-data-safely). # How are 'Similar Companies' identified? Source: https://docs.thedatacity.com/our-data/faqs/similar-companies You can find the most similar companies in the 'Full Info' of a selected company in either EXPLORE or ANALYSE. What does this mean, and how should it be used? We offer two methods for finding similar companies: **Semantic Similarity** and **Composite Similarity**. **Semantic Similarity** leverages cutting-edge LLM-based methods to identifying similarity between companies based on the text on their websites. More detail can be found [here](https://thedatacity.com/blog/building-the-industry-engine-similarity-score-update/). **Composite Similarity** combines the approach of the **Semantic Similarity** with a measure of similarity also based on structured business characteristics, such as sector, location, and employee count. **Quick Guide** * **Semantic** - broader exploration, more cross-sector matches. Useful for finding companies when building a Smart List, etc. * **Composite** - sector-aware recommendations, structural alignment. Useful when you need matches that share fundamental business attributes, i.e. when company size and location are important filters or signals, or you're looking for true industry peers or competitors. Try both and compare if you're unsure. Both return results in the same format. The Similar companies tab in Full Info, with the Semantic and Composite toggle highlighted **Please note**: This methodology is not the same as our classification engine, which we use to build our RTICs. It does not generate a comprehensive list of a sector, instead it only identifies the companies which are most similar to the given company # Why are there duplicate companies? Source: https://docs.thedatacity.com/our-data/faqs/why-are-there-duplicate-companies No companies are duplicated in our database. What may appear as duplicates are actually multiple distinct entities.
The single source of our input data is [Companies House](/our-data/third-party-data/companies-house). Companies will often register multiple entities of what seems like the same company. We hold group structure data to identify where companies are linked to one another in a parent-child relationship. **Tesco as an example** These are two of [Tesco](https://www.tesco.com/) entities: image png Jan 09 2024 03 36 35 6979 PM They are registered as separate entities on Companies House even though you could refer to them as the same company. As you can see in the screenshot below, TESCO STORES LIMITED is a child (or sub-child) of the ultimate parent, TESCO PLC: image png Jan 09 2024 04 26 59 3690 PM We are currently researching a method to effectively collapse all child companies under the ultimate parent company within the platform. You can remove duplicate companies using "Remove subsidiary companies": Screenshot 2025-03-04 at 15.49.36 You can read more about removing subsidiary companies [here](/our-data/using-our-data/what-are-and-how-to-remove-subsidiary-companies). # Why do all companies not have RTICs? Source: https://docs.thedatacity.com/our-data/faqs/why-do-all-companies-not-have-rtics Find out why some companies in our platform are not classified in an RTIC. We use website text data to classify companies into [RTICs](/our-data/rtics/what-are-rtics). Not all companies on the platform have an RTIC associated with them. These are the most common reasons as to why: * We have not yet matched a company in [Companies House](/our-data/third-party-data/companies-house) with a website. We can only obtain text data after successfully linking a company from Companies House to its website. If we do not have website text data for a company, we won't be able to classify it into an RTIC. We have >4.5 million individual companies on the platform, of which >1.2m have a URL match. * We built an RTIC before a company was founded or changed its website text. If we build an RTIC before the foundation of a company or a company's change of processes/technologies, it won't be in the RTIC until the next iteration. * Some companies do not work in the emergent sectors we target. This especially applies to companies working in foundational industries. * Some companies may use relevant processes or technologies but the way they describe them is not similar enough to the companies selected as training data. * They are false negatives and we missed them. We acknowledge some companies' classification falls through the cracks. When we use the term "false negative" we mean the algorithm classified a company as not being relevant for the sector when it should be. Our [quality assurance process](/our-data/rtics/how-do-we-build-rtics) is designed to recover genuine companies that might otherwise be missed, and this is regularly fixed in later iterations of RTIC building and with the help of industry experts and our users. If you find a company that has not been included in an RTIC you can report it on the platform. Reports are what trigger a [Hotfix](/our-data/rtics/updating-rtics#types-of-rtic-update) — a correction released outside the regular update cycle. You can follow the steps shown in the image below: Image Image **RTICs**: Find out more about RTICs and how they are built in our [What are RTICs? guide](/our-data/rtics/what-are-rtics). # Company locations Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/company-locations We use Companies House, company websites, and third-party data to provide company locations — including how we identify verified operating addresses and filter out non-genuine sites. Our operating address definition has changed. In our June 2026 release, operating addresses are evidence-backed locations where we have a signal that a company trades from that site. A location can now be both a registered office and an operating address — companies that operate from their registered office are no longer excluded. We've also introduced a data-driven blacklist to filter virtual offices and formation agents, with exemptions for companies with website evidence of genuine presence. ## Where location data comes from Our location data comes from three sources: * Companies House (registered address) * Company websites * Creditsafe (external provider) ### Companies House For the overwhelming majority of companies, a registered address is provided via Companies House. We use this postcode as the **registered address**. Companies are required to keep their own addresses up-to-date. ### Company websites We extract additional postcodes from the text on company websites. We only look for postcodes on specific pages that are likely to contain correct location information, such as `www.example.com/locations`. The postcodes extracted from websites must match a standard UK postcode format. We then validate these against the ONS Postcode Directory. ### Creditsafe We use additional postcodes provided by Creditsafe. Creditsafe use an external provider for trading addresses. ## Verified operating addresses A location is classified as a **verified operating address** when there is evidence that a company conducts business from that site. This evidence can come from the company's own website or from Creditsafe trading-address data, and helps distinguish genuine trading locations from addresses used only for registration or correspondence. A registered office is not treated as an operating address simply because it is registered with Companies House. When we find evidence that the company trades from that site, the same location can be both the registered office and a verified operating address. Many companies, particularly SMEs, operate out of their registered office. ### How we filter non-genuine operating addresses Some postcodes, particularly those associated with virtual offices, serviced address providers, and professional formation agents, contain an unusually high number of registered companies. To prevent these from appearing as genuine operating locations, we apply a blacklisting rule. A postcode is a candidate for blacklisting if it is an operating address and meets one of the following criteria: * **High density**: The postcode falls within the top 0.5% by registered company frequency. * **Professional services**: The postcode is flagged by Creditsafe as associated with solicitors or accountants, and the company's SIC code matches a relevant professional services classification. ### Exemptions Even if a postcode is blacklisted, an operating address is retained if: * **Website evidence**: The company explicitly lists the address on its own website. * **SIC self-evidence**: The company is itself a solicitor or accountant operating from its registered office. For the most extreme outlier postcodes, website evidence alone is not sufficient to override the blacklist — the address must also be the company's registered office. ## Filtering companies by location When you filter a company list by a location (a local authority, a postcode and radius, or a custom area), we return any company with at least one address in that area. By default this considers **both registered and operating addresses**, so a company can appear on the strength of a single branch in the area even when its registered office is elsewhere. You can narrow this: * **Registered address only**: Match companies whose Companies House registered office is in the area. * **Operating address only**: Match companies with a verified operating presence in the area (an address found on the company's website or via Creditsafe, and not blacklisted), wherever they are registered. A *genuine* operating location is one that is an operating address and is **not** blacklisted. Blacklisting does not remove the operating-address flag; the two are separate, so a virtual-office or formation-agent postcode can be both an operating address and blacklisted. ### Telling addresses apart in a download A download lists **every** address we hold for each company that matched, not only the address that fell inside your filter. So a download can include a company's addresses outside your area. These columns distinguish them: * **Is Registered Postcode**: The Companies House registered office. * **Is Operating Address**: There is evidence the company trades from this site. * **Is Blacklisted**: A likely virtual office or formation agent. Combine with **Is Operating Address** to find genuine operating locations (operating and not blacklisted). * **Is Subsidiary Address**: The address belongs to a subsidiary of the company rather than the company itself. ## Important notes If multiple companies are matched to the same website, they will have the same postcodes extracted from that website. A registered address is *not necessarily* the location of a company's head office — but it can be, and our system now correctly identifies these cases. # Company sizes Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/company-sizes Our company sizes are based on the UK's legislative definition on company sizes.
Here are our company size definitions: | **Company Size** | **Annual Turnover (£)** | **Number of Employees** | **Total Assets** | | ---------------- | ----------------------- | ----------------------- | ------------------ | | Micro | Less than 1m | Less than 10 | Less than £850,000 | | Small | Less than 15m | Less than 50 | Less than £7.5m | | Medium | Less than 54m | Less than 250 | Less than £27m | | Large | More than 54m | More than 250 | More than £27m | Our company sizes are based on the UK’s legislation definition on [company sizes](https://www.legislation.gov.uk/ukpga/2006/46/part/15). We also have a category for no employees or turnover. A given company must meet **at least two** of the above conditions to be categorised accordingly. For example, a medium company may have 20 employees (small), but has £20m in turnover (medium) and £10m in assets (medium). A company must have a registered postcode to be categorised into a company size bracket. # Estimated turnover, employees and growth Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/estimated-turnover-employees-growth We estimate turnover and employees for companies where we have enough data. We calculate growth rates to provide these estimates.
### Intro Our data on company employee counts and company turnovers is provided by CreditSafe and is based on financial reportings to Companies House. We get data, per company, per year. There is often no data or missing data for all or some years. Employee count data is more common than turnover data. Since there is a lag in financial reporting, we always use estimated employees and estimated turnover for the current year's values (this also helps to address missing data). Where we cannot estimate these values, we do not report them. We have developed our own methods for estimating company growth rates even when only limited data is available, for example, where employee counts are reported infrequently or have not been reported recently. We only estimate values where we have enough data to do so reliably. We use our company growth rates to estimate employees and to estimate turnover. You will see our best estimates for current turnover and employee count in the company summaries for roughly half of all businesses. Those businesses without an estimate are very likely to have zero employees. You can expect these estimates to change monthly as these estimates update as we receive more data. ### Growth in our UI The Growth tab shows our estimates of company employees and turnover by year: best estimate employee growth percentage per year and best estimate turnover growth percentage per year. The best way for you to get a feel for our estimation algorithm is to look at the graphs in this growth tab for a few companies that you know well. Screenshot 2025-03-04 at 15.31.13 If a company has reported employee count for three years or more we fit an exponential curve to those years and use this to calculate an annual employee growth rate. We do the same for turnover. Growth rate does not refer to year on year growth, but rather the average growth rate from the curve we fit. A year on year growth rate would not be able to handle missing data as well. Across the product we default to using the employee based growth rate, for example, in filtering. This is to account for inflation. Where a company has never reported employee count we assume zero employees. Where a company has never reported turnover we assume zero turnover. Where a company has reported employee count for only one or two years we estimate an employee count that starts in the first year an employee count is reported and continues at the level of the most recent year up to 2024 (the final year of our projections). **Note**: Companies much more frequently report employee count than turnover. In the case of a company reporting only employee counts we estimate turnover based on the average turnover per employee for that company’s SIC codes. ### Caveats * We constrain projections for sensible results. * The method here works very well for the vast majority of companies. But some companies do very strange things with their annual accounts and these edge case can affect aggregate results especially when those companies are very large. You can read more about how we're addressing this [here](/our-data/proprietary-data/what-are-companies-with-potential-anomalies). # Gross Value Added (GVA) Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/gva-data What GVA data does The Data City have? How should I use this data? What caveats are there? **Please use this metric with caution.** We are happy to talk through your particular use case should you wish. [Get in touch to find out more](mailto:support@thedatacity.com). The Data City estimates Gross Value Added (GVA) at company and RTIC level. This is available in ANALYSE and [EXPLORE](/using-industry-engine/tools/explore). In EXPLORE, you will find GVA information in the financials tab of a company profile. In ANALYSE, you will find GVA data in the analysis summary box. image png Jan 09 2024 11 51 01 5098 AM GVA data in ANALYSE ### How is GVA estimated? We have produced an estimated GVA measure at the company level using official GVA and employment data.This is an estimate of a company's UK GVA, if they have overseas operations. 1. The ONS produces GVA for 104 categories that can be matched to SIC codes. Similarly, the Business Register and Employment Survey (BRES) publishes sectoral employee data that can also be related to SIC codes. With this information, we have produced a measure of average GVA contribution per employee per SIC code. 2. The Data City have employee and SIC information for each company. Our employment data comes from Companies House, but we use profiles data from Lightcast to better estimate how many employees are UK based. It is therefore possible to estimate a GVA value per company by multiplying the company's UK employee count (sourced from Data City) by the average GVA employee contribution associated with that same company's SIC. 3. A company can have multiple SIC codes. Where this is the case, we split employees equally across the SIC codes and multiply each share of employees by the GVA per employee for each SIC code. Expect GVA per employee values to range from £50,000 to £200,000 in magnitude. ### How to use this data? When using GVA data for an RTIC or a list of companies, we strongly recommend reviewing the RTIC or list. You should check that the largest companies, by employee count, are appropriate for your analysis. * For large companies, is this employee count in line with what you would expect for this company - is the employee count accurate? * You may also want to consider whether large companies primarily operate in this sector, or whether their UK operations are not focused on this sector. You may want to remove companies that do not have significant UK operations, or if only a small part of their business falls within the selected sector or RTIC. We try and estimate the UK GVA of companies using data on where employees are located from Lightcast. We have a large coverage for this data, but it's important to note not all companies have this data point therefore it should be used with caution still. We cannot estimate GVA for a company if we don't know which sector (SIC) they're in. There is a small proportion of companies that do not have SICs. Not all companies have GVA data therefore estimated GVA per employee is calculated by dividing the total UK-adjusted GVA (sum of each company's EstimatedGVA multiplied by their UK employee proportion) by the total UK employees from companies that have GVA data available. These values are presented in the pop-up box. > Σ(EstimatedGVA × UKProportion) /Σ(BestEstimateUKEmployees\*) > > \*where companies have EstimatedGVA > 0 You might want to consider the impact that selecting the "exclude all ultimately foreign companies" filter has on the GVA value. Note: this will not remove UK based multinational firms that have large global employment, such as BP. **Note**: You should also be aware that our estimate is of GVA per full-time employee. GVA might be overestimated in sectors which are more likely to have part-time employees, for example, in recruiting agencies. # Industrial Strategy Classifications Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/industrial-strategy-classifications The UK government's Industrial Strategy identifies eight sectors — the IS-8 — as the drivers of UK economic growth over the next decade. The Data City has built precise, data-driven definitions for each, designed to be tracked and evaluated over time. This page explains how those definitions work, why they differ from traditional classification approaches, and what evidence underpins them. *** ## The problem with SIC codes Standard Industrial Classification (SIC) codes are the conventional way of categorising UK businesses. They have three well-documented limitations for Industrial Strategy analysis: **They are backwards-looking.** SIC codes were last updated in 2007. Sectors like AI, quantum computing, or clean energy cannot be reliably identified using them — companies in these sectors typically register under broad codes like "IT consultancy activities". **They only capture primary activity.** A company's SIC code reflects what it mainly does. Dual-use businesses — common in defence and fintech — are systematically miscategorised. A financial firm using AI has a different SIC code from a tech firm providing financial services, even though both are FinTech companies. **Self-selection introduces noise.** Businesses choose their own SIC codes, and many choose incorrectly or not at all. Around 26,000 active companies declare "activities of head offices" — a vague catch-all that obscures what they actually do. The government acknowledges these limitations explicitly. For three of the IS-8 sectors (Clean Energy, Defence, and Digital and Technologies), it states that no reliable SIC-based definition exists. *** ## How The Data City defines the IS-8 The Data City uses two proprietary classification systems to overcome the limitations of SIC. ### Real-Time Industrial Classifications (RTICs) RTICs identify companies operating in sectors that SIC codes cannot capture. They are built using machine learning trained on company website text, combined with expert input from government departments, industry bodies and academics. RTICs are live — they update continuously as companies' activities evolve — and explainable, meaning analysts can inspect why a company has been included or excluded. Frontier sectors within each IS-8 sector are mapped to dedicated RTICs. Many of these have been co-designed directly with government, including DSIT, Innovate UK, and DCMS. ### Real-Time SIC Codes (RSICs) RSICs address the self-selection problem. For sectors where SIC codes exist but are unreliable, RSICs use website text and machine learning to assign companies up to four more accurate SIC codes, correcting for vague declarations and outdated classifications. RSICs are used in sectors like Creative Industries, Financial Services and Professional and Business Services, where the underlying SIC framework is sound but its application by businesses is not. *** ## Confidence ratings The Data City infers a confidence rating for each IS-8 sector based on how well existing government definitions capture the sector's activities. | Rating | Sectors | What it means | | ---------- | ------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------- | | **Low** | Clean Energy, Defence, Digital and Technologies | Government acknowledges no reliable SIC definition exists. Definition relies entirely on RTICs. | | **Medium** | Advanced Manufacturing | A SIC-based proxy exists but misses significant portions of the sector. RTICs provide a more precise view. | | **High** | Creative Industries, Financial Services, Life Sciences, Professional and Business Services | SIC codes broadly capture the sector. RSICs correct misclassifications; RTICs extend coverage to frontier areas. | *** ## The IS-8 sectors ### IS01 — Advanced Manufacturing Covers medium-high technology firms across advanced materials, aerospace, agritech, automotive manufacturing, batteries and space. SIC codes are used by government as a proxy, but analysis shows they identify only 4% of companies in the sector as innovative, compared to 12% under The Data City's RTIC-based definition — a three-fold difference. Both approaches produce a broadly similar overall business count (\~25,000 companies), indicating that the RTIC definition is capturing the right population with greater precision. Frontier RTICs were developed in partnership with DSIT and Innovate UK. Three were co-created directly with government teams. *** ### IS02 — Clean Energy Industries Covers carbon capture, heat pumps, hydrogen, nuclear fission, nuclear fusion, and offshore and onshore wind. The government acknowledges that SIC codes are too restrictive for this sector and currently relies on the Low Carbon and Renewable Energy Economy (LCREE) survey. Survey-based estimates carry an inherent margin of error, making it difficult to attribute changes in sector size to specific policies with confidence. The Data City's RTIC-based definition provides a company-level view that can be consistently monitored over time without survey uncertainty. The Net Zero RTIC underpins this approach and has been adopted in published academic research, including work that informed the Skidmore Review. *** ### IS03 — Creative Industries Covers advertising and marketing, film and TV, music, performing and visual arts, and video games. Also includes **CreaTech** — the 4,879 companies at the intersection of creative industries and digital technology. The government's Creative Industries definition dates to 1998 and is well-established. The Data City aligns with it but uses RSICs to correct misclassifications. Our business count is broadly consistent with official data. CreaTech is defined as companies operating in both the creative industries (excluding software development) and the Digital and Technologies sector. The government is committed to better capturing this subsector through future SIC revisions. *** ### IS04 — Defence Covers companies in land, sea, air, cyber and dual-use defence technologies. The government acknowledges SIC codes are insufficient for this sector, primarily because defence businesses commonly serve both civil and defence markets. A company's SIC code records its primary activity, making dual-use firms invisible in SIC-based analysis. The Data City's methodology combines RSICs and a dedicated Defence RTIC. Company website text captures dual-use activity that a single SIC code cannot. This definition is being expanded in partnership with ADS Group. *** ### IS05 — Digital and Technologies Covers artificial intelligence, cybersecurity, engineering biology, quantum technologies, semiconductors and advanced connectivity. DSIT has acknowledged that SIC codes cannot define this sector with the granularity modern policy requires. Around half of companies in this sector carry SIC codes that fall outside the digital economy as DSIT defines it, meaning a SIC-based approach would miss them entirely. The Data City's definition relies entirely on RTICs. The RTIC-based approach is well-established: it underpins the government's own revised methodology for measuring the UK digital economy, developed with DSIT, Cambridge Econometrics and The Innovation and Research Caucus and published in July 2025. The specific RTICs used here differ from that publication but share the same methodological foundation. RTICs covering AI, Cyber, Quantum and Advanced Connectivity were co-designed directly with government departments. *** ### IS06 — Financial Services Covers asset management, capital markets, FinTech, insurance and reinsurance, and sustainable finance. Financial Services is one of the highest-confidence IS-8 sectors. SIC codes capture most of the sector well; RSICs correct systematic misclassifications in complex group structures where companies declare head office SIC codes, ensuring the sector is accurately sized. Where SIC falls short is at the frontier. FinTech is captured through a dedicated RTIC developed with Innovate Finance, and Sustainable Finance through the Green Finance RTIC, developed with WPI Economics. *** ### IS07 — Life Sciences Covers biopharma and medtech, alongside broader life sciences activity in medical devices, diagnostics and research. The government advocates for the definition created by the Office for Life Sciences (OLS), which publishes an open dataset of companies in Life Sciences frontier sectors. The Data City's definition aligns with this. RSICs correct common misclassifications, and RTICs extend coverage to frontier areas. This makes Life Sciences one of the best-evidenced IS-8 sectors, with an open, shared evidence base already in place. *** ### IS08 — Professional and Business Services Covers accounting and audit, legal services and management consultancy, alongside a wide range of B2B professional activity. Professional and Business Services is a high-confidence sector. RSICs correct the misclassifications that affect SIC-based analysis — particularly the \~26,000 active companies that declare "activities of head offices" but are actually in management consultancy, IT services or other professional activities — ensuring the sector is accurately represented. *** ## Further reading Full methodology, sector comparisons and business count data are available in The Data City's discussion paper: \[Open Sourcing the Industrial Strategy]\([https://thedatacity.com/reports/open-sourcing-the-industrial-strategy/](https://thedatacity.com/reports/open-sourcing-the-industrial-strategy/)) # Location Quotients Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/location-quotients What are location quotients? How do I interpret them? What do I need to know about the data?
### What are location quotients? Location quotients are a measure of relative concentration of an industry in an area.They are useful for comparing areas of different sizes. Traditional location analysis sometimes overlooks the size of the area. Location quotients account for this. ### How are they calculated? Location quotients compare the local presence of an industry with the national presence of the industry. I.e. they measure the presence or size of the industry against what would be expected for an area of this size, based on the national average for the industry. The formula for calculating employee based location quotients is included at the bottom of the page, as an example, to illustrate how location quotients are calculated. ### Example They help us to answer questions such as "Which industries are overrepresented in Leeds?" and “Where is industry X overrepresented?”. Leeds is likely to have a smaller absolute count of businesses than London, but after accounting for the sizes of the geographical economies, location quotients could identify Leeds as having the greater relative concentration of businesses in a specific sector. ### How do I interpret location quotients? Location quotients are a unit-less measure. A value of 1 means that the proportion of companies in industry X in an area is the same as the proportion found nationally. This means you are equally as likely to find industry X in an area as you would across the nation. A value of 2 means you are twice as likely to find industry X in an area as you would across the nation. Conversely, a value of 0.5 means you are half as likely to find industry X in an areas as you would across the nation. Location quotients are found on ANALYSE, under the locations tab. image png Jan 10 2024 11 32 12 2938 AM Above: The top 5 local authorities where Agency Market companies are overrepresented. ### What should I know about the data? The Data City provide location quotients for business count, employees and turnover. Location quotients using employees and turnover are subject to [the wider caveats of our employee and turnover data](/our-data/using-our-data/using-our-data-safely#creditsafe-employee-data-incorrect). 1. Similarly, local authorities that have a very small business base are more prone to very large location quotient values. For example the Isles of Scilly has 82 businesses. Imagine that across the UK that 5% of UK businesses are restaurants. If the Isles of Scilly had 30 restaurants, that would not be a particularly high number of restaurants in absolute terms. In relative terms this would be a large proportion of the overall business base (37%). This would indicate a strong overrepresentation of restaurants in Isles of Scilly, with a location quotient of 7.4. **Some care is required in using location quotients for particularly small local authorities.** 2. Location data can be subject to biases, such as the registered office effect. UK companies of all sizes often register in London despite having minimal *real* activity in London. Large global companies are more likely to be registered in London than other regions. This results in oddities, such as London having large location quotients for tobacco and mining companies. 3. Additionally, multinational companies' employee counts, as reported in their annual accounts, often report [all global employees](/our-data/using-our-data/using-our-data-safely#creditsafe-employee-data-global). This will have an impact on location quotients. #### Location quotient formula example - employees image png Jan 10 2024 03 38 17 8334 PM # Scale-up definition Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/scale-ups Our scale-up definition is the OECD's definition of a scale-up, applied to filed employment data. This page sets out the exact calculation, step by step. **Definition** Our scale-up flag looks at a single growth window: it ends at the company's most recent filed accounts and starts at the most recent filed year at least three years earlier. Within that window the company needs: * 10 or more employees at the start of the window * at least 4 filed years of employment data with 10 or more employees * no rise of more than tenfold between consecutive filed figures * annualised employment growth of 20% or more per year from start to end This follows the Eurostat-OECD definition of a high-growth enterprise, the basis of the term "scale-up", measured on employment over the most recent three years of filed data. The rest of this page sets out exactly how we apply it, including the edge cases, so you can reproduce any company's flag from its filed accounts. ## The data behind the flag * **Filed accounts only.** Employee figures come from companies' filed annual accounts. Each figure is attached to the year the accounts were made up to, giving at most one employment figure per filed year. Our estimated and projected employment series are never used for the scale-up flag. * **Employment, not turnover, by decision.** The OECD definition allows growth to be measured by employees or by turnover. We apply the employment measure only: employee counts are disclosed far more consistently in UK filings than turnover, which smaller companies often do not file. A company scaling revenue on a flat headcount will not be flagged. * **Anomaly filtering.** A figure only counts if it is greater than zero and has not been flagged by our anomaly detection (for example an implausible headcount for the size of the business). Anomalous figures are skipped entirely; they never influence the flag. * **Missing figures are skipped, not zero.** Roughly 4 in 10 filed accounts carry no employee figure at all, most commonly because the filing does not include a captured headcount. A year without a figure is not treated as zero employees; the window simply starts or ends at the nearest year that has one. * **Growth is growth in the filed headcount, however it arose.** We cannot fully distinguish organic hiring from acquisitions, intra-group staff transfers, or changes in how a group allocates employees between its entities; the OECD guidance acknowledges the same limitation. The tenfold cap below removes the worst of it: mis-transcribed figures and switches to consolidated group accounts arrive as huge single-year steps, while 99% of genuine qualifiers never step more than about eightfold between filings. A large step under the cap is still worth checking against the company's accounts before reading it as organic scaling. * **Each release uses the six most recent years of accounts.** The flag is recalculated with every data release; we hold the six most recent years of filed accounts (currently accounts made up to 2020 onwards). A company's flag can therefore change between releases as new accounts arrive or old years leave the held history. ## The calculation, step by step 1. Take the company's reliable filed employment figures (positive, non-anomalous), ordered by year. 2. The window **ends at the most recent** of those figures. 3. The window **starts at the most recent filed year at least three years earlier**. If no filed year is that old, the company cannot qualify yet. 4. The company is a scale-up if all four of these hold: * the start-year figure is 10 or more employees * at least 4 filed years inside the window (start and end inclusive) have 10 or more employees * no figure inside the window is more than ten times the previous filed figure * annualised growth across the window is at least 20% per year Annualised growth is the compound rate between the two window endpoints: $$ \text{annualised growth} = (\text{end employees} \, / \, \text{start employees})^{1/n} - 1 $$ where $n$ is the window length in years (end year minus start year). A company qualifies when this is at least 0.20. No rounding is applied before the comparison, and the years between the endpoints do not enter the growth formula; they matter only for the four-filed-years requirement. The window is fixed by data availability alone: it is always the shortest one the filings allow. The start moves to an older year only when nearer years carry no figure, a filed start is never skipped in search of a better growth rate, and a start below 10 employees fails rather than reaching further back. There is exactly one window to check per company, which keeps the flag focused on recent growth and makes it straightforward to reproduce. ## Worked examples These are real companies, with employee counts exactly as filed in their annual accounts and held in our July 2026 release. You can look any of them up on the platform by company number. Their figures, and in some cases their flags, will change as new accounts arrive. **Steady growth qualifies.** Principle Estate Services Limited (11056986): | Year | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 | | --------- | ---- | ---- | ---- | ---- | ---- | ---- | | Employees | 19 | 30 | 44 | 54 | 68 | 99 | The window runs from 2022, the most recent filed year at least three years before the 2025 accounts. It starts at 44 (10+), has 4 filed 10+ years, and grows $(99/44)^{1/3} - 1 = 31.0\%$ per year. Scale-up. **Growth must be recent.** Olfasense UK Ltd (02900894): | Year | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 | | --------- | ---- | ---- | ---- | ---- | ---- | ---- | | Employees | 22 | 22 | 46 | 44 | 47 | 48 | The window runs from 2022 to 2025 and annualises at $(48/46)^{1/3} - 1 = 1.4\%$ per year: not a scale-up. The company more than doubled between 2021 and 2022, and a window drawn from 2021 would average $(48/22)^{1/4} - 1 = 21.5\%$ per year, but 2022 is a filed year, so the window starts there. An early jump cannot carry a recent plateau. **The flag follows the window as new accounts arrive.** UD Restaurants Ltd (10515301): | Year | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 | | --------- | ---- | ---- | ---- | ---- | ---- | ---- | | Employees | 19 | 26 | 48 | 40 | 41 | 45 | When the 2023 accounts were the latest, the window ran from 2020 and annualised at $(40/19)^{1/3} - 1 = 28.2\%$ per year: a scale-up. With the 2024 accounts the window moved to 2021 and fell to $(41/26)^{1/3} - 1 = 16.4\%$. With the 2025 accounts it moved to 2022 and fell to $(45/48)^{1/3} - 1 = -2.1\%$. The flag dropped as soon as the growth stopped being recent. **Years without figures are bridged.** TSC Kent Ltd (10853210): | Year | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 | | --------- | --------- | ---- | --------- | ---- | ---- | ---- | | Employees | no figure | 16 | no figure | 37 | 36 | 38 | Three years before the 2025 accounts is 2022, which carries no figure, so the window starts at the next older filed year, 2021. It starts at 16 (10+), has 4 filed 10+ years (2021, 2023, 2024, 2025), and grows $(38/16)^{1/4} - 1 = 24.1\%$ per year. Scale-up. The missing year still counts towards elapsed time (we divide over 4 years, not 3 observations), so bridging never inflates a growth rate. **Bridging stops at the nearest filed year.** Royal London Asset Management Limited (02244297): | Year | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 | | --------- | ---- | ---- | --------- | --------- | ---- | ---- | | Employees | 221 | 384 | no figure | no figure | 534 | 591 | The 2022 and 2023 accounts carry no employee figure, so the window starts at 2021 and annualises at $(591/384)^{1/4} - 1 = 11.4\%$ per year: not a scale-up. A window from 2020 would average $(591/221)^{1/5} - 1 = 21.7\%$, but 2021 is a filed year and is never skipped in search of a better rate. **Implausible steps are rejected.** A real series, anonymised: a national charity that employs around 5,000 people, whose early years were transcribed wrongly at source: | Year | 2020 | 2021 | 2022 | 2023 | 2024 | | --------- | ---- | ---- | ---- | ---- | ---- | | Employees | 50 | 56 | 5193 | 5165 | 5163 | The window from 2021 to 2024 would annualise at $(5163/56)^{1/3} - 1 = 350\%$ per year, but the 2021 to 2022 step is a 93-fold rise: far beyond anything hiring can do, and the signature of a data error or a switch to consolidated group accounts. Any rise of more than tenfold between consecutive filed figures inside the window disqualifies the company. Steps under the cap pass: 99% of genuine scale-ups never exceed about eightfold. **Reaching 10 employees only recently is not enough.** Alexa Capital Limited (10759666): | Year | 2020 | 2021 | 2022 | 2023 | 2024 | | --------- | ---- | ---- | ---- | ---- | ---- | | Employees | 3 | 3 | 6 | 8 | 10 | The window starts at 2021, which has 3 employees: below the 10-employee floor, so the company is not a scale-up, however fast it is growing. This is the OECD's own threshold, which stops very small bases producing inflated growth rates. The floor applies at the actual window start; a start below 10 is never bridged past. ## Reproducing the flag from an export Platform exports that include the year-by-year financial history (`CompanyFinancialsCreditSafe`) contain everything the flag is computed from: 1. Take the `Reported_Numberofemployees` figures by year, dropping years where the figure is missing or zero, or where `DeclaredEmployeesAnomalous` is true. 2. The window ends at the latest remaining year and starts at the most recent remaining year at least three years earlier. 3. Check the four criteria: 10 or more employees at the start, four 10-or-more years inside the window, no figure more than ten times the previous filed figure inside the window, and annualised growth of at least 0.20 between the two endpoints, using the formula above. If a recomputation disagrees, the usual causes are skipping the anomaly filter in step 1 or comparing across releases: the platform recalculates with every data release, and both figures and flags move as new accounts arrive and old years leave the six-year history. ## Background The definition of an OECD scale-up company is what we have implemented on the platform. This is the OECD definition ([source](https://committees.parliament.uk/writtenevidence/109251/pdf/)): "*All enterprises with average annualised growth greater than 20% per annum, over a three-year period should be considered as high-growth enterprises. Growth can be measured by the number of employees or by turnover.*" Our general growth-rate metric and company size definitions are different: those do use estimated and projected figures. You can read more about those estimates and average annual growth rates [here](https://thedatacity.com/blog/focusing-on-company-growth/). The scale-up filter is located within the growth tab: image png Mar 04 2025 03 56 32 2701 PM Using the OECD's definition to determine a company's growth stage enhances the ease of making global comparisons. The OECD does not provide a definition for a start-up. We appreciate there are many different definitions of company size and company growth stages, and which one is right for you will depend on the purpose of your analysis. Our platform still allows for custom definitions using the filters bar, particularly within the financial tab: image png Mar 04 2025 03 57 10 6041 PM # Website Matching Source: https://docs.thedatacity.com/our-data/key-data-and-definitions/website-matching Assigning websites to companies is a foundation of what we do at The Data City. The more accurate and precise our website matching is, the better our Real-Time Industrial Classifications (RTICs), Real-Time Standard Industrial Classifications (RSICs), and your Smart Lists are. ### The short version From a range of sources, we assemble a set of potential websites for each company. We call these *candidates*. We then scrape up to 25 pages of each candidate website. Finally, we apply a machine learning model to select the best candidate for each company. This model is trained on manually checked website matches for thousands of companies collected over the last decade. We measure the success of our model continuously. The two most important metrics are: * *Accuracy*: how often we pick the correct website, or no website if the company does not have one. * *Precision*: how often we pick the correct website, or no website (regardless of whether the company does actually have a website). **Accuracy** tells us how likely a website match that you see in our product is to be correct. **Precision** is a broader measure. It considers that sometimes we won’t match small or newly incorporated companies with basic websites to that website. In our latest releases our model regularly exceeds 90% accuracy, and 97% precision. ### The long version V6 of our industry engine represented our biggest ever step forward in website matching. We switched from a logical scoring method to a machine learning model trained on top of year's worth of manually collected data. You can read about this in more detail in our blog post *[Industry Engine V6: What's new?](https://thedatacity.com/blog/industry-engine-v6-whats-new/)* # Environmental, Social and corporate Governance (ESG) Source: https://docs.thedatacity.com/our-data/proprietary-data/esg-statements We identify sub-pages of company websites which likely relate to ESG ## Why we provide ESG data ESG refers to a set of standards used to measure an organisation's environmental and social impact. The increasing importance of ESG made us curious about what we could contribute. Since we scrape company websites, it is possible for us to indentify pages of a company's website that are likely to contain ESG statements. ## How we indentify ESG statements When we scrape a company's website, we collect information from up to 25 pages of the website. We compare each URL that we scrape to a list of ESG keywords. If the URL contains any of these words, we assume that the web content for that sub-page relates to ESG. A few examples of the words we're looking for are: *sustainable*, *environment* and *gender*. In addition, we check the internal links on each scraped page for any further matches to our list of ESG keywords. ## Where to find ESG data On a company page, click the `ESG (beta)` tab. If we have identified any ESG statement pages for that company, they will appear here. ESG example ## Next steps This data is currently released under beta version status. That means it could contain some errors. We like to ship early and often, and we will continue to improve this data. # Gender Data Source: https://docs.thedatacity.com/our-data/proprietary-data/gender-data What gender data is available? How do The Data City estimate the founders of companies? What do I need to know to use this data? **New in the June 2026 update** — gender leadership is now a single, mutually-exclusive category per company (six in total), applied separately to directors and founders, with a symmetric two-thirds threshold for a "majority". Every company is placed in exactly one category, so the counts aggregate cleanly. ### What is available and how are founders identified? The Data City's platform can analyse the leaders and founders of a company. Leaders are the active directors (officers) listed on Companies House. We place every company in one of six mutually-exclusive gender categories based on its active directors — and, separately, on its founders (see below). We identify founders by looking at directors appointed **within 2 years and 28 days** of a company's incorporation. We add a 28-day buffer to account for standard administrative reporting delays permitted under UK corporate law. Specifically, the [Companies House "14+14" rule](https://www.gov.uk/guidance/people-with-significant-control-pscs) grants a company up to 14 days to update its internal register of People with Significant Control (PSC) following an appointment, and an additional 14 days to formally submit this information to the public register. Founders are specifically a **person with significant control** and they are *still* an **active director** at the company. A [person of significant control](https://www.gov.uk/guidance/people-with-significant-control-pscs) is defined by Government. In short it means a person likely has voting rights, shares, or a controlling influence in a company. Companies are required to declare persons of significant control. It is possible to have information protected, however. ### How we categorise leadership and founders Gender is assigned from each officer's Companies House title (e.g. Mr, Mrs). Titles without a gender (e.g. Dr, Prof) count as **unknown** and are still included in the total. The Data City **does not** use machine learning, and never uses names, to estimate gender. Each company is placed in exactly one of six categories, based on the share of active directors with a female-gendered title (`X`) or a male-gendered title (`Y`): | Category | Definition | | ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **All women led** | `X = 100%` — every active director has a female-gendered title | | **Majority women led** | `66.7% < X < 100%` — more than two-thirds, but not all | | **Mixed led** | `33.3% ≤ X ≤ 66.7%` — between a third and two-thirds have a female-gendered title | | **Majority men led** | `66.7% < Y < 100%` — more than two-thirds, but not all, have a male-gendered title | | **All men led** | `Y = 100%` — every active director has a male-gendered title | | **Uncertain** | none of the above — including companies with no active directors, and boards where unknown-gender titles mean the category can't be determined without guessing | The percentages are rounded for display; the boundaries are applied as exact thirds, so there is no rounding drift. A clear majority (more than two-thirds, but not all) is only possible on a team of four or more, so most companies fall into *All*, *Mixed* or *Uncertain* rather than a *Majority* category. The same six categories are applied separately to a company's **founders** (using the share of founders with each title). Because the categories are mutually exclusive, every company sits in exactly one — so you can aggregate them safely (for example, "women-led businesses are X% of companies"). A company is **women led** when it is *All women led* or *Majority women led* — i.e. more than two-thirds of its active directors have a female-gendered title. **Why "Uncertain"?** If a company has, say, two women, four men, and one director with an unknown-gender title, that final title could tip the company into either *Mixed led* or *Majority men led*. Because we never guess gender from a name, the honest answer is *Uncertain*. In some instances we are not able to identify founders, or identify the genders of founders. Firstly, we rely on persons of significant control from Companies House. Persons of significant control is legislation that was introduced in 2015 and our ability to identify founders before then is reduced. An example is included in [appendix one](#appendix-one). In addition, sometimes the true founders of a company are not listed as persons of significant control, because of their initial level of equity in a company not meeting the threshold for significant control. Lastly, we are not able to identify the gender of a founder (or leader) where they have used a unisex title — these companies fall into the *Uncertain* category. In both leader and founder data, gender is based on the declared titles of officers on Companies House. Data City **does not** use machine learning to estimate the gender. You should be careful when using this data to compare summary statistics between men and women led (or founded) businesses. We recommend you [remove outliers](/our-data/proprietary-data/what-are-companies-with-potential-anomalies) from your analysis. ### Where is the data? #### ANALYSE In ANALYSE, you can find gender data in the analysis summary box and in the company details panel. In the company details panel, the **Gender** section shows two 100% stacked bars — the founder gender mix and the director gender mix — each split into the six categories above. This communicates the proportion of companies in your list that are women led, men led, mixed, or uncertain, for both founders and directors. #### EXPLORE In EXPLORE, you can find gender data in the people tab of each company page, under the **Woman led statistics** header. This includes the women founder and women officer counts, whether the company is women led, and the company's director and founder gender categories. #### Filtering In ANALYSE and EXPLORE you can filter companies by any of the six categories — separately for directors (leaders) and for founders. Ticking more than one category in a group returns companies in any of them. These options are available under the company filter. In [November 2023](https://thedatacity.com/blog/new-founder-gender-data-in-platform/#:~:text=The%20women%20led%20statistics%20here,a%20majority%20women%20led%20business.) we wrote an article detailing definitions, coverage and drawbacks of this approach. ### Appendix one: ANN SUMMERS LTD. (01034349) As per Ann Summers' website: > 10th December 1971, the first Ann Summers shop opened in Marble Arch. Their Executive Chair, Jacqueline Gold, joins her dad's business as an intern with a brilliant idea - Tupperware parties, but for Ann Summers Product. The Ann Summers party is born. On Companies House, their incorporation date is [10 December 1971](https://find-and-update.company-information.service.gov.uk/company/01034349). However, on Companies House, their earliest officer (director) is David Gold who was appointed "before 1991":\ image png Sep 17 2025 12 28 42 9112 PMThere are two issues here for founder analysis using Companies House data: 1. We do not know the exact year David was appointed as a director. 2. The incorporation date is 20 years before the director appointment date. So although it's quite clear from the website, it's unclear using the data hence why we are unable to identify any founders regardless of gender: Image # Innovation Score Source: https://docs.thedatacity.com/our-data/proprietary-data/innovation-score What does the Innovation score represent, and how is it calculated?
### The short version: Our Innovation indicator uses a proprietary Machine Learning model to estimate how innovative every company in our database is based on their websites, where available. In the absence of any better method, we use R\&D intensity as a proxy for innovation. image png Dec 20 2023 11 28 01 2716 AM The model is trained on 980 companies with known R\&D intensities (R\&D £ expenditure per employee) and applied to 1.6 million companies in the UK to estimate whether they are innovative or not. A 3-star rating system is used to indicate our confidence in the estimation. ### The long version: Defining and measuring "innovation" is difficult. Data which indicates whether a company is innovative or not does not exist for all UK companies, and predicting unknown innovation is also tricky. The Central Bureau of Statistics of the Netherlands (CBS) have shown that [the website text of a company can accurately predict its innovativeness score.](https://www.cbs.nl/-/media/innovatie/using-website-texts-todetect-innovative-companies.pdf) But obtaining the training data to replicate this in the UK is not straightforward. Where proxy data which could be used to estimate unknown company innovation does exist, it is either private (see [the ONS UK Innovation Survey](https://www.ons.gov.uk/surveys/informationforbusinesses/businesssurveys/ukinnovationsurvey)), or difficult to capture. R\&D spending, a proxy for innovation, is not required in annually reported accounts, and is rarely voluntarily reported. Where it is reported, it is regularly marked improperly, rendering the field in the machine-readable XBRL format accounts unreadable. After parsing 1.5TB of machine-readable accounts in XBRL format and experimenting with OCR at scale, we have managed to capture R\&D spending data for 980 UK registered companies. These companies operate across all regions of the UK and all industrial sectors, and cover a wide range of R\&D spending and business sizes. From this we have calculated R\&D intensity (R\&D spending £ per employee). Combined with company website text, this provides a solid source of training data from which we have developed a Machine Learning method to estimate if a company is significantly more likely to be highly innovative, based only on the content of its website. A measure of 0-3 stars is then applied based on our confidence in the innovation likelihood predicted (0 stars - not innovative; 1 star - innovative, low confidence; 2 star - innovative, medium confidence; 3 stars - innovative, high confidence). Though we do not recommend users use the innovation indicators to build lists, this star rating system allows users to filter for innovative companies within the entire company database, or within specific lists/RTICs. **Filtering companies by innovation score is included in both our EXPLORE and ANALYSE platforms.** #### The Innovation Score appears as a number and not a category. What is this? If you're using the API, you will see the raw innovation score. To convert the raw score into the confidence rating mentioned above, use the following logic: 3 star = score >= 3 2 star = 1.5 >= score \< 3 1 star = 0 \< score > 1.5 # Linked companies Source: https://docs.thedatacity.com/our-data/proprietary-data/linked-companies ## What are linked companies? Linked companies data is an extension of our group structure data. Sets of companies can form “groups” that are connected, but are registered as separate companies for financial, legal and organisational reasons. Group structure data explains how the sets of companies are connected, the structure of their connections, and the directors that connect them. Creditsafe uses a rules-based criteria to identify companies that form these groups. However, sometimes companies can be “linked” together, but are not marked as being part of the same group structure by Creditsafe because they do not meet their criteria. By eye, we can see that these companies are linked in some way; usually it is because they have similar (or identical) company names, websites, directors, shareholders, registered addresses, or SIC codes and more. Here is an example. The company [**A Shade Greener Limited**](https://products.thedatacity.com/companypage/?company_id=06922318) is not part of a group structure. It is its own ultimate parent company. A Shade Greener Group Structure Our linked companies tab shows that there are in fact 63 companies linked to **A Shade Greener Limited**. All of these share the same postcode, and many have similar names. A Shade Greener Linked Companies For example, [A Shade Greener Member LLP](https://products.thedatacity.com/companypage/?company_id=OC386772). Both companies are registered at the same address *STERLING HOUSE MAPLE COURT, MAPLE ROAD, S75 3DP* and have the same person of significant control *Mr Stewart James Davies* born in 1951. We think these companies are likely to be part of the same organisation, even though our group structure data does not show this link. When these type of companies are *not* considered to be in a group structure, we can end up “double-counting” companies that are effectively one economic unit. This can make economic analysis more challenging by inflating the figures. ## How do we identify Linked Companies? We use a machine learning model trained on hundreds of thousands of data points to estimate the probability that any two companies are linked. We consider companies to be linked if they exceed 97% probability. The data behind this model includes * company names * company addresses * company websites * director names * director addresses * director birthdates * persons of significant control (shareholders) and more. These data points are used to find potential links. We then train our model by providing it pairs of companies that are extremely likely to be linked. ## Where can I find Linked Companies data? You can see a company's *linked companies* by clicking on the "Linked Companies" tab on a company page. We currently have around 460k "clusters" or "groups" of linked companies, which contain approximately 1.5million companies in total. As of Industry Engine v6.1, this data is released under "beta" version status, which means it could contain some errors. We are actively working to keep improving this data. ## How we're using Linked Companies data When counting companies in a sector, multiple registered companies can represent the same economic entity, for example, if they share a director, location, or website. Linked Companies identifies these relationships. Distinct Brands already aimed to solve this by counting the likely unique economic entities rather than raw company counts. It's now powered by Linked Companies data: where a group of linked companies exists, only the earliest-incorporated one counts as a Distinct Brand. This makes Distinct Brands a more accurate measure of the true number of independent companies in your analysis. Image # What are companies with potential anomalies? Source: https://docs.thedatacity.com/our-data/proprietary-data/what-are-companies-with-potential-anomalies The Data City has identified companies with accounts that are likely misrepresented. We will answer: Why have we added this feature? What are some examples of anomalies in accounts? What is the basis for the predicted anomalies?
### Why have we added this feature? Our data covers all active companies registered at Companies House. As well as a breadth of data, we have depth of data. We have the ability to drill down to company level financials. To offer this depth of data, Companies House data uses the submission of financial accounts by each company. This is a mandated process. However, self-declared (and especially unaudited) company accounts can contain mistakes. Companies House are not responsible for verifying company accounts: *"We carry out basic checks on documents received to make sure that they have been fully completed and signed, but we do not have the statutory power or capability to verify the accuracy of the information that companies send to us."* \~[Companies House](https://resources.companieshouse.gov.uk/serviceInformation.shtml) Outliers affect less than 0.05% of our companies but their impact, by nature, can be large. Identifying possible outliers will allow our subscribers to review and remove the very small number of companies that have bad data. ### What are some examples of anomalies in accounts? Below is a **non-exhaustive** list of anomalies that can occur in financial accounts. #### Companies can report another financial variable as their employees. When a company is filling in their accounts, in some instances, they will copy values from another field. For example, [GILLARDS FARMS LIMITED](https://find-and-update.company-information.service.gov.uk/company/00981261/filing-history) in their 2021 accounts report assets as their number of employees. You can see this in the two images below: image png Mar 15 2024 11 50 27 0155 AM *** image png Mar 15 2024 11 50 49 2594 AM #### Companies can report the year as their number of employees. For example, [ABBOTT & ABBOTT LIMITED](https://find-and-update.company-information.service.gov.uk/company/09150841/filing-history) revised their 2016 employees as '2016', in their 2017 accounts. image png Mar 15 2024 11 49 48 0296 AM #### Companies can report their wage costs as the number of employees For example, [AR CARS CLUB LIMITED](https://find-and-update.company-information.service.gov.uk/company/10798168/filing-history) in their 2017 accounts report director renumeration as the number of employees. image png Mar 15 2024 11 49 00 1090 AM In a small number of cases, anomalies are introduced by parsing of company accounts. The number of examples of this is very low, with a general accuracy of over 99.9%. The model has been trained to identify these outliers too. In addition, we are working with our data provider to improve the parsing process, which will reduce the frequency of the outliers further in the future. The examples mentioned above are where financial accounts are misleading. In training the model to identify unusual financials, we have also identified companies that correctly have unusual financials. In particular, the model will also identify companies that can have high levels of employments with low levels of resources. Specific examples of these include recruitment or healthcare agencies. In these companies, employees are added to a company's books, but it is another company that is funding salaries through their sales. The model will identify companies likely using signficant temporary or part-time employment, if it appears the companies financials are not sufficient to support that level of employment. Removing agencies, or companies with temporary workers, will be beneficial for any analysis using Gross Value Added or turnover. Our calculations of GVA rely explicitly on the number of employees referring to the number of full-time employees. To estimate turnover, we rely implicitly on the assumption that each employee is a full-time employee. ### What is the basis for predicted anomalies? We trained a model to predict whether a company has anomalous financials. To do this, we started with known examples of anomalies. We used the model to predict anomalies, validating the predictions, and incorporating these into a training set. We completed this iteration over 20 times. The validation process was manual. For thousands of companies, this involved inspecting their accounts on Companies House and understanding where the reported values had come from. We now have a training set of over 10,000 companies and we will continue to review predictions, incorporating new kinds of outliers, if and when we find them. If you are aware of outliers that we are missing, please get [in touch](mailto:andrew.purdy@thedatacity.com). # What are Distinct Brands? Source: https://docs.thedatacity.com/our-data/proprietary-data/what-are-distinct-brands Long-time users of our data will know that group structure has always been a challenge. Behind what looks like a single business, there can be layers of subsidiaries, brands and parent organisations shaping activity, employment and revenue. And there’s no consistent way in which this happens.  A group can contain many different brands. In the case of Tesco, this includes the overarching Tesco PLC brand, but also Dunhumby – Tesco’s analytics subsidiary – and Tesco Finance.  “Distinct brands” has been created to help you quickly find the meaningful companies within a group. For reliability, companies that file dormant accounts, or appear inactive, cannot be considered distinct brands.  ## Where can I find Distinct Brands? To better answer the question of ‘how many companies are there in a sector or list?’, we’ve brought distinct brands to analyse. Image The image above shows the number of distinct brands in the Net Zero sector. We think there are 20,500 meaningful companies in the Net Zero sector, across 28,000 registered companies. You can also find flags for distinct brands in downloads from explore, and on a company's group structure tab. # What are RSICs and how do we ensure data quality? Source: https://docs.thedatacity.com/our-data/proprietary-data/what-are-rsics Companies can misreport their activities, by selecting the wrong SIC code. We have solved this issue. RTICs are great for the emerging economy. For the foundational economy, where there is more likely to be an appropriate SIC, an issue remains. A company can choose the wrong SIC code, or the SIC code they've selected does not match their activities. We fixed this issue. Real-Time Standard Industrial Classifications (RSICs) use machine learning and a company's website text to better classify company's activities. For example, Shell PLC is a large energy company. Because they are large, the only SIC code they file at Companies House is activities of head offices. A better description of their activities is provided through RSICs (and RTICs): extraction of crude petroleum and natural gas, mineral oil refining, and the wholesale of fuels and petroleum products. Shell PLC in Industry Engine, showing one filed SIC code alongside six RSICs RSICs follow the same structure as SICs. A reminder, our RSICs: * Fill gaps where SIC codes are missing * Correct inaccuracies in existing SIC codes * Add granularity where SIC codes are vague You can read more about RSICs [here](https://thedatacity.com/blog/sic-codes-fixed-introducing-real-time-sic-codes-rsics/). ### Data Quality To ensure quality we focus on *methodological integrity*: **Trust in the methodology** Primarily, we’ve built trust in our RSICs *within* the RSIC methodology itself. We do this through three distinct layers: 1. **Evidence, not prediction:** We treat classification as an evidence problem, not a prediction problem. RSICs are not arbitrary predictions. Instead, we evaluate the *empirical likelihood* of a classification based on the company’s website text, and what we uniquely understand about companies in each sector. If the data doesn't support the code, we don't assign it. 2. **Coherence Filtering:** This validation layer which rejects codes that lack alignment with the company's specific niche. This allows us to distinguish between a company *mentioning* a topic and actually *doing* it. We identify this distinction, and we classify appropriately. 3. **Specificity:** We also penalise generic classifications. Broad, "catch-all" codes are rarely useful for decision-making, so we deprioritise them in favour of precise definitions. Companies spread across more of the classification instead of piling into a handful of catch-all codes. **Trust in transparency** Unlike black box AI models where the logic is hidden, our RSIC system is built on transparency. Every classification is traceable back to the specific evidence that supports it. The framework is auditable, and we remain in control. **Quality in everything** Quality RSICs rely on quality inputs. By prioritising quality in everything, beginning with high-fidelity website matching and cutting-edge website text analysis, we build trust at every step of the pipeline. That includes knowing when an input isn't good enough. Not every page we find behind a company's website is really a website: some are bot checks, cookie walls, holding pages or errors. A language model will use those perfectly happily, and the result looks well-formed while telling you nothing true. So we check every website before we use it. Where we can't produce RSICs we trust, a company gets no RSIC rather than a misleading one. # How do we build RTICs? Source: https://docs.thedatacity.com/our-data/rtics/how-do-we-build-rtics How an RTIC is built: The Data City's approach to classifying real economies in real time. An [RTIC (Real-Time Industrial Classification)](/our-data/rtics/what-are-rtics) is The Data City's classification of UK companies into precisely defined sectors and sub-sectors. Standard [SIC codes are often too vague, or too out of date](/our-data/rtics/what-is-the-difference-between-rtics-sic-codes), to capture emerging and fast-moving industries — RTICs are built to describe what companies actually do, based on the evidence they publish about themselves. ### How it works Every RTIC is produced through a rigorous, multi-stage methodology that combines our proprietary AI technology with expert analyst judgement at every step: 1. **Define** — the sector is given a rigorous written definition — its scope, its boundaries, and real companies that exemplify it — developed with input from industry and academic experts wherever possible. 2. **Classify** — our proprietary AI identifies the companies that genuinely belong in the sector, assessing each against the sector's definition using the evidence of what that company actually says about itself. Our technology draws on a continuously refreshed view of hundreds of thousands of [matched UK company websites](/our-data/key-data-and-definitions/website-matching). 3. **Quality-assure** — every list goes through multiple independent layers of AI-assisted quality assurance, designed to catch false positives, recover genuine companies that might otherwise be missed, and sense-check the sector's overall shape and statistics before anything is published. 4. **Sign off** — analysts review the results throughout, and a manager formally signs off every RTIC before publication. No AI output goes live on its own. These stages run as a continuous loop, not a one-way pipeline: when real companies test a sector's boundaries, the definition is refined and the affected work re-run — for as long as it takes to get the sector right. Because classification is grounded in website evidence, a company needs a [matched website](/our-data/key-data-and-definitions/website-matching) to be classified into an RTIC. See [why not all companies have RTICs](/our-data/faqs/why-do-all-companies-not-have-rtics). ### Kept current, not built once Published RTICs don't stand still. Newly identified UK companies are continuously assessed against every published sector, and an analyst reviews every proposed addition before it joins a live list. Periodic full updates refresh the sector definition and company list together — and each update comes with a report explaining what changed and why, so you can always understand how your list has evolved. Three types of update keep an RTIC current: an annual **Full Update**, a monthly **Maintenance Update**, and ad hoc **Hotfixes** raised from reports on the platform. Each is reflected in the RTIC's version number. See [RTIC update types and versioning](/our-data/rtics/updating-rtics). ### What you can rely on * **Human oversight, always.** An analyst reviews, and can override, every AI output. Ambiguous and borderline companies are decided by expert judgement, not by an algorithm alone. * **Evidence, not guesswork.** Classification decisions are grounded in what companies actually publish about themselves. Where evidence can't be found, a company is flagged for review — never guessed at. * **Full traceability.** Every decision is versioned and auditable: any classification can be traced back to the sector definition and evidence that produced it. * **Honest accuracy.** Where we quote a confidence or accuracy figure, it comes with the reasoning behind it — we show our working, not just a headline number. A detailed description of our methodology is available on request — email [support@thedatacity.com](mailto:support@thedatacity.com). # RTIC update types and versioning Source: https://docs.thedatacity.com/our-data/rtics/updating-rtics The three types of update that keep RTICs current, and how an RTIC's version number reflects its update history. A quick-reference overview of how [RTICs](/our-data/rtics/what-are-rtics) are kept up to date, and how that update history is reflected in an RTIC's version number. ### Types of RTIC update RTICs are kept accurate and relevant through three types of update: | Update type | Cadence | What it covers | | ---------------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | | **Full Update** | Annual | A full refresh of the RTIC vertical: taxonomy review, updated training set, and a rebuilt company list. | | **Maintenance Update** | Monthly | New companies are classified into existing verticals using our proprietary AI, then reviewed before release. | | **Hotfix** | Ad hoc | Released in response to mismatch or suggestion reports raised on the platform; corrects a small number of companies outside the regular cycle. | Hotfixes are how your reports reach the live data. If you spot a company that's missing from an RTIC, or classified into one it doesn't belong in, [report it on the platform](/our-data/faqs/why-do-all-companies-not-have-rtics) — corrections are released without waiting for the next scheduled update. ### Versioning Every update is reflected in the RTIC's version number, using semantic versioning: Major.Minor.Patch (for example, 1.3.0). | Version segment | Maps to | Meaning | | ----------------- | ------------------ | -------------------------------------------------------------------------------------------------------------- | | **Major** (X.0.0) | Full Update | A full refresh of the RTIC vertical, typically once a year; may include reclassification and taxonomy changes. | | **Minor** (0.X.0) | Maintenance Update | Monthly classification of new or updated companies into existing verticals; reviewed before release. | | **Patch** (0.0.X) | Hotfix | Ad hoc correction for a small number of companies, raised via mismatch/suggestion reports. | **Example:** V 3.0.4 — this RTIC vertical has undergone three Full Updates, and four Hotfixes. A detailed description of our update process is available on request — email [support@thedatacity.com](mailto:support@thedatacity.com). # What are RTICs? Source: https://docs.thedatacity.com/our-data/rtics/what-are-rtics Learn what RTICs are, why you can trust them and how they are used. ### What is an RTIC? Real-Time Industrial Classifications (RTICs) are The Data City's classification of UK companies into precisely defined sectors and sub-sectors, built with our proprietary AI technology and grounded in the evidence companies publish about themselves. RTICs are a modern classification approach, unlike the traditional SIC system that relies on predetermined and static categories. They are based on how companies describe themselves on their websites. RTICs are exclusively frontier sector classifications — sectors and activities that sit outside SIC's coverage altogether. They're not a replacement for SIC codes; for correcting or extending classification within SIC's existing scope, see [RSICs](/our-data/proprietary-data/what-are-rsics). The output is a dataset that gathers companies working in the same field. In this dataset, you will find all the data available for the companies in an RTIC. The data can be explored on the platform or downloaded. **How are RTIC different from SIC codes?** Find out more about the differences between the two classification systems [here](/our-data/rtics/what-is-the-difference-between-rtics-sic-codes) Untitled design (50) ### Why can you trust RTICs? Every RTIC combines our proprietary AI technology with expert analyst judgement: an analyst reviews, and can override, every AI output, classification decisions are grounded in what companies actually publish about themselves, and every decision is versioned and auditable. See [how we build RTICs](/our-data/rtics/how-do-we-build-rtics) for the full methodology. RTICs are also built alongside experts in their respective fields. As an example, our various 'Space' RTICs (space economy, space energy, in-orbit space manufacturing etc.) were built in collaboration with the Satellite Applications Catapult. They helped us both build the taxonomy, as well as check which companies belonged in the training set. Some of our RTICs were built with our own internal experts. For example, Software Development was built in collaboration with our development team. **Kept current**: Published RTICs are continuously maintained — newly identified companies are assessed against every published sector, and periodic full updates refresh the definition and list together. See [how we build RTICs](/our-data/rtics/how-do-we-build-rtics#kept-current-not-built-once). ### How can you use RTICs? RTICs provide unique insights into emerging economies that SIC codes cannot. View RTICs at a top level to get a glimpse into the sector as a whole, or combine them with filters to get a more refined view. RTICs can be found under the RTICs section in the platform. Use location filters to get regional insights, or growth filters to find high-growth companies and much more.