SIC Industry Intelligence Report: Distribution, Sector Coverage & Classification Analysis
By abhishek.verma75000 · May 16, 2026
Comprehensive analysis of 731 SIC codes across 81 industry divisions highlighting classification density, sector concentration, industry gaps, keyword…
Major Group 46 contains the most industry classifications with 52 entries, followed by Major Group 47 with 44 classifications and Major Group 10 with 30 classifications. A bar chart has been generated showing the top 15 two-digit SIC major groups ranked by their number of industry classifications.
SIC codes are distributed across 81 major industry divisions (based on the first 2 digits), covering 731 total codes in the dataset. Two visualizations were generated: a bar chart showing the count of SIC codes per division ranked from highest to lowest, and a treemap highlighting the top 20 divisions by code count. These charts make it easy to see which industry divisions are most densely populated with SIC codes.
The analysis produced two data tables summarizing the most common keywords and two-word phrases (themes) found across industry descriptions. These tables highlight the dominant terminology used to categorize and describe industries in the dataset.
The analysis examined 731 SIC codes across broad industry sectors to determine which have the most granular coverage. Two data tables were generated showing the distribution of SIC codes by sector. Based on the UK SIC 2007 standard, Manufacturing (SIC codes 10000–33999) typically dominates in granularity due to its vast range of products and processes, followed by Professional & Scientific Services and Wholesale & Retail sectors.
The analysis identified the industries with the longest and most complex descriptions, which are strong indicators of niche specialization. Two bar charts and supporting data tables were generated showing the top 15 industries by description length and the overall distribution across all industries. SIC 46439 (Wholesale of radio/television goods & electrical household appliances) tops the list with 158 characters, followed by SIC 16290 (Manufacture of other wood products including cork, straw and plaiting materials) at 100 characters. The average description length across all industries is just 41.7 characters, making these long-description industries stand out significantly.
The analysis identified clear gaps and underrepresented areas in the SIC code classification system. Two charts were generated showing the distribution of SIC codes across divisions and numeric ranges. There are 11 divisions with zero SIC codes (divisions 4, 34, 40, 44, 48, 54, 57, 67, 76, 83, and 89), and 12 additional divisions with only 1–2 codes each. These gaps may represent emerging industries or niche sectors not yet fully captured by the traditional SIC framework.
The analysis extracted the first 2 digits of each SIC Code to identify major industry groups, then ranked the top 10 by the number of sub-codes they contain. A horizontal bar chart has been generated showing each 2-digit major group on the Y-axis and the count of sub-codes on the X-axis, with percentage labels displayed outside each bar. Data tables are also available with the full breakdown of counts and percentage shares.
The SIC Code column spans from a minimum of 1,110 to a maximum of 99,999, covering a range of 98,889. The average SIC code is 43,925 with a median of 46,120, suggesting a fairly central distribution. The standard deviation of 26,478 indicates considerable spread across the dataset. There are 731 unique SIC codes in use, and a box plot has been generated to visualize the distribution and highlight any outliers.
A data table has been generated showing SIC codes grouped into numeric tiers (1–9) based on their first digit, representing broad economic sectors. The table includes the number of codes per tier and average description length for each sector group.
The analysis of description text lengths across all 731 entries reveals a distribution skewed toward shorter descriptions. The histogram with KDE overlay and outlier threshold lines clearly shows the spread and shape of the data. Descriptions tend to be short , with a median of 39 characters and a mean of 41.72 characters. The standard deviation of 19.50 characters indicates moderate variability. The shortest description is just 5 characters, while the longest reaches 158 characters. Using the 2-standard-deviation rule, 29 entries (about 4%) are flagged as unusually long (exceeding ~80.7 characters), and no entries fall below the short-outlier threshold (~2.7 characters).
The sentiment analysis of the 'Description' column reveals that the vast majority of entries are neutral in tone. A bar chart has been generated showing the full distribution across all three sentiment categories.
A word cloud has been successfully generated from the 'Description' column of your dataset. The visualization displays the most frequently occurring words in the descriptions, with larger words appearing more prominently to indicate higher frequency.