Open Data Index Analysis – Countries, Scores and Data Availability
By abhishek.verma75000 · March 26, 2026
Data analysis and visualization of global open data performance using the datastam.ai
Datasets marked isOpen=1 have an average Score of 92.96, while those marked isOpen=0 average just 28.48 — a difference of 64.48 points. A bar chart has been generated to visually compare these two groups.
The analysis identified the top 10 countries with the highest average open data scores across all variables. A bar chart has been generated showing each country's average score, making it easy to compare performance at a glance. The visualization ranks countries from highest to lowest average score, giving a clear picture of which nations lead in open data transparency and accessibility.
The bar chart shows the percentage of 'isOpen' across all data variables. ODB.2013.D9 has the highest openness rate at 20.93%, while ODB.2013.D7 has the lowest at just 3.49%. Overall, openness rates across variables are quite low, with even the top variable falling below 21%.
The UK leads all countries with 100% of its datasets having a cMachineReadable value greater than zero. A bar chart has been generated showing the top 20 countries ranked by this percentage, making it easy to compare machine readability rates across nations.
Out of 86 countries analyzed, 49 (57%) have NO datasets with open licensing (fLicense 0). The bar chart shows the proportion of licensed datasets per country, sorted from lowest to highest. Countries like Malaysia, Hungary, Turkey, India, South Africa, Ireland, Israel, Jamaica, Jordan, Kazakhstan, Venezuela, Korea, and many more have zero open-licensed datasets.
The analysis produced a bar chart showing the average score for each dataset variable (type), along with supporting data tables. The chart displays average scores with error bars representing variability, making it easy to see which variables score highest and which are most consistent. The most consistently well-published datasets are those with both high average scores and low standard deviation — these appear at the top of the consistency ranking in the generated tables.
A bar chart was generated showing how many countries offer bulk download (dBulk 0) for each variable in the dataset. The visualization uses a color scale from red (fewer countries) to green (more countries), making it easy to spot which datasets lag behind and which are widely available.
The bar chart shows how strongly each of the six sub-indicators (aExists, bAvailable, cMachineReadable, dBulk, eFree, fLicense) correlates with the overall Score. All six sub-indicators have a positive Pearson correlation with Score, meaning higher values in each indicator are associated with higher scores. The visualization ranks them from strongest to weakest correlation, making it easy to see which indicators matter most for achieving high scores.
There is a strong positive relationship between gUpdated values and both overall Score and isOpen status. The correlation between gUpdated and Score is 0.745 (strong), while the correlation with isOpen is 0.387 (moderate). Two visualizations were generated: a box plot showing score distributions by gUpdated value and open status, and a dual-axis chart comparing average scores and open rates across gUpdated values.
The analysis identifies which countries have the highest proportion of datasets with linked data (jLinked 0). A bar chart has been generated showing the top 20 countries ranked by this proportion, filtered to include only countries with at least 3 datasets for statistical reliability.
The 10 countries with the lowest average open data scores — indicating the biggest open data gaps — are Mali (3.33), Myanmar (5.0), Haiti (5.33), Cameroon (6.33), Nigeria (8.67), Yemen (9.33), Botswana (9.67), Senegal (10.33), Sierra Leone (10.67), and Zambia (10.67). A horizontal bar chart and data table have been generated to visualize these results clearly.
Two bar charts were generated to reveal which countries and variables most frequently charge for data access (eFree=0), creating barriers for users. The first chart shows the top 15 countries with the highest number of paid data records, and the second highlights the top 15 variables that are most commonly behind a paywall.
A bar chart has been generated showing the top 20 most frequently reported data formats found in the cFormats column across all countries and variables. The chart ranks each format by how often it appears, giving a clear picture of which data formats dominate the dataset.
The analysis shows the percentage of assessed variables classified as open (isOpen=1) for each country. A bar chart displays the top 20 countries ranked by their open variable percentage, and supporting data tables provide the detailed breakdown per country.
The analysis produced a scatter plot comparing average hSustainable and iDiscoverable scores across all countries in the dataset. Each point on the chart represents a country, making it easy to spot which nations score high on both dimensions simultaneously — those appearing in the upper-right corner of the chart excel at both metrics.
The analysis identified 62 Country-Variable combinations where the Score equals 0, revealing critical data availability gaps across the dataset. Two bar charts were generated showing which countries and variables have the most gaps, along with detailed data tables for deeper exploration.
There are 15 datasets that achieve a perfect Score of 100. These span 7 countries: France, Germany, Netherlands, New Zealand, Norway, Sweden, UK, and the US. The UK leads with 4 perfect-scoring datasets, followed by New Zealand and the US with 2 each. A bar chart has been generated showing the distribution of perfect scores by country, along with detailed tables listing each country-variable combination.
The analysis identified which responsible agencies have the highest and lowest average scores. Data tables have been generated showing the agencies with at least 3 entries, ranked by their average scores. The results are split into top performers and bottom performers, giving a clear picture of which agencies score best and worst.
The sentiment analysis of the 'aDetails' column reveals that the majority of entries are neutral in tone. A bar chart has been generated showing the full sentiment distribution across all records.