Conflict Data Quality Report: Missing Notes and Format Inconsistencies

By abhishek.verma75000 · July 29, 2026

This analysis examines a dataset of 11 conflict records, revealing data quality issues that affect research reliability. Five records lack supporting…

A bar chart and table have been generated showing the number of unique entities associated with each conflict category, ranked from highest to lowest. This visualization makes it easy to see which categories involve the broadest range of distinct entities.

The analysis produced a heatmap showing how conflict categories are distributed across different time periods. Darker red cells indicate periods where a category appears more frequently, making it easy to spot categories that are concentrated in specific eras versus those spread across many periods.

5 records in your dataset are missing supporting Notes documentation. These gaps are concentrated in the 'Trend' category, with related missing entries under the 'Region' period grouping. A chart comparing documented vs. undocumented records by category has been generated to help you visualize where the gaps are.

Entity representation varies noticeably across the 7 time periods identified in the data, which together include 10 distinct entities. The period labeled 'Region' stands out with the most diversity, containing 5 different entities. Overall, 'Conflict deaths' was the most frequently recorded entity, appearing in 2 records. This suggests that different time periods tend to focus on specific entity types, reflecting shifts in data emphasis or terminology over time.

Looking across all categories and time periods, 'Conflict deaths' is the most frequently occurring entity, appearing 2 times, followed closely by 'Armed conflict deaths', 'Conflict', 'Disease & starvation', and 'UCDP; geoBoundaries', each appearing once. In terms of reach, 'Conflict deaths' also stands out by spanning 2 distinct periods within 1 category, showing it's tracked consistently over time. Overall, the dataset contains 10 unique entities, and the results are fairly evenly spread since most entities only appear once.

Looking at the Description field across categories, a few clear themes stand out. Many entries reference death and fatality counts, reflecting data about casualties from conflicts. Another common thread involves time and reporting periods, such as annual, quarterly, or preliminary data labels, showing that many descriptions clarify how current or complete the data is. There are also references to conflict and fighting, tying into the overall subject of armed conflict data. A bar chart breaks down the most frequently used words in the Description field, and a second chart shows how many descriptions fall under each category, helping to see where this descriptive detail is concentrated.

The dataset was broken down by the Entity column, counting how many records belong to each distinct entity (such as 'Armed conflict deaths' vs 'Conflict deaths') and calculating what percentage each represents of the total. A bar chart and data table were generated showing the count and proportion for each entity category.

The dataset contains 11 rows and 5 relevant columns: Period, Category, Entity, Description, and Notes. Entity has the most variety with 10 unique values, while Category has the least with 5. Only the Notes column has missing data, with 5 out of 11 entries blank; all other columns are fully complete.

The analysis examined the Period column and found that most rows (8 out of 11) don't fit the common patterns of a single year, year range, or quarter format, meaning they fall into an 'Other/Unknown' category. This suggests the Period column has significant formatting inconsistencies, mixing several different styles like year ranges (e.g., '1989–2026'), single years (e.g., '2025'), and quarters (e.g., '2026 Q1'). A bar chart was generated showing the breakdown of these format patterns, along with data tables highlighting the specific rows and their detected formats, so you can review exactly which entries deviate from expectations.

The analysis organized the dataset's rows to compare how specific each Period entry is (full date ranges, single years, or quarters) against the status noted in the Notes column (such as 'Complete year' or 'Incomplete'). While the detailed statistical breakdown encountered a technical hiccup, the underlying data tables were successfully generated and are available for review, showing the Period, Notes, and related fields for each row in your dataset.

The sentiment analysis of the 'Description' text found that most entries (10 out of 11, or 90.9%) are neutral in tone, with only 1 entry (9.1%) showing negative sentiment and none classified as positive. The average sentiment polarity score was -0.011, indicating an overall tone that's essentially balanced with a very slight negative lean.