Arxiv Ai Ml Papers

By shrijeetverma13 · May 13, 2026

This dataset captures 7,701 peer-reviewed and preprint research papers submitted to arXiv.org between September 2025 and April 2026 — the most active and…

The analysis identified the most prolific first authors in the dataset. Jon-Paul Cacioli tops the list with 11 papers, followed by Dongxin Guo with 9 papers and Xue Qin with 7 papers. A horizontal bar chart has been generated showing the top 15 first authors ranked by paper count.

The analysis examined 7,701 AI/ML paper submissions spanning 107 unique primary categories. Two visualizations were generated: a bar chart showing the top 10 primary research categories by paper count, and a donut pie chart displaying each category's share of total submissions. These charts reveal which research areas dominate the landscape of AI/ML academic publishing.

Paper submissions show a dramatic upward trend from September 2025 through April 2026. Starting with just 25 submissions in September 2025, the volume grew significantly to reach a peak of 6,031 papers in April 2026. A line chart has been generated showing this monthly progression, along with supporting data tables.

The analysis compares the number of authors between large collaborations and typical papers, with two visualizations (a box plot and a histogram) generated to illustrate the stark differences. Typical papers make up the vast majority of the dataset (7,684 papers) and cluster tightly around 3–6 authors, with a mean of 4.81 and a median of 4. Large collaborations, though rare (only 17 papers), have dramatically more authors — ranging from 55 to 547 — with a mean of 146 and a median of 88 authors.

Out of 7,701 total papers, 839 (about 10.9%) have been updated at least once. For those that do get updated, the typical revision happens very quickly — the median update lag is just 5 days after initial submission. Two visualizations were generated: a pie chart showing the split between updated and non-updated papers, and a histogram showing the distribution of update lag times.

The analysis examined which category combinations appear most frequently in papers that span multiple research areas. Two data tables were generated showing the top category pair combinations, revealing which AI/ML subfields most commonly appear together in the same paper.

The analysis successfully produced two visualizations breaking down how papers span multiple research areas. A bar chart shows the percentage distribution of papers by number of categories (1, 2, 3, 4, and 5+), and a horizontal bar chart highlights the top 10 most cross-disciplinary primary categories ranked by their average number of categories per paper.

The analysis compared 7,458 papers without journal references against 243 papers with journal references across four dimensions. A grouped bar chart and supporting tables were generated to visualize the differences. Interestingly, non-journal papers tend to be slightly longer (avg 179.9 words vs 173.6) and have more authors (5.15 vs 4.31), while journal-referenced papers span slightly more categories on average (2.14 vs 2.07). Large collaborations are rare in both groups but slightly more common in non-journal papers (0.2% vs 0%). A second chart shows the top fields publishing in journals, led by cs.CV, cs.LG, cs.NE, cs.AI, and cs.CL.

The analysis successfully computed monthly submission trends and generated a multi-line chart showing how key metrics evolve over time. The chart displays total papers submitted, average number of authors, average abstract length, and the proportion of updated papers — all plotted across months. High-volume months (more than 1.5 standard deviations above the mean) are highlighted with red star markers.

The analysis of abstract length and word count distributions reveals well-behaved, roughly symmetric distributions with very high correlation. Two visualizations were generated: a dual histogram showing both distributions side by side, and a scatter plot confirming their near-perfect linear relationship. Statistical tables with full descriptive stats and percentiles are also available.

The sentiment analysis of the 'abstract' column reveals that the majority of abstracts are neutral in tone, with a significant portion leaning positive. A bar chart has been generated showing the full sentiment distribution across all records.

A word cloud has been generated from the 'abstract' column of your dataset. The word cloud displays the most frequently occurring words in the abstracts, with larger words appearing more often in the text.