LLM Price Performance Tracker (March 2026)
By shrijeetverma13 · April 1, 2026
The most comprehensive cross-provider snapshot of the large language model API landscape as of March 31, 2026. Data source:
The bar chart and table show the top 10 AI models offering the best intelligence per dollar, along with their respective providers. The horizontal bar chart visually ranks each model from highest to lowest value efficiency, color-coded by provider so you can quickly spot which companies dominate the value rankings.
The analysis produced a dual-axis chart and supporting tables showing the top 15 providers by number of models, alongside their average composite benchmark scores. The bar chart displays model counts per provider, while the red line overlay tracks average benchmark performance — making it easy to spot which providers offer both quantity and quality.
A data table has been generated showing the models that sit on the efficiency frontier — those offering the highest composite benchmark scores at each cost level. These are the models that provide the best performance-per-dollar trade-off, meaning no other model beats them on both dimensions simultaneously.
The chart and data tables show how average composite benchmark scores and blended costs per 1M tokens have changed across release years. A dual-axis line chart was generated, making it easy to compare both metrics side by side over time. The visualization clearly tracks the trajectory of model performance and pricing as newer models have been released.
There is essentially no meaningful tradeoff between output speed and intelligence across the models analyzed. The correlation is r = -0.053, which is extremely weak, meaning faster models are not noticeably less intelligent, and smarter models are not noticeably slower. A scatter plot and supporting data tables were generated to visualize this relationship.
The analysis identified the top 10 models that excel across both the AA Coding Index and LiveCodeBench benchmarks, with their associated costs. A bar chart was generated comparing each model's AA Coding Index score, LiveCodeBench score (scaled ×50 for comparison), and blended cost per 1M tokens overlaid as a red line. A data table is also available showing each model's provider, scores, and pricing tier.
There's a moderate positive correlation (0.432) between parameter count and composite benchmark scores, meaning larger models tend to score higher — but many small models punch well above their weight. The scatter plot shows all models colored by how much they exceed (green) or fall below (red) the expected trendline, with top small overperformers highlighted as stars. Among models with 9B parameters or fewer, K-EXAONE (Reasoning) by LG AI Research leads the pack with a composite score of 57.52 — nearly 20 points above what its size would predict. OpenAI's GPT-5 nano variants and Alibaba's Qwen3 8B also significantly outperform expectations.
The analysis compares average composite benchmark scores and output speeds across different pricing tiers. A scatter plot and summary tables were generated showing how each tier performs on both dimensions, letting you see whether higher-priced tiers deliver better quality and/or faster responses.
The analysis compared 302 open-source and 151 closed-source models across three dimensions. Closed-source models lead on benchmark performance (51.37 vs 40.38) and speed (132.4 vs 113.0 tokens/sec), but open-source models are dramatically cheaper ($0.77 vs $7.96 per 1M tokens — over 10x less). A scatter plot and comparison charts were generated to visualize these trade-offs.
The analysis produced two visualizations showing which models lead on math benchmarks and how they compare on cost. A scatter plot maps math performance (average of AIME 2025 and MATH 500) against blended cost per 1M tokens, color-coded by pricing tier — making it easy to spot cost-effective leaders. A horizontal bar chart shows the top 10 models side-by-side on both benchmarks individually.
The analysis identified models with the fastest time-to-first-token (TTFT) while maintaining an above-median composite benchmark score (≥ 47.2). A bar chart was generated showing the top 20 fastest qualifying models, color-coded by their composite benchmark scores using a Viridis color scale. You can hover over each bar to see provider, benchmark score, and pricing tier details.
A table has been generated showing each provider's single best model ranked by price-performance ratio. You can explore the results in the data table to see which model from each provider offers the greatest value relative to its cost.
Two visualizations were generated to compare how providers differ in their input vs. output cost ratios. The scatter plot shows each provider's average input cost vs. average output cost per 1M tokens, with bubble size representing the number of models. The bar chart ranks providers by their average output-to-input cost ratio, making it easy to spot which providers charge proportionally more for output tokens.
Paid models clearly outperform free-tier models on the composite benchmark. The analysis compared 16 free models against 299 paid models, with paid models averaging 46.6 vs. 35.2 for free — an 11.4-point gap. Two visualizations were generated: a box plot showing the full distribution comparison and a bar chart highlighting the top 15 free models by provider.
A data table has been generated showing 2026-released open-source models ranked by their price-performance ratio. The table highlights each model's name, provider, price-performance ratio score, composite benchmark performance, and blended cost per 1 million tokens — giving you a comprehensive view of which models deliver the most value for their cost.
Here's the best model for each budget level to maximize composite benchmark score:
• $1/1M tokens : Grok 3 mini Reasoning (high) — Composite Score: 68.0, Cost: $0.35/1M tokens
• $5/1M tokens : GPT-5 (high) — Composite Score: 74.5, Cost: $3.44/1M tokens
• $20/1M tokens : GPT-5 (high) — Composite Score: 74.5, Cost: $3.44/1M tokens
Two data tables were generated summarizing these results. GPT-5 is the top performer at both the $5 and $20 budget tiers, delivering a composite score of 74.5 at just $3.44/1M tokens — well under both limits. At the tighter $1 budget, Grok 3 mini Reasoning (high) is the best choice with a solid score of 68.0 at only $0.35/1M tokens.