Balanced Medical Dataset
By shrijeetverma13 · May 19, 2026
The fastest way to prototype, benchmark, or fine-tune a medical LLM without spinning up massive compute. MedInstruct-5K is a carefully sampled 5,000-pair…
Great news — your dataset has no duplicate text entries! Out of 10,000 total text entries analyzed, zero are duplicates (0.00%), meaning every text value in the dataset is unique. Two bar charts were generated showing duplicate percentages and counts by source, confirming that all sources are free of repetition.
The analysis examined text lengths across 10,000 records and produced two visualizations: a histogram showing the distribution of text lengths by source, and a box plot comparing text lengths across different sources. The overall mean text length is 1,189 characters, while the median is 904 characters, indicating a right-skewed distribution where some longer texts pull the average up. Text lengths range from a minimum of 129 characters to a maximum of 5,817 characters.
The analysis successfully produced visualizations showing how average text length varies across different sources. A bar chart displays the average text length (in characters) for each source, ranked from longest to shortest, and a box plot reveals the full distribution of text lengths per source — highlighting spread, outliers, and variability.
The dataset contains 10,000 records spread across 7 distinct sources. The top contributor is 'chatdoctor' with 3,500 records, making up 35% of the entire dataset. Three visualizations were generated to help you explore this: a bar chart showing record counts per source, a donut pie chart showing the percentage distribution, and a box plot comparing text length distributions across sources.
Every source in the dataset provides completely unique text content with no overlapping material between sources. All 7 sources achieved a perfect 100% uniqueness score, meaning there are no duplicate texts shared across any of the sources. Two charts were generated: a stacked bar chart showing unique vs. overlapping content counts per source, and a percentage bar chart confirming 100% uniqueness across all sources.
The analysis identified the most frequent keywords appearing in text across each source. Two data tables were generated showing the top terms per source, filtered to remove common stopwords and short words, focusing on meaningful medical and clinical terminology.
The analysis successfully computed question complexity metrics across the three medical sources (meddialog, icliniq, and chatdoctor) and generated a grouped bar chart comparing average word count, average sentence count, and average words-per-sentence ratio for each source. The chart visually ranks the sources by complexity, making it easy to compare how verbose or structured questions are across platforms.
The analysis examined the text column to identify dialogue-style content (containing keywords like 'Doctor:' or 'Answer:') versus single-question formats. The results show that 100% of the texts in the dataset are dialogue-style, meaning every entry contains multi-turn conversation markers. A stacked bar chart has been generated showing the format breakdown by source, and data tables are available with the percentage splits per source.
A summary table was generated capturing text statistics for the dataset. The table includes overall and per-source breakdowns covering record counts, vocabulary size, median word counts, 25th/75th percentiles, and the percentage of texts containing numeric values.
The analysis identified texts that fall more than 3 standard deviations from the mean character length within each source. A total of 149 outlier texts were detected across all 7 sources. Two visualizations were generated: a bar chart showing outlier counts per source and a box plot displaying the full text length distribution with outlier points highlighted. Summary statistics tables are also available showing mean and standard deviation per source.
The sentiment analysis of your text data reveals that the majority of entries are neutral, with a generally positive lean overall. A bar chart has been generated showing the full sentiment distribution across all 10,000 records.
A word cloud has been successfully generated from the 'text' column of your dataset. The visualization displays the most frequently occurring words, with larger text indicating higher frequency. This gives you a quick visual overview of the most prominent terms in your text data.