A massive disparity in global genomic databases shows South Asians represent less than 1% of participants, threatening the accuracy of AI-driven personalized medicine for millions.

  • South Asians account for less than 1% of participants in global genome-wide association studies.
  • AI models trained on European-heavy data are less accurate for South Asian health risks.
  • The data gap risks exacerbating health inequalities in diabetes and cardiovascular disease.

As Artificial Intelligence (AI) and machine learning revolutionize medicine by enabling early disease detection and personalized treatments, a critical flaw has emerged: a lack of diversity in the underlying data. South Asian populations, despite their vast numbers, are significantly underrepresented in the global health databases that fuel these breakthroughs.

The Massive Representation Gap

The disparity is stark. According to the NHGRI-EBI GWAS Catalogue, between 2005 and 2025, over 86% of participants in human genome-wide association studies were of European ancestry. In contrast, South Asians accounted for less than 1%. This means the 'reference maps' for modern biology are being drawn using a very limited palette of human genetics.

Why This Matters

BozokMedia analysis shows that this isn't just a statistical error; it is a direct threat to public health. South Asians face disproportionately higher rates of type 2 diabetes, cardiovascular disease, and asthma compared to those of European descent. When AI models are trained predominantly on European genetic data, they fail to accurately predict risks or suggest effective treatments for South Asian individuals.

"More than 20% of the world is being neglected in multi-modal data integration, denying them the human right to attain maximal possible health." — Bhramar Mukherjee, Yale School of Public Health.

For instance, polygenic risk scores—tools used to estimate an individual's genetic predisposition to diseases—have proven to be less accurate when applied to South Asian populations. As geneticist Shweta Ramdas notes, most predictions regarding gene expression are inferred from European datasets, leaving a massive blind spot in understanding disease mechanisms relevant to South Asians.

Historical Background

Historically, biomedical research has been heavily skewed toward Western populations. Integrated biobanks like the U.K. Biobank have been instrumental in drug development and clinical guidelines, but their composition reflects a long-standing trend of prioritizing European-descended cohorts. This historical bias has created a cumulative disadvantage for non-European populations in the era of precision medicine.

MetricEuropean AncestrySouth Asian Ancestry
Study Participation (%)>86%<1%
Primary Health RisksStandard BaselineHigh Diabetes/Cardio Risk
AI Model AccuracyHighPotentially Low/Unverified
Did You Know?: In India, the number of people living with diabetes is projected to skyrocket to 125 million by the year 2045.

Frequently Asked Questions

1. How does the data gap affect individual treatment?
It can lead to inaccurate risk assessments and treatments that are not optimized for the specific genetic makeup of South Asian patients.

2. What is being done to fix this?
South Asian scientists are actively working to build their own indigenous genomic tools and datasets to ensure equitable healthcare.