A massive disparity in global genomic databases shows South Asians represent less than 1% of participants, threatening the accuracy of AI-driven personalized medicine for millions.
- South Asians account for less than 1% of participants in global genome-wide association studies.
- AI models trained on European-heavy data are less accurate for South Asian health risks.
- The data gap risks exacerbating health inequalities in diabetes and cardiovascular disease.
As Artificial Intelligence (AI) and machine learning revolutionize medicine by enabling early disease detection and personalized treatments, a critical flaw has emerged: a lack of diversity in the underlying data. South Asian populations, despite their vast numbers, are significantly underrepresented in the global health databases that fuel these breakthroughs.
The Massive Representation Gap
The disparity is stark. According to the NHGRI-EBI GWAS Catalogue, between 2005 and 2025, over 86% of participants in human genome-wide association studies were of European ancestry. In contrast, South Asians accounted for less than 1%. This means the 'reference maps' for modern biology are being drawn using a very limited palette of human genetics.
Why This Matters
BozokMedia analysis shows that this isn't just a statistical error; it is a direct threat to public health. South Asians face disproportionately higher rates of type 2 diabetes, cardiovascular disease, and asthma compared to those of European descent. When AI models are trained predominantly on European genetic data, they fail to accurately predict risks or suggest effective treatments for South Asian individuals.
"More than 20% of the world is being neglected in multi-modal data integration, denying them the human right to attain maximal possible health." — Bhramar Mukherjee, Yale School of Public Health.
For instance, polygenic risk scores—tools used to estimate an individual's genetic predisposition to diseases—have proven to be less accurate when applied to South Asian populations. As geneticist Shweta Ramdas notes, most predictions regarding gene expression are inferred from European datasets, leaving a massive blind spot in understanding disease mechanisms relevant to South Asians.
Historical Background
Historically, biomedical research has been heavily skewed toward Western populations. Integrated biobanks like the U.K. Biobank have been instrumental in drug development and clinical guidelines, but their composition reflects a long-standing trend of prioritizing European-descended cohorts. This historical bias has created a cumulative disadvantage for non-European populations in the era of precision medicine.
| Metric | European Ancestry | South Asian Ancestry |
|---|---|---|
| Study Participation (%) | >86% | <1% |
| Primary Health Risks | Standard Baseline | High Diabetes/Cardio Risk |
| AI Model Accuracy | High | Potentially Low/Unverified |
Frequently Asked Questions
1. How does the data gap affect individual treatment?
It can lead to inaccurate risk assessments and treatments that are not optimized for the specific genetic makeup of South Asian patients.
2. What is being done to fix this?
South Asian scientists are actively working to build their own indigenous genomic tools and datasets to ensure equitable healthcare.