Bengal is the eastern end of the South Asian cline and the western end of something else. The Ganges and Brahmaputra delta has been a meeting place for people arriving from the Gangetic plain to the west, the hills of the northeast and mainland Southeast Asia. A rule before the numbers. Ancestry is not identity. Being Bengali is a matter of language, culture, religion and family, on both sides of the border, and none of that is measured by a genome. This post describes the ancient populations that Bengali genomes resemble. It does not describe who anyone is.
The short answer: Bengali genomes are mostly South Asian in the usual sense, a blend of Iranian farmer-related and Ancient Ancestral South Indian (AASI) ancestry with a limited Steppe-related share, and they sit toward the AASI-rich end of the cline. On top of that they carry an East or Southeast Asian related layer, typically in the range of a tenth to a fifth, that arrived with Austroasiatic and later Tibeto-Burman speaking groups. No ancient genome from Bengal itself has been published.
The South Asian base#
The framework for the subcontinent comes from Narasimhan et al. (2019) and Reich et al. (2009). Present-day South Asians fall along a cline between two constructed poles. The Ancestral North Indian (ANI) pole is a blend of an Iranian farmer-related stream and Steppe pastoralist ancestry; the Ancestral South Indian (ASI) pole is the same Iranian farmer-related stream blended with more AASI, the deeply divergent lineage native to the subcontinent. The Iranian farmer-related and AASI streams mixed first, in the Indus region before 3000 BC (see the Indus Periphery); Steppe ancestry entered after 2000 BC. The South Asian qpAdm guide walks through the proxies.
On that cline, Bengalis sit toward the ASI end. Their AASI-related share is substantially higher than in Punjab or the Hindu Kush, and their Steppe-related share is correspondingly small, in the low single digits to around a tenth depending on the model and the community. Steppe ancestry thins with distance from the northwest, and the delta is a long way from the Swat valley.
The East and Southeast Asian layer#
What sets Bengal apart is a component that most other South Asian populations lack or carry only in traces. The 1000 Genomes BEB sample of Bengalis from Dhaka was the first widely used Bengali reference, and every analysis of it finds an East Asian related component that the Gujarati, Punjabi, Telugu and Tamil samples do not share, usually between a tenth and a fifth.
Its origin is mainland Southeast Asian. Chaubey et al. (2011), studying Austroasiatic-speaking groups of India such as the Munda, found that their East Asian related ancestry traced to Southeast Asia and was strongly male-biased, carried by Y-chromosome haplogroup O2a alongside almost entirely South Asian maternal lineages. Tätte et al. (2019) extended that work to Bengal and showed that the Bengali component has the same Southeast Asian affinity, with admixture dates within the past few thousand years and at least one episode within the first millennium AD. Basu et al. (2016) separately identified Ancestral Austroasiatic and Ancestral Tibeto-Burman components in India, both present in Bengal.
Two waves are therefore in play. The older is Austroasiatic: Munda-related populations who brought the O2a lineage from Southeast Asia and were absorbed by the farming population of the delta. The later is Tibeto-Burman: the peoples of the northeastern hills and the Chittagong Hill Tracts, whose contact with the plains continued into historical times. The Bengali component is a blend of both, weighted differently by district.
A dated timeline for Bengal#
- Before 5000 BC. AASI-related foragers occupy the subcontinent. No ancient genome represents them directly.
- About 4700 to 3000 BC. Iranian farmer-related and AASI-related people mix in the Indus region. The blend spreads east with farming over the following two millennia.
- About 2000 to 1000 BC. Steppe MLBA ancestry enters the northwest and dilutes eastward along the Gangetic plain. The Bengal delta is settled by farming communities whose genomes are unknown.
- Roughly 2000 BC to 500 AD. Austroasiatic-speaking populations carrying Southeast Asian ancestry and haplogroup O2a spread into eastern India. Dating is method-dependent; the Bengali admixture signal includes episodes well within this window.
- First millennium AD onward. Tibeto-Burman speaking groups from the northeast contribute a further East Asian related layer, especially in northern and eastern Bengal.
- Present day. The BEB sample and later Bengali cohorts place Bengal at the eastern edge of the South Asian cline with a distinct Southeast Asian shift.
Regional variation#
Bengal is not one gene pool. Samples from Sylhet and Chittagong carry more of the East Asian related component than western districts, consistent with proximity to the hills; West Bengal shades toward Bihar and Odisha with more of the ordinary ANI-ASI profile. Endogamous communities within Bengal differ through drift and marriage patterns rather than different sources. None of this variation is a ranking, and none of it maps onto religion: Muslim and Hindu Bengalis draw from the same regional ancestry.
What the Y-DNA and mitochondrial DNA add#
Paternal lineages in Bengal are more mixed than in most of South Asia. R1a, the lineage that arrived with Steppe pastoralists, is found in roughly a fifth to a third of Bengali men in the surveys of Sengupta et al. (2006) and the 1000 Genomes Y-chromosome study. H, the deep native South Asian lineage, is common. O2a, the Austroasiatic-associated lineage from Southeast Asia, varies from a few percent to well over a tenth by district, and it is the clearest single marker of the eastern layer. J2, L and R2 are present at lower frequency.
Maternal lineages are overwhelmingly South Asian. Branches of haplogroup M account for around two thirds of Bengali maternal lines, with West Eurasian lineages and the South Asian branches of R and U making up most of the rest. East Asian mitochondrial lineages such as D, F and G appear at low frequency, far below the autosomal East Asian share. That contrast is the sex bias Chaubey and colleagues documented: the Southeast Asian contribution came more through men than women.
Limitations#
- There is no ancient genome from Bengal. The nearest published ancient individuals come from the Swat valley, Rakhigarhi in Haryana, Roopkund in the Himalaya and Southeast Asia, all hundreds to thousands of kilometres away.
- AASI has no ancient sample either. The Andamanese Onge stand in for it, and the Bengali AASI share is the largest single source of model uncertainty.
- The Southeast Asian source is not pinned down. Modern Austroasiatic speakers and ancient genomes from Vietnam, Laos and Thailand are all used as proxies, with slightly different results.
- Sampling is uneven. The BEB sample comes from Dhaka; Sylhet, Chittagong and much of West Bengal are represented by far fewer individuals.
- Language, religion and identity are not in the genome. A Southeast Asian component does not make anyone Austroasiatic, and a Steppe share does not make anyone Indo-Aryan.
Frequently asked questions about Bengali DNA#
Do Bengalis have East Asian ancestry?#
Yes, at a meaningful minority share. Analyses of the 1000 Genomes Bengali sample and later cohorts find an East or Southeast Asian related component of roughly a tenth to a fifth. It traces to mainland Southeast Asia and arrived with Austroasiatic speaking groups, with a later Tibeto-Burman contribution from the northeastern hills. It is stronger in eastern districts than in the west.
How much Steppe ancestry do Bengalis have?#
Little. Bengal is at the far eastern end of the South Asian cline, and the Steppe pastoralist ancestry that arrived in the northwest after 2000 BC thins with distance. Published models put the Bengali Steppe-related share between the low single digits and around a tenth, well below Punjab or the Hindu Kush.
Are there ancient genomes from Bengal?#
No. As of this writing no ancient genome from Bangladesh or West Bengal has been published. The closest ancient references are from the Swat valley, Rakhigarhi, Roopkund and mainland Southeast Asia, so Ancient Matches for a Bengali genome report shared segments with individuals from those regions, not from Bengal itself.
Are Bangladeshis and West Bengalis genetically different?#
Not as populations. Both draw from the same regional ancestry, and religion does not track any genetic boundary. There is a gradient: the East or Southeast Asian related component is highest in the east near the hills and lowest in the west toward Bihar and Odisha, and West Bengal shades toward the ordinary ANI-ASI profile. Local endogamy adds structure on top of that.
Why does my ancestry test give Bengalis a "Southeast Asian" percentage?#
Because the East Asian related layer in Bengal really is Southeast Asian in origin, and consumer references have no Bengal-specific ancient source. Depending on the panel, the same ancestry may be labelled Southeast Asian, Chinese, Burmese or Tibetan. A formal model with a named source and a standard error tells you the share; the label is a convention.
How to model Bengali ancestry with qpAdm and Global25#
A Bengali genome needs four streams in a qpAdm model, and Ancestrify's catalog carries an ancient source for each. The distal model uses Iranian Neolithic Farmer (8000 - 5000 BC), Ancestral South Indian (10000 - 2000 BC), Western Steppe Herder (5000 - 2800 BC) and Northeast Asian (6000 - 2000 BC). The fourth source is the one that makes Bengal different: a Bengali genome modelled without it will usually fail, and one modelled with it should show a significant weight near a tenth to a fifth. The proximal model swaps in regional references: Peninsular South Indian (BC 600 - 600 AD) as the Indian base and Gandharan Swat (BC 200 - 300 AD) or Loebanr Swat Iron Age (1200 - 800 BC) for the Steppe-carrying northwestern share, with Northeast Asian again for the eastern layer.
One honest limit: the catalog's Northeast Asian source is broad. It measures the East Asian related stream with a real error bar, but it cannot say whether that stream is Austroasiatic or Tibeto-Burman in origin, because the two are too similar at this depth for qpAdm to separate. The Global25 analysis, with its worldwide reference set of modern and ancient populations from Southeast Asia and the Himalaya, is the complementary step for that question, and Ancient Matches reports shared segments with individual ancient genomes from the wider region. All three are part of the ancestry service.
Sources and further reading#
- Narasimhan, V. M. et al. (2019). The formation of human populations in South and Central Asia. Science, 365, eaat7487.
- Chaubey, G. et al. (2011). Population genetic structure in Indian Austroasiatic speakers: the role of landscape barriers and sex-specific admixture. Molecular Biology and Evolution, 28, 1013 to 1024.
- Tätte, K. et al. (2019). The genetic legacy of continental scale admixture between Indians and Southeast Asians. Scientific Reports, 9, 3818.
- Basu, A., Sarkar-Roy, N. and Majumder, P. P. (2016). Genomic reconstruction of the history of extant populations of India reveals five distinct ancestral components and a complex structure. PNAS, 113, 1594 to 1599.
- Reich, D. et al. (2009). Reconstructing Indian population history. Nature, 461, 489 to 494.
Editorial note: the hero artwork in this article was generated with AI as conceptual illustration. It does not reproduce a scientific figure, an ancient individual, a real site or a measured migration route.



