Pakistan Genomic Resource
>
Pakistan Genomic Resource

South Asia’s Largest National-Scale Biobank

The Pakistan Genomic Resource (PGR) is a landmark national genomics initiative led by the Center for Non-Communicable Diseases (CNCD), dedicated to advancing biomedical research globally and enabling the future of precision medicine in Pakistan.


Under the leadership of Professor Dr. Danish Saleheen, Principal Investigator of the program, PGR is building one of the largest and most comprehensive genomic resources in South Asia. The initiative combines genomic, clinical, biochemical and epidemiological data to better understand the genetic, biochemical and environmental factors that influence health and disease across Pakistan’s diverse populations.

A Resource Built at National Scale

PGR began as a focused cardiovascular case-control cohort and has scaled into nationwide genomic infrastructure. Enrollment is active and ongoing, expanding both statistical power and the catalogue of genetic variation with every new participant.

300,000+ Participants enrolled across Pakistan, and growing
50 sites 23 cities
Across Pakistan
6,476+ Genes with a homozygous loss-of-function carrier identified
1 in 4 Participants yields a gene knockout — far above outbred biobanks
Map of PGR enrollment sites across Pakistan

Enrolled Under Dozens of Phenotypes

PGR is a nested case-control resource. Participants are recruited across multiple disease areas and paired with rich clinical data, making every genetic finding immediately interpretable against real physiology.

Cardiometabolic diseases

Myocardial infarction, hypertension, CAD, Stroke, Takkayasu, diabetes mellitus, lipid disorders, and obesity, recruited at national scale across hospital systems.

Renal & hepatic disease

Chronic kidney disease and NAFLD cohorts, paired with biomarkers including eGFR, ALT and AST.

Neurological & sensory

Including Parkinsons, Dementia Phenotyping for neurodegenerative risk and sensory traits.

Reproductive & developmental

Fertility history and developmental phenotypes captured across generations within consenting families.

Musculoskeletal & autoimmune/rheumatologic disease

Rheumatoid arthritis, Osteoarthritis, Systemic lupus erythematosus, Ankylosing spondylitis, Psoriatic arthritis, Scleroderma and immune-pathway phenotypes, reflecting Pakistan’s distinct disease burden and selection history.

Respiratory disease

Asthma, Chronic obstructive pulmonary disease

Ophthalmic disorders

Diabetic retinopathy, Glaucoma, Keratoconus, High myopia, Age-related macular degeneration

Healthy controls

A deep healthy-control base recruited alongside cases, essential for distinguishing disease-driving variants from background variation.

How Participants are Enrolled

PGR is a nested case-control resource, recruiting both patients and healthy controls directly from the point of care across public and private hospitals in Pakistan. To minimize selection bias, enrollment spans both urban and suburban hospitals that also serve surrounding rural populations.

01

Hospital-based identification

Participants are identified at the time of hospital presentation. Cases are identified according to predefined inclusion and exclusion criteria; phenotypes are defined from confirmed clinical diagnoses made by primary physicians, supported by laboratory investigations and standardized diagnostic criteria.

02

Control selection

Control individuals are selected from the same hospitals as cases, typically unrelated or distantly related case attendants who meet control criteria — a design that preserves comparability while still capturing family structure.

03

Written and oral consent

Participants are enrolled following diagnostic confirmation, after both written and oral informed consent are obtained. The CNCD institutional review board and the National Bioethics Committee for Research, Pakistan approved the study.

04

Baseline clinical workup

Each participant undergoes standardized anthropometry, blood pressure measurement, and a baseline blood draw, with data and specimens logged into a central database for downstream sequencing and biomarker analysis.

Enrollment Questionnaire

Following consent, trained research staff administer questionnaires on past medical and family history in the participant’s local language — the structured context that turns a genetic finding into an interpretable one.

Demographics & identity

Sex is self-reported, alongside place of birth and self-reported ethnicity — the basis for PGR’s population genetic analyses across 14 ethnic groups.

Consanguinity

30.6% of participants self-report parents who are first cousins — a single questionnaire item strongly predictive of carrying a homozygous loss-of-function variant.

Medical history

Confirmed clinical diagnoses across cardiometabolic, renal, hepatic, neurological, and infectious disease, established by primary physicians and supported by laboratory investigations.

Family history

Past medical and family history is recorded during recall-by-genotype (RBG) interviews, administered to both the proband and consenting family members by trained research staff.

Lifestyle factors

Tobacco usage (ever versus never) and years of formal education are captured, allowing genetic associations to be distinguished from behavioural and socioeconomic confounders.

Recall-by-genotype consent

Probands are recontacted under CNCD’s IRB approval; recruitment proceeds until all potentially recontactable individuals either consent or decline to participate.

What’s Measured at Enrollment

Every PGR participant completes a standardized baseline visit, generating a uniform clinical and biological record alongside their genetic data, the foundation that makes later genotype-phenotype discovery possible.

Insights on PGR Data

Beyond recruitment and sequencing, PGR’s value comes from a structured programme of population-genetic, statistical, and translational analyses applied to the cohort. The sections below summarize the major analysis studies described in the recent flagship Nature publication.

01

Population genetic architecture

Fixation index (FST) calculations, unsupervised admixture analysis (ADMIXTURE, k = 17), and principal component analysis place PGR’s 14 self-reported ethnicities relative to 1000 Genomes and Human Genome Diversity Project reference populations, characterizing Pakistan’s South and Central Asian genetic structure.

02

Autozygosity (runs of homozygosity)

Runs of homozygosity (ROH) are called to calculate FROH for every participant. Mean FROH in PGR (0.034) falls between the theoretical expectations for first-cousin (0.0625) and second-cousin (0.016) parental relatedness, consistent with self-reported consanguinity.

03

Biomarker burden association analyses

Burden tests (regenie) correlate predicted loss-of-function and damaging missense variants against 46 circulating biomarkers across more than 500,000 gene-phenotype tests, validating known associations (e.g. APOC3-triglycerides, PCSK9-LDL-C) and revealing new directional findings (e.g. TIMD4, HBB).

04

FROH-phenotype association analyses

Genome-wide autozygosity is tested against 16 quantitative and 18 binary clinical phenotypes, with sensitivity analyses stratified by ethnicity, sex, age, and self-reported consanguinity to rule out confounding.

05

homLoF gene discovery & subsampling

Across all sequenced participants, 6,476 genes show at least one homozygous loss-of-function (homLoF) carrier. Subsampling experiments benchmark PGR’s discovery rate against each gnomAD population at matched sample sizes, and project whether discovery has reached saturation.

06

Gene-set constraint (LoF intolerance)

homLoF odds ratios and homozygote-depletion statistics are calculated across more than 100 gene sets, essentiality screens, ClinGen disease genes, drug targets, tissue-expression categories, and GSEA pathways, to map which biological pathways are essential for human life.

07

Recall-by-genotype (RBG) studies

Probands carrying therapeutically relevant homLoF variants, and their consenting family members, are recontacted for deep phenotyping — including echocardiography, cold-pressor testing, and Sanger-sequencing confirmation of zygosity — enabling family-level genotype-phenotype analysis.

08

In vitro functional characterization

Select variants (for example, in RXFP1) are functionally validated outside the cohort: variant constructs are expressed in HEK293 cells, with protein expression and receptor signalling (cAMP response to relaxin-2) measured to confirm loss-of-function status.

Recall-by-genotype: Returning to the Family

Every participant consents at enrollment to future call-back. When a knockout carrier is identified, PGR can return — not just to that individual, but to their consenting relatives — turning a single rare finding into a family-level study.

Why This Matters for Discovery

Chart showing uniqueness of PGR in identification of human knockouts

Discover Biology That Exists Nowhere Else

Whether you’re validating a drug target, studying a rare phenotype, or building the next chapter of South Asian genomics. PGR’s growing resource, structured methodology, and recall-by-genotype capability are built to be worked with.