The Pakistan Genomic Resource (PGR) is a landmark national genomics initiative led by the Center for Non-Communicable Diseases (CNCD), dedicated to advancing biomedical research globally and enabling the future of precision medicine in Pakistan.
Under the leadership of Professor Dr. Danish Saleheen, Principal Investigator of the program, PGR is building one of the largest and most comprehensive genomic resources in South Asia. The initiative combines genomic, clinical, biochemical and epidemiological data to better understand the genetic, biochemical and environmental factors that influence health and disease across Pakistan’s diverse populations.
PGR began as a focused cardiovascular case-control cohort and has scaled into nationwide genomic infrastructure. Enrollment is active and ongoing, expanding both statistical power and the catalogue of genetic variation with every new participant.
PGR is a nested case-control resource. Participants are recruited across multiple disease areas and paired with rich clinical data, making every genetic finding immediately interpretable against real physiology.
Myocardial infarction, hypertension, CAD, Stroke, Takkayasu, diabetes mellitus, lipid disorders, and obesity, recruited at national scale across hospital systems.
Chronic kidney disease and NAFLD cohorts, paired with biomarkers including eGFR, ALT and AST.
Including Parkinsons, Dementia Phenotyping for neurodegenerative risk and sensory traits.
Fertility history and developmental phenotypes captured across generations within consenting families.
Rheumatoid arthritis, Osteoarthritis, Systemic lupus erythematosus, Ankylosing spondylitis, Psoriatic arthritis, Scleroderma and immune-pathway phenotypes, reflecting Pakistan’s distinct disease burden and selection history.
Asthma, Chronic obstructive pulmonary disease
Diabetic retinopathy, Glaucoma, Keratoconus, High myopia, Age-related macular degeneration
A deep healthy-control base recruited alongside cases, essential for distinguishing disease-driving variants from background variation.
PGR is a nested case-control resource, recruiting both patients and healthy controls directly from the point of care across public and private hospitals in Pakistan. To minimize selection bias, enrollment spans both urban and suburban hospitals that also serve surrounding rural populations.
Participants are identified at the time of hospital presentation. Cases are identified according to predefined inclusion and exclusion criteria; phenotypes are defined from confirmed clinical diagnoses made by primary physicians, supported by laboratory investigations and standardized diagnostic criteria.
Control individuals are selected from the same hospitals as cases, typically unrelated or distantly related case attendants who meet control criteria — a design that preserves comparability while still capturing family structure.
Participants are enrolled following diagnostic confirmation, after both written and oral informed consent are obtained. The CNCD institutional review board and the National Bioethics Committee for Research, Pakistan approved the study.
Each participant undergoes standardized anthropometry, blood pressure measurement, and a baseline blood draw, with data and specimens logged into a central database for downstream sequencing and biomarker analysis.
Following consent, trained research staff administer questionnaires on past medical and family history in the participant’s local language — the structured context that turns a genetic finding into an interpretable one.
Sex is self-reported, alongside place of birth and self-reported ethnicity — the basis for PGR’s population genetic analyses across 14 ethnic groups.
30.6% of participants self-report parents who are first cousins — a single questionnaire item strongly predictive of carrying a homozygous loss-of-function variant.
Confirmed clinical diagnoses across cardiometabolic, renal, hepatic, neurological, and infectious disease, established by primary physicians and supported by laboratory investigations.
Past medical and family history is recorded during recall-by-genotype (RBG) interviews, administered to both the proband and consenting family members by trained research staff.
Tobacco usage (ever versus never) and years of formal education are captured, allowing genetic associations to be distinguished from behavioural and socioeconomic confounders.
Probands are recontacted under CNCD’s IRB approval; recruitment proceeds until all potentially recontactable individuals either consent or decline to participate.
Every PGR participant completes a standardized baseline visit, generating a uniform clinical and biological record alongside their genetic data, the foundation that makes later genotype-phenotype discovery possible.
Beyond recruitment and sequencing, PGR’s value comes from a structured programme of population-genetic, statistical, and translational analyses applied to the cohort. The sections below summarize the major analysis studies described in the recent flagship Nature publication.
Fixation index (FST) calculations, unsupervised admixture analysis (ADMIXTURE, k = 17), and principal component analysis place PGR’s 14 self-reported ethnicities relative to 1000 Genomes and Human Genome Diversity Project reference populations, characterizing Pakistan’s South and Central Asian genetic structure.
Runs of homozygosity (ROH) are called to calculate FROH for every participant. Mean FROH in PGR (0.034) falls between the theoretical expectations for first-cousin (0.0625) and second-cousin (0.016) parental relatedness, consistent with self-reported consanguinity.
Burden tests (regenie) correlate predicted loss-of-function and damaging missense variants against 46 circulating biomarkers across more than 500,000 gene-phenotype tests, validating known associations (e.g. APOC3-triglycerides, PCSK9-LDL-C) and revealing new directional findings (e.g. TIMD4, HBB).
Genome-wide autozygosity is tested against 16 quantitative and 18 binary clinical phenotypes, with sensitivity analyses stratified by ethnicity, sex, age, and self-reported consanguinity to rule out confounding.
Across all sequenced participants, 6,476 genes show at least one homozygous loss-of-function (homLoF) carrier. Subsampling experiments benchmark PGR’s discovery rate against each gnomAD population at matched sample sizes, and project whether discovery has reached saturation.
homLoF odds ratios and homozygote-depletion statistics are calculated across more than 100 gene sets, essentiality screens, ClinGen disease genes, drug targets, tissue-expression categories, and GSEA pathways, to map which biological pathways are essential for human life.
Probands carrying therapeutically relevant homLoF variants, and their consenting family members, are recontacted for deep phenotyping — including echocardiography, cold-pressor testing, and Sanger-sequencing confirmation of zygosity — enabling family-level genotype-phenotype analysis.
Select variants (for example, in RXFP1) are functionally validated outside the cohort: variant constructs are expressed in HEK293 cells, with protein expression and receptor signalling (cAMP response to relaxin-2) measured to confirm loss-of-function status.
Every participant consents at enrollment to future call-back. When a knockout carrier is identified, PGR can return — not just to that individual, but to their consenting relatives — turning a single rare finding into a family-level study.
Whether you’re validating a drug target, studying a rare phenotype, or building the next chapter of South Asian genomics. PGR’s growing resource, structured methodology, and recall-by-genotype capability are built to be worked with.