Repository series · ncbi-prjeb5139
A systematic analysis of two whole genomes and thirteen exomes from Saudi Arabian tribes
The Kuwaiti population is composed of three genetic subgroups of Persian, Saudi Arabian tribe and Bedouin origin. The Saudi Arabian tribe subgroup traces its origin to the Najd region of Saudi Arabia. By sequencing two whole genomes and thirteen exomes from this subgroup at high coverage (>40X), we identify 4,950,724 Single Nucleotide Polymorphisms (SNPs), 515,802 indels and 39,762 structural variations. Of the total identified variants, 10,098 (8.3%) exomic SNPs, 139,923 (2.9%) non-exomic SNPs, 5,256 (54.3%) exomic indels, and 374,959 (74.08%) non-exomic indels are ‘novel’. Up to 8,070 (79.9%) novel biallelic exomic SNPs are seen in low frequency (minor allele frequency < 5%). We observe 5,462 known and 1,004 novel potentially deleterious nonsynonymous SNPs. Allele Frequencies of common SNPs derived using the 15 exomes is significantly correlated with those derived using genotype data from a larger cohort of 48 individuals (Pearson correlation coefficient, 0.91; p <2.2x10-16). A set of 2,485 SNPs show significantly different allele frequencies as compared to populations from other continents. Two notable variants having risk alleles seen in high frequencies in the Saudi Arabian tribe subgroup are: a nonsynonymous deleterious SNP (rs2108622, CYP4F2 gene) associated with warfarin dosage levels required to elicit normal anticoagulant response; and a 3’UTR SNP (rs6151429, ARSA gene) associated with Metachromatic Leukodystrophy. Hemoglobin Riyadh variant (identified for the first time in a Saudi Arabian woman) is observed in the presented exome data. The profile of the mitochondrial haplogroups derived from the 15 individuals is consistent with the haplogroup diversity seen in Saudi Arabian natives, who are believed to have received substantial gene flow from Africa and eastern provenance. We present the first genome resource for designing genetic studies in Saudi Arabian tribe subgroup.
The accession and its GCC connection are verified. Registration alone does not establish that the broader research programme remains active.
01 / Project overview
What the record establishes.
- Geographic scope
- Saudi Arabia connection indexed in BioProject metadata
- Project type
- Repository project
- Research domain
- Human health and population genomics
- Years
- 2014–
- Lifecycle status
- repository_recorded
- Status basis
- Registered in NCBI BioProject on 2014/02/21; operational lifecycle is not asserted.
- Status evidence date
- 2014-02-21
- Scale
- 1 BioProject accession grouped by matching submitter, date, data type and narrative.
02 / Organizations and population
Who and what the project connects.
- Lead organizations
- DASMAN DIABETES INSTITUTE
- Partner organizations
- Not stated
- Organism / population
- Homo sapiens
03 / Data and access
What exists and how it can be reached.
Data types
- Other
- Sequencing
Data access
Public repository metadata with linked data where supplied by the submitter
Identifiers
- BioProject
PRJEB5139
04 / Evidence and provenance
Why the record is included.
Inclusion basis
Exact country-name match in authoritative NCBI BioProject metadata; repeated submissions are grouped into one Atlas series.
Editorial note
Repository verification confirms the accession and regional connection, not whether the broader research programme remains active.
Sources
- primary record source Verified 2026-08-15
Release v0.2.0
