← Back to Nuqta Atlas

Repository series · ncbi-prjeb5139

A systematic analysis of two whole genomes and thirteen exomes from Saudi Arabian tribes

Saudi ArabiaHuman health and population genomicsRepository record

The Kuwaiti population is composed of three genetic subgroups of Persian, Saudi Arabian tribe and Bedouin origin. The Saudi Arabian tribe subgroup traces its origin to the Najd region of Saudi Arabia. By sequencing two whole genomes and thirteen exomes from this subgroup at high coverage (>40X), we identify 4,950,724 Single Nucleotide Polymorphisms (SNPs), 515,802 indels and 39,762 structural variations. Of the total identified variants, 10,098 (8.3%) exomic SNPs, 139,923 (2.9%) non-exomic SNPs, 5,256 (54.3%) exomic indels, and 374,959 (74.08%) non-exomic indels are ‘novel’. Up to 8,070 (79.9%) novel biallelic exomic SNPs are seen in low frequency (minor allele frequency < 5%). We observe 5,462 known and 1,004 novel potentially deleterious nonsynonymous SNPs. Allele Frequencies of common SNPs derived using the 15 exomes is significantly correlated with those derived using genotype data from a larger cohort of 48 individuals (Pearson correlation coefficient, 0.91; p <2.2x10-16). A set of 2,485 SNPs show significantly different allele frequencies as compared to populations from other continents. Two notable variants having risk alleles seen in high frequencies in the Saudi Arabian tribe subgroup are: a nonsynonymous deleterious SNP (rs2108622, CYP4F2 gene) associated with warfarin dosage levels required to elicit normal anticoagulant response; and a 3’UTR SNP (rs6151429, ARSA gene) associated with Metachromatic Leukodystrophy. Hemoglobin Riyadh variant (identified for the first time in a Saudi Arabian woman) is observed in the presented exome data. The profile of the mitochondrial haplogroups derived from the 15 individuals is consistent with the haplogroup diversity seen in Saudi Arabian natives, who are believed to have received substantial gene flow from Africa and eastern provenance. We present the first genome resource for designing genetic studies in Saudi Arabian tribe subgroup.

Repository interpretation

The accession and its GCC connection are verified. Registration alone does not establish that the broader research programme remains active.

01 / Project overview

What the record establishes.

Geographic scope
Saudi Arabia connection indexed in BioProject metadata
Project type
Repository project
Research domain
Human health and population genomics
Years
2014–
Lifecycle status
repository_recorded
Status basis
Registered in NCBI BioProject on 2014/02/21; operational lifecycle is not asserted.
Status evidence date
2014-02-21
Scale
1 BioProject accession grouped by matching submitter, date, data type and narrative.

02 / Organizations and population

Who and what the project connects.

Lead organizations
DASMAN DIABETES INSTITUTE
Partner organizations
Not stated
Organism / population
Homo sapiens

03 / Data and access

What exists and how it can be reached.

Data types

  • Other
  • Sequencing

Data access

Public repository metadata with linked data where supplied by the submitter

Identifiers

  • BioProjectPRJEB5139

04 / Evidence and provenance

Why the record is included.

Inclusion basis

Exact country-name match in authoritative NCBI BioProject metadata; repeated submissions are grouped into one Atlas series.

Editorial note

Repository verification confirms the accession and regional connection, not whether the broader research programme remains active.

Sources

  1. primary record source Verified 2026-08-15

Release v0.2.0

Take the public record with you.

Projects CSVSources CSVIdentifiers CSV