Skip to content
← Theoretical 1

Q25 - Nucleotide Diversity and Selective Sweeps

Theoretical 1 Real exam question - full text reproduced under IBO's CC BY-NC-SA 4.0 license

The average pairwise difference between a pair of sequences in a population is called nucleotide diversity, which is a simple measure of the genetic diversity of a population.

Figure 1 describes 3 alleles (A, B, C) in terms of their frequency in a given population (column 2). Column 3 provides the nucleotide sequences: for alleles B and C only the sequence differences from allele A are provided while positions with the same nucleotide as in A are marked with an asterisk “*”.

Table showing three alleles. Allele A (frequency 0.5) has the sequence GTATA GACTAGATCTACATGTAGCTGACTGCATATGCTGT. Allele B (frequency 0.3) differs from A at 4 positions. Allele C (frequency 0.2) differs from A at 2 positions. Figure 1. Three alleles with their frequencies and nucleotide sequences.

Nucleotide diversity can be calculated using the formula: pi = sum over all pairs (2 x pi x pj x dij), where pi and pj are frequencies of two alleles, and dij is the difference between them.

Figure 2 plots nucleotide diversity (y-axis) in 300 bp sliding windows along the sequence of gene Y (x-axis) for a modern cultivated crop plant (dashed line 2) and its wild ancestor (solid line 1). Below the plot, the structure of the gene is shown: boxes indicate exons, with grey areas being protein-coding regions and white areas being untranslated regions.

Plot showing nucleotide diversity along gene Y. The wild ancestor (solid line 1) shows high diversity in the upstream/promoter region that drops through the intron and rises again in the coding exons. The crop (dashed line 2) shows near-zero diversity in the upstream region but similar diversity to the wild type in the coding regions. Below the plot: A = Intron, B = Exon 1 (small, with grey coding region), C = Exon 2 (large, with grey coding region). Figure 2. Nucleotide diversity along gene Y. Line 1 (solid) = wild ancestor, Line 2 (dashed) = cultivated crop. A = Intron, B = Exon 1, C = Exon 2.

It has been concluded from the results in Figure 2 that a mutation in this locus was positively selected during the domestication process.

Q25.1. Calculate the nucleotide diversity in this population. Give the correct answer to 3 decimal places.
Q25.2. Cheetahs were found to have significantly lower genome-wide genetic diversity than other mammals. Select the likeliest explanation.
Q25.3. What is the most likely effect of the positively selected mutation in gene Y during domestication?
Q25.4.1. This data suggests purifying selection is acting more strongly on this locus in the wild ancestor than in the domesticated crop.
Q25.4.2. This data suggests, that in the wild ancestor, purifying selection is acting more strongly on the sequence upstream of exon 1 compared to the intron of gene Y.
Q25.4.3. No recombination occurred in the analysed region during the domestication process.

Question reproduced from IBO 2023, Theoretical Paper 1, licensed under CC BY-NC-SA 4.0 - attributed to the International Biology Olympiad. Open the full exam PDF