Biochem 501 – Section IV
Problem Set 1
“Nucleic Acid Structure, Function and Replication”
(Lectures 1-3).
I. Exploiting nucleotide structure for drugs.
1. Nucleoside analogues are a class of drugs widely used to treat diseases like cancer and
viruses. Two examples of such drugs in use are shown below:
H
For each drug, name the nucleobase present.
Zidovudine has a Thymine.
Gemcitabine has a Cytosine.
2. These drugs exert their influence when they are incorporated into DNA during genome
replication. Knowing this, explain why nucleoside kinases – enzymes that can transfer
phosphates to nucleosides – are needed for these nucleosides to have a pharmacological
effect.
Note that as nucleoside drugs these compounds lack 5’ phosphates. Because nucleotides (in
particular nucleoside triphosphates) are the actual substrates for incorporation by
polymerases, Zidovudine and Gemcitabine must first be converted into a nucleotide form to be
incorporated. Nucleoside kinases do this, allowing the drugs to have their impact.
Interestingly, as a consequence of this one mechanism of resistance to nucleoside drugs by
cancer or viruses is to decrease the activity of nucleoside kinases, thus preventing the
poisonous analogues from being incorporated into nucleic acids.
3. For each of the drugs, following incorporation of the nucleoside analogue into the growing
strand during DNA synthesis, is it possible for the DNA chain be further extended? Explain.
Remember that the key position to consider for extension of a nucleic acid is the 3’ position.
This is ordinarily where the 3’OH is that is used as the nucleophile to attack the phosphate on
the incoming nucleotide triphosphate.
In the case of Zidovudine, the 3’ position contains an azide (N3). This is not a 3’OH and
thus the chain would not be able to extend.
In Gemcitabine, although there are fluorines at the 2’ position of the ribose, the 3’ position still
contains an OH. Thus we would expect that if gemcitabine is incorporated, additional
nucleotides could still be added.
Interestingly, with respect to the efficacy of gemcitabine as a drug, the ability of the chain to be
extended beyond the initial incorporation makes it more difficult to be removed by DNA
polymerase 3’->5’ exonuclease activity compared to nucleoside analogues that block extension
entirely.
4. Gemcitabine is metabolized into 2′,2′-difluoro-deoxyuridine before it is incorporated into the
growing chain. The sugar pucker conformations for this derivative of gemcitabine are shown
below:
Based on the equilibrium arrow shown above, is the dominant conformer C3’endo or C2’endo?
Looking back at the slides we see that when the 3’ position is “up” like on the left, the sugar is
said to be in the C3’endo conformation. Because the equilibrium arrow implies that this
conformation dominates over the other, we would expect the dominant conformer to be C3’
endo.
Do you expect the backbone conformation to be A-form or B-form in character?
C3’ endo conformations place the phosphates of the backbone closer together. Thus we
would expect the backbone to have A-form character rather than B-form character.
II. Structure, Stability and Disruption of RNA Hairpins.
Recall that single-stranded nucleic acids can also form base pairing interactions with
themselves to produce secondary structure. Two examples of a structure that commonly forms
in single-stranded RNAs called a “hairpin” are shown below:
1. Which structure do you predict to be more stable? Explain your reasoning.
Hairpin A is made up of 13 A/U basepairs, while Hairpin B is made up of 13 G/C basepairs.
Because the structures are otherwise identical this means that Hairpin A is held together by 26
H-bonds (A/U makes 2-H bonds) while Hairpin B is held together by 39 H-bonds This mean
that hairpin B is more stable than hairpin A.
2. What enzymatic activity is needed to unwind these structures?
Helicases are the activity that is responsible for unwinding/melting duplex nucleic acid
structures.
3. There are many examples in biology in which a change in RNA secondary structure is related
to a change in activity of the RNA molecule. Consider two structures of the RNA shown below:
Which of the two structures is more stable? Explain.
The hairpin structure on the left contains 13 A/U base pairs adding 26 H-bonds to the
structure compared to the structure on the right. Thus we would expect the equilibrium would
favor the structure on the left under these conditions.
4. Now consider the system when a “trigger oligo” is included:
Which of the two structures is more stable? Explain.
The structure on the left still contains 13 A/U base pairs giving 26 H-bonds to that structure.
However with the trigger oligo included in the structure on the right, we now have 6 G/C base
pairs and 7 A/U base pairs, for a total of (3*6+2*7=18+14=32) 32 H-bonds. This suggests
that the structure on the right is now more stable than the structure on the left.
5. RNase V1 is a nuclease enzyme that specifically cleaves the phosphodiester backbone of
double-stranded regions of RNA. Would you expect the sequence region highlighted in yellow
to be cleaved less, the same, or more by RNAse V1 if the trigger oligo is included in the
reaction? Explain.
Because RNase V1 specifically cleaves double-stranded RNA, we need to see whether the
yellow-region of sequence in the structure is more double-stranded or less double-stranded
under each condition.
In the absence of trigger oligo, the hairpin structure is favored and a portion of the yellow
sequence Is double-stranded as it is making base-pairing interactions in the hairpin structure.
In the presence of trigger oligo, the non-hairpin structure is favored, and the yellow
sequence is not forming base-pairing interactions with any sequence now.
Thus we would expect the amount of yellow sequence cleaved by RNase V1 to decrease
if trigger oligo is included in the reaction.
III. Twisted Stuff.
Recall that the natural twist of a B-form dsDNA is given by:
Lk0= (# of base-pairs) / (10.9 base-pairs per turn)
1. Compute the natural twist, Lk0 for a dsDNA fragment that is 10900 base-pairs long.
Plug it in: 10900 base pairs / (10.9 base-pairs per turn) = 1000 turns
2. The observed topology for a DNA structure is given by the linking number:
Lk=Twist+Writhe
For each of the 5 observed topologies below, indicate whether it is underwound or overwound
with respect to the natural twist of the 10900 base-pair fragment.
A
B
C
D
E
Tw
1100
900
1200
600
1000
Wr
-100
200
-205
400
50
Remember that being underwound means the observed Lk (Lk=Tw+Wr) is less than Lk0, and
being overwound means the observed Lk is > Lk0. When Lk=Lk0 we say the DNA is relaxed.
We calculated Lk0 to be 1000, so we just compute the Lk for each condition and compare to
Lk0.
A: 1100-100=1000.
1000=1000,
so DNA is relaxed.
B: 900+200=1100.
1100>1000,
so DNA is overwound
C: 1200-205=995.
995<1000,
so DNA is underwound
D: 600+400=1000.
1000=1000,
so DNA is relaxed.
E: 1000+50=1050.
1050>1000,
so DNA is overwound.
3. Name the enzymatic activity that is required to change the topology of a DNA structure.
Topoisomerases are the enzymes that change DNA topology (I.e. can change the linking
number of the DNA structure).
4. Many of the above enzymes do not require ATP because they perform topological changes
that move the structure’s linking number closer to Lk0. In contrast, the bacterial enzyme Gyrase
spends large amounts of ATP to underwind genomic DNA so that Lk<<< Lk 0. What structural
advantage does this afford the cell?
Underwinding the DNA leads to negative supercoiling which helps package the DNA into
a smaller, more compact size.
To see why, note that if the DNA is underwound then Lk <<< Lk0. In principle any combination
of Twist and Writhe that sum to this underwound Lk value would be acceptable structures.
However, since DNA prefers a B-form structure, ideally the Twist component of Lk would still be
Lk0. However, to keep the linking number the same with Lk<Lk0 and Tw=Lk0, negative writhe
must be introduced into the structure. This negative writhe corresponds to negative supercoiling
which compacts the structure.
For example, if our Lk0 is 1000 and we underwind the DNA so that the Lk is 900. Then to
maintain DNA in a B-form structure we would have:
Lk=Tw+Wr
Lk=Lk0+Wr
Substitute in the values
900=1000+Wr
Wr=-100.
Thus underwound DNA will tend to become negatively supercoiled.
IV. Origin Stories.
DNA replication initiates from origin sequences in the genome.
The badger genome is ~1 billion base pairs long.
1. If DNA replication occurs at a rate of roughly 100 base pairs / second, how long would it take
to copy the badger genome if there were only a single origin sequence?
Simple calculation:
1,000,000,000 / (100 base pairs / second) = 10,000,000 seconds
That’s like 115 days, yikes!
2. Assuming all origins operate independently and are well distributed throughout the genome,
how many origins would be required to reduce the badger genome duplication time to less than
24 hours?
If we assume that the replication rate scales with the number of origins then we just need to do
some simple algebra to solve this.
Note that 24 hours is 86400 seconds.
1,000,000,000 / (100 bp/second * N number of origins) < 86400 seconds.
Rearrange to solve for N:
N> 1,000,000,000 / (100*86400)
N > 116 origins
The actual number of origins in a mammalian genome (where cells typically divide once every
24 hours) typically ranges from 60-150 so this back of the envelope calculate seems to make
sense.
3. An example of an origin sequence from E. coli is shown below.
CTATTTATTTAGAGATCTGTTCTATTGTGATCTCTTATTAGGATCGCACTGCCCTGTGG
ATAACAAGGATCCGGCTTTTAAGATCAACAACCTGGAAAGGATCATTAACTGTGAATGA
TCGGTGATCCTGGACCGTATAAGCTGGGATCAGAATGAGGGGTTATACACAACTCAAAA
ACTGAACAACAGTTGTTCTTTGGATAACTACCGGTTGATCCAAGCTTCC
This sequence is made up of 40% GC. The average GC% for the E. coli genome is 51%.
Explain why this helps make the origin sequence a useful place to initiate DNA replication.
Recall that for DNA polymerase to begin initiating from an origin of replication, the two strands
at the origin need to be “melted” and unwound by helicases. To melt DNA hydrogen
bonds between the two stands will be broken, and so the more A/T and less C/G content
the origin has, the easier it is for this origin to be unwound and initiate replication. Thus,
it is natural for an origin sequence to be more A/T rich than the rest of the genome.
V. Fork DNA.
Consider the replication fork at the exact moment in time shown below.
Suppose enzyme inhibitors are added to the reaction at the exact moment in time shown above
which can immediately inhibit 100% of the specified activity only. For each inhibitors below,
draw the expected nucleic acid fragments you would find at the end of the reaction.
1. Primase Inhibitor.
If primase is inhibited, no new primers can be synthesized but replication can otherwise
proceed.
2. Helicase Inhibitor.
Without helicase activity, replication cannot proceed past the dsDNA portion of the replication
fork.
3. Pol III inhibitor.
If Pol III is inhibited our main polymerase for DNA synthesis will be stopped. Pol I however can
still replace RNA primers that were put down w/ DNA.
4. Ligase inhibitor.
If ligase is inhibited, our Okazaki fragments will get made but not ligated together. The lagging
strand will end up being broken up and remain fragmented.
5. Pol I inhibitor.
If Pol I is inhibited, RNA primers will not be converted to DNA. Since DNA ligase ligates DNA to
DNA, the fragments will not be stitched together either.