DNA and RNA Structure, Replication, Repair, Transcription, RNA Processing, Genetic Code
1. DNA Structure
Deoxyribonucleic acid (DNA) is the molecule that carries the genetic instructions for the development, functioning, growth, and reproduction of all known organisms and many viruses. It's a double helix, a structure that looks like a twisted ladder. The building blocks of DNA are called nucleotides. Each nucleotide consists of three parts: a deoxyribose sugar, a phosphate group, and a nitrogenous base. There are four types of nitrogenous bases in DNA: Adenine (A), Guanine (G), Cytosine (C), and Thymine (T).
The "sides" of the ladder are formed by alternating sugar and phosphate groups. The "rungs" of the ladder are formed by pairs of nitrogenous bases. These bases pair up in a specific way: Adenine always pairs with Thymine (A-T), and Guanine always pairs with Cytosine (G-C). This is known as complementary base pairing. The two strands of the DNA double helix run in opposite directions, a property called antiparallel. One strand runs in the 5' to 3' direction, and the other runs in the 3' to 5' direction. The 5' and 3' refer to the numbering of carbon atoms in the deoxyribose sugar molecule.
The structure of DNA was famously elucidated by James Watson and Francis Crick in 1953, based on the X-ray diffraction data produced by Rosalind Franklin and Maurice Wilkins. This discovery earned Watson, Crick, and Wilkins the Nobel Prize in Physiology or Medicine in 1962.
All Things good, Good Cookies.
A pairs with T; G pairs with C.
2. RNA Structure
Ribonucleic acid (RNA) is a nucleic acid essential in various biological roles in coding, decoding, regulation, and expression of genes. RNA differs from DNA in three main ways: it is usually single-stranded, it contains the sugar ribose instead of deoxyribose, and it uses the nitrogenous base Uracil (U) instead of Thymine (T). Like DNA, RNA has four bases: Adenine (A), Guanine (G), Cytosine (C), and Uracil (U). Adenine pairs with Uracil (A-U), and Guanine pairs with Cytosine (G-C).
There are several types of RNA, each with a specific function:
- Messenger RNA (mRNA): Carries genetic information from DNA in the nucleus to the ribosome in the cytoplasm, where it directs protein synthesis.
- Transfer RNA (tRNA): Acts as an adapter molecule, bringing specific amino acids to the ribosome during protein synthesis, matching them to the codons on the mRNA.
- Ribosomal RNA (rRNA): A structural and catalytic component of ribosomes, the cellular machinery responsible for protein synthesis.
- Small nuclear RNA (snRNA): Involved in RNA splicing and other nuclear processes.
- MicroRNA (miRNA) and small interfering RNA (siRNA): Involved in gene silencing and post-transcriptional regulation.
3. DNA Replication
DNA replication is the biological process of producing two identical replicas of DNA from one original DNA molecule. This process is fundamental to cell division and heredity. Replication is semi-conservative, meaning each new DNA molecule consists of one original (parental) strand and one newly synthesized strand.
The process begins at specific sites called origins of replication. Enzymes are crucial for replication:
- Helicase: Unwinds the DNA double helix by breaking the hydrogen bonds between complementary bases, creating a replication fork.
- Single-strand binding proteins (SSBs): Bind to the separated strands to prevent them from re-annealing.
- Topoisomerase (or DNA gyrase): Relieves the torsional strain caused by unwinding the DNA.
- Primase: Synthesizes a short RNA primer, which provides a free 3'-OH group for DNA polymerase to start adding nucleotides.
- DNA Polymerase: The main enzyme responsible for synthesizing new DNA strands. It reads the template strand and adds complementary nucleotides in the 5' to 3' direction. There are different types of DNA polymerases with specific roles (e.g., DNA Pol III for elongation, DNA Pol I for primer removal and gap filling).
- Ligase: Joins the Okazaki fragments on the lagging strand into a continuous DNA molecule.
Replication proceeds in two main directions from the origin:
- Leading Strand: Synthesized continuously in the 5' to 3' direction, following the replication fork.
- Lagging Strand: Synthesized discontinuously in short fragments called Okazaki fragments, also in the 5' to 3' direction, but in the opposite direction to the movement of the replication fork. Each Okazaki fragment requires its own RNA primer.
After the new DNA strands are synthesized, the RNA primers are removed and replaced with DNA nucleotides by DNA polymerase I, and the Okazaki fragments are joined by DNA ligase. The process results in two identical DNA molecules, each with one old and one new strand.
Helper Starts The Process, Polymerase Leads.
Helicase, SSBs, Topoisomerase, Primase, Polymerase, Ligase.
4. DNA Repair
DNA repair mechanisms are essential for maintaining the integrity of the genome. DNA is constantly subjected to damage from various sources, including environmental mutagens (like UV radiation and chemicals) and errors during DNA replication. If unrepaired, this damage can lead to mutations, which can cause diseases like cancer.
There are several major DNA repair pathways:
- Direct Reversal: Some specific types of damage can be directly reversed by enzymes. For example, photolyase uses visible light energy to break the covalent bonds in pyrimidine dimers caused by UV radiation.
- Base Excision Repair (BER): This pathway removes and replaces a single damaged or modified base. A DNA glycosylase recognizes and removes the damaged base, creating an apurinic/apyrimidinic (AP) site. AP endonuclease then cleaves the DNA backbone at the AP site, and DNA polymerase and ligase fill the gap and seal the nick.
- Nucleotide Excision Repair (NER): This pathway repairs bulky, helix-distorting lesions, such as pyrimidine dimers and chemically modified bases. It involves recognizing the lesion, excising a short oligonucleotide segment containing the damaged base, synthesizing a new segment using the undamaged strand as a template, and ligating the ends.
- Mismatch Repair (MMR): This system corrects errors that escape the proofreading activity of DNA polymerase, such as mismatched bases or small insertions/deletions. MMR proteins identify the mismatch, distinguish the newly synthesized strand (which contains the error) from the template strand, remove a segment of the new strand containing the mismatch, and resynthesize the correct sequence.
- Double-Strand Break Repair (DSBR): DNA double-strand breaks are particularly dangerous. Two major pathways exist:
- Non-homologous End Joining (NHEJ): This is the primary pathway in eukaryotes. It directly ligates the broken ends, often resulting in small insertions or deletions at the break site.
- Homologous Recombination (HR): This pathway uses an undamaged homologous chromosome or sister chromatid as a template to accurately repair the break. It is more accurate than NHEJ but is typically active only during the S and G2 phases of the cell cycle when a sister chromatid is available.
DNA repair is like a cellular "spell check" and "undo" button for genetic information, preventing mutations that could lead to disease.
5. Transcription
Transcription is the process of synthesizing an RNA molecule from a DNA template. It is the first step in gene expression, where the genetic information encoded in DNA is copied into a complementary RNA sequence. This process occurs in the nucleus of eukaryotic cells and in the cytoplasm of prokaryotic cells.
The key enzyme in transcription is RNA Polymerase. This enzyme binds to a specific region on the DNA called the promoter, which signals the start of a gene. Transcription involves three main stages:
- Initiation: RNA polymerase binds to the promoter region of the DNA. In prokaryotes, a sigma factor helps RNA polymerase recognize and bind to the promoter. In eukaryotes, transcription factors are required for RNA polymerase to bind to the promoter. Once bound, RNA polymerase unwinds a short segment of the DNA double helix.
- Elongation: RNA polymerase moves along the DNA template strand, reading the sequence and synthesizing a complementary RNA strand in the 5' to 3' direction. It uses ribonucleoside triphosphates (ATP, UTP, CTP, GTP) as substrates. As RNA polymerase moves, it unwinds the DNA ahead of it and rewinds the DNA behind it.
- Termination: Transcription stops when RNA polymerase encounters a termination signal on the DNA template. In prokaryotes, termination can occur via two mechanisms: Rho-dependent termination (involving the Rho protein) or Rho-independent termination (forming a hairpin loop in the RNA). In eukaryotes, termination mechanisms are more complex and often involve specific protein factors.
The resulting RNA molecule is called a primary transcript or pre-mRNA in eukaryotes. In prokaryotes, the transcribed RNA can often function immediately. The sequence of the RNA molecule is complementary to the template strand of the DNA and is almost identical to the coding strand (sense strand), with Uracil replacing Thymine.
Think of DNA as a master cookbook in a library (nucleus). Transcription is like making a photocopy (RNA) of a specific recipe (gene) to take to the kitchen (cytoplasm) to prepare the dish (protein).
6. RNA Processing (in Eukaryotes)
In eukaryotic cells, the primary RNA transcript (pre-mRNA) undergoes several modifications before it can be translated into protein. This process is called RNA processing and occurs in the nucleus. The main steps are:
- 5' Capping: A modified guanine nucleotide (7-methylguanosine cap) is added to the 5' end of the pre-mRNA molecule. This cap protects the mRNA from degradation by enzymes, helps in its export from the nucleus, and is recognized by ribosomes for initiation of translation.
- 3' Polyadenylation: A tail of about 50 to 250 adenine nucleotides (poly-A tail) is added to the 3' end of the pre-mRNA. This tail also protects the mRNA from degradation, aids in its export from the nucleus, and plays a role in translation initiation.
- Splicing: Eukaryotic genes often contain non-coding regions called introns interspersed among coding regions called exons. During splicing, the introns are removed, and the exons are joined together to form a continuous coding sequence. This process is carried out by a complex of proteins and small nuclear RNAs (snRNAs) called the spliceosome.
Alternative Splicing: A remarkable feature of eukaryotic gene expression is alternative splicing. This allows a single gene to produce multiple different mRNA molecules, and consequently, multiple different proteins, by varying which exons are included or excluded in the final transcript. This significantly increases the coding potential of the genome.
After processing, the mature mRNA molecule is exported from the nucleus to the cytoplasm, where it serves as the template for protein synthesis.
Cap at 5', Tail at 3', Splice the middle.
Capping, Tail (Polyadenylation), Splicing.
7. Genetic Code
The genetic code is the set of rules by which information encoded in genetic material (DNA or RNA sequences) is translated into proteins (amino acid sequences) by living cells. It is a nearly universal code, meaning it is the same for almost all organisms.
The genetic code is read in triplets of nucleotides called codons. Each codon specifies either a particular amino acid or a start/stop signal for translation.
- There are 64 possible codons (4 bases raised to the power of 3 positions: 43 = 64).
- Of these, 61 codons specify one of the 20 common amino acids.
- 3 codons (UAA, UAG, UGA) are stop codons, signaling the termination of translation.
- One codon (AUG) typically codes for the amino acid Methionine and also serves as the start codon, initiating translation.
The genetic code has several important properties:
- Triplet Nature: It is read in groups of three nucleotides.
- Universality: It is nearly the same in all organisms, from bacteria to humans. Minor variations exist, for example, in mitochondria.
- Degeneracy or Redundancy: Most amino acids are specified by more than one codon. This is because there are 61 codons for 20 amino acids. For example, Leucine is specified by six different codons. The first two bases of a codon are usually critical, while the third base (wobble position) can often vary without changing the amino acid specified.
- Non-overlapping: Codons are read sequentially without overlap. For instance, in the sequence AUG GCC UCA, the codons are AUG, GCC, and UCA, not AUG, UGC, GCC, CCU, etc.
- Specificity: Each codon specifies only one amino acid or a stop signal.
Understanding the genetic code is crucial for comprehending how genes are expressed and how proteins are synthesized. It forms the basis of molecular genetics and biotechnology.
| First Position | Second Position | Third Position | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| U | C | A | G | |||||||
| U | Phe (F) / Leu (L) | U | U – UUU, UUC (Phe) | C – UCU, UCC, UCA, UCG (Ser) | A – UAU, UAC (Tyr) | G – UGU, UGC (Cys) | U | |||
| A – UAA (Stop), UAG (Stop) | G – UGA (Stop), UGG (Trp) | |||||||||
| C | Leu (L) | U – CUU, CUC, CUA, CUG (Leu) | C – CCU, CCC, CCA, CCG (Pro) | A – CAU, CAC (His) | G – CGU, CGC, CGA, CGG (Arg) | U | ||||
| A – CAA, CAG (Gln) | G – CGA, CGG (Arg) | |||||||||
| A | Ile (I) / Met (M) | U – AUU, AUC, AUA (Ile) | C – ACU, ACC, ACA, ACG (Thr) | A – AAU, AAC (Asn) | G – AGU, AGC (Ser) | U | ||||
| A – AAA, AAG (Lys) | G – AGG (Arg) | |||||||||
| G | Val (V) | U – GUU, GUC, GUA, GUG (Val) | C – GCU, GCC, GCA, GCG (Ala) | A – GAU, GAC (Asp) | G – GGU, GGC, GGA, GGG (Gly) | U | ||||
| A – GAA, GAG (Glu) | G – GGG (Gly) | |||||||||
Focus on the first two bases for the general amino acid group. The third base often provides redundancy. Remember the start codon (AUG - Met) and the three stop codons (UAA, UAG, UGA).