# INTRODUCTION

## A Bioinformatics Guide by the Armenian Bioinformatics Institute

Contemporary medicine, pharmaceuticals, and biotechnology have become increasingly data-driven fields, where the analysis of large-scale biological data serves as a cornerstone of progress and innovation. Revolutionary technologies, such as high-throughput sequencing, have transformed bio-related research and industries. However, these advances have also highlighted a critical gap: the shortage of specialists—bioinformaticians—who can convert vast datasets into actionable insights. Training experts in bioinformatics and genomics is therefore essential for advancing life sciences research.

To address this need, the Armenian Bioinformatics Institute (ABI, <https://abi.am>), in collaboration with the Bioinformatics Group at the Institute of Molecular Biology (IMB) NAS RA, has initiated an annual series of summer schools in genome bioinformatics. To extend the reach of these efforts, we have compiled the materials from these schools into this guide.

This guide provides you with the foundational knowledge and skills needed to embark on your first bioinformatics project. While this is not an exhaustive resource, it serves as a starting point, equipping you with the tools and inspiration to pursue further education in this dynamic field.

## Who Is The Guide For?

This guide is designed for biologists or data/computer scientists eager to embark on their journey in Bioinformatics and Genomics. However, if you do not fit into these categories but are still interested in exploring the fascinating world of science, we encourage you to challenge yourself and complete the guide. Lack of prior knowledge in either discipline should not hinder your progress (see details below).

The guide encompasses fundamental aspects of Molecular Biology, Programming in R and Python, basics of Unix command-line, Statistics, Experimental Techniques in Genetics and Genomics, Bioinformatics algorithms, Sequencing data analysis and Functional Genomics.&#x20;

## Advanced access and certification <a href="#advanced-access-and-certification" id="advanced-access-and-certification"></a>

We provide free access to this guide for anyone interested in starting their journey in bioinformatics. If you find our work valuable, we encourage you to support the Armenian Bioinformatics Institute with a donation. As a non-profit organization, your contributions help us continue developing training materials and resources.

To donate, visit:

{% embed url="<https://donate.abi.am>" %}

Thank you for supporting our mission to advance bioinformatics education!

## Credits

This guide includes both original content and links to external resources. The original content is authored by:&#x20;

* Created by the ABI team&#x20;
  * Final editing by: Maria Nikoghosyan and Lilit Nersisyan, PhD
* Original content in specific modules created by:&#x20;
  * Aleksey Kurnosov, PhD (Molecular Biology)
  * Vahan Huroyan, PhD (Statistics)
  * Susanna Avagyan (Single-cell gene expression data analysis)

## Follow us

Follow us on social media and stay updated on upcoming events, workshops, and new opportunities.

* [ABI website](https://abi.am)
* [Facebook](https://www.facebook.com/abi.arm.bio)
* [LinkedIn](https://www.linkedin.com/company/76327403/admin/dashboard/)
* [X (Twitter)](https://x.com/arm_abi)


# How to use the guide

This guide consists of several modules that you can take either in parallel or in order. Some modules are optional, depending on your background. Below are some tips to help you organize your time around these modules.

#### Molecular Biology

This module is designed for beginners in biology. If you're already familiar with the content, feel free to skip most of the sections. For those with a more advanced biology background, additional reading materials are provided that you may find useful.

#### Programming

* **Python**: This module is ideal if you're new to programming.
* **R**: If you've just mastered Python and don't want to feel overwhelmed, you may skip this module. The only place you'll need R in this guide is for a couple of practicals in the **Functional Genomics** module (e.g., gene expression data analysis).\
  However, if you already know Python or prefer R, you can take this module. Some practicals in the **Bioinformatics Algorithms** module are explained using Python, but the same tasks can be easily implemented in R if you're more comfortable with it.

#### Statistics

Statistics can be overwhelming for beginners, requiring time, patience, and persistence for the material to sink in. It might feel frustrating to focus on statistics when you're eager to start coding for bioinformatics. If that’s the case, you could lighten the mood by taking the **Bioinformatics Algorithms** module first or in parallel with statistics.

On the other hand, if you choose to tackle statistics first, you'll be well-prepared with both coding and statistical thinking by the time you reach the **Bioinformatics Algorithms** and **Functional Genomics** modules, which will make those easier to master.

You can also choose to follow the theory and practicals in order or mix them up and take them in parallel—practicing each concept in the programming language you prefer, whether that’s Python or R.

#### Bioinformatics Algorithms

This module provides an engaging introduction to bioinformatics, teaching you how to code and process genetic sequences. You can choose to take this module before or after statistics, or even in parallel. However, it's recommended to take it before the **Functional Genomics** module, as it will introduce you to the algorithmic bases of the command-line tools that you'll be using for alignment and other algorithms.

#### Sequencing Data Analysis and Functional Genomics

This module is your first real introduction to computational biology. While it’s not all-encompassing (many data modalities and analysis techniques, such as epigenomics and metagenomics, are not included), by the end of this module, you will have the foundation to explore further topics and resources on your own.

With that, you're ready to dive in! Enjoy the journey of learning and self-improvement.


# MOLECULAR BIOLOGY

## Introduction

This guide begins with a foundational introduction to molecular biology. You will learn about the cellular composition of life, how genetic material is encoded in DNA, and how it is transcribed into RNA and translated into proteins—the building blocks of life.

Understanding molecular biology is not only essential for following this guide but also serves to ignite your interest in the life sciences. While it is possible to perform bioinformatics tasks without a deep knowledge of biology, a great bioinformatician or computational biologist benefits immensely from understanding molecular biology. It provides insights into the challenges you are addressing, enables meaningful interpretation of data, and fosters the ability to formulate your own hypotheses. The more you delve into biology, the more engaging and rewarding bioinformatics becomes.<br>

### Credits

The main content of this module was written and edited by **Aleksey Kurnosov, PhD**.\
Additional material is provided to complement the core content.<br>

### Advanced material

This module offers the foundational knowledge necessary to begin exploring bioinformatics. For a deeper understanding of molecular biology, we recommend the book:

Alberts B, Johnson A, Lewis J, Raff M, Roberts K, and Walter P. *Molecular Biology of the Cell*. 6th Edition. Garland Science, 2014.

To access this book, please contact **<info@abi.am>**.


# The Cell


# Cells and Their Organelles

When using a microscope to magnify a section of tissue from any organism, you will observe that it is composed of units known as **cells**. The cell is a basic structural and functional unit of an organism.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F8RFqSREeXTDkkJJ1a69d%2FFig_1_%20cells.png?alt=media&amp;token=c671a2fc-cb78-49b6-aa85-97f57323f7f2" alt=""><figcaption><p>Cells are visible in the epithelium tissue of the mammalian gut (<em>left</em>) and onion root hair tip (<em>right</em>)<br>Image sources (from left to right):<br>Berkshire Community College Bioscience Image Library - Epithelial Tissues: Mucous Glands in Simple Columnar Epithelium, CC0, <a href="https://commons.wikimedia.org/w/index.php?curid=70159512">https://commons.wikimedia.org/w/index.php?curid=70159512</a><br>Johnathan Gutierrez, CC BY-SA 4.0, <a href="https://creativecommons.org/licenses/by-sa/4.0">https://creativecommons.org/licenses/by-sa/4.0</a>  </p></figcaption></figure>

Multicellular organisms, such as animals and plants, exhibit a complex organisation. For instance, the human body is comprised of approximately 30 trillion cells. In contrast, some organisms, like amoebae, consist of a single cell.\
Cells come in various shapes and sizes, each equipped with specialised functions that contribute to the overall functioning of living systems. Different functions of a cell can be assigned to specialised compartments called **organelles**.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2Fzs2XGnrm2dzPGk44dpFc%2FFig_2_%20cells.png?alt=media&amp;token=f03114f4-4853-4d51-aa55-27d4f35c8696" alt=""><figcaption><p>Some types of mammalian cells<br>Image sources (from left to right):<br>UC Regents Davis campus - http://brainmaps.org, CC BY 3.0, <a href="https://commons.wikimedia.org/w/index.php?curid=22012513">https://commons.wikimedia.org/w/index.php?curid=22012513</a><br>Electron Microscopy Facility at The National Cancer Institute at Frederick (NCI-Frederick) - [1], Public Domain, <a href="https://commons.wikimedia.org/w/index.php?curid=407197">https://commons.wikimedia.org/w/index.php?curid=407197</a><br>CC BY-SA 3.0, <a href="https://commons.wikimedia.org/w/index.php?curid=600755">https://commons.wikimedia.org/w/index.php?curid=600755</a><br>SubtleGuest at English Wikipedia, CC BY 2.5, <a href="https://commons.wikimedia.org/w/index.php?curid=19688924">https://commons.wikimedia.org/w/index.php?curid=19688924</a></p></figcaption></figure>

Cells can be categorised as **eukaryotic** or **prokaryotic**. Two large groups of organisms, **bacteria** and **archaea** have prokaryotic cells, while all other organisms have eukaryotic cells. Take a look at the diagram below and try to identify two major similarities and one difference between prokaryotic and eukaryotic cells.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FPunV1c20h8k3PvLNGUyh%2FFig_6_eukaryotic_prokaryotic_cells.png?alt=media&amp;token=7c5a5c4b-4d14-422a-ad51-b95a691bef4f" alt="" width="540"><figcaption><p>Bacterial (prokaryotic) cell on the left and animal (eukaryotic) cell on the right.<br>Image source: Biorender.com</p></figcaption></figure>

All cells share common characteristics, one of the most notable being the presence of a **plasma membrane** and **ribosomes**. The cell membrane serves to control what goes in and out of the cell. The ribosomes perform the synthesis of proteins.

The major difference is the presence of a **nucleus** and other membrane-enclosed organelles in eukaryotic cells. Membranes separate such organelles from the rest of the cell in order to maintain conditions (e.g. pH or concentration of some molecules) optimised to perform particular tasks. A jelly-like substance surrounding the organelles is called **cytosol**.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FglPmtmdFAeuZXyNHnMEa%2FFig_3_eukaryotic_cell.png?alt=media&amp;token=914cf52a-e6c4-4024-9159-e000f1bfc6be" alt=""><figcaption><p>Animal cell<br>Image source: Biorender.com</p></figcaption></figure>

The functions of major organelles that can be found in most eukaryotic cells are described in the table below.

<table data-full-width="true"><thead><tr><th width="157">Organelle</th><th>Function</th></tr></thead><tbody><tr><td>Nucleus</td><td>Stores most of the cell’s genetic material</td></tr><tr><td>Mitochondrion</td><td>The site of aerobic respiration, the process that produces adenosine triphosphate (ATP) – a molecule that cells use as a major source of energy</td></tr><tr><td>Ribosome</td><td>Makes proteins</td></tr><tr><td>Rough endoplasmic reticulum</td><td>Is covered by ribosomes. Transports proteins to other parts of the cell</td></tr><tr><td>Smooth endoplasmic reticulum</td><td>Makes lipids, stores calcium ions, breaks toxins</td></tr><tr><td>Golgi apparatus</td><td>Modifies proteins and packages them into vesicles</td></tr><tr><td>Lysosome</td><td>Digests molecules</td></tr><tr><td>Cytoskeleton</td><td>A network of protein filaments that defines the cell shape, enables the cell movement and intracellular transport</td></tr></tbody></table>

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FF1BKBrpeRRT4G1apV8jU%2FFig_4_plant_cell.png?alt=media&amp;token=a57793a8-5836-43d1-afb2-874e9419e0c1" alt="" width="540"><figcaption><p>Plant cell<br>Image source: Biorender.com</p></figcaption></figure>

**Plant cells** have some organelles distinguishing them from animal cells.

<table><thead><tr><th width="201">Plant cell organelle</th><th>Function</th></tr></thead><tbody><tr><td>Permanent vacuole </td><td>Contains the cell sap which stores degradative enzymes, waste products and inorganic ions</td></tr><tr><td>Chloroplast</td><td>Glucose is synthesised here in course of photosynthesis</td></tr><tr><td>Cell wall</td><td>A rigid structure surrounding the cell which protects the cell and defines its shape</td></tr></tbody></table>

**Prokaryotic cells** have no nucleus and other membrane-bound organelles. Many prokaryotes are surrounded by two membranes (gram-negative bacteria).

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2Fk2H0odBA9hsdkUjPj1YS%2FFig_5_bacterial_cell.png?alt=media&amp;token=17caff75-4ee8-4c28-80eb-bf51597d6a76" alt="" width="540"><figcaption><p>Bacterial cell<br>Image source: Biorender.com</p></figcaption></figure>

<table><thead><tr><th width="219">Bacterial cell organelle</th><th>Function</th></tr></thead><tbody><tr><td>Nucleoid</td><td>A region where the bacterial chromosome is located</td></tr><tr><td>Plasmid</td><td>Small circular DNA that often carries genes for antibiotic resistance</td></tr><tr><td>Flagellum</td><td>An organelle responsible for movement</td></tr></tbody></table>

## Summary

The video below provides a summary of the topic and some additional details:

{% embed url="<https://www.youtube.com/watch?v=URUJD5NEXC8>" %}


# Cell Specialisation

In multicellular organisms, functions are typically divided between specialised cell types. Cells of similar origin working together are organised into **tissues**. Despite having the same genetic material, the cells of an organism acquire distinct structural and molecular characteristics that enable them to perform their specific role. This process is known as cellular **differentiation**. During differentiation, different genes are selectively switched on and off, leading to the development of specialised cell types.

To illustrate the structural diversity that cells can achieve through differentiation, let us take a look at several examples from hundreds of specialised cell types of the human organism.&#x20;

The respiratory epithelium is a specialised tissue which lines the lung airways preventing the entry of pathogens and dust particles into the organism. Its predominant cell type is the ciliated cells. Ciliated cells are adapted to move mucus, a viscous fluid that traps pathogens and foreign particles, towards the pharynx. These elongated cells are characterised by the presence of numerous hair-like structures called cilia, which extend from their surface. The coordinated beating of cilia propels the mucus along the airways, clearing it from the lungs.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FCl2W1EUR5x0ZqT8smn6K%2FFig_63_lung_epithelium.jpg?alt=media&amp;token=056b710c-b22b-49fd-85df-1451fd8e4319" alt="" width="563"><figcaption><p>Light microscopy image of lung epithelium. Goblet cells secrete mucus and ciliated cells propel the mucus by the beating of cilia<br>Image source:<br>OpenStax College - Anatomy &#x26; Physiology, Connexions Web site. http://cnx.org/content/col11496/1.6/, Jun 19, 2013., CC BY 3.0, https://commons.wikimedia.org/w/index.php?curid=30148359</p></figcaption></figure>

Another example of highly specialised cells is macrophages. Macrophages, a subtype of white blood cells, engulf and digest pathogens, such as bacteria, as well as dying cells. Macrophages are capable of moving towards their targets by extending long flexible processes called pseudopodia. Upon encountering a pathogen, a macrophage initiates phagocytosis, ingesting the invader into a vesicle for subsequent digestion.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2Fjm0PMVrMWEH1rOnvITTR%2FFig_64_macrophage.jpg?alt=media&amp;token=ffbf626d-9bcd-4a07-a0cc-24d07620c791" alt="" width="375"><figcaption><p>Light microscopy image of macrophage stretching pseudopodia to engulf two particles<br>Image source:<br>Obli - Transferred from en.wikipedia to Commons., CC BY-SA 2.0, https://commons.wikimedia.org/w/index.php?curid=635700</p></figcaption></figure>


# Quiz 1

{% embed url="<https://docs.google.com/forms/d/e/1FAIpQLSeUEQggUsq8ju2MJOzY3Wc9M0gBGNraUghsFdD9Z69mMg6Vzw/viewform?usp=sharing>" %}


# Biological Molecules

The majority of biological molecules are **carbon**-based. Each carbon atom has the ability to form four **bonds**, typically with other carbon, hydrogen, oxygen, nitrogen, sulfur or phosphorus atoms.

When carbon atoms bond together, they create the **carbon skeleton**. The shape of this carbon skeleton, along with the positioning of other **chemical groups** defines the unique properties of a molecule.

The picture below provides an example of a biological molecule – serotonin. Serotonin is a signalling molecule used by neurons to communicate with each other.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FGPzDVtM3DVaLLyoRodPA%2FFig_7_serotonin.png?alt=media&amp;token=3ea361b9-b3e8-458f-81ae-e6992a8957cb" alt="" width="375"><figcaption><p>Serotonin is a biological molecule. Note that in this molecule carbon atoms form single or double bonds with other atoms.<br>Image source: Molview.org</p></figcaption></figure>

## What Are Polymers? <a href="#what-are-macromolecules" id="what-are-macromolecules"></a>

Proteins, nucleic acids and many of the carbohydrates are **polymers** – long molecules made of multiple repeating units called **monomers**. Monomers can be either identical (e.g. in cellulose) or similar (e.g. in DNA).

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2Fj6B9VMGrOFWLqGWnoAt4%2FFig_8_polymer.png?alt=media&amp;token=28890894-101d-4981-9f5f-dd41f2a6d57e" alt="" width="188"><figcaption><p>Formation and degradation of polymers<br>Image source:<br>Christinelmiller - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=98001922</p></figcaption></figure>

## Major Classes of Biological Molecules  <a href="#the-bare-minimum" id="the-bare-minimum"></a>

If you already have a biological background or do not want to indulge in the details of biological molecules' structure and functions, then consider the below information just enough to move on to the next chapter.

The four major classes of biological molecules are carbohydrates, proteins, nucleic acids (DNA and RNA) and lipids.

| Class                       | Function                                                                                              |
| --------------------------- | ----------------------------------------------------------------------------------------------------- |
| Carbohydrates               | Serve as an energy source and as structural components                                                |
| Proteins                    | Catalyse biochemical reactions, transport molecules, have structural, signalling and other structures |
| Nucleic acids (DNA and RNA) | Store and carry genetic information                                                                   |
| Lipids                      | Store energy, make up cell membranes and serve as hormones                                            |


# Carbohydrates

Carbohydrates are a group of organic molecules encompassing sugars (monosaccharides and disaccharides) and their polymers. Most carbohydrates have a general formula of Cn(H2O)m.

The nutritional sources of carbohydrates include table sugar (a common name for sucrose) and starchy products such as potatoes and cereals.&#x20;

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F7rM1CtjQP1gGIXk686m8%2FFig_20_table_sugar_and_starch.png?alt=media&amp;token=cdbffe15-06c9-4c24-b165-85a4a2a4cc9d" alt=""><figcaption><p>Table sugar and starch – the major dietary sources of carbohydrates<br>Image sources:<br>Emilian Robert Vicol from Com. Balanesti, Romania - White-Sugar-Crystals_91637-480x360, CC BY 2.0, https://commons.wikimedia.org/w/index.php?curid=38382827<br>MoRntashamL GOLD - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=125835813</p></figcaption></figure>

## Monosaccharides

Monosaccharides are simple sugars typically having a general formula of (CH2O)n. The number of carbon atoms in monosaccharides ranges from 3 to 7. The common example is a six-carbon sugar – **glucose**, with a general formula of C6H12O6. Like most monosaccharides, glucose adopts a ring structure in aqueous solutions.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2Fbk8LqGAXma9RoAmzAYbs%2FFig_9_glucose.png?alt=media&amp;token=952bb96e-b2eb-4a30-8749-179e839667ba" alt="" width="194"><figcaption><p>Glucose</p></figcaption></figure>

Glucose is the major source of energy for the cell. The release of energy from monosaccharides occurs through cellular **respiration**, a process that involves breaking them down into smaller molecules.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FT4VskuTOwt0DeRGHE61B%2FFig_10_respiration.png?alt=media&amp;token=7e4cf410-b40e-434b-bf73-a2349b533a56" alt=""><figcaption><p>Equation summarising cellular respiration<br>Image source:<br>BChristinelmiller - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=98216277</p></figcaption></figure>

Another function of monosaccharides is to serve as building material for other kinds of molecules, such as amino acids.

## Dehydration synthesis

Two monosaccharide units can be joined together in a dehydration reaction.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FqVkaiOhyioBFADLimaNg%2FFig_11_dehydration.png?alt=media&amp;token=0ecb46ee-5dc0-4332-90f7-caf0d2a5a001" alt=""><figcaption><p>Two glucose units are joined via a dehydration reaction to form a disaccharide called maltose<br>Image source:<br>CNX OpenStax - http://cnx.org/contents/GFy_h8cu@10.53:rZudN6XP@2/Introduction, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=49922728</p></figcaption></figure>

The most common **disaccharides** formed by the dehydration reaction are sucrose ("table sugar"), maltose and lactose. The addition of more sugar units to the chain results in the synthesis of polysaccharides.

## Polysaccharides

Polysaccharides are large polymers consisting of hundreds to thousands of monosaccharide units. Polysaccharides serve either energy storage or structural functions.

Both plant and animal cells use polymers made of glucose monomers to store energy. Glucose can be quickly retrieved from these polymers by a reaction called **hydrolysis.** Unlike glucose, these polysaccharides are insoluble, resulting in compact storage.

The plant storage polysaccharide is **starch**. Starch consists of two types of molecules: **amylose**, a linear molecule, and **amylopectin**, a more complex, branched molecule. Amylose allows a slower release of glucose, while amylopectin can be broken down faster due to more free ends.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FOSQl2VpQpZ2FyfZnX5pO%2FFig_12_starch.png?alt=media&amp;token=5ffa565c-c4d2-43a8-b5e8-57db29ffe9a1" alt=""><figcaption><p>Starch consists of two types of molecules<br>Image source: Biorender.com</p></figcaption></figure>

The animal storage polysaccharide **glycogen** is similar to amylopectin but is more branched. Glycogen is predominantly stored in liver and muscle cells, though other cells also contain a limited amount of glycogen.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FC9CpXrGGWEPaD8R0oGMi%2FFig_13_glycogen.png?alt=media&amp;token=f580b1d0-a88f-4b3f-8ccf-bf89456d5f61" alt="" width="540"><figcaption><p>Glycogen has plenty of branching points<br>Image source: Biorender.com</p></figcaption></figure>

**Cellulose**, another polymer of glucose, is the most abundant structural polysaccharide in the world. Numerous linear and unbranched cellulose molecules align and are connected by weak electrostatic attractions known as **hydrogen bonds**. Aggregated cellulose molecules, forming structures called microfibrils, constitute a major component of plant cell walls.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2Falj4QCQywAVEQBjrCgI2%2FFig_14_cellulose.jpeg?alt=media&amp;token=c5107d9e-cb62-4fcb-bb43-7fa97a4983e1" alt=""><figcaption><p>Cellulose molecules make up microfibrils in the plant cell wall</p></figcaption></figure>

&#x20;

## Summary

The videos below provide a summary of the topic and some additional details:

{% embed url="<https://youtu.be/-Aj5BTnz-v0>" %}

{% embed url="<https://youtu.be/FEAXI5XeJ4M>" %}

{% embed url="<https://youtu.be/SOQyiM6V3RQ>" %}


# Lipids

Lipids are a diverse group of biological molecules that do not mix with water. The key groups of lipids encompass fats, phospholipids, steroids and waxes. Lipids serve energy storage functions, are the major components of biological membranes, and act as precursors for hormones.

Many lipids can be synthesised by human cells. However, some essential lipids such as omega-3 and omega-6 fatty acids have to be obtained from diet.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FfXK0pLvf7uYk6qptWSuy%2FFig_21_lipids_diet.png?alt=media&amp;token=000a13ae-a181-4cd3-ac55-00c526ebff74" alt="" width="375"><figcaption><p>Some dietary sources of essential lipids</p></figcaption></figure>

## Fats

A typical fat molecule is a **triglyceride** consisting of a **glycerol** part and three **fatty acid** tails. The long fatty acid tails provide fats with their hydrophobic properties.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F8Yb0dX2VbGbaTVjAUCQE%2FFig_15_triglyceride.jpeg?alt=media&amp;token=987712e2-d444-4830-bf2d-f145182271cc" alt=""><figcaption><p>Fats are synthesised by dehydration synthesis<br>Image source:<br>OpenStax College - Anatomy &#x26; Physiology, Connexions Web site. http://cnx.org/content/col11496/1.6/, Jun 19, 2013., CC BY 3.0, https://commons.wikimedia.org/w/index.php?curid=30131156</p></figcaption></figure>

In animals, fats serve as long-term energy storage molecules. Fats are mainly deposited in **adipose tissue** cells that provide thermal insulation and cushion the body's organs.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FIuJxUhC45q7wvS0LsYes%2FFig_16_adipose_tissue.jpeg?alt=media&amp;token=cd83e585-0569-4b33-9def-b74fc75f7da2" alt=""><figcaption><p>Fat is stored in adipose tissue cells<br>Image source:<br>OpenStax College - Anatomy &#x26; Physiology, Connexions Web site. http://cnx.org/content/col11496/1.6/, Jun 19, 2013., CC BY 3.0, https://commons.wikimedia.org/w/index.php?curid=30131287</p></figcaption></figure>

## Phospholipids

Phospholipids are the key components of cell membranes. Phospholipids are structurally similar to fats but only have two fatty acid tails. Instead of the third tail, phospholipids have a negatively charged **phosphate group**. This arrangement results in a dual nature: a strongly hydrophobic part ("tails") that repels water and a hydrophilic part ("head") that attracts water.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FQxrDiJ0YmF99MYRcD2hg%2FFig_17_phospholipid.jpeg?alt=media&amp;token=135ccc7b-b909-4db9-94ba-e63d4fae9d89" alt="" width="255"><figcaption><p>The structure of a phospholipid<br>Image source:<br>OpenStax - https://cnx.org/contents/FPtK1zmh@8.25:fEI3C8Ot@10/Preface, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=30131167</p></figcaption></figure>

In an aqueous environment, phospholipids tend to form a **bilayer**: the hydrophilic heads face outward, interacting with water, while the hydrophobic tails are concealed within. Phospholipid bilayers comprise the plasma membrane and the internal membranes of eukaryotic cells.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F2gLHQGC0zrPKGd7JtgUc%2FFig_18_phospholipid_bilayer.jpeg?alt=media&amp;token=a30c89c6-60bd-4af1-8152-df2dfe59ce73" alt=""><figcaption><p>Phospholipid bilayer<br>Image source:<br>OpenStax - https://cnx.org/contents/FPtK1zmh@8.25:fEI3C8Ot@10/Preface, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=30131169</p></figcaption></figure>

## Steroids

Steroids are lipid molecules that have four fused hydrocarbon rings. A  typical steroid is **cholesterol** – the component of animal cell membranes and a precursor of such hormones as estradiol and testosterone.&#x20;

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FjWLnNSlWpJ8NtRPrtOIt%2FFig_19_cholesterol.png?alt=media&amp;token=844a1990-fed9-46cb-9963-cf35462167ae" alt="" width="375"><figcaption><p>Cholesterol</p></figcaption></figure>

***

## Summary

The videos below provide the summary of the topic and some additional details:

{% embed url="<https://youtu.be/OpyTJbzA7Fk>" %}

{% embed url="<https://youtu.be/O9lL2KStW9s>" %}

{% embed url="<https://youtu.be/Ezp8F7XJHWE>" %}


# Nucleic Acids (DNA and RNA)

### Introduction

Nucleic acids are polymers utilised by cells as information molecules. The genetic information is stored in the form of **deoxyribonucleic acid** (**DNA**) providing instructions for the synthesis of proteins needed to build and maintain functioning cells, tissues, and organisms. Notably, some viruses employ **ribonucleic acid** (**RNA**) for storing genetic information. Within cells, RNA serves to carry the instructions from the DNA to the protein-synthesising machines, known as ribosomes.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FtfywDSz0ZBJBhj9xOT6T%2FFig_22_Transcription-translation_Summary.jpg?alt=media&amp;token=abd7e06f-7509-48da-b4c4-20114de442a5" alt="" width="274"><figcaption><p>DNA is the template for the synthesis of RNA which is decoded by a ribosome to make a protein<br>Image source:<br>OpenStax - https://cnx.org/contents/FPtK1zmh@8.25:fEI3C8Ot@10/Preface, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=30131214</p></figcaption></figure>

The choice of nucleic acids as information molecules is determined by their ability to drive their self-replication. The **replication** of DNA before cell division enables each cell to have its copy of the genetic material.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FA972QUbZ1oe641HfHHTs%2FFig_23_replication.jpg?alt=media&amp;token=e8b02add-c4fc-4ec6-874a-984c1207b575" alt="" width="150"><figcaption><p>The structure of DNA enables a replication mechanism in which two identical copies of the parental molecule are produced<br>Image source:<br>Genomics Education Programme - Semi conservative replication of DNA, CC BY 2.0, https://commons.wikimedia.org/w/index.php?curid=50542923</p></figcaption></figure>

## Nucleic acids are made of nucleotides&#x20;

Nucleic acids are often referred to as **polynucleotides** since they consist of monomers called nucleotides.&#x20;

Each **nucleotide** comprises three components: a five-carbon sugar, a phosphate group and a nitrogenous base. In RNA, the sugar is **ribose**, while DNA incorporates **deoxyribose**. The nitrogenous bases are attributed to one of two families: purines and pyrimidines. **Purines**, distinguished by their larger size and two fused rings, include **adenine** (**A**) and **guanine** (**G**). **Pyrimidines** contain one ring and include **cytosine** (**C**), **thymine** (**T**) and **uracil** (**U**). Thymine is exclusive to DNA, while uracil is only found in RNA.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F9VHyGsQBAnF1DplpQCYk%2FFig_24_nucleic_acid.png?alt=media&amp;token=66dac95d-165a-4495-8a8a-40fdc902a603" alt=""><figcaption><p>Nucleotides make up nucleic acids<br>(a) Each nucleotide contains a pentose sugar, a phosphate and a nitrogenous base<br>(b) Nitrogenous bases<br>(c) Fice-carbon sugars of DNA and RNA<br>Image source:<br>OpenStax College - Anatomy &#x26; Physiology, Connexions Web site. http://cnx.org/content/col11496/1.6/, Jun 19, 2013., CC BY 3.0, https://commons.wikimedia.org/w/index.php?curid=30131164</p></figcaption></figure>

## DNA and RNA structure

In a polynucleotide, nucleotides are joined by **phosphodiester linkages** formed through condensation reactions. This bonding creates the **sugar-phosphate backbone**, characterized by an alternating pattern of sugars and phosphates. The two ends of a polynucleotide molecule feature distinct chemical groups. The end with a phosphate group is termed the **5' end**, while the end with a hydroxyl group is called the **3' end**.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FJ7Z6JXIAm5Pmi0P2tkC9%2FFig_25_DNA_structure.png?alt=media&amp;token=0c2273bc-2d89-4186-a2e4-7d6670231ccf" alt="" width="375"><figcaption><p>DNA structure<br>Image source:<br>Zephyris - Own work, CC BY-SA 3.0, https://commons.wikimedia.org/w/index.php?curid=15027555</p></figcaption></figure>

A DNA molecule is comprised of two polynucleotide strands that intertwine, forming a **double helix**. The strands are referred to as antiparallel as they run in opposite 5' to 3' directions. The nitrogenous bases are positioned inwardly towards the double helix from the sugar-phosphate backbone. Thus, the nitrogenous bases of the two strands face each other forming **complementary base pairs**. In DNA, adenine (A) always pairs with thymine (T), while guanine (G) pairs with cytosine (C). The complementary bases are held together by **hydrogen bonds**. There are two hydrogen bonds between A and T, and three hydrogen bonds between G and C.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FoXLOnJ2Xp1TaoEKQ13iH%2FFig_26_tRNA.png?alt=media&amp;token=3b998db2-b37e-46fd-8427-97e694d599cf" alt="" width="301"><figcaption><p>Left: the regions of complementarity in tRNA<br>Right: tRNA 3D shape<br>Image source:<br>Kyle Schneider (SchneiderKD) (Transfered by BQmUB2010090/Original uploaded by SchneiderKD) - Schneider KD (Original uploaded on en.wikipedia), Public Domain, https://commons.wikimedia.org/w/index.php?curid=12309266</p></figcaption></figure>

The majority of RNA molecules are **single-stranded**. However, the complementary base-pairing between regions of the same RNA molecule can make it fold into a complex three-dimensional shape. A striking example is the 3D L-shape of transfer RNA (tRNA) molecules that serve to carry amino acids to the ribosome.

##

## Summary

The videos below provide the summary of the topic and some additional details:

{% embed url="<https://youtu.be/AmOO4j0E408>" %}

{% embed url="<https://youtu.be/jUUJSOM1ihU>" %}


# Quiz 2

[Quiz 2. Carbohydrates, lipids and nucleic acids](https://docs.google.com/forms/d/e/1FAIpQLScto_Lpc7PGTt956xSxoB2BPxyFNhIOEgLGZZWdgz6lSP18pQ/viewform?usp=sf_link)


# Proteins

## Introduction

Proteins play a pivotal role in virtually every cellular function. They are responsible for molecular transport, provide structural support, and participate in defence mechanisms. Proteins, known as **enzymes**, act as catalysts, speeding up biochemical reactions. Other proteins are responsible for intercellular communications.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FTldi3gsgvJpXs7rdozKL%2FFig_27_proteins_functions.png?alt=media&amp;token=47e1ac05-8d7c-41d0-8e3a-646a97a675f9" alt="" width="540"><figcaption><p>Functions of proteins<br>Image source: Biorender.com</p></figcaption></figure>

## Proteins are made of amino acids

A polymer formed by the linkage of amino acids is termed a **polypeptide** and is called a **protein** when it serves a biological function. Any compound containing both an amino group (-NH2) and a carboxylic group (-COOH) can be broadly termed an **amino acid**. Yet, it is specifically the α-amino acids, where the amino and carboxylic groups are attached to the same carbon atom, that are the building blocks of proteins. The **side chain**, which can as well be called the R group, is variable among amino acids.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FdrdR8uhLdY7SpT18ax2L%2FFig_28_amino_acid.png?alt=media&amp;token=e07ae412-b57c-44a7-a2e6-0bf9972dae22" alt="" width="375"><figcaption><p>Amino acid structure<br>Image source:<br>OpenStax College - Anatomy &#x26; Physiology, Connexions Web site. http://cnx.org/content/col11496/1.6/, Jun 19, 2013., CC BY 3.0, https://commons.wikimedia.org/w/index.php?curid=30131160</p></figcaption></figure>

A total of 22 amino acids can be incorporated into polypeptides by ribosomes. Notably, among these, 20 amino acids are considered "canonical" and are universally present in all organisms. The distinctive properties of a particular amino acid are defined by its side chain, which can be categorised as negatively charged (acidic), positively charged (basic), nonpolar (hydrophobic), or polar (hydrophilic).&#x20;

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FHIJQMyDMPBY2fAq018Or%2FFig_29_amino_acids.png?alt=media&amp;token=95c1e06c-1b67-4ee3-b120-b6e0d3a17f77" alt="" width="563"><figcaption><p>Proteinogenic amino acids<br>Image source:<br>CNX OpenStax - http://cnx.org/contents/GFy_h8cu@10.53:rZudN6XP@2/Introduction, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=49923700</p></figcaption></figure>

A linkage between a carboxyl group of one amino acid and an amino group of another amino acid is called a **peptide bond**.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F3CeJs9tN544fJEEEOjiz%2FFig_30_peptide_bond.svg?alt=media&amp;token=a6e723e5-f60f-4779-b9f3-795681a09bd5" alt="" width="375"><figcaption><p>Formation of a peptide bond by a dehydration reaction<br></p></figcaption></figure>

## Levels of protein structure

Following synthesis by a ribosome, a polypeptide undergoes a folding process, dictated by the sequence of amino acids. The resulting complex shape of the protein and the physical properties of its surface govern its capability to interact with other molecules and determine its function. The structure of proteins can be classified into four levels: primary, secondary, tertiary, and, in some cases, quaternary.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FxxWC4ZuFEMOh8Ujg00zp%2FFig_31_levels_of_structure.jpg?alt=media&amp;token=b48ab740-1184-49c2-8200-fc62decc845f" alt="" width="563"><figcaption><p>Levels of protein structure<br>Image source:<br>CNX OpenStax - https://cnx.org/contents/5CvTdmJL@4.4, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=53712842</p></figcaption></figure>

The **primary structure** is the sequence of amino acids in a peptide. The primary structure of a protein is defined by the genetic information encoded in the DNA.&#x20;

The **secondary structure** is a regular local conformation of the polypeptide backbone stabilised by hydrogen bonds between carboxyl oxygen atoms and hydrogen atoms attached to nitrogen atoms. Two common types of secondary structures include the spiral-shaped **alpha helix** and the **beta-pleated sheet**, the latter formed by two or several parallel polypeptide segments.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FedDQ10Sb1IHlMY4NVw1z%2FFig_32_secondary_structure.jpg?alt=media&amp;token=4542c4e1-bab6-4d26-aca5-d23edc2a54c9" alt="" width="563"><figcaption><p>Secondary structures: alpha helix and beta-pleated sheet<br>Image source:<br>CNX OpenStax - https://cnx.org/contents/5CvTdmJL@4.4, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=53712849</p></figcaption></figure>

The **tertiary structure** refers to the overall three-dimensional shape of a polypeptide and is stabilised by interactions between amino acids' R groups. These interactions include hydrogen bonds, ionic bonds, disulfide bridges formed between cysteine residues, van der Waals forces, and hydrophobic interactions. The latter entails the grouping of hydrophobic amino acid residues at the core of a protein globule to minimise contact with water, while hydrophilic amino acid residues predominantly occupy the surface, exposed to the aqueous environment.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2Fl4Xo56fpEa41lQvusR27%2FFig_33_tertiary_structure.jpg?alt=media&amp;token=c75c23cd-9383-4148-914c-409f04c7667e" alt="" width="563"><figcaption><p>Interactions stabilising the tertiary structure<br>Image source:<br>CNX OpenStax - https://cnx.org/contents/5CvTdmJL@4.4, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=53712850</p></figcaption></figure>

The **quaternary structure** represents the overall structure of a protein, formed by the arrangement of two or several polypeptides referred to as subunits. A classic example is haemoglobin consisting of two α globin and two β globin subunits, each carrying a non-protein haem group.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FeHZFxPzoIFLAG0hMGwww%2FFig_34_haemoglobin.png?alt=media&amp;token=8803e010-a525-44ea-8312-8693bdda3d78" alt="" width="375"><figcaption><p>Haemoglobin structure. α globin and β globin subunits are shown in red and blue respectively. Haem groups are shown in green.<br>Image source:<br>Zephyris at the English-language Wikipedia, CC BY-SA 3.0, https://commons.wikimedia.org/w/index.php?curid=2300973</p></figcaption></figure>

Overall, the diversity of functions performed by proteins is reflected by the diversity of their shapes.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FmrTcA6myEuoiWwCl0Wb6%2FFig_80_protein_shapes.png?alt=media&amp;token=b885eec5-b457-4afd-bb39-59a76ea92413" alt=""><figcaption><p>Examples of proteins' 3D shapes<br>Image source:<br>David Goodsell, https://book.bionumbers.org/how-big-is-the-average-protein/, with changes</p></figcaption></figure>


# Catalysis of Biological Reactions

Numerous chemical reactions within cells would proceed at unfeasibly slow rates under the conditions in which biological organisms exist. This sluggishness arises from the fundamental nature of chemical reactions, entailing breaking and forming bonds, which requires reactant molecules to convert to unstable transition states. The energy that needs to be absorbed from the environment to reach the transition state is termed **activation energy**.&#x20;

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FdrL7vXaZKiMnkeX62XfG%2FFig_35_activation_energy.png?alt=media&amp;token=1493a2f5-a003-491c-93a2-791fd28a8d2e" alt="" width="375"><figcaption><p>Activation energy of a reaction with and without enzyme<br>Image source:<br>Microbialmatt - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=126786827</p></figcaption></figure>

**Catalysts** are substances that accelerate reactions without being consumed and without altering the reaction equilibrium. In biological systems, these catalysts are primarily proteins known as **enzymes**. Enzymes facilitate chemical reactions by lowering the activation energy required for the reaction to occur.

In enzyme-catalyzed reactions, the molecules undergoing transformation are referred to as **substrates**. The substrate binds to a specific region on the enzyme known as the **active site**, forming an **enzyme-substrate complex**. The active site's shape and physical properties are tailored to accommodate only particular substrates, which explains the specificity of enzymes to their substrates.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F2TVkEXf0YrOHzMXtyqAq%2FFig_36_enzyme.png?alt=media&amp;token=0380a56c-1576-4d78-8489-906f456aae17" alt="" width="563"><figcaption><p>The induced fit model. Substrate binding makes the enzyme undergo a conformational change</p></figcaption></figure>

During catalysis, the substrate remains bound to the active site, undergoing conversion into products. This interaction is characterised by an **induced fit** model, where the active site's shape is not rigid but rather adjusts upon substrate binding. Weak interactions between chemical groups on the enzyme and substrate cause the active site to conform around the substrate.

Enzymes employ diverse mechanisms to reduce activation energy in chemical reactions. One approach involves properly orienting substrate molecules within the active site, facilitating bond formation. Alternatively, enzymes may distort substrate molecules, promoting bond-breaking. Some enzymes create a microenvironment within their active sites that favours specific reactions. For instance, the presence of acidic amino acid residues can lower pH, increasing the likelihood of hydrogen ion transfer to the substrate. Additionally, certain enzymes utilise amino acid residues in their active sites to form transient covalent bonds with substrates as part of the reaction process.


# Quiz 3

[Quiz 3. Proteins and Catalysis of Biological Reactions](https://docs.google.com/forms/d/e/1FAIpQLScNFY1tZemK8pPXojRZaW951wtl4m0SmFkWoL3Si_RBtT0Hwg/viewform?usp=sf_link)


# Information Flow in the Cell


# DNA Replication

During cell division, each daughter cell requires a copy of the genetic information from the parental cell, necessitating DNA replication prior to division. The fundamental principles of DNA replication are dictated by the structure of DNA, which comprises two complementary strands. According to the **semi-conservative** replication model, each strand of the DNA serves as a template for the synthesis of a new strand. As a result, after replication, each daughter DNA molecule contains one original (old) strand and one newly synthesised strand.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FrIjsZILWQDEtz13x9s7Y%2FFig_37_semiconservative.png?alt=media&amp;token=b85e6b42-56b1-4250-84f4-42b32f3582bc" alt="" width="223"><figcaption><p>Semi-conservative model of DNA replication: each strand of the DNA double helix serves as a template for the synthesis of a new strand<br>Image source:<br>Eunice Laurent - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=110374906</p></figcaption></figure>

While the major principles of DNA replication are similar in eukaryotes and prokaryotes, this chapter will focus on the process as it occurs in bacteria, such as *Escherichia coli*, due to its relative simplicity.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FsdGRh6y8fXVSmCCy3dSq%2FFig_38_origin.jpeg?alt=media&amp;token=248ebe9a-de73-448f-bae2-e64a0073efd2" alt="" width="563"><figcaption><p>Origin of replication in bacteria</p></figcaption></figure>

DNA synthesis initiates at specific sites known as **replication origins**. Most bacteria possess a single circular chromosome with only one replication origin. Upon the separation of DNA strands at the origin, a **replication bubble** forms. Replication starts simultaneously in two directions from the origin, leading to the formation of two **replication forks**. These forks move away from each other, performing DNA synthesis in their respective directions.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FlNpZts39IEJdxJNWPVst%2FFig_39_DNA_synthesis_reaction.png?alt=media&amp;token=d5240704-7efd-491d-b933-ae9e0c7641c4" alt="" width="563"><figcaption><p>Addition of nucleotides to the growing DNA chain<br>Image source:<br>Untitled image adapted from Madeleine Price Ball. "DNA polymerase". Accessed February 22, 2024. <a href="https://www.khanacademy.org/science/ap-biology/gene-expression-and-regulation/replication/a/molecular-mechanism-of-dna-replication">https://www.khanacademy.org/science/ap-biology/gene-expression-and-regulation/replication/a/molecular-mechanism-of-dna-replication</a>.</p></figcaption></figure>

At the replication fork, DNA synthesis is catalyzed by enzymes known as **DNA polymerases**. These polymerases require a single-stranded DNA template, which remains in the newly synthesized DNA as the original DNA strand. The substrates for the DNA polymerase reaction include the elongating new DNA strand and free nucleotides in the form of deoxyribonucleoside triphosphates. These nucleotides serve as the building blocks for the growing DNA strand, with each nucleotide being incorporated into the elongating DNA chain in a template-directed manner.

At the replication fork, the two template DNA strands are anti-parallel, with one running in the 5'-to-3' direction and the other in the 3'-to-5' direction. However, DNA polymerases can only elongate new DNA strands in the 5'-to-3' direction. Consequently, one DNA strand, known as the **leading strand**, can be synthesised continuously towards the replication fork.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FBYt5Q5dniYbZvPlHDLuP%2FFig_40_Okazaki_fragments.svg?alt=media&amp;token=0de9fa04-941d-4681-9ff3-57ecdd7e0335" alt="" width="364"><figcaption><p>DNA replication fork: leading and lagging strands<br>Image source:<br>Andi Schmitt - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=99996272</p></figcaption></figure>

The other strand, termed the **lagging strand**, must be synthesised in a discontinuous manner. DNA polymerase moves away from the replication fork, reading the template, and synthesises short fragments called **Okazaki fragments**. After each fragment is synthesised, the DNA polymerase detaches from the template and attaches to a new segment of single-stranded DNA cleared by the progressing replication fork. In bacteria, Okazaki fragments are typically 1000-2000 nucleotides long, whereas in eukaryotes, they are approximately 10 times shorter.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FGAJiJB73xCp8RyeSClkU%2FFig_41_DNA_repilcation.png?alt=media&amp;token=bcb9ba80-d1de-44ba-a933-c14b63bbed54" alt="" width="561"><figcaption><p>Summary of DNA replication in bacteria</p></figcaption></figure>

In addition to DNA polymerase, DNA replication involves numerous other proteins. The movement of the replication fork is facilitated by **helicase**, which unwinds the double helix, separating the template strands from each other.&#x20;

Immediately following unwinding, the template strands are bound by **single-strand DNA binding** (**SSB**) **proteins**, preventing them from reannealing. This ensures the accessibility of the template strands for replication.&#x20;

The unwinding of the double helix by helicase generates supercoiling, excessive twisting, in front of the replication fork. This supercoiling is relieved by another enzyme called **topoisomerase**.

DNA replication initiation poses a challenge as DNA polymerase cannot initiate the synthesis of a new polynucleotide; it can only elongate an existing one. To overcome this, another enzyme, **primase**, is required. Primase synthesises a short stretch of RNA, termed a **primer**, to initiate DNA synthesis. The RNA primer serves as a starting point for DNA polymerase, which then adds deoxynucleotides to extend it, synthesising the new DNA strand.

DNA replication involves two types of DNA polymerases to overcome the challenge of primer removal and replacement. **DNA polymerase III** is responsible for extending the RNA primer by adding deoxynucleotides. However, it must detach from the template when it encounters another Okazaki fragment's primer. At this point, another enzyme, **DNA polymerase I**, takes over. DNA polymerase I removes the RNA nucleotides of the primer and replaces them with DNA nucleotides, filling in the gap left behind by the RNA primer.

Finally, the Okazaki fragments need to be joined into a single continuous DNA strand by the enzyme **DNA ligase**.&#x20;

The video below provides a summary of the DNA replication process with some additional details.

{% embed url="<https://www.youtube.com/watch?v=0Ha9nppnwOc>" %}
DNA replication summary
{% endembed %}


# Gene Expression: Transcription

## Introduction

DNA sequences known as **genes** encode proteins or functional RNAs. However, even protein-coding genes do not directly drive protein synthesis but act through the process of **gene expression**. Gene expression includes two main processes:

1. Transcription: RNA synthesis occurs on the DNA template.
2. Translation: protein synthesis takes place on the RNA template.

In eukaryotic cells, RNA molecules undergo several processing steps before translation can occur.

## <mark style="color:purple;">Question on the topic</mark>

The nucleotide sequence of a DNA strand serving as a template for transcription is given below.\
5'-AACCTGACAGA-3'

What is the nucleotide sequence of the RNA produced on this template?

{% hint style="info" %}
Note that nucleic acid sequences should be written in the 5'-3' direction.
{% endhint %}

<details>

<summary>Answer</summary>

UCUGUCAGGUU

</details>

## RNA polymerase

The main enzyme involved in transcription is **RNA polymerase** which catalyses the synthesis of RNA using the DNA template. Similar to DNA polymerase, RNA polymerase elongates RNA in the 5'-to-3' direction. However, unlike DNA polymerase, RNA polymerase does not require a primer to initiate polynucleotide synthesis.

In bacteria, a single RNA polymerase enzyme is responsible for transcription. Eukaryotes typically possess three RNA polymerases, each responsible for transcribing different types of genes. However, all eukaryotic protein-coding genes are transcribed by RNA polymerase II.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FyxyY2GL8TWJLxEQ3LHbe%2FFig_42_RNA_pol_II.png?alt=media&amp;token=c3d4f7d7-b170-4ebb-8d88-1f1199da94a9" alt="" width="375"><figcaption><p>Eukaryotic RNA polymerase II<br>Image source:<br>Litvinanna - Own work using https://www.rcsb.org/structure/1WCM, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=77523301</p></figcaption></figure>

The RNA produced on a protein-coding gene is referred to as messenger RNA (mRNA). Transcription of other genes results in non-coding RNAs, which include ribosomal RNAs (rRNAs), transfer RNAs (tRNAs), small nuclear RNAs (snRNAs), microRNAs (miRNAs) and several other RNA types.

## Initiation of Transcription

To start transcription, RNA polymerase must bind to a specific DNA sequence known as a **promoter**, which marks the start site for RNA synthesis. Most, though not all, promoters are located **upstream** from (before) the transcription start site.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FjQOSImFn7NXIhGzFpZTg%2FFig_43_bacterial_promoter.jpeg?alt=media&amp;token=70eb4caa-b236-411a-85c4-76eeb43a3c8e" alt="" width="408"><figcaption><p>Image source:<br>Connie Rye, Robert Wise, Vladimir Jurukovski, Jean DeSaix, Jung Choi, Yael Avissar - Biology, Oct 21, 2016, https://openstax.org/books/biology/pages/15-2-prokaryotic-transcription</p></figcaption></figure>

Transcription initiation in bacteria usually requires just the RNA polymerase and one additional protein called σ factor. Promoters, though variable, typically contain similar sequences located at positions -10 and -35 upstream from the transcription start site. The initiation of transcription requires the unwinding of a short region of DNA which is stabilised by σ factor. This open DNA region is termed a **transcription bubble**. Once the transcription bubble is formed, RNA polymerase initiates RNA synthesis. Just one of the DNA strands called the template strand is used for transcription.

In eukaryotes, the binding of the promoter by RNA polymerase requires the assistance of multiple additional proteins known as **transcription factors**. A typical promoter recognised by RNA polymerase II contains a conserved nucleotide sequence rich in adenines and thymines.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FoLsGStqJR5PviGzL8dVm%2FFig_44_eukaryotic_promoter.jpeg?alt=media&amp;token=ea0b4e5d-748b-4c61-b0b1-c2995c8b54d8" alt="" width="272"><figcaption><p>RNA polymerase II and transcription factors required to bind the promoter<br>Image source:<br>ByCNX OpenStax - http://cnx.org/contents/GFy_h8cu@10.53:rZudN6XP@2/Introduction, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=49929297</p></figcaption></figure>

## Elongation

The RNA polymerase moves downstream from the transcription start site adding nucleotides to the growing RNA molecule. The elongated RNA grows in the 3' direction.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FOJ0doBpqDe1hHr2XbGgd%2FSynthese_eines_RNA-Molek%C3%BCls_mit_dem_Enzym_RNA-Polymerase-1.jpg?alt=media&amp;token=0f3732ee-fdc1-4a36-a33e-813a561257b6" alt=""><figcaption><p>RNA elongation Image source: Rosti1985 - Own work, CC BY-SA 4.0, https://commons.wikimedia.org/w/index.php?curid=89613003</p></figcaption></figure>

Multiple RNA polymerase enzymes can transcribe the same gene at the same time to increase RNA production.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FfGwnBkHXpmUPV5deGWHN%2FFig_46_transcription_microscopy.jpg?alt=media&amp;token=d09cb9c5-728d-407a-af07-1e72c83e6dbc" alt="" width="375"><figcaption><p>Multiple RNA polymerase enzymes transcribe three closely located genes. "Begin" indicates the 5' end of the coding strand of DNA, where new RNA synthesis begins; "end" indicates the 3' end, where the transcripts are almost complete. RNA polymerase molecules are seen on DNA as dots.<br>Image source:<br>CC BY-SA 3.0, https://commons.wikimedia.org/w/index.php?curid=1830046</p></figcaption></figure>

## Transcription Termination

Termination in bacteria occurs when RNA polymerase encounters a termination signal. In most cases, termination involves the formation of a "hairpin" structure which results in polymerase detaching from the transcript.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FX0WaFuEKTnJ8YODABM3V%2FFig_47_termination.png?alt=media&amp;token=ce09d26b-3768-4466-9862-b867406ea145" alt="" width="563"><figcaption><p>Transcription termination in bacteria<br>Image source:<br>Untitled image. "Stages of transcription". Accessed February 29, 2024.<br>https://www.khanacademy.org/science/biology/gene-expression-central-dogma/transcription-of-dna-into-rna/a/stages-of-transcription </p></figcaption></figure>

In eukaryotes, the process of transcription termination is more sophisticated and is not yet fully understood. The enzyme RNA polymerase encounters a specific sequence called polyadenylation signal which marks the end of the transcript and recruits an additional enzyme that cleaves the newly synthesised RNA. Despite the cleavage of the RNA, RNA polymerase may continue transcription until it is eventually pushed away from the DNA by enzymes responsible for degrading the non-functional RNA.

## Summary

The video below summarises the transcription and provides some additional details:

{% embed url="<https://www.youtube.com/watch?v=XzVXhemtwmA>" %}


# Gene Expression: RNA Processing

In eukaryotes, the transcription of a gene by RNA polymerase II produces **pre-mRNA** which has to undergo several processing steps to become a fully functional mRNA. These steps include modifications of the transcript's 5' and 3' ends and the excision of some internal parts of the RNA molecule.&#x20;

## RNA Capping

Soon after the start of transcription by RNA polymerase, a modified guanine nucleotide is added to the 5' end of the growing RNA molecule.&#x20;

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F9JKGwkqFH2Hpw18XDsLs%2FFig_48_cap.png?alt=media&amp;token=20489194-00c4-421a-ab79-59579fd1ea2a" alt=""><figcaption><p>The structure of the 5' cap<br>Image source:<br>Zephyris - English Wikipedia, CC BY-SA 3.0, https://commons.wikimedia.org/w/index.php?curid=1379696<br></p></figcaption></figure>

This modification, known as a **5' cap**, serves several functions:

1. Protection from degradation: the 5' cap protects the pre-mRNA from degradation by nucleases.&#x20;
2. Nuclear export signal: the presence of the 5' cap serves as one of the molecular signals indicating that the RNA is mature and ready to be exported from the nucleus to the cytoplasm. This ensures that only fully processed mRNA molecules are transported to the site of protein synthesis.
3. Initiation of translation: in eukaryotes, the 5' cap is essential for the ribosome to bind to the mRNA and initiate protein synthesis.

## RNA Splicing

Most eukaryotic protein-coding genes contain two major types of segments: coding segments called **exons** and non-coding sequences called **introns**. During transcription by RNA polymerase II, both exons and introns are included in the pre-mRNA transcript.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FMZKX1aq5mWnL25pLgtHb%2FFig_49_rna_splicing.jpg?alt=media&amp;token=2625533a-e060-4c4d-bc4c-18ca7a87a561" alt="" width="544"><figcaption><p>RNA splicing: introns are excised and exons are fused<br>Image source:<br>Genomics Education Programme - Splice donor sites, CC BY 2.0, https://commons.wikimedia.org/w/index.php?curid=50543001</p></figcaption></figure>

However, prior to translation, introns must be removed from the pre-mRNA and exons must be joined together in a process called **RNA splicing**. Splicing is catalysed by a large molecular complex known as the **spliceosome**, which consists of several proteins and small nuclear RNAs (snRNAs). The spliceosome recognises specific sequences at the boundaries between exons and introns, called splice sites. Splicing results in the production of mature mRNA molecules containing only the sequences necessary for protein synthesis.

While the precise role of introns and splicing in eukaryotic organisms is still not fully understood, there are several potential advantages associated with splicing:

1. Regulatory elements in introns: some introns contain regulatory elements that can modulate gene expression.&#x20;
2. Exon-intron structure and protein evolution: the exon-intron structure of genes may facilitate the faster evolution of new proteins. It has been observed that many protein domains, which are functional units of proteins capable of folding into stable tertiary structures, are encoded by single exons. Additionally, proteins often consist of similar sets of domains. The alternation of exons with introns allows for **exon shuffling** through recombination events. This process can generate new combinations of exons, giving rise to novel genes that encode functionally active proteins, thus contributing to protein evolution and diversification.
3. Alternative splicing for protein diversity: RNA splicing enables the generation of diverse protein isoforms from a single gene through a mechanism known as **alternative splicing**. Different exons can be included or excluded from mature mRNA transcripts, leading to the production of multiple protein isoforms with distinct functions.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FomKBysLdY0h7VK0j9YcM%2FFig_50_alternative_splicing.gif?alt=media&amp;token=af186c3b-ea5f-4b80-873a-0a8b501cefc8" alt="" width="563"><figcaption><p>Alternative splicing results in different protein isoforms<br>Image source:<br>National Human Genome Research Institute - http://www.genome.gov/Images/EdKit/bio2j_large.gif, Public Domain, https://commons.wikimedia.org/w/index.php?curid=2132737</p></figcaption></figure>

## Polyadenylation

As RNA polymerase II progresses along the gene during transcription, it eventually encounters a specific sequence known as the **polyadenylation signal**, which is transcribed into the growing pre-mRNA. The canonical polyadenylation signal sequence is AAUAAA.

The polyadenylation signal is recognised by cleavage proteins that cut the transcript a short distance downstream from the signal. This cleavage event results in the release of the pre-mRNA.

An enzyme called poly-A polymerase (PAP) adds a string of adenine nucleotides, known as the **poly-A tail**, to the 3' end of the pre-mRNA. Typically, about 200 adenine nucleotides are added to create the poly-A tail.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F210H75Q7WnpYGHslL7wL%2FFig_51_polyadenylation.png?alt=media&amp;token=846b4a66-4ea0-4525-b7a3-0ee2778bffd7" alt="" width="375"><figcaption><p>Polyadenylation of RNA<br>Image source:<br>Zephyris (en Wikipedia user) - English Wikipedia, CC BY-SA 3.0, https://commons.wikimedia.org/w/index.php?curid=1379635</p></figcaption></figure>

Similar to the 5' cap, the poly-A tail protects mRNA from degradation and is required for mRNA export from the nucleus.

## Summary

The figure below summarises RNA processing in eukaryotic cells.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FUfAQFcs8D2MMoGDsTVmE%2FFig_52_rna_processing.png?alt=media&amp;token=c6793e13-fffe-4ef9-849a-641ab33af8b3" alt="" width="540"><figcaption><p>RNA processing in eukaryotic cells<br>Image source: Biorender.com</p></figcaption></figure>

More details on RNA splicing can be found in the following videos.

{% embed url="<https://www.youtube.com/watch?v=aVgwr0QpYNE>" %}

{% embed url="<https://www.youtube.com/watch?v=YgmoHtLGb5c>" %}


# Quiz 4

[Quiz 4. DNA Replication, Transcription and RNA Processing](https://docs.google.com/forms/d/e/1FAIpQLScgLMZkvQ6kCunnS6CdG4SyoEIdxaFYwFr8KmmuIt8N5yqOzw/viewform?usp=sharing)


# Chromatin and Chromosomes

DNA carrying genetic information is packaged in structures called **chromosomes**. Bacteria typically have a single chromosome which contains one circular DNA, while eukaryotic genomes are divided into many chromosomes. A eukaryotic chromosome contains a linear DNA molecule associated with numerous proteins classified as **histones** and non-histone proteins. This complex of DNA and proteins is referred to as **chromatin**.

Packaging DNA into chromatin fibres achieves a remarkably high level of compaction. For instance, approximately 205 cm of DNA present in every human cell is enclosed within a nucleus about 10 micrometres in diameter.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F8gXcOvQDVMJkZIAAK7GH%2FFig_65_beads_on_a_string.jpg?alt=media&amp;token=a34fc967-e4b8-4e21-8cd6-2a59b08f7df1" alt="" width="375"><figcaption><p>Electron microscopy of chromatin. Nucleosomes remind beads on a string<br>Image source:<br>Oak Ridge National Laboratory - ORNL History, CC BY 2.0, https://commons.wikimedia.org/w/index.php?curid=90681196</p></figcaption></figure>

The primary level of chromatin compaction involves the formation of **nucleosomes**, which are protein complexes composed of histones. When observed under an electron microscope, chromatin gently extracted from the nuclei resembles beads on a string. Each "bead" represents a nucleosome, around which DNA is wrapped, while the "string" connecting them is linker DNA.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2Fq9qZfx5hwZvhDbIFNavK%2FFig_66_chromatin.jpg?alt=media&amp;token=e27276ab-86c6-48ac-9cb4-858166b37a3d" alt="" width="450"><figcaption><p>Chromatin structure<br>Image source:<br>Darryl Leja, NHGRI - https://medlineplus.gov/genetics/condition/adnp-syndrome/#causes, Public Domain, https://commons.wikimedia.org/w/index.php?curid=130950339</p></figcaption></figure>

A nucleosome consists of a protein octamer comprising eight histones, each of which has a tail extending from the nucleosome core. These tails undergo various chemical modifications that can alter chromatin structure and regulate gene expression.


# Regulation of Gene Expression

As we have noted earlier, cells in multicellular organisms can express different sets of genes defining cells' specific functions and structural features. Expression can be regulated at any step leading from a gene to a protein and even at the proceeding steps of protein inactivation and degradation. We have previously outlined the RNA processing control resulting in diverse RNA splicing products, however, for most genes, the control at the transcription level is of primary importance.

## Chromatin regulation

The regulation of gene expression through altering chromatin structure determines the accessibility of genes and their promoters to the transcription machinery. Genes situated in dense chromatin regions, known as **heterochromatin**, are typically inaccessible for transcription. In contrast, genes located in decondensed chromatin regions (**euchromatin**), can be transcribed, with the level of transcription influenced by factors such as nucleosome positioning, histone modifications, and modifications of nucleotides in the DNA.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FXIjwB1Af7ck6NwbU5WAC%2FFig_67_histone_acetylation.png?alt=media&amp;token=b41cf302-61d5-48fa-b87f-f70b77529491" alt="" width="479"><figcaption><p>Histone acetylation results in opening up chromosome<br>Image source:<br>Eslaminejad, Mohamadreza &#x26; Fani, Nesa &#x26; Shahhoseini, Maryam. (2013). Epigenetic Regulation of Osteogenic and Chondrogenic Differentiation of Mesenchymal Stem Cells in Culture. Cell Journal. 15. 1-10.</p></figcaption></figure>

One common histone modification that influences chromatin structure is **histone acetylation**, where acetyl groups are added to lysine side chains, removing their positive charge. This modification reduces the affinity of histones for DNA, resulting in chromatin opening and increased accessibility of DNA for transcription-related proteins, thus enhancing gene expression.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F93EdyLXcweegh2Uh5TRD%2FFig_68_DNA_methylation.jpg?alt=media&amp;token=59a07572-e741-4ee6-b638-44247c1d93a7" alt="" width="563"><figcaption><p>DNA methylation silences chromatin</p></figcaption></figure>

A prevalent nucleotide modification in eukaryotes is **cytosine methylation**. Methylated genes are typically inactive, and their demethylation leads to gene activation and increased expression.

DNA methylation and histone modifications, despite being reversible, can be transmitted to daughter cells during cell division. This type of inheritance, which involves the transmission of information that is not encoded in the DNA sequence itself but rather in the modifications of DNA and histones, is called **epigenetic inheritance**.

## Transcription factors

The initiation of transcription of all protein-coding genes in eukaryotes relies on a set of proteins known as general **transcription factors**, which bind to the promoter region. However, for genes with tissue-specific expression patterns, additional transcription factors are required. These specialised transcription factors bind to specific DNA sequences and either facilitate or block the binding of RNA polymerase, thus regulating transcription initiation.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FHLAHNWJB3M2V7UW6X1Io%2FFig_69_transcription_activators.jpg?alt=media&amp;token=22467f23-6101-4ac2-8964-d8a69703d151" alt="" width="408"><figcaption><p>Transcription activators bind the enhancer to promote the expression of a gene<br>Image source:<br>CNX OpenStax - http://cnx.org/contents/GFy_h8cu@10.53:rZudN6XP@2/Introduction, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=49929334</p></figcaption></figure>

Transcription factors that enhance transcription are called **activators**, and they bind to **enhancer** sequences in the DNA. Conversely, transcription factors that inhibit transcription are termed **repressors**, and they bind to **silencer** sequences.

Each transcription factor typically regulates multiple genes. However, activation of a particular gene's expression depends on the specific combination of transcription factors present (in the case of activators) or absent (in the case of repressors). This intricate network of transcription factors ensures precise control over gene expression in response to cellular signals and developmental cues.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FMDOmdPW6ncAMA3TRlT1I%2FFig_70_transcription_regulation.png?alt=media&amp;token=940235c2-afe1-42ec-b66c-52997a430306" alt="" width="563"><figcaption><p>A combination of transcription factors is often required to initiate gene expression<br>Image source:<br>Untitled image. "Combinatorial control". Accessed April 17, 2024.<br>https://www.khanacademy.org/science/biology/gene-expression-central-dogma/translation-polypeptides/a/the-stages-of-translation</p></figcaption></figure>

## Activation of transcription factors

Alterations in gene expression within a cell can occur through various mechanisms in response to numerous internal and external stimuli. One common example is the cell's response to a chemical stimulus, such as a hormone.

In this scenario, a hormone typically binds to a receptor located on the cell's surface. This binding event induces a conformational change in the receptor, initiating a cascade of molecular events within the cell known as a **signal transduction pathway**. Eventually, this pathway leads to the activation of specific transcription factors.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FyVRZZD0YPJd2idjC0a2C%2FFig_71_signal_pathway.jpg?alt=media&amp;token=c449c613-0c23-45db-8d7c-7d7acd650ab1" alt="" width="375"><figcaption><p>Activation of a transcription factor through a signal transduction pathway<br>Image source:<br>Untitled image. "Transcription and DNA-Protein Binding". Accessed April 17, 2024.<br>https://biologicalmodeling.org/motifs/transcription<br> Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 </p></figcaption></figure>

Once activated, these transcription factors modulate gene expression by binding to the regulatory regions of their target genes. This binding alters the expression levels of these genes, thereby implementing the response to the initial stimulus.

## Transcriptome

[One of the ways to measure gene expression is the quantification of different RNA molecules present in a sample using sequencing-based approaches. The whole set of RNA molecules in a tissue sample or an organism is called the transcriptome.](#user-content-fn-1)[^1]

[^1]: Not sure we need it really


# Quiz 5

[Quiz 5. Chromatin, Regulation of Gene Expression, and Translation](https://docs.google.com/forms/d/e/1FAIpQLSdtxxTWV89MBEVOrEZ-mhAtI9luUx5-YYV-yoV7aRPfyBl7aA/viewform?usp=sharing)


# The Genetic Code

In mRNA, nucleotides are used to code for the amino acids that make up proteins. However, with only four different nucleotides available, it's not possible for each nucleotide to directly correspond to one of the twenty amino acids. Instead, nucleotides are combined into "words" called **codons**.

If codons contained only two nucleotides, there would be 16 possible combinations ($$4^{2}$$). However, this would not be sufficient to code for all 20 amino acids. To provide enough codons for twenty amino acids, the **triplet code** is used, where each codon consists of three nucleotides.

With three nucleotides per codon, there are 64 possible combinations ($$4^{3}$$), which is more than enough to code for the 20 amino acids. This **redundancy** in the genetic code means that some amino acids are encoded by more than one codon. For example, the amino acid leucine is coded for by six different codons.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FvB6pdO83fKz5BFuKSl5P%2FFig_53_genetic_code.png?alt=media&amp;token=90e0a6ca-5f7f-473a-ac99-2f0282b2ae0f" alt="" width="432"><figcaption><p>The triplet code<br>Image source: Biorender.com</p></figcaption></figure>

While most codons in the genetic code serve the primary function of coding for amino acids, four codons have special functions. The **start codon** (AUG) not only codes for the amino acid methionine but also serves as the initiation signal for translation, marking the point at which the ribosome begins protein synthesis. Also, there are three **stop codons** in the genetic code: UAA, UAG, and UGA. Unlike other codons, stop codons do not specify any amino acids. Instead, they signal the termination of translation, indicating to the ribosome that the protein synthesis should conclude.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F75sPy1vmbHGa1r2JlZKh%2FFig_55_genetic_code.png?alt=media&amp;token=cc6d1dd0-f528-4213-b1bc-cca248d04009" alt="" width="540"><figcaption><p>The codon table<br>Image source: Biorender.com</p></figcaption></figure>

## <mark style="color:purple;">Question on the topic</mark>

The nucleotide sequence of a fragment of a short RNA molecule is given below. Find the start codon and write the sequence of amino acids in a polypeptide encoded by this RNA using the codon table.

ACUGCAAAUGGAUGCCGGAUUAUGGAGUUAAGAUGUGC

<details>

<summary>Answer</summary>

Met-Asp-Ala-Gly-Leu-Trp-Ser

</details>

During translation, the ribosome reads mRNA in triplets, each coding for one amino acid. Theoretically, the ribosome could read each mRNA molecule using three different **reading frames**, depending on the nucleotide from which translation starts.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FTlPDPPbzzVnqNjoFpieZ%2FFig_54_reading_frame.png?alt=media&amp;token=84f33a83-decb-4bc7-a721-54cf85ba990f" alt="" width="375"><figcaption><p>Three reading frames can be used for the polypeptide synthesis<br>Image source:<br>Untitled image. "Reading frames". Accessed March 11, 2024.<br>https://www.khanacademy.org/science/biology/gene-expression-central-dogma/transcription-of-dna-into-rna/a/stages-of-transcription </p></figcaption></figure>

However, in the majority of cases, only one of the reading frames is utilised for translation by the ribosome. The choice of the reading frame is determined by the start codon. Once the ribosome encounters the start codon, it establishes the correct reading frame and begins translating the mRNA into a polypeptide chain. Subsequent codons are then read in the same frame.

A striking property of the genetic code is that it is **universal** for all organisms. The same codons encode the same amino acids in virtually all living organisms, from bacteria to plants to animals. This universality suggests that the genetic code is evolutionarily ancient and was likely established in the common ancestor of all currently existing organisms.&#x20;

Minor deviations from the universal genetic code only exist in compact genomes. For example, mitochondria, which possess their separate translation machinery, exhibit slight variations in the genetic code compared to the nuclear genome.


# Gene Expression: Translation

## Overview of Translation

Translation is the process by which the genetic information encoded in mRNA molecules is used to synthesise polypeptides. During translation, the mRNA's nucleotide sequence is decoded into a sequence of amino acids.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FRwY6V8WFZJxvnCVTbQdd%2FFig_56_translation_overview.png?alt=media&amp;token=585d460e-e30e-42bd-818c-293eb32dddfa" alt="" width="563"><figcaption><p>Translation of mRNA by a ribosome<br>Image source:<br>LadyofHats - Own work using:nih.govmiami.edumolecularassembler.comufl.eduphschool.comuic.eduyoutube.comyoutube.com, Public Domain, https://commons.wikimedia.org/w/index.php?curid=4889777</p></figcaption></figure>

Nucleotides in the mRNA are read in sets of three, known as codons. Each codon corresponds to a specific amino acid or serves as a signal to start or stop protein synthesis. These codons are recognised by adaptor molecules called transfer RNAs (tRNAs).

Transfer RNAs carry specific amino acids to the ribosome, a large macromolecular complex where protein synthesis occurs. The ribosome catalyses the formation of peptide bonds between adjacent amino acids, linking them together to form a growing polypeptide chain.

## Transfer RNAs

A tRNA is a relatively short RNA molecule, typically about 80 nucleotides long, that recognises a specific codon in the mRNA and matches it with an appropriate amino acid.

The tRNA molecule folds into a complex structure due to the formation of four double-helical segments held together by intramolecular hydrogen bonds. When drawn in one plane, the tRNA resembles a cloverleaf, while its 3D structure is L-shaped.&#x20;

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FMAROIiiuvAOjt7AoOz2V%2FFig_57_tRNA.jpeg?alt=media&amp;token=21781fc3-7ec0-45a2-b9bc-04fbb28a9d66" alt="" width="563"><figcaption><p>(a) The two-dimensional tRNA structure  resembles a cloverleaf<br>(b) The three-dimensional tRNA structure is L-shaped<br>Image modified from:<br>CNX OpenStax - https://cnx.org/contents/5CvTdmJL@4.4, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=53713134</p></figcaption></figure>

At the 3' end of the L-shaped molecule, an amino acid is covalently attached to the tRNA through an ester bond, forming an aminoacyl-tRNA molecule. The specific amino acid attached to the tRNA corresponds to the codon recognised by the tRNA during translation. The loop at the tip of the L-shaped tRNA contains the **anticodon**, a sequence of three nucleotides complementary to a specific codon in the mRNA.

Although there are 61 codons coding for amino acids, the number of distinct tRNA molecules does not necessarily match this count. Instead, the actual number of tRNAs varies among species, with some bacteria having as few as 31 tRNAs and humans possessing 48 tRNAs.&#x20;

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FO1sWTwn53awVTM7YTPgT%2FFig_58_wobble.webp?alt=media&amp;token=bbc69edd-5b63-40e1-bd76-e86f3f08d8f2" alt="" width="375"><figcaption><p>Normal and wobble base-pairing</p></figcaption></figure>

The reason for this discrepancy is that the structure of some tRNA molecules allows them to bend the rules of complementary base-pairing in the third position of a codon – a phenomenon known as **wobble** base pairing. In those cases, accurate codon-anticodon recognition is required only for the first two nucleotide pairs while the third nucleotide pair can be mismatched. For instance, in bacteria, a guanine nucleotide in the third position of a codon can pair with either cytosine or uracil. The wobble base-pairing allows one tRNA to recognise several codons despite differences in their nucleotide sequences.

## Ribosome Structure and Function

Ribosomes are large molecular machines that catalyse protein synthesis. They are composed of ribosomal RNAs (rRNAs) associated with ribosomal proteins, and while the exact number of rRNAs and proteins may vary across different organisms, the general organisation of the ribosome remains rather conservative. A ribosome comprises a **large** and a **small subunit**, that come together around the mRNA and separate after the protein synthesis is complete so that they can be re-used.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FeqgAhly5CqoAZSHBXU9I%2FFig_59_ribosome_structure.png?alt=media&amp;token=6255ef21-e8bb-47aa-bce4-4c7ee80e4acb" alt="" width="375"><figcaption><p>The large ribosomal subunit (red) and small ribosomal subunit (blue). For the large subunit, the rRNAs are shown in dark red and orange-red, and the ribosomal proteins are shown in pink. For the small subunit, the rRNA is shown in dark blue and the ribosomal proteins in light blue).<br>Image source:<br>Vossman - Own work, CC BY-SA 3.0, https://commons.wikimedia.org/w/index.php?curid=6865434</p></figcaption></figure>

The ribosome has three binding sites for tRNAs, namely the A (aminoacyl), P (peptidyl), and E (exit) sites. Transfer RNAs move sequentially from one site to another during translation. The **A site** is responsible for binding the tRNA carrying the next amino acid to be added to the growing polypeptide chain. The **P site** holds the tRNA, which carries the growing polypeptide chain. The **E site** facilitates the exit of tRNA molecules that have already participated in translation.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F4dK2QbmO6VJLX4I1isBB%2FFig_60_tRNA_binding_sites.png?alt=media&amp;token=3b94935c-f288-4e0f-a074-97204329fe58" alt="" width="563"><figcaption><p>The tRNA binding sites in the ribosome: A (aminoacyl), P (peptidyl), and E (exit) sites<br>Image source:<br>Untitled image. "Ribosome". Accessed March 18, 2024.<br>https://www.khanacademy.org/science/biology/gene-expression-central-dogma/translation-polypeptides/a/the-stages-of-translation</p></figcaption></figure>

The catalytic activity in the peptide bond formation is performed by the rRNA component of the ribosome. This enzymatic activity makes ribosomes ribozymes—catalytically active RNA molecules.

Unlike prokaryotic ribosomes, in eukaryotes, some ribosomes synthesise proteins when docked on the internal membranous structures of the cell (i.e. on the rough endoplasmic reticulum). In that case, a produced protein is compartmentalised to this membranous structure, can be inserted into cell membranes, or destined for secretion.

## The Process of Translation

Translation involves three major stages: initiation, elongation and termination.

### Initiation of translation

At the initiation stage, the small subunit of the ribosome first binds to a specialised **initiator tRNA**. This initiator tRNA carries the amino acid methionine (or formyl-methionine in bacteria). The initiator tRNA has a unique ability to bind directly to the P site of the ribosome without initially passing through the A site.

In prokaryotes, translation initiation is guided by the presence of a sequence known as the **Shine-Dalgarno sequence**, located upstream of the start codon on the mRNA. This sequence helps position the initiator tRNA precisely at the start codon. In eukaryotes, the small ribosomal subunit, with the initiator tRNA already on board, associates with the 5' cap of the mRNA first. The ribosome then scans along the mRNA until it encounters the start codon.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FQSoeSd8wnu1m6TecQa1Z%2FFig_61_translation_initiation.png?alt=media&amp;token=76ac7fe2-b894-4e98-91b5-b9007564ff39" alt="" width="563"><figcaption><p>Initiation of translation in prokaryotes<br>Image source:<br>Untitled image. "Initiation of translation". Accessed March 18, 2024.<br>https://www.khanacademy.org/science/biology/gene-expression-central-dogma/translation-polypeptides/a/the-stages-of-translation</p></figcaption></figure>

Once the initiator tRNA is correctly positioned at the start codon, the large ribosomal subunit joins the complex, forming a complete ribosome. The A site is now ready to receive the next aminoacyl-tRNA and the translation can start.

Translation initiation in both prokaryotes and eukaryotes requires the assistance of specific protein factors. The process also requires the expense of energy in the form of guanosine triphosphate (GTP).

### Elongation

After the translation initiation is complete, the elongation cycle begins, during which amino acids are added to the growing polypeptide chain. This process proceeds from the amino group end (N-terminus) towards the carboxyl end (C-terminus) of the polypeptide.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2Fy3tXfCorUMUwno2x2Y6r%2FFig_62_translation_elongation.png?alt=media&amp;token=1c208187-b953-4a22-bbac-8718af320adc" alt="" width="563"><figcaption><p>The elongation cycle<br>Image source:<br>Untitled image. "The elongation cycle". Accessed March 19, 2024.<br>https://www.khanacademy.org/science/biology/gene-expression-central-dogma/translation-polypeptides/a/the-stages-of-translation</p></figcaption></figure>

The elongation cycle involves several steps assisted by a set of specific protein factors:

1. **Codon recognition**. In this step, an aminoacyl-tRNA enters the A site and binds the codon if codon-anticodon recognition occurs properly. To ensure accurate tRNA selection, a molecule of GTP is expended.
2. **Formation of the peptide bond**. The growing polypeptide chain, held by the peptidyl-tRNA located in the P site, is transferred to the amino acid attached to the aminoacyl-tRNA in the A site.
3. **Translocation**. In this step, the peptidyl-tRNA moves from the A site to the P site. Simultaneously, the discharged tRNA in the P site moves to the E site and vacates the ribosome. The mRNA, bound to the tRNA, also moves, bringing the next codon into the A site for the next cycle of elongation. The translocation step requires the hydrolysis of one GTP molecule.

These steps repeat iteratively until a stop codon is encountered on the mRNA, signalling the termination of translation.

### Termination

Translation terminates when the ribosome encounters a stop codon. The stop codons are recognised by no tRNA but can be bound by a protein called release factor which mimics the tRNA shape to occupy the A site. The release factors catalyse the hydrolytic reaction which cleaves the polypeptide from its tRNA in the P site thus allowing the polypeptide to be released from the ribosome.

## Summary

The video below summarises the translation process and provides additional details on protein factors:

{% embed url="<https://www.youtube.com/watch?v=8Hsz_Vmcy-Y>" %}


# Cell Cycle and Cell Division

The cell cycle is a sequence of events occurring within a cell encompassing growth, development, and division. It consists of two primary phases:

1. **Interphase**: the cell grows, performs DNA replication and prepares for the division.
2. **Mitotic (M) phase** involves the separation of the cell's genetic material and cytoplasm to produce two daughter cells.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2Fk2Ry9RD1bA1bxmDp5ifr%2FFig_72_cell_cycle.jpg?alt=media&amp;token=93d55343-1677-408f-aea3-9e95ae38bcee" alt="" width="563"><figcaption><p>The cell cycle<br>Image source:<br>CNX OpenStax - http://cnx.org/contents/GFy_h8cu@10.53:rZudN6XP@2/Introduction, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=49926267</p></figcaption></figure>

## Interphase

Interphase is further divided into three sub-phases:

1. **G1 (Gap 1) phase**: the cell grows, increases the number of organelles, and synthesises proteins.
2. **S (Synthesis) phase**: DNA replication occurs, resulting in the duplication of the genetic material.
3. **G2 (Gap 2) phase**: the cell performs additional protein synthesis and growth which prepares it for mitosis.

Many cells of a multicellular organism quit the replicative cell cycle and after the G1 phase switch to a non-proliferative G0 phase.

## Mitosis and cytokinesis

After DNA replication, each chromosome consists of two **chromatids** containing identical copies of the chromosome's DNA. During mitosis chromosomes are split and the chromatids are directed to the opposite poles of the cell.&#x20;

In animal cells, mitosis is facilitated by organelles called **centrosomes** which serve as microtubule-organising centres. **Microtubules** are filaments that consist of multiple protein subunits and can move chromosomes.

The microtubules of the so-called **mitotic spindle** extend from the centrosomes and attach to regions of the chromosomes called **centromeres**. The attachment occurs via protein complexes called **kinetochores**, which assemble at the centromeres.

As the mitotic spindle forms and microtubules attach to kinetochores, they exert forces on the chromosomes, pulling the chromatids apart and guiding their movement to opposite poles of the cell. This ensures that each daughter cell receives an identical set of chromosomes during cell division.

Mitosis includes five stages: prophase, prometaphase, metaphase, anaphase and telophase. Cytokinesis is the process of cytoplasm division which usually starts in the late telophase.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2Fghfwj2YwSMf26oP3HIjG%2FFig_73_mitosis.jpg?alt=media&amp;token=b4707ae5-7d49-456d-9dda-37246ddfd440" alt=""><figcaption><p>Stages of mitosis and cytokinesis<br>Image source:<br>OpenStax - https://cnx.org/contents/FPtK1zmh@8.25:fEI3C8Ot@10/Preface, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=30131217</p></figcaption></figure>


# Quiz 6

[Quiz 6. Cell Cycle and Cell Division](https://docs.google.com/forms/d/e/1FAIpQLSc5XbiGM2nud-01W0ntqimIODQyzrp83D5ptDe0NdKzK5VerQ/viewform?usp=sharing)


# Mutations and Variations


# Point mutations

**Mutations** are changes in the genetic information of a cell, arising from various sources, including errors during DNA replication and repair. While some mutations lead to large-scale changes in the genome, such as chromosomal translocations, others, known as **point mutations**, involve alterations of a single base pair in the DNA. Point mutations are the most common type of mutation and are the focus of this chapter.

Mutations can occur in either germline cells, which give rise to gametes and can be passed on to offspring, or somatic cells, which make up the body but are not involved in reproduction. **Germline mutations** can be inherited by future generations, while **somatic mutations** are only passed on to descendant cells during cell division. Accumulation of somatic mutations over time can lead to the development of cancer.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FILoTP5Dgwm0yId3FSeP7%2FFig_76_germline_somatic_mutations.jpg?alt=media&amp;token=85bc69ae-4a3a-4890-9d34-0f1057754f2a" alt="" width="375"><figcaption><p>Germline and somatic mutations<br>Image source:<br>Prostate cancer genotyping for risk stratification and precision treatment - Scientific Figure on ResearchGate, https://www.researchgate.net/figure/Fundamental-difference-between-germline-and-somatic-mutations-wwwlearncolontownorg_fig3_380066207, CC BY 4.0<br></p></figcaption></figure>

Most mutations in the human genome occur in non-coding DNA regions and thus have no or little impact on the phenotype. Mutations causing significant consequences usually occur in the genes or their regulatory regions.&#x20;

## Nucleotide substitutions

**Substitution** is a type of mutation in which one nucleotide in the DNA sequence is replaced by another (e.g. adenine is substituted by guanine).

These substitutions can have varying effects on the resulting protein. **Silent mutations** occur when the substitution does not alter the amino acid sequence of the protein due to the redundancy of the genetic code. In other words, the codon affected by a silent mutation still codes for the same amino acid, so the protein's function remains unchanged.

**Missense mutations** result in the substitution of one amino acid for another in the protein sequence. Depending on the location and type of the alteration, missense mutations vary in their effect. Some missense mutations may only cause subtle changes in protein function, while others can significantly alter the protein's structure or activity e.g. when a mutation changes the shape of an enzyme's active site.

**Nonsense mutations** are substitutions that change a codon coding for an amino acid into a stop codon. This premature termination of translation leads to the production of a truncated protein, which is often non-functional.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F6bcrqPmEsbf4LhahRc24%2FFig_74_mutations.jpg?alt=media&amp;token=001bdee1-0300-4b68-8cd3-af283d135b1a" alt="" width="408"><figcaption><p>Consequences of nucleotide substitutions<br>Image source:<br>CNX OpenStax - http://cnx.org/contents/GFy_h8cu@10.53:rZudN6XP@2/Introduction, CC BY 4.0, https://commons.wikimedia.org/w/index.php?curid=49926491<br></p></figcaption></figure>

## Deletions and insertions

The effects of adding (**insertion**) or removing (**deletion**) a nucleotide from the DNA sequence in a coding region are often detrimental. Because the genetic code is read in sets of three nucleotides (codons), the addition or deletion of a single nucleotide causes a **frameshift**. Deletions and insertions change the reading frame and thus the whole sequence of codons downstream from the mutation. This results in totally different amino acids and, since a stop codon usually appears close to a frameshift mutation, in a truncated polypeptide.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FQLN3GJfHsxjIzNmV0yTh%2FFig_75_frameshift_mutations.jpg?alt=media&amp;token=28fccc7e-e157-4c5a-b1db-38ac22532729" alt=""><figcaption><p>Frameshift mutations result in a wrong amino acid sequence and truncated polypeptide<br>Image source:<br>National Human Genome Research Institute – https://www.genome.gov/genetics-glossary/Frameshift-Mutation, CC BY 4.0 </p></figcaption></figure>


# Genotype-Phenotype Interactions

Mutations are the primary source of genetic variation within a species, creating variants of nucleotide sequences in a gene called **alleles**. While many alleles do not significantly affect gene function, some result in distinct phenotypic traits.

A classic example of allele variation is observed in pea plants, particularly in genes determining flower colour. In these plants, a gene encoding a protein which controls the production of a purple pigment called anthocyanin has two alleles. One allele codes for a functional protein, leading to purple flowers, while the other allele is faulty, resulting in no pigment production and thus white flowers.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2Flx1no57FrlLYH8sSXLsJ%2FFig_77_peas_pigment.png?alt=media&amp;token=64d31528-cbbc-47b8-af00-a0bb1839143a" alt="" width="563"><figcaption><p>Alleles are variants of a nucleotide sequence in a gene</p></figcaption></figure>

Most multicellular organisms are **diploid**, meaning they possess two sets of chromosomes in their cells. Consequently, genes are typically represented by two alleles in a cell. If both alleles are identical, the organism is termed **homozygous** for that gene. Conversely, if the alleles differ, the organism is termed **heterozygous**.

In a heterozygous organism, if only one of the alleles affects the observable characteristics it is called **dominant**, while the second allele is called **recessive**. For instance, in pea plants with heterozygous alleles for flower colour, where one allele codes for a functional protein and the other for a non-functional protein, the purple flower trait is dominant.

Interactions between alleles can often be more nuanced than a binary dominant-recessive relationship. For example, human blood groups are determined by a gene with three alleles $$I$$$$I^{A}$$, $$I$$$$I^{B}$$, and $$i$$. This gene governs the presence of a specific carbohydrate on the surface of red blood cells, with each allele resulting in a different structure of this carbohydrate. Alleles $$I^{A}$$and $$I^{B}$$ lead to the production of carbohydrate structures A and B, respectively, while allele $$i$$ results in the absence of the carbohydrate. Both  $$I^{A}$$and $$I^{B}$$ are dominant over allele $$i$$, meaning individuals with genotypes $$I^{A}i$$ and $$I^{B}i$$ will have blood A and B, respectively. Only individuals with the genotype $$ii$$ will have no carbohydrate on their red blood cells and will have the blood group O. However, individuals with the genotype  $$I^{A}I^{B}$$ will exhibit blood group AB, indicating the presence of both carbohydrate structures. This type of allele interaction, where both alleles contribute to the phenotype in a heterozygous individual, is termed **codominance**.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FB8oYLOJururIqyIDWuM2%2FFig_78_blod_groups.png?alt=media&amp;token=c1e6ece8-85b4-44cf-8678-553355efbe18" alt="" width="563"><figcaption><p>Blood groups<br>Image source:<br>OpenStax College - Anatomy &#x26; Physiology, Connexions Web site. http://cnx.org/content/col11496/1.6/, Jun 19, 2013., CC BY 3.0, https://commons.wikimedia.org/w/index.php?curid=30148183 with changes</p></figcaption></figure>

Examples of phenotypes discussed above demonstrate discontinuous variation, which means that every individual can be assigned to a particular group according to its phenotype. However, plenty of traits, such as human body mass or skin tone are distributed continuously, in other words, an individual can have any characteristic value within a certain range. Such characteristics are commonly **polygenic**, which means, they are controlled by multiple genes, and/or are influenced by the environment.

Skin tone is polygenic, meaning it's controlled by multiple genes. At least 150 genes are known to contribute to skin colour, and even a simplified model considering just 3 genes, each with 2 alleles, can illustrate polygenic inheritance and its effect on phenotype. In this model, an individual who is homozygous recessive for all three genes would have a very light skin tone, while a homozygous dominant individual would have very dark skin. The combination of different alleles of these genes produces the wide spectrum of skin tones observed in human populations.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FVyxo5v3UZ0ilJB1ss541%2FFig_79_skin_tones.jpg?alt=media&amp;token=b9787f2d-e0af-4233-941d-42bd8767bc6c" alt=""><figcaption><p>A simplified model of polygenic inheritance of skin colour</p></figcaption></figure>


# Quiz 7

[Quiz 7. Mutations and Variation](https://docs.google.com/forms/d/e/1FAIpQLSd5i_RVHmPmrwTIn-rLUwW_pmeZ_CAq2GbCZeunL4BSDGo1aQ/viewform?usp=sharing)


# PROGRAMMING

Programming is a fundamental skill in bioinformatics, enabling you to make sense of complex datasets. After preprocessing your sequencing data in a bash environment, you’ll often end up with files like count matrices that need further analysis and visualization. This is where programming becomes your most valuable tool, helping you transform raw data into meaningful insights.

For basic analysis and visualization, you can use almost any programming language. However, R and Python dominate the bioinformatics landscape, as many widely used packages are written in these languages. As a result, most bioinformaticians work with R, Python, or often both.

Beyond data analysis, bioinformatics can also involve developing software tools or packages for others in the field. For these tasks, adopting an Object-Oriented Programming (OOP) approach can be more effective than writing basic scripts. Python and Java are popular choices for such projects. However, this guide will focus on helping you get started with scripting in R and Python, leaving the complexities of OOP for another day.

If you’re new to programming, starting with Python is a great idea. Its straightforward syntax and versatility make it beginner-friendly. Once you’re comfortable with Python, picking up R will feel much easier. For now, focus on one language to build your confidence and avoid feeling overwhelmed.


# Python for Genomics

Get started with Python by taking the following course from Johns Hopkins University:&#x20;

{% embed url="<https://www.coursera.org/learn/python-genomics?specialization=genomic-data-science>" %}

With this course, you will learn:&#x20;

* Gain a foundational understanding of Python programming, focusing on genomic data analysis.
* Learn to use Jupyter Notebooks for interactive coding, data visualization, and documentation.
* Master essential data structures like lists and dictionaries, as well as control flow mechanisms such as loops and conditionals.
* Develop skills to read, write, and parse files in formats commonly used in genomics.
* Explore the Biopython library for handling and analyzing biological data.
* Apply Python programming techniques to solve real-world genomic data science problems.

### Access hints

To access this course, you’ll need to create a Coursera account. Financial aid is available, allowing you to take the course for free. The application process is simple and straightforward, so don’t hesitate—just go for it!


# R programming (optional)

R is a powerful programming language widely used in bioinformatics for data analysis, statistical computing, and data visualization. With its extensive range of libraries and tools designed specifically for biological data, R is an essential tool for anyone working in genomics, transcriptomics, or other areas of bioinformatics. By learning R, you'll be able to efficiently manipulate and analyze complex biological datasets, perform statistical analyses, and visualize results in meaningful ways. This course will introduce you to the fundamentals of R programming, providing a solid foundation for tackling real-world bioinformatics problems.<br>

{% embed url="<https://www.coursera.org/learn/r-programming/>" %}

This course will help you:&#x20;

* Set up R and RStudio to create a robust programming environment.
* Understand core programming concepts such as data types, control structures, and functions in R.
* Learn to read, write, and manipulate data, including handling missing values and subsetting data structures.
* Create and debug custom functions to streamline your data analysis workflow.
* Enhance the efficiency of your code through profiling and optimization techniques.
* Organize and comment your code effectively to ensure clarity and reproducibility.

### Access hints

To access this course, you’ll need to create a Coursera account. Financial aid is available, allowing you to take the course for free. The application process is simple and straightforward, so don’t hesitate—just go for it!\ <br>


# STATISTICS: THEORY

STATISTICS IS THE ART OF LEARNING FROM DATA.

Statistics is a crucial component of bioinformatics because it allows researchers to make sense of complex biological data and draw meaningful conclusions. In bioinformatics, data is often noisy and high-dimensional, and statistical methods help to identify patterns, test hypotheses, and make predictions with a high degree of confidence. Without a strong statistical foundation, it would be difficult to extract reliable insights from genomic, transcriptomic, or other biological datasets.

This module offers a comprehensive introduction to statistical learning, covering key concepts and methods that are widely used in data science and bioinformatics. It will further enhance your ability to understand and apply statistical techniques to complex datasets.

To ensure that you can effectively use these concepts, we’ve supplemented later modules with practical courses where you’ll perform statistical analyses. These practical sessions will build on the theory and programming basics you’ve learned earlier, providing hands-on experience in analyzing real biological data.

### Credits

This module contains original content developed by Vahan Huroyan, PhD.&#x20;

### Additional material (optional)

If you'd like to delve deeper into the foundations of statistics, you can also take this course from Stanford University:

{% embed url="<https://www.coursera.org/learn/stanford-statistics>" %}


# Introduction to Probability

Probability Theory is a fundamental branch of mathematics concerned with quantifying the likelihood of events. Probability answers the question: *"How likely is it that a particular outcome will happen?”*&#x20;

This concept is expressed as a number between 0 and 1, where 0 indicates impossibility of an event to occur and 1 indicates certainty of an event happening.

In probability theory, a probability space includes three key components:&#x20;

1. **Sample Space (S):** the set of all possible outcomes of a particular experiment
2. **Sigma-Algebra (F):** representing subsets of the sample space for which probabilities are defined
3. **Probability Function (P):** assigning probabilities to these subsets.&#x20;

The sample space, encompasses all possible outcomes of a random experiment. The sigma-algebra consists of subsets of the sample space, ensuring that certain properties hold true under probability measures, such as closure under complementation and countable unions. Lastly, the probability function assigns probabilities to these subsets, reflecting the likelihood of different outcomes or events occurring within the sample space. These 3 components form the foundational framework for analyzing uncertainty and randomness in various statistical contexts.

### <mark style="color:blue;">Example 1: Flipping a Coin</mark>

**Sample Space (S)**: The set of all possible outcomes.

* $$S={Heads,Tails}$$

**Sigma-Algebra (F)**: The collection of all subsets of the sample space including the empty set.

* $$F = { \emptyset, {\text{Heads}}, {\text{Tails}}, {\text{Heads}, \text{Tails}} }$$

**Probability Function (P)**: Assigns a probability to each subset in the sigma-algebra.

* $$P(∅)=0$$
* $$P({\text{Heads}}) = \frac{1}{2}$$
* $$P({\text{Tails}}) = \frac{1}{2}$$
* $$P({\text{Heads, Tails}}) = {1}$$

### <mark style="color:blue;">Example 2: Rolling a Die</mark>

**Sample Space (S)**: The set of all possible outcomes.

* $$S = { 1, 2, 3, 4, 5, 6 }$$

**Sigma-Algebra (F)**: The collection of all subsets of the sample space including the empty set.

* $$F = { \emptyset, {1}, {2}, \ldots, {6}, {1, 2}, {1, 3}, \ldots, {1, 2, 3, 4, 5, 6} }$$

**Probability Function (P)**: Assigns a probability to each subset in the sigma-algebra.

* $$P(\emptyset) = 0$$
* $$P({1}) = \frac{1}{6}$$
* $$P({2}) = \frac{1}{6}$$​
* $$\vdots$$
* $$P({1, 2}) = \frac{2}{6}$$
* $$\vdots$$
* $$P({1, 2, 3, 4, 5, 6}) = 1$$

### <mark style="color:blue;">Example 3: Detecting the Presence of a Gene Variant</mark>

**Sample Space (S)**: The set of all possible genotypes at a particular genetic locus.

* $$S = { \text{AA}, \text{Aa}, \text{aa} }$$

**Sigma-Algebra (F)**: The collection of all subsets of the sample space.

* $$F = { \emptyset, {\text{AA}}, {\text{Aa}}, {\text{aa}}, {\text{AA}, \text{Aa}}, {\text{AA}, \text{aa}}, {\text{Aa}, \text{aa}}, {\text{AA}, \text{Aa}, \text{aa}} }$$

**Probability Function (P)**: Assigns a probability to each subset in the sigma-algebra, typically based on population genetics models such as Hardy-Weinberg equilibrium.

* $$P(\emptyset) = 0$$
* $$P({\text{AA}}) = p^2$$
* $$P({\text{Aa}}) = 2pq$$
* $$P({\text{aa}}) = q^2$$
* $$P({\text{AA}, \text{Aa}}) = p^2 + 2pq$$
* $$P({\text{AA}, \text{aa}}) = p^2 + q^2$$
* $$P({\text{Aa}, \text{aa}}) = 2pq + q^2$$
* $$P({\text{AA}, \text{Aa}, \text{aa}}) = 1$$

where $$p$$ is the frequency of the dominant allele $$A$$ and $$q$$ is the frequency of the recessive allele $$a$$ in the population, with $$p+q=1$$.


# Conditional Probability

Conditional probability is used when two or more events are not independent. This means the likelihood of one event is influenced by whether another event occurred. It asks, "If we know A has happened, what's the chance of B also happening?" It's calculated by multiplying the probability of the preceding event by the updated probability of the succeeding, or conditional, event.

Two events are said to be independent if one event occurring does not affect the probability that the other event will occur. However, if one event occurring or not does affect the likelihood that the other event will happen, the two events are said to be dependent, for example, a company's stock price increasing after reporting higher-than-expected earnings. If events are independent, then the probability of some event B is not contingent on what happens with event A, for example, an increase in Apple's shares and a drop in wheat prices.

Conditional probability is often written as the "probability of B *given* A" and notated as $$P(B|A)$$.

$$
P(B \mid A) = \frac{P(A \cap B)}{P(A)}
$$

{% embed url="<https://www.youtube.com/watch?v=_IgyaD7vOOA>" %}


# Independent Events

In probability theory, two events are considered independent if the occurrence of one event does not affect the probability of the occurrence of the other event. In other words, knowing that one event has occurred does not change the likelihood of the other event occurring.

For two events A and B, they are independent if and only if the probability of both events occurring together (the intersection of A and B) is equal to the product of their individual probabilities. Mathematically, this is expressed as:

$$
P(A \cap B) = P(A) \times P(B)
$$

If this equation holds true, then $$A$$ and $$B$$ are independent events.

For example, consider the events of flipping a coin and rolling a die. Let event $$A$$ be getting heads on the coin flip, and event $$B$$ be rolling a 4 on the die. The probability of getting heads on the coin is $$0.5$$, and the probability of rolling a 4 on the die is $$\frac{1}{6}$$​. Since these two events do not influence each other, they are independent. Therefore, the probability of both events happening together is:

$$P(A \cap B) = P(A) \times P(B) = 0.5 \times \frac{1}{6} = \frac{1}{12}$$

This confirms that the events are independent because the probability of their intersection is equal to the product of their individual probabilities.

{% embed url="<https://www.youtube.com/watch?v=1wuRV5z0PPE>" %}


# Random Variables

A key component of the theory of probability is the concept of random variables, which are functions that assign real numbers to each possible outcome of a random phenomenon. For instance, in the context of rolling a six-sided die, the values of the random variable could be represented by integers 1 through 6, each corresponding to one of the possible outcomes of the roll. Or the temperature on any given day can be considered a random variable because it depends on various unpredictable factors. In other words, a random variable is a function that maps from sample space to real numbers.&#x20;

**Daily Life Example**:

* **Coin Toss Outcome**: Let $$X$$ be a random variable representing the outcome of a coin toss, where $$X=1$$for heads and $$X=0$$ for tails. Here, $$X$$ is a random variable because its value is determined by a random process (the coin toss).

**Biology Example**:

* **Presence of a Genetic Marker**: Let $$Y$$ be a random variable representing the presence (1) or absence (0) of a particular genetic marker in an individual. The value of $$Y$$ is determined by a random genetic process.

We consider discrete and continuous random variables separately.&#x20;

### Discrete Random Variables

Discrete random variables represent outcomes that can be counted or enumerated, often taking on integer values. Examples include the outcomes of rolling a die or the number of heads obtained in a series of coin flips.

**Daily Life Example**:

* **Number of Emails Received**: Let $$X$$ be the number of emails you receive in a day. Possible values of $$X$$ are $$0, 1, 2, 3$$, and so on, making it a discrete random variable.

**Biology Example**:

* **Number of Offspring**: Let $$Y$$ be the number of offspring produced by a particular organism. The possible values are countable integers $$(e.g., 0, 1, 2, 3)$$.

### Continuous Random Variables

Continuous random variables are characterized by outcomes that can take on any value within a certain range, typically over an interval of real numbers. Common examples of continuous random variables include measurements such as height, weight, or time intervals.

**Daily Life Example**:

* **Daily Rainfall**: Let $$X$$ be the amount of rainfall in a day measured in millimeters. Since rainfall can take any non-negative value within a continuous range, $$X$$ is a continuous random variable.

**Biology Example**:

* **Blood Glucose Levels**: Let $$Y$$ be the blood glucose level of an individual measured in mg/dL. Since blood glucose can vary continuously, Y a continuous random variable.


# Independent, Dependent and Controlled Variables

### Independent and Dependent Random Variables

**Independent Random Variables:**

* **Definition:** Random variables $$X$$ and $$Y$$ are independent if the occurrence of one does not affect the probability distribution of the other.
* **Example:** Let $$X$$ be the outcome of rolling a fair six-sided die, and $$Y$$ be the outcome of flipping a fair coin. These events are independent since the outcome of the die roll does not influence the coin flip.

**Dependent Random Variables:**

* **Definition:** Random variables $$X$$ and $$Y$$ are dependent if the occurrence of one affects the probability distribution of the other.
* **Example:** Let $$X$$ be the amount of rainfall on a given day, and $$Y$$ be the water level in a nearby river. These variables are dependent because more rainfall likely increases the river's water level.

In a science experiment, there are three main types of variables, besides independent and dependent variables there is also the controlled variables.

The **independent variable** is the one that the researcher deliberately changes or manipulates. Its purpose is to observe the effect it has on another variable. For example, in an experiment testing the effect of light on plant growth, the amount of light is the independent variable.

The **dependent variable** is the one that is measured or observed in response to changes in the independent variable. Its purpose is to assess the effect of the independent variable. In the same plant growth experiment, the height of the plants is the dependent variable.

**Controlled variables** are variables that are kept constant throughout the experiment to ensure that the test results are reliable. The purpose of controlling these variables is to make sure that any observed changes in the dependent variable are due to the manipulation of the independent variable alone. In the plant growth experiment, controlled variables might include the type of plant, the amount of water given, the soil type, and the ambient temperature.

{% embed url="<https://www.youtube.com/watch?v=UvCSzl2IDbY>" %}


# Data distribution PMF, PDF, CDF

### Probability Mass Functions (PMFs)

Probability distributions are essential tools in understanding the behavior of random variables. For discrete random variables, we use Probability Mass Functions (PMFs), which assign probabilities to each possible outcome. The sum of these probabilities always equals 1, providing a complete description of the distribution. PMFs allow us to determine the likelihood of observing a specific value of the random variable.&#x20;

#### **1. PMF of a Fair Six-Sided Die**

A fair six-sided die has outcomes {1, 2, 3, 4, 5, 6}. Each outcome has an equal probability of occurring.

$$P(X = x) =  \begin{cases}  \frac{1}{6} & \text{if } x \in {1, 2, 3, 4, 5, 6} \ 0 & \text{otherwise} \end{cases}$$&#x20;

**2.** **PMF of a Biased Coin**

&#x20;A biased coin has a 70% chance of landing on heads (H) and a 30% chance of landing on tails (T).&#x20;

&#x20;$$P(X = x) = \begin{cases} 0.7 & x = \text{H} \ 0.3 & x = \text{T} \ 0 & \text{otherwise} \end{cases}$$

### Probability Density Functions (PDFs)

Conversely, for continuous random variables, Probability Density Functions (PDFs) are employed. Unlike PMFs, PDFs indicate the likelihood of the variable falling within a particular range. While the PDF itself doesn't provide probabilities directly, the area under the curve within a range represents the probability of the variable falling within that range. Understanding the shape and behavior of PDFs is crucial for analyzing continuous probability distributions.&#x20;

#### **Uniform Distribution Example**:&#x20;

The time it takes for a bus to arrive at a bus stop, assuming buses arrive at a regular interval. If buses arrive every 10 minutes, and you arrive at the bus stop at a random time, the waiting time $$X$$ can be modeled as a uniform distribution between 0 and 10 minutes.

​ $$f(x) =  \begin{cases}  \frac{1}{10} & 0 \le x \le 10 \ 0 & \text{otherwise} \end{cases}$$

### Cumulative Distribution Function (CDF)

Both PMFs and PDFs are complemented by the Cumulative Distribution Function (CDF). The CDF provides a comprehensive view of the probability distribution by specifying the probability that the variable takes on a value less than or equal to a given value. It accumulates the probabilities of all possible outcomes, starting at zero and approaching one as the variable's value increases. The CDF is indispensable for calculating probabilities, making statistical inferences, and understanding the behavior of random variables across their entire range. By understanding PMFs, PDFs, and CDFs, analysts can effectively model and analyze data in various fields, enabling informed decision-making and predictions.

#### **CDF of a Uniform Distribution**

**Example**: The time you wait for a bus that arrives every 10 minutes uniformly.

The CDF of a uniform distribution over $$\[a,b]$$ is:

​ $$F(x) =  \begin{cases}  0 & x < a \ \frac{x - a}{b - a} & a \le x \le b \ 1 & x > b \end{cases}$$

For $$a=0 \ and \ b=10$$:&#x20;

$$F(x) =  \begin{cases}  0 & x < 0 \ \frac{x}{10} & 0 \le x \le 10 \ 1 & x > 10 \end{cases}$$

####

{% embed url="<https://youtu.be/YXLVjCKVP7U>" %}


# Mean, Variance of a Random Variable

Mean and variance are essential statistical measures that provide a concise summary of a random variable's distribution.

### Mean

The mean, which is also referred to as the expected value or the expectation of a random variable, is a measure of central tendency in probability distributions. For discrete random variables, it is calculated by summing the product of each possible value of the variable and its corresponding probability from the PMF. Similarly, for continuous random variables, the mean is obtained by integrating the product of each possible value and its probability density from the PDF. The mean provides insight into the average value or central tendency of the distribution and is a fundamental metric for understanding the behavior of random variables. The expectation formula for discrete and continuous random variable $$X$$ is given below respectively:

$$
E(X) = \mu = \sum\_{x} x P(X = x) \ E(X) = \mu = \int\_{-\infty}^{\infty} x f(x) , dx
$$

#### Examples:

* **Income Distribution**:
  * **Context**: In economics, the mean income of a population is crucial for understanding the economic well-being of individuals and households.
  * **Mean Interpretation**: The mean income provides an average income level across a population, aiding in policy-making, taxation, and economic planning.
* **Gene Expression Levels**:
  * **Context**: In genetics and molecular biology, gene expression levels can be measured across samples.
  * **Mean Interpretation**: The mean gene expression level indicates the average amount of mRNA or protein produced by a gene in a population of cells or organisms.

### &#x20;Variance

Variance measures the spread or dispersion of a probability distribution around its mean. It quantifies the degree to which values deviate from the expected value. In the case of discrete random variables, variance is computed by summing the squared differences between each value and the mean, weighted by their respective probabilities from the PMF. For continuous random variables, variance is obtained by integrating the squared differences between each value and the mean, weighted by their probability densities from the PDF. Variance is a crucial measure in assessing the variability and uncertainty within a probability distribution. The variance formula for discrete and continuous random variable $$X$$ is given below respectively:

$$
\text{Var}(X) = \sum\_{x} (x - \mu)^2 P(X = x) \ \text{Var}(X) = \int\_{-\infty}^{\infty} (x - \mu)^2 f(x) , dx
$$

**Example: Cell Growth Rates**

* **Random Variable**: $$W$$ represents the growth rate of cells under specific conditions.
* **Variance Interpretation**: The variance of $$W$$, $$Var(W)$$, measures how growth rates vary among cells.
* **Biological Context**:
  * **Mean**: $$E(W)$$ could be the average growth rate observed in a cell culture.
  * **Variance**: A higher variance $$Var(W)$$ indicates greater variability in cell growth rates, which could be influenced by genetic mutations, nutrient availability, or environmental factors.

The relationship between the mean (expected value) and variance of a random variable 𝑋 X can be expressed as follows:

$$
\text{Var}(X) = E(X^2) - \[E(X)]^2
$$

Try to prove it!

### Quantiles

Quantiles divide a probability distribution into equal-sized intervals, each containing a specified percentage of the data. Common quantiles include the median (50th percentile), quartiles (25th, 50th, and 75th percentiles), and percentiles (any arbitrary percentile). Quantiles provide valuable insights into the distribution of data, allowing for comparisons, summarizations, and identification of extreme values. They are particularly useful in understanding the spread and shape of a distribution, aiding in decision-making and risk assessment.

{% embed url="<https://www.youtube.com/watch?v=IFKQLDmRK0Y>" %}


# Some Common Distributions

### Uniform distribution for discrete random variable

A uniform distribution for a discrete random variable signifies a scenario where each possible outcome has an equal probability of occurring. This implies that the probability mass function (PMF) assigns the same value to each outcome within a finite set.&#x20;

For instance, consider rolling a fair six-sided die; if each face has an equal chance of appearing, this would lead to a uniform distribution among the six possible outcomes.&#x20;

Mathematically, if X represents the discrete random variable, and the outcomes are&#x20;

$$x₁, x₂, ..., xₙ$$, then the PMF is given by the following formula

$$P(X = xᵢ) = \frac{1}{n} \ for\  all \ i$$,&#x20;

where n is the total number of possible outcomes.

{% embed url="<https://youtu.be/cyIEhL92wiw>" %}

### Uniform distribution for continuous random variable

A uniform distribution for a continuous random variable describes a situation where the probability density function (PDF) remains constant within a specified interval. The PDF of a continuous uniform distribution is constant, indicating a constant probability density across the interval. Mathematically, if $$X$$is a continuous random variable distributed uniformly over the interval $$\[a, b]$$, then the probability density function is&#x20;

$$f(x) = \frac{1}{(b - a)} \ for \ a ≤ x ≤ b.$$

{% embed url="<https://www.youtube.com/watch?t=90s&v=-qt8CPIadWQ>" %}

### Binomial distribution

The binomial distribution  models the number of successes in a fixed number of independent Bernoulli trials, where each trial has the same probability of success, denoted by p. Commonly used in scenarios such as coin flips or product defect rates, the binomial distribution's probability mass function (PMF) calculates the likelihood of observing a specific number of successes in a fixed number of trials. The key parameters are the number of trials $$(n)$$ and the probability of success $$(p)$$on each trial.&#x20;

Mathematically, if X represents the number of successes in n trials, then the probability mass function is:

&#x20;$$P(X = k) = \binom{n}{k} \cdot p^k \cdot (1-p)^{n-k}$$, &#x20;

where $$\binom{n}{k}$$ denotes the binomial coefficient.

P.S. if you're unfamiliar with statistical tests, feel free to skip the tests part of the video for now and return to it later :)

{% embed url="<https://www.youtube.com/watch?v=J8jNoF-K8E8>" %}

### Normal distribution

The normal distribution, also known as the Gaussian distribution, is the most common distribution in statistics. The normal distribution is characterized by a symmetric bell-shaped curve. In a normal distribution, the distribution is entirely defined by its mean (μ) and standard deviation (σ). Many natural phenomena, tend to follow a normal distribution.&#x20;

Its probability density function is described by the formula&#x20;

&#x20;$$f(x) = \frac{1}{\sigma \sqrt{2 \pi}} , e^{-\frac{(x - \mu)^2}{2 \sigma^2}}$$

{% embed url="<https://www.youtube.com/watch?v=rzFX5NWojp0>" %}

If you feel comfortable with mathematics and want to watch something more advanced consider this video⬇️

{% embed url="<https://youtu.be/cy8r7WSuT1I?si=_rpauV3IuBOFOxD2>" %}

### Student's t-Distribution

The Student's t-distribution (often referred to simply as the t-distribution) is a probability distribution that arises in statistical inference when the sample size is small or when the population variance is unknown. It is widely used in hypothesis testing, confidence interval estimation, and linear regression analysis. The shape of the t-distribution depends on a parameter known as degrees of freedom, which is related to the sample size. As the degrees of freedom increase, the t-distribution approaches the standard normal distribution. However, for smaller sample sizes, the t-distribution has heavier tails, making it more robust against outliers compared to the normal distribution. Understanding the t-distribution is crucial for conducting accurate statistical analyses, especially when dealing with small samples or unknown population variances.<br>

{% embed url="<https://www.youtube.com/watch?t=278s&v=T0xRanwAIiI>" %}

### Poisson Distribution

The Poisson distribution is employed to model the number of events occurring within a fixed interval of time (or space), given a known average rate of occurrence. It is particularly useful in scenarios where events happen independently and at a constant average rate, such as the number of customer arrivals in a queue or the number of emails received per day. The Poisson distribution's probability mass function calculates the likelihood of observing a specific number of events in a given interval. Its key parameter is lambda, representing the average rate of occurrence. Mathematically, if X represents the number of events occurring in a fixed interval, then the probability mass function is:&#x20;

$$P(X = k) = \frac{e^{-\lambda} \lambda^k}{k!}$$

where k is a non-negative integer.

{% embed url="<https://www.youtube.com/watch?t=1s&v=jmqZG6roVqU>" %}


# Exploratory Statistics: Mean, Median, Quantiles, Variance/SD

Exploratory statistics serves as the initial stage of data analysis, focusing on summarizing and understanding the essential characteristics of a dataset. It involves employing various descriptive statistics to gain insights into the distribution, central tendency, and variability of the data. Key measures include the mean, median, quantiles, variance, and standard deviation, each providing unique perspectives on the dataset.&#x20;

<mark style="background-color:blue;">It's important to note that these measures differ from the parameters of a random variable's theoretical distribution, such as the population mean or variance. Instead, they are derived from observed data and offer a practical summary of the sample at hand.</mark>

<br>

### Mean

The mean, or average, is perhaps the most commonly used measure of central tendency in exploratory statistics. It represents the sum of all data values divided by the total number of observations. The mean provides a single numerical summary of the data's central location, indicating the typical value around which the observations tend to cluster. While sensitive to extreme values (outliers), the mean offers a straightforward interpretation and is often utilized in various analytical contexts.&#x20;

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F8NS2lrUaIr8WDmtzepLc%2Fimage.png?alt=media&amp;token=11307d1e-5383-4005-8588-546cc9d30a4c" alt="" width="346"><figcaption></figcaption></figure>

### Median

Unlike the mean, which is influenced by extreme values, the median represents the middle value of a dataset when arranged in ascending or descending order. It is robust to outliers and provides a measure of central tendency that is less affected by extreme observations. The median is particularly useful when the dataset contains skewed distributions or when there are concerns about the influence of outliers on the mean. It offers a more robust representation of the typical value, especially in scenarios where the data is not symmetrically distributed.&#x20;

### Mode

For a discrete random variable, the mode is the value with the highest probability mass function (PMF). For a continuous random variable, it refers to the peak of the probability density function (PDF). The mode can be directly derived from the distribution's parameters (mean, variance, etc.) without needing a sample. In a sample (a set of observed data points), the mode is the most frequently occurring value. It's derived from the data and represents the value that appears most frequently.

{% embed url="<https://www.youtube.com/watch?v=k3aKKasOmIw>" %}

### Quantiles

Quantiles divide a dataset into equal-sized portions, providing insight into the distribution of data across various percentiles. Common examples include quartiles (dividing the data into four parts) and percentiles (dividing the data into hundred parts). Quantiles help identify the spread and variability of the data, facilitating comparisons and understanding of data distributions. They are particularly useful for assessing the relative position of individual observations within a dataset and for identifying potential outliers or extreme values. <br>

### Variance and Standard Deviation

Variance and standard deviation quantify the spread or dispersion of data points around the mean. Variance measures the average squared deviation of each data point from the mean, providing a measure of the overall variability within the dataset. Standard deviation, the square root of the variance, offers a more interpretable measure by providing the spread of data in the same units as the original data. Together, variance and standard deviation offer insights into the degree of variability within the dataset, aiding in understanding the distribution's shape and characteristics. They are fundamental measures in exploratory statistics, providing valuable information about the data's variability and distribution.<br>

{% embed url="<https://www.youtube.com/watch?t=836s&v=SzZ6GpcfoQY>" %}


# Data Visualization

Whenever you have the opportunity to visualize your data, take it. While numbers and statistics can convey a lot of information, visual representations often provide a clearer and more intuitive understanding of the data. People tend to grasp concepts more easily when they see them visually, making it easier to identify patterns, trends, and outliers. Visualizations can transform complex datasets into accessible insights, facilitating better communication and more informed decision-making.

#### Histogram

A histogram is a graphical representation of the distribution of a dataset. It divides the data into bins (or intervals) and displays the frequency (or count) of data points in each bin. The x-axis represents the bins, while the y-axis shows the frequency of data points in each bin. Histograms are useful for understanding the shape of the data distribution, identifying skewness, and detecting any potential outliers. They are commonly used for continuous data.

{% embed url="<https://www.youtube.com/watch?v=AndS0RLdxtk>" %}

#### Barplot

A barplot, or bar chart, is a graph that represents categorical data with rectangular bars. Each bar's length or height corresponds to the frequency or proportion of the category it represents. The x-axis displays the categories, while the y-axis shows the frequency or proportion of each category. Barplots are useful for comparing different categories and visualizing the distribution of categorical data. Unlike histograms, barplots are used for discrete data.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F1hG9wXnfVEQEcNjeybcW%2Fimage.png?alt=media&amp;token=5a403557-7fb4-4d62-83d3-cb592bc0b1fb" alt="" width="375"><figcaption></figcaption></figure>

#### QQ Plot

A QQ plot, or quantile-quantile plot, is a graphical tool to assess whether a dataset follows a specific theoretical distribution, often the normal distribution. It plots the quantiles of the dataset against the quantiles of the theoretical distribution. If the points lie approximately along a straight line, it suggests that the data follows the theoretical distribution. QQ plots are useful for checking the normality assumption and identifying deviations from the expected distribution.

{% embed url="<https://youtu.be/okjYjClSjOg?si=f-5S7hWC1De5Y14g>" %}

#### Boxplot

A boxplot, also known as a whisker plot, is a standardized way of displaying the distribution of data based on a five-number summary: minimum, first quartile (Q1), median, third quartile (Q3), and maximum. The box shows the interquartile range (IQR) between Q1 and Q3, while a line inside the box indicates the median. Whiskers extend from the box to the minimum and maximum values within 1.5 times the IQR. Points outside this range are considered outliers and are plotted individually. Boxplots are useful for identifying the central tendency, dispersion, and outliers in a dataset.

{% embed url="<https://www.youtube.com/watch?v=fHLhBnmwUM0>" %}

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2F0BWAFw26OzXy3opxcVd3%2Fimage.png?alt=media&amp;token=a1c6a513-a865-4315-a973-d3ef308a6fc5" alt="" width="368"><figcaption></figcaption></figure>

#### Scatter Plot

A scatter plot is a type of data visualization that displays the relationship between two quantitative variables. Each point on the plot represents an observation in the dataset, with the position on the x-axis corresponding to the value of one variable and the position on the y-axis corresponding to the value of the other variable. Scatter plots are useful for identifying correlations, trends, and potential outliers within the data. They are often used to determine whether there is a linear or non-linear relationship between the two variables. Additionally, scatter plots can help reveal clusters or patterns that might suggest further investigation.


# Confidence Intervals

Confidence intervals are a statistical concept used to estimate the range of values within which we expect a population parameter, such as the mean or proportion, to lie. They provide a measure of uncertainty around an estimated statistic based on sample data.

1. **Purpose**:
   * **Estimate Precision**: Confidence intervals help quantify the uncertainty in our estimates of population parameters derived from sample data.
   * **Inferential Tool**: They provide a range of plausible values for the parameter, allowing us to make inferences about the population.
2. **Construction**:
   * **Sample Data**: Start with a sample from the population and compute a sample statistic (e.g., mean, proportion).
   * **Distribution Assumptions**: Underlying assumptions about the population distribution (e.g., normality for means, binomial for proportions) guide the calculation.
   * **Formula**: Typically constructed as $$Estimate±Margin \ of \ Error$$, where the margin of error accounts for variability and is based on the standard error of the statistic.
3. **Interpretation**:
   * **Confidence Level**: Often expressed as a percentage (e.g., 95%, 99%). It represents the probability that the confidence interval includes the true population parameter if the sampling and estimation process were repeated many times.
   * **Example**: A 95% confidence interval suggests that if we were to take 100 different samples and compute confidence intervals for each, approximately 95 of those intervals would contain the true population parameter.
4. **Factors Influencing Width**:
   * **Sample Size**: Larger samples generally result in narrower confidence intervals because they provide more precise estimates of the population parameter.
   * **Variability**: Higher variability in the data results in wider intervals, as it increases the uncertainty in estimating the parameter.

#### Practical Use:

* **Decision Making**: Confidence intervals aid in making informed decisions by providing a range of plausible values for a population parameter.
* **Comparisons**: They allow comparisons between groups or over time, assessing whether differences are statistically significant.

{% embed url="<https://www.youtube.com/watch?v=TqOeMYtOc1w>" %}


# Comparison tests, p-value, z-score

Comparison tests are statistical methods used to determine if there are significant differences between groups, they are crucial tools in inferential statistics, allowing researchers to draw conclusions about the populations from which their samples are drawn, and to make informed decisions based on statistical evidence.

### Student's t-test

The Student's t-test is used to compare the means of two groups and is most effective when the data is normally distributed and the variances are equal.&#x20;

### Non-parametric tests

Non-parametric tests, such as the Mann-Whitney U test or Kruskal-Wallis test, are alternatives to t-tests that do not assume a normal distribution and are used when data does not meet parametric test assumptions.&#x20;

{% embed url="<https://www.youtube.com/watch?v=l86wEhUzkY4>" %}

### ANOVA

ANOVA (Analysis of Variance) is used to compare the means of three or more groups. It assesses the overall variance to determine if there is at least one significant difference between group means.&#x20;

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FAhPZIahuriPhE4O7I549%2Fimage.png?alt=media&amp;token=2f58cf84-153d-43b1-ab88-ae59115632a2" alt="" width="282"><figcaption></figcaption></figure>

{% embed url="<https://youtu.be/EFdlFoHI_0I?si=0LGwN8tvJkdZ80a->" %}

{% embed url="<https://www.youtube.com/watch?t=1s&v=j9ZPMlVHJVs>" %}

{% embed url="<https://www.youtube.com/watch?v=Xg8_iSkJpAE>" %}

### P-values

P-values are a measure of the strength of evidence against the null hypothesis in statistical tests. A p-value indicates the probability of obtaining test results at least as extreme as the results observed, assuming that the null hypothesis is true. In hypothesis testing, a low p-value (typically ≤ 0.05) suggests that the null hypothesis can be rejected, indicating that there is a statistically significant difference between groups. Conversely, a high p-value suggests that there is insufficient evidence to reject the null hypothesis, implying that any observed differences may be due to random chance.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FdvKRQNlSQbAzrrGOUwm0%2Fimage.png?alt=media&amp;token=e1b63df4-41dc-468e-b53a-bd892e11a012" alt="" width="282"><figcaption></figcaption></figure>

{% embed url="<https://www.youtube.com/watch?v=vemZtEM63GY>" %}

{% embed url="<https://www.youtube.com/watch?v=JQc3yx0-Q9E>" %}

### Z-score

A z-score, also known as a standard score, is a statistical measurement that describes a data point's relation to the mean of a group of values. It is expressed in terms of standard deviations from the mean.

The z-score of a data point xxx is calculated using the formula:

$$z = \frac{x - \mu}{\sigma}$$

where $$x$$ is the value of the data point, $$\mu$$ is the mean of the dataset, and $$\sigma$$ is the standard deviation of the dataset.

A positive z-score indicates that the data point is above the mean, while a negative z-score indicates that the data point is below the mean. The magnitude of the z-score reflects the number of standard deviations the data point is from the mean. A larger absolute value indicates the data point is further from the mean.

Z-scores are used for standardization, making different datasets comparable by converting data into a common scale. They also help identify outliers, as data points with z-scores beyond a certain threshold (commonly ±2 or ±3) are considered unusual. Additionally, in a normal distribution, z-scores correspond to probabilities, helping determine the likelihood of a data point occurring within a certain range.

For example, suppose we have a dataset with a mean (μ) of 50 and a standard deviation (σ) of 10. For a data point x = 70, the z-score is calculated as:

$$z = \frac{70 - 50}{10} = \frac{20}{10} = 2$$

This z-score of 2 means that the data point is 2 standard deviations above the mean.

{% embed url="<https://www.youtube.com/watch?v=5ABpqVSx33I>" %}


# Multiple test correction: Bonferroni, FDR

Multiple test correction methods are used to control the increased risk of Type I errors (false positives) when performing multiple statistical tests simultaneously. Multiple testing corrections adjust p-values derived from multiple statistical tests to correct for occurrence of false positives.

Watch this video to take a closer look at Type I and Type II errors:

{% embed url="<https://youtu.be/7mE-K_w1v90?si=SqIcoQWWE4Q_W2JA>" %}

### The Bonferroni correction

The Bonferroni correction is a straightforward and conservative approach that adjusts the significance level by dividing it by the number of tests performed. This reduces the likelihood of false positives but can also increase the risk of Type II errors (false negatives) due to its stringent nature.&#x20;

{% embed url="<https://www.youtube.com/watch?t=389s&v=HLzS5wPqWR0>" %}

### The False Discovery Rate

The False Discovery Rate (FDR) method, such as the Benjamini-Hochberg procedure, is a less conservative approach that controls the expected proportion of false positives among the rejected hypotheses. FDR methods are more powerful than the Bonferroni correction, allowing for the identification of more true positives while still controlling for false discoveries. This makes FDR particularly useful in large-scale testing scenarios, such as genomic studies, where a large number of tests are conducted simultaneously.

{% embed url="<https://www.youtube.com/watch?v=K8LQSvtjcEo>" %}

<br>


# Regression & Correlation

### Regression&#x20;

Regression analysis is a statistical method used to examine the relationship between a dependent variable and one or more independent variables (here, "independent" not in the statistical sense of being uncorrelated). The goal of regression is to model the expected value of the dependent variable based on the values of the independent variables.&#x20;

Linear regression, the most common type, fits a straight line (hyperplane to be precise) to the data points that best represents the relationship between the variables. This line (hyperplane), known as the regression line, can be used to make predictions. Regression analysis helps in understanding how the typical value of the dependent variable changes when any one of the independent variables is varied, while the other independent variables are held constant.

{% embed url="<https://www.youtube.com/watch?v=7ArmBVF2dCs>" %}

### Correlation

Correlation is a statistical measure that describes the strength and direction of a relationship between two variables. The correlation coefficient, typically represented by r, ranges from -1 to 1. A value of 1 indicates a perfect positive correlation, meaning that as one variable increases, the other also increases proportionally.  A value of -1 indicates a perfect negative correlation, where one variable increases as the other decreases. A value of 0 indicates no correlation, implying that there is no linear relationship between the variables. Correlation is useful for identifying and quantifying the degree to which two variables are related, but it does not imply causation.

{% embed url="<https://www.youtube.com/watch?list=PLblh5JKOoLUIzaEkCLIUxQFjPIlapw8nU&v=PaFPbb66DxQ>" %}

{% embed url="<https://www.youtube.com/watch?v=xZ_z8KWkhXE>" %}


# Dimentionality Reduction


# PCA (Principal Component Analysis)

Principal Component Analysis (PCA) is a statistical technique used for dimensionality reduction. It transforms a large set of variables into a smaller one that still contains most of the information in the original set. Here's a plain text explanation:

#### What is PCA?

PCA identifies patterns in data by finding the directions (principal components) along which the variance of the data is maximized. It projects the data onto these principal components to reduce the number of dimensions while retaining the most important features of the data.

#### Why Use PCA?

* **Reduce Complexity:** By reducing the number of features, PCA simplifies the dataset, making it easier to visualize and analyze.
* **Remove Noise:** PCA can help filter out noise from the data, improving the performance of machine learning algorithms.
* **Prevent Overfitting:** With fewer features, models are less likely to overfit, especially when dealing with small datasets.

#### Example Application

Imagine you have a dataset with 100 features, and you want to reduce it to 2 dimensions for visualization:

1. **Standardize the Data:** Ensure all features have the same scale.
2. **Compute Covariance Matrix:** Understand how features vary together.
3. **Eigenvalues and Eigenvectors:** Determine the principal components.
4. **Select Principal Components:** Choose the top 2 components.
5. **Transform the Data:** Project data onto the new 2D space.

{% embed url="<https://www.youtube.com/watch?v=HMOI_lkzW08>" %}

{% embed url="<https://www.youtube.com/watch?t=308s&v=FgakZw6K1QQ>" %}


# t-SNE (t-Distributed Stochastic Neighbor Embedding)

**Overview:**

* **t-SNE** is a non-linear dimensionality reduction technique primarily used for visualization.
* It converts high-dimensional Euclidean distances into conditional probabilities that represent similarities.
* The algorithm minimizes the Kullback-Leibler divergence between these probability distributions in the high-dimensional and low-dimensional space.

**Key Characteristics:**

* Effective at creating a visual representation of complex data, revealing clusters and patterns.
* Sensitive to parameters such as perplexity and learning rate.
* Computationally intensive, especially for large datasets.

**Applications:**

* Visualizing high-dimensional data such as images, text, and gene expression data.
* Exploring and understanding the structure of the data.

{% embed url="<https://www.youtube.com/watch?v=NEaUSP4YerM>" %}


# UMAP (Uniform Manifold Approximation and Projection)

**Overview:**

* **UMAP** is another non-linear dimensionality reduction technique that focuses on preserving both the local and global structure of the data.
* It constructs a high-dimensional graph and then optimizes a low-dimensional graph to be as similar as possible to the high-dimensional one.

**Key Characteristics:**

* Generally faster and more scalable than t-SNE, making it suitable for larger datasets.
* Often produces more meaningful global structure in the low-dimensional representation.
* Less sensitive to hyperparameters compared to t-SNE, with only a few parameters to tune (n\_neighbors and min\_dist).

**Applications:**

* Similar to t-SNE, UMAP is used for visualizing high-dimensional data in fields like genomics, image analysis, and natural language processing.
* It is also used as a preprocessing step for clustering and classification algorithms.

{% embed url="<https://www.youtube.com/watch?v=eN0wFzBA4Sc>" %}

{% embed url="<https://www.youtube.com/watch?v=jth4kEvJ3P8>" %}

### Comparison

| Feature            | PCA                                      | t-SNE                                          | UMAP                                           |
| ------------------ | ---------------------------------------- | ---------------------------------------------- | ---------------------------------------------- |
| **Algorithm Type** | Linear                                   | Non-linear                                     | Non-linear                                     |
| **Parameters**     | None                                     | Perplexity, learning rate                      | n\_neighbors, min\_dist                        |
| **Scalability**    | Efficient, handles large datasets        | Computationally intensive                      | Fast, scalable                                 |
| **Output**         | Linear combinations of original features | 2D or 3D embedding for visualization           | 2D or 3D embedding for visualization           |
| **Strengths**      | Simple, fast, captures variance          | Reveals clusters, good for visualization       | Fast, captures both local and global structure |
| **Weaknesses**     | Only captures linear relationships       | Computationally expensive, parameter-sensitive | Slightly complex, needs parameter tuning       |


# QUIZ

{% embed url="<https://forms.gle/m8hzrtTxFjEwnj9y8>" %}


# STATISTICS & PROGRAMMING

### Statistics and Python

This module will help you bridge the gap between theory and practice by applying the statistical concepts you’ve learned and using your coding skills to make statistical inferences from real-world examples.

If Python is the language you're most comfortable with, we recommend checking out this course from the University of Michigan:

{% embed url="<https://www.coursera.org/learn/inferential-statistical-analysis-python>" %}

### Statistics and R

Otherwise, if R is your language of choice, we recommend the following course from HarvardX on the EdX platform:

{% embed url="<https://www.edx.org/learn/r-programming/harvard-university-statistics-and-r>" %}


# BIOINFORMATICS ALGORITHMS

We will be designing and analyzing algorithms for bioinformatics. This should be useful to you whether you will be building novel methods and tools, or executing existing tools on new data.


# Introduction

Please try to prepare these before the class. If you cannot do something, don't worry and just come to the teacher in some break.

* Make an account on rosalind.info and solve the first problem [DNA](https://rosalind.info/problems/dna/)
* Bring two good bioinformatics papers that you are already familiar with
* Bring your laptop so that you can write code on your favorite programming language


# DNA strings and sequencing file formats

In this section, you will follow a few videos to revise your knowledge about the nature of DNA strings, the basics of working with genomes and reads, and the different file formats. After you're done with the videos, you will find a few problems to solve and check your knowledge and programming skills. The practical videos are explained in Python, but feel free to use any language of your choice to reproduce the same algorithm.

## DNA strings and sequencing

### Why study DNA sequencing and computational genomics?

{% embed url="<https://www.youtube.com/watch?v=hpb-mH-yjLc&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=1>" %}

### DNA: the molecule, the string

{% embed url="<https://www.youtube.com/watch?v=s1j-DuYJFr0&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=3>" %}

### String definitions and Python syntax

{% embed url="<https://www.youtube.com/watch?v=xmFR9VXzG7c&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=4>" %}

### Practical: String basics

{% embed url="<https://www.youtube.com/watch?v=YA2lBLz40SM&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=5>" %}

### Practical: Manipulating DNA strings

{% embed url="<https://www.youtube.com/watch?v=DphFe1vFJmU&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=6>" %}

### Practical: Downloading and reading a genome

{% embed url="<https://www.youtube.com/watch?v=FEvaKOZDDBk&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=7>" %}

### How DNA is copied

{% embed url="<https://www.youtube.com/watch?v=vaYxqrKn7Pk&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=8>" %}

### Sequencing by Synthesis

{% embed url="<https://www.youtube.com/watch?v=IzXQVwWYFv4&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=9>" %}

### Base calling and sequencing errors

{% embed url="<https://www.youtube.com/watch?v=U4QnpciIJhM&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=10>" %}

## File formats to store DNA reads

### Reads in FASTQ format

{% embed url="<https://www.youtube.com/watch?v=SZ6suqu-eLA&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=11>" %}

### Practical: Working with sequencing reads

{% embed url="<https://www.youtube.com/watch?v=YXVheiWAN5M&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=12>" %}

### Practical: Analyzing reads by position

{% embed url="<https://www.youtube.com/watch?v=7Nc04PFysEU&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=13>" %}

### Sequencers give pieces to genomic puzzles

{% embed url="<https://www.youtube.com/watch?v=X307VyAeHfI&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=14>" %}

## Problems to solve

To solve the following problems you need to [create an account on the Rosalind website](https://rosalind.info/accounts/register/) and [log in](https://rosalind.info/accounts/login/). On this site, you will find instructions on how to solve the problems and submit your answers. Please note that you need to solve the problems in the specified order. The problems will open up for you to solve only after you have unlocked the previous problems.&#x20;

Now, try to solve these problems after completing all the videos above.\
Some of them may be challenging :)

1. [Counting DNA Nucleotides](https://rosalind.info/problems/dna/)
2. [Transcribing DNA into RNA](https://rosalind.info/problems/rna/)
3. [Complementing a Strand of DNA](https://rosalind.info/problems/revc/)
4. [Computing GC Content](https://rosalind.info/problems/gc/)
5. [Translating RNA into Protein](https://rosalind.info/problems/prot/)

## Congratulations!

If you made it here, then congratulations! You have successfully completed this section. Move to the next portion of the guide with the arrow buttons below.


# Read alignment: exact matching

Read alignment, also known as sequence alignment, refers to the process of mapping or aligning short sequence reads obtained from high-throughput sequencing technologies to a reference genome or transcriptome. The goal is to determine where each read likely originated from in the reference sequence. This information is important for many downstream applications, such as variant calling, i.e. identifying differences between the reads and the reference genome,  and transcriptome analysis, i.e. identifying which genes are being transcribed and at what levels.

There are numerous algorithms designed for aligning reads to a genome. This section introduces fundamental exact matching algorithms, which specifically locate exact occurrences of a given string of characters within a larger body of text or identify patterns in text.

### Read alignment and why it's hard

{% embed url="<https://www.youtube.com/watch?v=PMGstYcBgTY&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=15>" %}

### Naive exact matching

{% embed url="<https://www.youtube.com/watch?v=KUbsdGm3G7s&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=16>" %}

### Practical: Matching artificial reads

{% embed url="<https://www.youtube.com/watch?v=ep91JWd6fs0&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=17>" %}

### Practical: Matching real reads

{% embed url="<https://www.youtube.com/watch?v=SFYpw87lHWQ&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=18>" %} <br>
{% endembed %}

### Boyer-Moore basics

{% embed url="<https://www.youtube.com/watch?v=4Xyhb72LCX4&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=19>" %}

### Boyer-Moore: putting it all together

{% embed url="<https://www.youtube.com/watch?v=Wj606N0IAsw&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=20>" %}

### Diversion: Repetitive elements

{% embed url="<https://www.youtube.com/watch?v=ZwHCRb_y7vA&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=21>" %}

### Practical: Implementing Boyer-Moore

{% embed url="<https://www.youtube.com/watch?v=CT1lQN73UMs&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=22>" %}

## Problems to solve

Try to solve these problems after completing the section.\
Some of them may be challenging :)

1. [Finding a Motif in DNA](https://rosalind.info/problems/subs/)
2. [Open Reading Frames](https://rosalind.info/problems/orf/)
3. [RNA Splicing](https://rosalind.info/problems/splc/)

## Congratulations !

If you made it here, then congratulations! You have successfully completed this section. Move to the next portion of the guide with the arrow buttons below.


# Indexing before alignment

Indexing is an essential step in many bioinformatics applications, as it can greatly reduce the computational time and resources required for sequence alignment. It allows the alignment algorithm to quickly locate the query sequences in the reference genome, without having to search the entire genome for matches. In this section you will learn how indexing works.<br>

### Preprocessing

{% embed url="<https://www.youtube.com/watch?v=HGVQi5xX44M&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=23>" %}

### Indexing and k-mer indexes

{% embed url="<https://www.youtube.com/watch?v=2UsmUgJtwAI&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=24>" %}

### Ordered structures for indexing

{% embed url="<https://www.youtube.com/watch?v=CA0_CxKe_SY&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=25>" %}

### Hash tables for indexing

{% embed url="<https://www.youtube.com/watch?v=X2OTyGTi82M&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=26>" %}

### Practical: Implementing a k-mer index

{% embed url="<https://www.youtube.com/watch?v=LUcAXi2TijM&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=27>" %}

### Variations on k-mer indexes

{% embed url="<https://www.youtube.com/watch?v=My_sw_Rf_4U&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=28>" %}

### Genome indexes used in research

{% embed url="<https://www.youtube.com/watch?v=bfnFgcB_Nlw&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=29>" %}

## Problem to solve

Try to solve this problem after completing the section.

1. [Finding a Spliced Motif](https://rosalind.info/problems/sseq/)

## Congratulations!

If you made it here, then congratulations! You have successfully completed this section. Move to the next portion of the guide with the arrow buttons below.


# Read alignment: approximate matching

Often, reads do not exactly match the reference genome, owing to natural genomic variations or errors introduced during sequencing. In such instances, approximate matching algorithms facilitate identifying similarities between the query and reference sequences, enabling the alignment of reads. This approach is especially useful when working with organisms that are closely related, as their genomes may have evolved somewhat differently over time.

In this section you will learn how approximate matching algorithms work.

### Approximate matching, Hamming and edit distance

{% embed url="<https://www.youtube.com/watch?v=MCvHeW13DsE&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=30>" %}

### Pigeonhole principle

{% embed url="<https://www.youtube.com/watch?v=duatMINEGgE&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=31>" %}

### Practical: Implementing the pigeonhole principle

{% embed url="<https://www.youtube.com/watch?v=9M8aNFgwNG0&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=32>" %}

### Solving the edit distance problem

{% embed url="<https://www.youtube.com/watch?v=8Q2IEIY2pDU&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=33>" %}

### Using dynamic programming for edit distance

{% embed url="<https://www.youtube.com/watch?v=0KzWq118UNI&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=34>" %}

### Practical: Implementing dynamic programming for edit distance

{% embed url="<https://www.youtube.com/watch?v=Xg6uyW9Bscs&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=35>" %}

### Edit distance for approximate matching

{% embed url="<https://www.youtube.com/watch?v=NjfNZzJiu8o&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=36>" %}

## Problems to solve

Try to solve these problems after completing the section.

1. [Counting Point Mutations](https://rosalind.info/problems/hamm/)
2. [Finding a Shared Motif](https://rosalind.info/problems/lcsm/)
3. [Enumerating Gene Orders](https://rosalind.info/problems/perm/)
4. [Enumerating k-mers Lexicographically](https://rosalind.info/problems/lexf/)

If these were too easy for you, try unlocking the following set of **advanced problems**

1. [Longest Increasing Subsequence](https://rosalind.info/problems/lgis/)
2. [k-Mer Composition](https://rosalind.info/problems/kmer/)
3. [Finding a Shared Spliced Motif](https://rosalind.info/problems/lcsq/)
4. [Edit Distance](https://rosalind.info/problems/edit/)
5. [Edit Distance Alignment](https://rosalind.info/problems/edta/)
6. [Counting Optimal Alignments](https://rosalind.info/problems/ctea/)

## Congratulations!

If you made it here, then congratulations! You have successfully completed this section. Move to the next portion of the guide with the arrow buttons below.


# Global and local alignment

There are two primary approaches to aligning two sequences. Global alignment involves aligning the entire length of both sequences, whereas local alignment focuses only on aligning regions that exhibit similarity. Global alignment is suitable for comparing sequences expected to be similar overall but may differ in some areas. In contrast, local alignment is effective for identifying similar segments within longer sequences that may otherwise differ substantially. Local alignments are commonly utilized to pinpoint functional or structural motifs within a sequence.

### Meet the family: global and local alignment

{% embed url="<https://www.youtube.com/watch?v=-bjSPP2v6_Q&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=37>" %}

### Practical: Implementing global alignment

{% embed url="<https://www.youtube.com/watch?v=BGV-hUoHF9k&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=38>" %}

### Read alignment in the field

{% embed url="<https://www.youtube.com/watch?v=u06cm08U0BQ&list=PL2mpR0RYFQsBiCWVJSvVAO3OJ2t7DzoHA&index=39>" %}

## Problems to solve (advanced)

Try to solve these problems after completing this section.

1. [Global Alignment with Scoring Matrix](https://rosalind.info/problems/glob/)
2. [Maximizing the Gap Symbols of an Optimal Alignment](https://rosalind.info/problems/mgap/)
3. [Global Alignment with Constant Gap Penalty](https://rosalind.info/problems/gcon/)
4. [Local Alignment with Scoring Matrix](https://rosalind.info/problems/loca/)

## Congratulations!

If you made it here, then congratulations! You have successfully completed this section. Move to the next portion of the guide with the arrow buttons below.


# NGS DATA ANALYSIS & FUNCTIONAL GENOMICS

## Introduction

We will explore Next-Generation Sequencing (NGS) data analysis, starting from raw data processing to final interpretation. We will begin by learning how to handle and analyze raw NGS data, understanding the main data formats used to store genomic information. We will cover Whole Genome (WG) data handling, focusing on the key steps of alignment and variant calling. By the end of this course, you will have a solid understanding of the complete workflow involved in NGS data analysis, including the tools and techniques essential for accurate and efficient genomic data processing.

## Why Do We Need NGS Data?

**1. Genome Assembly**\
Genome assembly involves reconstructing the complete sequence of an organism's DNA from small, overlapping fragments generated by sequencing. This is essential for studying organisms without a reference genome or for improving existing assemblies. NGS provides the high-throughput data needed to piece together these fragments accurately, enabling a comprehensive view of an organism's genome.

**2. Variant Detection**\
Variant detection identifies differences between an individual's genome and a reference genome, such as single nucleotide polymorphisms (SNPs), insertions, deletions, or structural variations. These variants are crucial for understanding genetic diversity, disease mechanisms, and traits of interest in both research and clinical settings. NGS's sensitivity and resolution make it ideal for uncovering even rare or complex variants.\
Learn more about genetic variation [here](https://www.ebi.ac.uk/training/online/courses/human-genetic-variation-introduction/what-is-genetic-variation/)

**3. Gene Expression Analysis**\
Gene expression analysis determines which genes are actively transcribed and at what levels under specific conditions or in different tissues. NGS, provides a precise and quantitative method for measuring gene expression. This allows researchers to identify gene regulation patterns, compare expression across samples, and study pathways involved in biological processes and diseases.

Beyond these, NGS has many more applications, such as epigenomics, metagenomics, and transcriptome profiling, making it an indispensable tool in modern biology.


# Experimental Techniques

## Learning outcomes:

* Have a fundamental knowledge in theoretical basics experimental techniques like PCR&#x20;
* Have a general understanding of the principles of sequencing technologies
* Understand the differences between the three generations of sequencing technologies


# Polymerase Chain Reaction

## Polymerase Chain Reaction&#x20;

Polymerase Chain Reaction (PCR) is a revolutionary molecular biology technique used to amplify specific DNA sequences. Developed in the 1980s, PCR enables researchers to produce millions of copies of a particular DNA segment quickly and efficiently. This process involves repeated cycles of heating and cooling to denature the DNA, anneal primers, and extend new DNA strands with the help of a DNA polymerase enzyme. PCR is widely used in various fields, including medical diagnostics, forensic science, and genetic research, making it an essential tool for DNA analysis and manipulation.

Watch the video below to understand how PCR works.

{% embed url="<https://www.khanacademy.org/science/ap-biology/gene-expression-and-regulation/biotechnology/v/the-polymerase-chain-reaction-pcr>" %}

## Gel electrophoresis

Gel electrophoresis is a laboratory technique used to separate mixtures of DNA, RNA, or proteins according to their size and charge. This method involves applying an electric current to a gel matrix, typically made of agarose or polyacrylamide. Molecules migrate through the gel at different rates, allowing for the separation and analysis of the components in the mixture.

Gel electrophoresis is commonly used to analyze the results of Polymerase Chain Reaction (PCR). PCR amplifies specific DNA sequences, producing a large number of copies of a particular DNA fragment. Gel electrophoresis allows researchers to verify the success of the PCR amplification and to determine the size of the amplified DNA fragments.

Watch the video below to understand how electrophoresis works.

{% embed url="<https://www.khanacademy.org/science/ap-biology/gene-expression-and-regulation/biotechnology/v/gel-electrophoresis-dna>" %}


# Sanger (first generation) Sequencing Technologies

**DNA Sequencing** is a way to figure out the order of nucleotides along a strand of DNA. **Sanger sequencing** represents the first generation of sequencing technologies, and is based on the process of **DNA replication**. Scientists make copies of DNA strands. Then they observe which nucleotides are added. This way the sequence of nucleotides can be seen.&#x20;

To learn more about this topic, navigate through [this link](https://letstalkscience.ca/educational-resources/backgrounders/sanger-sequencing), and completely review the content on the page. Return here once you finish.<br>

You can also watch this video, to better understand the concept of Sanger sequencing

{% embed url="<https://www.youtube.com/watch?v=-QIMkQ4E_wE>" %}

##


# Next (second) Generation Sequencing technologies

Next-generation sequencing (NGS) is a massively parallel sequencing technology that offers ultra-high throughput, scalability, and speed. The technology is used to determine the order of nucleotides in entire genomes or targeted regions of DNA or RNA. NGS has revolutionized the biological sciences, allowing labs to perform a wide variety of applications and study biological systems at a level never before possible.&#x20;

The third generation of sequencing technologies followed, with the power of sequencing long stretches of nucleic acids.

Let's start with the review of 2nd and 3rd Generation Sequencing technologies.&#x20;

Now, watch the following video explaining Illumina sequencing by synthesis in more detail.

{% embed url="<https://www.youtube.com/watch?v=mI0Fo9kaWqo>" %}


# The third generation of sequencing technologies

The third generation of sequencing technologies has enabled the generation of long stretches of contiguous sequence information from a single molecule of DNA. This can be useful for resolving complex genomic structures and accurately assembling large genomes.  Additionally, third-generation sequencing technologies may enable fast and cost-efficient applications, for example, in diagnostics.&#x20;

The common technologies are Single Molecule Real-Time (SMRT) sequencing by Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT) sequencing.&#x20;

Watch the video below to understand how Nanopore sequencing works.&#x20;

{% embed url="<https://youtu.be/E9-Rm5AoZGw>" %}

##


# The Linux Command-line

## What Is Command-line?

The **command-line**, also called command screen, or text interface is a user interface that's navigated by typing commands at prompts, instead of using a mouse.

## Why Should I Learn Command-Line?

Using a command-line, you can perform almost all the same tasks that can be done with a GUI,  but quicker and easier to automate and do remotely. Good knowledge of the command-line is an essential skill for any successful bioinformatician, as most of the common bioinformatics tools operate in the command-line environment.

## Learning outcomes

At the end of this week, you should:

* Know how to connect to the server and use the command-line&#x20;
* Be familiar with the Bash terminal & the Linux command-line
* Get familiar with sequencing and annotation file formats
* Get familiar with common command-line tools for genomics


# Connecting to the Server

For the following sections of our guide, you will need to have access a Linux command-line environment to practice with hands-on exercises. If your computer operates on a Unix system with more than 8GB RAM and 100GB storage space, you're in luck! Although real-world bioinformatics datasets require a much more space and memory for processing, you can still use your laptop for the practical exercises in this guide.&#x20;

If you don't have access to a Linux environment, there are two options you can try:&#x20;

1. You can install a Linux Virtual Machine (VM) on your computer. This requires patience, as VMs can be slow ;)&#x20;
2. You can apply for access to the ABI computing resources. This will give you access to a fast server, where we will also provide example datasets used in the tutorials that follow.&#x20;

### Installing a Virtual Machine

Register for the following course from Johns Hopkins University here:&#x20;

{% embed url="<https://www.coursera.org/learn/genomic-tools>" %}

Then access the **VMBox Download & Instructions** in the Week 1 materials.

#### Access hints

To access this course, you’ll need to create a Coursera account. Financial aid is available, allowing you to take the course for free. The application process is simple and straightforward, so don’t hesitate—just go for it!

### Accessing the ABI server&#x20;

#### Step 1: Generate a key pair

An SSH key pair is like a lock and key for secure online access. The **private key** is your secret key that you keep safe, and the **public key** is the lock you give to the server. When they match, the server knows it's you, allowing secure access without needing a password.

* If you have a Windows machine, follow [this tutorial](https://www.oracle.com/webfolder/technetwork/tutorials/obe/cloud/javaservice/JCS/JCS_SSH/create_sshkey.html) to generate your key pair. **Remember the location of your private and public keys!** You will provide your public key when requesting access. When granted access, you will need to locate your private key, so **be sure to remember where you stored it** for future use!&#x20;
* If you have a Unix Machine, you can use [this tutorial](https://docs.oracle.com/en/cloud/cloud-at-customer/occ-get-started/generate-ssh-key-pair.html#GUID-8B9E7FCB-CEA3-4FB3-BF1A-FD3406A2432F__GENERATINGANSSHKEYPAIRONUNIXANDUNIX-B329368E).

#### Step 2: Apply for access

Please, make a donation of at least 10,000 AMD at [donate.abi.am](https://donate.abi.am) as a token of gratitude and support to ABI. Then fill in [this form](https://docs.google.com/forms/d/e/1FAIpQLSfs-Tw8s1Eb9EbwHxicVRY1wQ4d_Z54Qgt3XyPX_YNW1BbgAA/viewform?usp=dialog) to request a month-long access. You can prolong your access by filling in another form as you go. You will need to provide your public key when filling the form, so have it ready.&#x20;


# The Linux Command-Line For Beginners

To learn the basics of the Linux command line, complete their [official tutorial](https://ubuntu.com/tutorials/command-line-for-beginners#1-overview), and return here afterward.

{% embed url="<https://ubuntu.com/tutorials/command-line-for-beginners>" %}

Make sure to complete the entire tutorial, as you will use the acquired knowledge in the future sections of the guide.


# The Bash Terminal

Lastly, in order to complete your journey of learning command-line, watch the video below to learn the basic commands of the Bash terminal, which you will use later in the guide.

{% embed url="<https://youtu.be/oxuRxtrO2Ag>" %}


# File formats, alignment, and genomic features

This module is based  the online  [Coursera course Command Line Tools for Genomic Data Science offered by Johns Hopkins University](https://www.coursera.org/learn/genomic-tools/home/welcome). In this module, we will explore the primary data formats used in genomic data analysis and learn how to use command-line tools to process and analyze them.

## Course material for genomic data analysis

### On the server

The example data materials are stored on our server, we will give a path once you open a user .

The files are organized by type and assignment. You need to copy them to your working directory to begin working with them.

**Note:** All upcoming file paths will be provided relative to this directory.

### On Zenodo

Additionally, the example data materials are stored on Zenodo. You can download them and use them locally on your computer.

{% embed url="<https://zenodo.org/records/14547079>" %}


# FASTA & FASTQ file formats

## FASTA and FASTQ data formats

FASTA and FASTQ are two fundamental file formats used in bioinformatics for storing nucleotide sequences.

**FASTA** is a simple text-based format that stores nucleotide or protein sequences. Each entry in a FASTA file begins with a header line starting with a '>' character, followed by the sequence identifier and optional description. The subsequent lines contain the sequence data. FASTA is widely used for sequence alignment and database searches.

**FASTQ** is an extension of the FASTA format that includes quality scores for each nucleotide. Each entry in a FASTQ file consists of four lines: a header line starting with '@' followed by the sequence identifier, a line with the sequence data, a '+' separator line, and a line with the quality scores. FASTQ is essential for storing and analyzing raw sequencing data, providing both sequence information and the associated quality metrics.

Follow the link below for more details about FASTA and FASTQ formats

{% embed url="<https://compgenomr.github.io/book/fasta-and-fastq-formats.html>" %}

<br>


# Basic Unix Commands for Genomics

Let's start by learning how to use basic command-line commands.

Register for the following course from Johns Hopkins University here:&#x20;

{% embed url="<https://www.coursera.org/learn/genomic-tools>" %}

#### Access hints

To access this course, you’ll need to create a Coursera account. Financial aid is available, allowing you to take the course for free. The application process is simple and straightforward, so don’t hesitate—just go for it!

### Get started with basic Unix commands

Watch the following videos to get started, then come back here to continue!

1. [Basic Unix Commands 10: Archiving Content](https://www.coursera.org/learn/genomic-tools/lecture/pPBJb/basic-unix-commands-10-archiving-content)
2. [Basic Unix Commands 11: Practical Exercises I](https://www.coursera.org/learn/genomic-tools/lecture/2s9ut/basic-unix-commands-11-practical-exercises-i)
3. [Basic Unix Commands 12: Practical Exercises II](https://www.coursera.org/learn/genomic-tools/lecture/VzyF7/basic-unix-commands-12-practical-exercises-ii)

## Materials

You can use the following files, which are the ones featured in the video materials.

The files are stored in the following directory on the server:

> Plants/


# Sequences and Genomic Features Part 1

This module explores key topics in genomic data science, starting with a molecular biology primer to review the basics of DNA, RNA, and proteins. It covers sequence representation, generation, and annotation to give biological meaning to raw data. You'll learn the fundamentals of sequence alignmen. Finally, it introduces tools for retrieving genomic features efficiently, enabling deeper insights into genome structure and function.

1. [Sequences and Genomic Features 1: Molecular Bio Primer](https://www.coursera.org/learn/genomic-tools/lecture/aN4oD/sequences-and-genomic-features-1-molecular-bio-primer)
2. [Sequences and Genomic Features 2: Sequence Representation and Generation](https://www.coursera.org/learn/genomic-tools/lecture/wnlsM/sequences-and-genomic-features-2-sequence-representation-and-generation)
3. [Sequences and Genomic Features 3: Annotation](https://www.coursera.org/learn/genomic-tools/lecture/s3pYM/sequences-and-genomic-features-3-annotation)
4. [Sequences and Genomic Features 4.1: Alignment I](https://www.coursera.org/learn/genomic-tools/lecture/m49pF/sequences-and-genomic-features-4-1-alignment-i)
5. [Sequences and Genomic Features 4.2: Alignment II](https://www.coursera.org/learn/genomic-tools/lecture/mu1aa/sequences-and-genomic-features-4-2-alignment-ii)
6. [Sequences and Genomic Features 5: Recreatng Sequences and Features](https://www.coursera.org/learn/genomic-tools/lecture/IMNuh/sequences-and-genomic-features-5-recreating-sequences-features)
7. [Sequences and Genomic Features 6: Genomic Feature Retrieval](https://www.coursera.org/learn/genomic-tools/lecture/Sv2sD/sequences-and-genomic-features-6-genomic-feature-retrieval)


# Sequences and Genomic Features Part 2: SAMtools

In the previous module, we learned how to perform an alignment to a reference genome and explored the resulting SAM/BAM files generated after alignment. We also covered the main properties of these file formats.

Now, let’s dive into SAMtools, a powerful suite of tools that allows us to efficiently manipulate and analyze these data files.

* [Sequences and Genomic Features 7: SAMtools I](https://www.coursera.org/learn/genomic-tools/lecture/6YQMM/sequences-and-genomic-features-7-samtools-i)
* [Sequences and Genomic Features 8: SAMtools II](https://www.coursera.org/learn/genomic-tools/lecture/GUgmY/sequences-and-genomic-features-8-samtools-ii)

## Materials

You can use the following files, which are similar to the examples presented in the video.&#x20;

The files are stored in the following directory on the server:

> sam\_bam/


# Sequences and Genomic Features Part 3: BEDtools

BED (Browser Extensible Data) files are simple text files that define genomic regions of interest, such as genes, exons, or regulatory elements. They consist of columns specifying chromosome, start position, end position, and optional annotations. BED files are widely used for tasks like visualizing regions in genome browsers, performing feature-based analyses, and comparing genomic intervals.

EDtools is a  set of utilities often referred to as the "Swiss Army knife" of genomics. It enables a wide range of tasks, including intersecting, merging, and analyzing genomic intervals. To deepen your understanding of BEDtools and its applications in bioinformatics, watch the following videos from the Coursera course and return here to continue.

* ​[Sequences and Genomic Features 9: BEDtools I](https://www.coursera.org/learn/genomic-tools/lecture/6SpUD/sequences-and-genomic-features-9-bedtools-i)​
* ​[Sequences and Genomic Features 10: BEDtools II](https://www.coursera.org/learn/genomic-tools/lecture/xcDQV/sequences-and-genomic-features-10-bedtools-ii)​

Now, read [the instructions](https://www.coursera.org/learn/genomic-tools/supplement/4zWiX/module-2-exam-instructions-important) and complete the [practice quiz](https://www.coursera.org/learn/genomic-tools/exam/c9MBd/module-2-exam).

### Materials

You can use the following files, which are similar to the examples presented in the video.&#x20;

The **bed** files are stored in the following directory on the server:

> bed/


# Genetic variations & variant calling


# Genomic Variations

Genomic variations are differences in the DNA sequence between individuals within a species. There are several different types of genomic variations, including single nucleotide polymorphisms (SNPs), insertions and deletions, copy number variations (CNVs), and structural variations.

SNPs are the most common type of genomic variation, and they occur when a single nucleotide (A, T, C, or G) in the DNA sequence is altered. These variations can have a range of effects, from having no impact on an organism's characteristics to causing significant changes.

Insertions and deletions are changes in the DNA sequence that involve the addition or removal of nucleotides. These variations can be small, involving just a few nucleotides, or large, involving thousands of nucleotides.

Copy number variations (CNVs) are changes in the number of copies of a particular section of DNA. These variations can have a range of effects, from having no impact on an organism's characteristics to causing significant changes.

Structural variations are changes in the structure of the DNA molecule itself. These variations can involve the rearrangement of large sections of DNA, and they can have a range of effects, from having no impact on an organism's characteristics to causing significant changes.

Overall, genomic variations are an important source of diversity within a species, and they can have a range of effects on an organism's characteristics and behavior.

Read the following chapter to gain a general understanding of the use of variant identification and analysis in bioinformatics and genomics.

{% embed url="<https://www.ebi.ac.uk/training/online/courses/human-genetic-variation-introduction/variant-identification-and-analysis/>" %}


# Alignment and variant detection: Practical

This module is dedicated to practicing alignment techniques. Here, you’ll apply what you’ve learned about aligning sequences to reference genomes and work hands-on with alignment tools and data formats like SAM/BAM.

1. ​[Alignment and Sequence Variation 1: Overview](https://www.coursera.org/learn/genomic-tools/lecture/KobfR/alignment-sequence-variation-1-overview)​
2. ​[Alignment and Sequence Variation 2: Alignment and Variant Detection Tools](https://www.coursera.org/learn/genomic-tools/lecture/6DT1v/alignment-sequence-variation-2-alignment-variant-detection-tools)​
3. ​[Alignment and Sequence Variation 3: VCF](https://www.coursera.org/learn/genomic-tools/lecture/pscrb/alignment-sequence-variation-3-vcf)​
4. ​[Alignment and Sequence Variation 4: Bowtie](https://www.coursera.org/learn/genomic-tools/lecture/aZDuX/alignment-sequence-variation-4-bowtie)​
5. ​[Alignment and Sequence Variation 5: BWA](https://www.coursera.org/learn/genomic-tools/lecture/naE7a/alignment-sequence-variation-5-bwa)

#### Materials

You can use the following files, which are similar to the examples presented in the video.&#x20;

For the alignment practice you can use the following **fastq** file and as a reference the **fasta** file

> \#Example fastq raw reads
>
> fastq/example.fastq&#x20;
>
> \#Example reference fasta files
>
> fasta/Sars\_cov\_2.fa

You can also independently review the example VCF files stored in the following folder. Additionally, there is [a page](https://www.ebi.ac.uk/training/online/courses/human-genetic-variation-introduction/variant-identification-and-analysis/understanding-vcf-format/) that explains the VCF data format in detail.

> vcf/


# Integrative Genomics Viewer

The Integrative Genomics Viewer (IGV) is a powerful and user-friendly visualization tool used in bioinformatics to explore and analyze large-scale genomic data. It supports a wide range of data types, including sequence alignments, variant calls, and gene annotations, making it an essential tool for researchers and clinicians. IGV provides interactive, high-performance visualization capabilities, allowing users to zoom in on specific genomic regions, view multiple data tracks, and easily identify patterns and anomalies. By offering an intuitive interface and robust visualization features, IGV facilitates the interpretation and understanding of complex genomic datasets.

Here is a link to the [IGV Desktop application](https://igv.org/doc/desktop/##_top)

Watch this video for a more detailed introduction to IGV.

{% embed url="<https://www.youtube.com/watch?v=E_G8z_2gTYM>" %}

### Materials

You can use the following **bam** and **gff** files, which are similar to the examples presented in the video.&#x20;

The files are stored in the following directory on the server:

> \#bam file
>
> sam\_bam/Sars\_cov\_2.bam
>
> \#gff file
>
> gtf\_gff/Sars\_cov\_2.gff

##


# Variant Calling with GATK

In this section, you will learn how to perform variant calling to identify single nucleotide polymorphisms (SNPs) and small insertions and deletions (indels) from NGS data using one of the most widely used tools. Below are the main steps involved in the variant calling pipeline.

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FDohkfZPGSPsyd8XqUMoh%2Fvariant_calling_workflow.png?alt=media&amp;token=d0d34f32-2009-4373-99cc-23decf3a2705" alt=""><figcaption><p>The photo is adapted form <a href="https://datacarpentry.github.io/wrangling-genomics/instructor/04-variant_calling.html">here</a></p></figcaption></figure>

## GATK

The Genome Analysis Toolkit (GATK) is a software suite created by the Broad Institute to help analyze DNA sequencing data. It is commonly used to find genetic variants, like single nucleotide changes (SNPs) and small insertions or deletions (indels), in genomic data. GATK includes tools to improve the accuracy of the data, such as fixing errors in base quality scores and aligning indels correctly. Its workflows are designed to make variant calling reliable and precise, making GATK a valuable tool for both research and clinical studies. Even if you're new to genomics, GATK's features and step-by-step guides make it easier to start working with your data.

Learn nore about GATK using following link

{% embed url="<https://gatk.broadinstitute.org/hc/en-us/articles/360036194592-Getting-started-with-GATK4>" %}

For an advanced understanding of the variant caller algorithm, you can watch the following video.

{% embed url="<https://www.youtube.com/watch?v=TwxzDHKLM58>" %}

## Practice

In order to perform variant calling yourself, you will use the [Variant Calling with GATK webpage ](https://learn.gencore.bio.nyu.edu/variant-calling/)and the presentation below as helping material.

### Materials

The files are stored in the following directory on the server.&#x20;

Copy it to your home directory

> variant\_calling\_files/

In the folder you will have several files:

> \#Reference fasta file
>
> variant\_calling\_files/ref\_genome/ecoli\_rel606.fasta
>
> \#Raw fastq files
>
> variant\_calling\_files/trimmed\_fastq\_small/SRR2584863\_1.trim.sub.fastq
>
> variant\_calling\_files/trimmed\_fastq\_small/SRR2584863\_2.trim.sub.fastq
>
> \#File with all codes
>
> variant\_calling\_files/variant\_calling\_script.md #which is a markdown file
>
> variant\_calling\_files/variant\_calling\_script.pdf&#x20;

Note! All the paths in the **variant\_calling\_script.md** codes are relative to the  v**ariant\_calling\_files** folder.&#x20;


# RNA Sequencing & Gene expression

## How and why we study the expression patterns of genes

RNA sequencing, also known as RNA-Seq, is a powerful technique used in the analysis of gene expression. RNA-Seq involves sequencing the RNA molecules present in a biological sample, which can provide valuable insights into the genes that are being expressed and the levels at which they are being expressed. This information can be used to better understand the role of genes in various biological processes and can help identify potential targets for therapeutic intervention.

Studies that explore gene expression are essential for understanding the underlying mechanisms of many biological processes. By analyzing the levels of gene expression in different tissues, researchers can gain insights into how genes are regulated and how they interact with each other to perform their specific functions. These studies can also provide important clues about the mechanisms of disease and can help identify potential therapeutic targets for a wide range of conditions.


# Gene expression and how we measure it

### What is gene expression

You may recall from the molecular biology sections of this guide, that during gene expression, the information in a gene is transcribed, or copied, into a molecule called RNA. The RNA molecule is then translated into a specific protein or other functional molecule. Gene expression is regulated at various steps in this process, allowing cells to turn genes on or off as needed to produce the proteins and other molecules they require.

Gene expression is an important aspect of how an organism's traits are inherited and how they can evolve over time. It is also a key factor in how cells and tissues develop and function.

Start by reading a short introduction about the RNA-sequencing data analysis and  gene expression.

{% embed url="<https://compgenomr.github.io/book/rnaseqanalysis.html>" %}

{% embed url="<https://compgenomr.github.io/book/what-is-gene-expression.html>" %}

### Measureing gene expression

There are several techniques that can be used to measure gene expression, including:

1. Northern blotting: This technique involves isolating RNA from cells or tissues and separating it based on size using electrophoresis. The separated RNA is then transferred to a membrane, and specific RNA molecules are detected using a labeled probe.
2. Reverse transcription polymerase chain reaction (RT-PCR): This technique involves using the enzyme reverse transcriptase to make a complementary DNA (cDNA) copy of RNA. The cDNA is then amplified using polymerase chain reaction (PCR), and the amount of amplified cDNA is measured.
3. Microarrays: A microarray is a glass slide or other substrate with thousands of tiny spots of DNA, each representing a different gene. RNA from cells or tissues is labeled with a fluorescent dye and added to the microarray. The amount of fluorescence at each spot is measured, and this can be used to determine the relative expression levels of different genes.
4. RNA sequencing: This technique involves sequentially reading the nucleotide bases in an RNA molecule, which allows for the precise measurement of the abundance of specific RNA molecules.
5. Protein expression: Gene expression can also be measured by analyzing the levels of the proteins encoded by specific genes. This can be done using techniques such as western blotting, immunoprecipitation, and mass spectrometry.

In this section, you will be introduced to NGS RNA sequencing data analysis for gene expression measurement. Visit this subchapters 8.3.1 to 8.3.3 to get a theoretical understanding of the pipeline. Optionally, you may try out the practical exercises with R as well. We will have a separate tutorial on how to perform normalization and differential expression analysis in the next sections.

{% embed url="<https://compgenomr.github.io/book/gene-expression-analysis-using-high-throughput-sequencing-technologies.html>" %}


# Gene expression quantification and normalization

Normalization is an important step in the analysis of gene expression data, as it helps to control for variations in the amount of RNA that is present in a sample. This is important because the amount of RNA in a sample can be affected by factors such as the health of the organism, the tissue being studied, and the experimental conditions.

Normalization allows researchers to compare gene expression levels between different samples and conditions in a meaningful way, by adjusting for any differences in the amount of RNA present. Without normalization, it would be difficult to accurately compare gene expression levels between samples, as differences in RNA levels could mask differences in gene expression.

Read through subchapter 8.3.4 of this textbook to understand what factors contribute to RNA abundance variances and how to address them (R practice parts are optional).&#x20;

{% embed url="<https://compgenomr.github.io/book/gene-expression-analysis-using-high-throughput-sequencing-technologies.html#quantification>" %}

Now, explore the different types of normalization methods. Includes practical exercises with R. &#x20;

{% embed url="<https://hbctraining.github.io/DGE_workshop/lessons/02_DGE_count_normalization.html>" %}


# Explorative analysis of gene expression

Now that you know how to normalize gene expression counts to make samples comparable, you can proceed to different ways of exploring the gene expression patterns in different samples.&#x20;

Go through the subchapter 8.3.6 of the textbook below. In case you'd like to give a try to the R practical, you'll have to follow the previous particles to have the dataset ready.

{% embed url="<https://compgenomr.github.io/book/gene-expression-analysis-using-high-throughput-sequencing-technologies.html#exploratory-analysis-of-the-read-count-table>" %}


# Differential expression analysis with DESeq2

Differential gene expression analysis is a technique used to identify genes that are differentially expressed, or turned on or off, between two or more biological conditions or samples. This can be used to identify genes that are specifically involved in a particular process or disease, or to understand the underlying molecular mechanisms of a biological system.

Once the differentially expressed genes have been identified, researchers can use various techniques, such as gene ontology analysis and pathway analysis, to further understand the biological functions and pathways in which these genes are involved. This information can be used to gain insights into the underlying mechanisms of a biological process or disease, and may ultimately lead to the development of new treatments or therapies.

There are several tools used to identify genes differentially expressed between different conditions or groups of samples. Here, we will explore DESeq2 (Differential Expression analysis for Sequencing). It uses statistical methods to analyze RNA-seq data and identify genes that are differentially expressed between two or more conditions or samples. It takes into account various sources of variability in the data, such as batch effects and technical noise, to accurately identify differentially expressed genes.

We are offering you two tutorials to run DESeq2 on two different samples.&#x20;

## DESeq2 tutorial #1 by Griffith lab

Use your local R Studio and replicate the pipeline described in this tutorial.&#x20;

{% embed url="<https://genviz.org/module-04-expression/0004/02/01/DifferentialExpression/>" %}

## DESeq2 tutorial #2 by Altuna Akalin

You can also try out this tutorial, where you will find additional explanation for the pipeline. Note that in order to be able to access the dataset for this tutorial you'll need to have run the previous R practicals of this textbook subchapter.

{% embed url="<https://compgenomr.github.io/book/gene-expression-analysis-using-high-throughput-sequencing-technologies.html#differential-expression-analysis>" %}


# Functional enrichment analysis

Differential expression analysis can give you a list of genes that differ between groups of samples. In order to understadn the functional consequences of the observed differential gene regulation, you need to undertake functional enrichment analysis steps that are often considered as systems level analyses.&#x20;

Functional enrichment analysis is a statistical method used to identify overrepresented biological pathways or functional categories in a set of genes. It is often used to gain insights into the biological functions and pathways that are potentially involved in a particular biological process or disease.

There are several types of functional enrichment analysis, including gene ontology (GO) enrichment analysis, and pathway enrichment analysis. GO terms are standardized definitions of gene functions that are organized into a hierarchy, allowing researchers to identify the broad categories of functions that are overrepresented in their data. Pathways are sets of interconnected genes that are involved in specific biological processes, and pathway enrichment analysis can be used to identify pathways that are potentially involved in a particular biological process or disease.

Read through the subchapter 8.3.8 to gain general understanding of how functional enrichment analysis is done.

{% embed url="<https://compgenomr.github.io/book/gene-expression-analysis-using-high-throughput-sequencing-technologies.html#functional-enrichment-analysis>" %}


# Single-cell Sequencing and Data Analysis

Single-cell sequencing is a revolutionary technology that allows researchers to explore the genetic and transcriptomic landscapes of individual cells. Unlike traditional bulk sequencing, which averages signals across millions of cells, single-cell sequencing provides a high-resolution view of cellular heterogeneity, uncovering rare cell types, states, and dynamic processes that would otherwise be obscured.\
\
There are many technologies used for sequencing single-cell RNA (scRNA-seq), with one of the more popular platforms being 10x Genomics. For a brief introduction on how it works, see below:\
\
[![](https://embed-ssl.wistia.com/deliveries/817c8c823157c5cd6791daa773576d13.jpg?image_play_button_size=2x\&image_crop_resized=960x540\&image_play_button=1\&image_play_button_color=54bbffe0)](https://www.10xgenomics.com/blog/single-cell-rna-seq-an-introductory-overview-and-tools-for-getting-started?wvideo=f75ht43w1q)

[Single cell RNA-seq: An introductory overview and tools for getting started - 10x Genomics](https://www.10xgenomics.com/blog/single-cell-rna-seq-an-introductory-overview-and-tools-for-getting-started?wvideo=f75ht43w1q)\
\
**How does this image relate to our topic?**

<figure><img src="https://3514673221-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FdDE3NSiCXcu5YQDQmgdU%2Fuploads%2FLkEbAYlsYSdKUKCJPmFt%2Fimage.png?alt=media&amp;token=4a21bcc4-deb0-43ca-a881-6491779d8d27" alt=""><figcaption></figcaption></figure>


# scRNA-seq Data Analysis Workflow

Analyzing single-cell sequencing data involves several key steps:

1. **Preprocessing:** This includes quality control, normalization, and filtering to ensure high-quality data.
2. **Dimensionality Reduction:** Techniques like PCA, t-SNE, and UMAP are used to reduce the complexity of the data and visualize the relationships between cells.
3. **Clustering:** Cells are grouped based on similar gene expression patterns, allowing for the identification of distinct cell types or states.
4. **Differential Expression Analysis:** Identifying genes that are differentially expressed between cell clusters or conditions to understand the underlying biology.

One of the most used packages for single-cell data analysis is [Seurat](https://satijalab.org/seurat/), which is a comprehensive **R** toolkit that facilitates the entire analysis workflow, from data preprocessing to advanced visualization and downstream analysis.\
A guided tutorial of a basic single-cell analysis worklow for a  10X Genomics dataset of Peripheral Blood Mononuclear Cells (PBMC) in Seurat can be found here:&#x20;

{% embed url="<https://satijalab.org/seurat/articles/pbmc3k_tutorial.html>" %}

A **Python** alternative to Seurat is [scanpy](https://scanpy.readthedocs.io/en/stable/index.html), a scalable toolkit for analyzing single-cell gene expression data built jointly with [anndata](https://anndata.readthedocs.io/).  You can find a scanpy tutorial on the same PBMC dataset can be found here:

{% embed url="<https://scanpy-tutorials.readthedocs.io/en/latest/pbmc3k.html>" %}


# scRNA-seq Data Visualization Methods

Visualizing single-cell RNA-seq (scRNA-seq) data is crucial for understanding the complex patterns and relationships within cellular populations. Effective visualization techniques help in identifying clusters, discovering new cell types, and interpreting differential expression results. \
\
We have already explored some visualization methods in the basic scRNA-seq data analysis workflow tutorials. For a deeper dive into various visualization techniques, you can follow the tutorials below.&#x20;

### Seurat

{% embed url="<https://satijalab.org/seurat/articles/visualization_vignette>" %}

### scanpy

{% embed url="<https://scanpy-tutorials.readthedocs.io/en/latest/plotting/core.html>" %}


# FINAL REMARKS

## Congratulations!&#x20;

You’ve done an amazing job working through all the sections of this guide. We hope it has been informative and helpful in your bioinformatics journey.<br>

### Provide feedback

We’d love to hear your thoughts! Please take a few minutes to let us know which modules you found most useful and where we could improve. Your feedback will help us enhance this guide for future users.

[Bioinformatics Guide 2024: Feedback](/)

### Support us

The Armenian Bioinformatics Institute (ABI) is a non-profit organization driven by a passionate team dedicated to advancing bioinformatics training in Armenia and beyond. Every member of our team works tirelessly to make this vision a reality. If you’d like to support our efforts, please consider making a donation. Your contribution helps us expand our programs and reach more people.

Thank you for supporting the future of bioinformatics!

{% embed url="<https://donate.abi.am>" %}

### Continue your career in bioinformatics

If you’ve made it this far, you’re well on your way in your bioinformatics career! We often have openings for internships and research positions in our labs, which we think you might find exciting.

Explore what’s happening in our research labs ([Nersisyan lab](https://www.abi.am/research/nersisyan-lab), [Binder lab](https://www.abi.am/research/binder-lab)), and feel free to contact lab leaders or ABI at **<info@abi.am>**. In your email, share your background, career goals, and why you’re interested in joining our team. We welcome applications from all around the world!

### **Stay Connected**

Follow us on social media and stay updated on upcoming events, workshops, and new opportunities.

* [ABI website](https://abi.am)
* [Facebook](https://www.facebook.com/abi.arm.bio)
* [LinkedIn](https://www.linkedin.com/company/76327403/admin/dashboard/)
* [X (Twitter)](https://x.com/arm_abi)

&#x20;


