0% found this document useful (0 votes)
5 views55 pages

Understanding Biology Through Computing

The document discusses the integration of computer technology in understanding biology, particularly through DNA and RNA sequencing. It highlights the significance of the Human Genome Project, the complexity of human genomes, and the advancements in single-cell RNA sequencing for disease research. Additionally, it emphasizes the importance of non-coding DNA and alternative splicing in genetic functions and diseases.

Uploaded by

ksym812
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views55 pages

Understanding Biology Through Computing

The document discusses the integration of computer technology in understanding biology, particularly through DNA and RNA sequencing. It highlights the significance of the Human Genome Project, the complexity of human genomes, and the advancements in single-cell RNA sequencing for disease research. Additionally, it emphasizes the importance of non-coding DNA and alternative splicing in genetic functions and diseases.

Uploaded by

ksym812
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

The Digital Cell: How to understand

biology using computer


Jihwan Park
GIST School of Life Sciences Associate Professor
Eyeoncell Co., Ltd. CEO
A 1987 American film
Director: Joe Dante, Executive producer: Steven Spielberg
Won an Oscar for Best Visual Effects
The sequencer is our inner space ship and new microscope
A large amount of information is in the individual cells of our body

2m long DNA in individual cells


X
37 trillion cells in our body

displayed at the Wellcome Collection, London


Understanding biology as a flow of information

DNA RNA Amino acids Protein


transcription
translation folding

DNA sequence of ASCL1

displayed at the Wellcome Collection, London


How to understand biology using computer
How to understand biology using computer

Cell types Phenotype

Brain diseases
Neurons

Genetic variants Muscle diseases


Muscle cells

Cardiac diseases
Cardiac cells

Obesity
Adipocytes
A neuron and a liver cell share the same genome
Chromatin structure and epigenetic regulation

Kidney Int. 2011 79(1):23-32.


1 GENOME Catalogue

> 200 EPIGENOME Operation manual

> 200 cell types


DNA sequencing- past

[Link]
breakthroughs-one-letter-at-a-time/
Credit: National Cancer Institute.
DNA sequencing- present
DNA sequencing- future?

Portable DNA sequencers are being


used to monitor epidemics, as shown
here during the Zika outbreak in Brazil.
University of Birmingham
What is the Human Genome Project?

 Decoding the Human Genome

 The DNA sequencing is used to determine the exact sequence of bases


in a DNA molecule (cost $3 billion, took 13 years)

 Aims of the project:


• to identify all genes in the human DNA
• determine the sequences of the 3 billion bases that make up
human DNA
• to obtain a physical map of human genome
• to know the function of genes
• To find out which sequence of DNA is responsible for genetic
disorders
Surprise #1: most of the genome is non-coding

• only 1.5% of the DNA codes for proteins, tRNAs, or rRNAs

• remaining 98.5% of the DNA is noncoding DNA


 transposable elements, telomeres, centromeres and
etc.
But non-coding does not mean non-functional

Development 2017 144: 2548-2559 Oncotarget. 2017; 8:48424-48435


Surprise #2: Human have a small number of protein-
coding genes

Organism # of Genome size # of protein-


chromosomes coding genes
Rice 12 389Mb ~37,000

Mouse 40 2.6Gb ~24,000

Human 46 3.2Gb ~21,000

Chimpanzee 48 2.7Gb ~21,000

Roundworm 12 97Mb ~19,000

Fruit Fly 8 137Mb ~15,000

Bacteria (E. coli) 1 4.6Mb ~3,200


The complexity of human genomes

• Alternative splicing
• Chemical modifications of proteins
• non-coding RNAs
• cis-regulatory sequences
Surprise #3: Human and chimp genomes differ only by 1.23%

Don’t worry, you share 99.9% of your DNA with them!


We are all essentially identical twins
but still there are many differences

0.1% of 3,2000,000,000 = 3.2 million/strand


x 2 strands

= 6.4 million differences!


How to understand biology using computer

Cell types Phenotype

Brain diseases
Neurons

Genetic variants Muscle diseases


Muscle cells

Cardiac diseases
Cardiac cells

Obesity
Adipocytes
RNA sequencing and data analysis
Process of quantifying RNA abundance by determining the precise order of nucleotides within a DNA strand
? ? ? ? ? ? ? ?? ? ? ? ? ? ? ? ? ? ? ?? ? ? ? ? ? ? ? ? ? ? ?? ? ? ? ? ? ? ? ? ? ? ?? ? ? ?

RNA ? ? ? ? ? ? ? ?? ? ? ? ? ? ? ? ? ? ? ?? ? ? ? ? ? ? ? ? ? ? ?? ? ? ?
? ? ? ? ? ? ? ?? ? ? ? ? ? ? ? ? ? ? ?? ? ? ?
? ? ? ? ? ? ? ?? ? ? ?

Seqeuncing

Sequencing
reads Alignment

Reference
genome

gene1 gene2 gene3 gene4 Gene


3 1 4 2 quantification

Gene expression patterns


Disease research using RNA sequencing

Normal Disease

Energy metabolism
Genes

Inflammation

Expression level

low high
Single cell RNA sequencing

G
Single cell data matrix

PT Cells 1 2 3 4 5
Gene 1 = 2 6 0 3 1
DCT Gene 2 = 0 0 4 5 3

Gene 3 = x
LOH

......
[Link] (Chapter 37) CD

G: glomerulus, PT: proximal tubule, DCT: distal convoluted tubule


LOH: loop of Henle, CD: collecting duct
Bulk RNA sequencing vs Single cell RNA sequencing

bulk RNA sequencing

Single cell RNA sequencing


How much RNA does a typical mammalian cell contain?

[Link]

• 360,000 mRNA molecules


• 10–30 pg total RNA
• 0.25-0.9 pg mRNA
Evolution of single cell RNAseq techniques

Svensson et al. 2018 Nature Protocol


Single cell RNA sequencing workflow
Single Cell Atlas of the Healthy Kidney (43,745 cells)

cluster nCells %
13 4 1 1,001 2.29 Endothelial cells
1
2 78 0.18 Podocytes
14
16 3 26,482 60.54 Proximal tubule epithelial cells

2 15 4 1,581 3.61 Loop of Henle


5 8,544 19.53 Distal tubule epithelial cells
12 6 6 870 1.99 Collecting duct/Principal cells
7 1,729 3.95 Collecting duct/Intercalated cells
3 8
8 110 0.25 Collecting duct
10 7 9 601 1.37 Novel cell type
9
10 549 1.26 Fibroblasts
11 228 0.52 Macrophages
12 74 0.17 Neutrophils
13 235 0.54 B lymphocytes
5 14 1,308 2.99 T lymphocytes
11 15 313 0.72 Lymphocytes/NK-cells
16 42 0.1 Novel cell type 2
The Human Cell Atlas project

The Scientist Magazine Domcke et al. 2020 Science


Single cell RNA sequencing for disease research

Normal Disease

Sandberg 2014 Nature Methods


Transcriptomics
Single cell long
(Science ‘18, Kidney Int. ‘19)
read sequencing
(under review) Technology Multi-omics

Cell lineage tracing Epigenomics


(Kidney Int. 2024) (Kidney Int. ’24)
Pipeline
development Spatial
(under revision) Single cell
transcriptomics
biology

Application

• Thyroid cancer (Sci. Adv. ’22, Adv. Sci. ‘24) • Kidney fibrosis (Cell Metabol. ‘20,
• Pan-cancer (Nat. Comm. ‘22) Kidney Int. ‘24)
• Acute myeloid leukemia (Mol. Cells ‘23) • Aging (JASN ‘24)
• Tumor microenvironment (Mil. Med. Res. ‘22) • Diabetes induced erectile dysfunction
(eLife ‘23)
Integration of single cell multi-omics data

• Biological and technical variation in genome/epigenome/transcriptome/proteome


• Identification of genetic or epigenetic effects on gene expression
• Pre-, post-transcriptional regulation

배서경, 김경대, 박지환 TiBMB 2020


The single-cell multi-omics studies

Single cell transcriptome profiling is highly scalable

Wen et al 2022 The Innovation


2세대 NGS – Short read sequencing

2세대 NGS (Illumina)


전체 NGS 시장의
약 70% 점유

최근까지 유전체 결정에 있어서 주된 data로 사용

진정한 인간유전체 결정 $1,000 시대

100~150 bp의 짧은(short-read) 서열을 제공


유전체를 조립하면 구조 적인 에러 발생 확률 높음
Third generation sequencing:
long-read, real-time, single molecule sequencing
Oxford Nanopore Technologies PacBio

Kraft and Kurth. Medizinische genetic 2019


 High throughput sequencing core  Single cell analysis
- Nanopore MinION, PromethION, PacBio Revio - 10x Genomics Single cell multi omics
(3rd generation sequencers) - MERSCOPE Spatial transcriptomics
- Elements AVITI NGS

High performance
computing servers
Short-read seq vs long read sequencing

Udine et al. Mol. Neurodegener. 2023


Solve more genetic diseases with long-read sequencing

L. Hickey [Link] 2020


The Functional Impact of Alternative Splicing in
Cancer and normal physiology

González et al 2017 Cell Reports Stevens and Oltean 2016 JASN


Long read sequencing for a single cell RNA sequencing libraries

• Cost-effective short-read sequencing can read only 3’ ends of mRNAs


• 3’ ends of mRNAs + cell barcodes  can quantify the gene expression levels

polyA tail

economic,
Accurate,
Short-read

Expensive,
Low accuracy
Ultra long

Alternative splicing
Somatic mutations within gene body
[Link]
Transposable elements
Ouro-seq: Circularization-based targeted scRNAseq

11
Ouro-seq can efficiently cover full-length transcripts

cDNAs with cell


barcode sequence

Before Ouro-seq

After Ouro-seq

43
unpublished
Comparison with Previous Methods - Cell Barcode containing reads

Proportion of biological
3' or 5' end identifying
molecules (>3,000bp)
0.8

0.6

0.4
Ouro-Seq applied
0.2 Previous studies
(16 studies, 2017-2023)
0

0.6
molecules (>3,000bp)
biological full-length
Proportion of true

0.4

0.2

0
1,000bp 1,500bp 2,000bp 2,000bp 3,000bp
Median molecule length (N50) satisfying the criteria
(all molecule sizes)

unpublished
Ouro-tools, a comprehensive toolkit for QC and Ouro-seq data analysis
Ouro-tools provides accurate TSS and TES mapping with
unreferenced Gs and unreferenced As

External PolyA
Unreferenced Gs

46
unpublished
Ouro-tools provides accurate TSS and TES mapping with
unreferenced Gs and unreferenced As

47
unpublished
Constructing a mouse alternative splicing atlas

Throughput
Platform
per sample
HiSeq 50Gbp
(short-read) (175M reads)
PromethION 150Gbp Total 3.4Tbp (1,653M reads)
(long-read) (50M reads)
48
unpublished
Ouro-seq distinguishes transcript isoform expression (Hnrnpf)

49
unpublished
Cell type specific exon switching patterns

50
unpublished
Ouro-seq distinguishes transcript isoform expression (Tpm1)
Ouro-seq distinguishes transcript isoform expression (Tpm1)
Cell type-specific splicing isoforms

NKCC2-F: the lowest binding affinities for Na+, K-, and Cl-
NKCC2-B: the highest binding affinity for Na+, K-, and Cl-
NKCC2-A: intermediate binding affinity for Na+, K-, and Cl-
Discovery of novel promoters
Acknowledgement
GIST
Chungnam National University Hospital
HyunSu An PhD candidate
Sin Young Choi, PhD Yea Eun Kang, MD, PhD
So-I Shin, PhD Koo Bon Seok, MD, PhD
Gyeong Dae Kim, PhD candidate Jin Man Kim, MD, PhD
Jawoon Yi, MS
KyuMin Park Undergraduate Da Hyun Kang, MD, PhD
Minho Eun, MS candidate
Donggun Kim, PhD candidate
Kyung Hee University Hospital
Seo-Gyeong Bae, MS
Ju-Young Moon, MD, PhD
Seoul National University Hospital
Ha Jeong Lee, MD, PhD Inha University Hospital
Chung-Ang University Jun-Kyu Suh, MD, PhD
Kyoung-Dong Kim, PhD Ji-Kan Ryu, MD, PhD

Sichuan University Chonnam National University Hospital


Han Luo, MD, PhD Jae-Sook Ahn, MD
Jingqiang Zhu, MD Hyeoung-Joon Kim, MD, PhD
Heng Xu, PhD
KNIH
Byeong-Sun Choi, PhD

You might also like