The Digital Cell: How to understand
biology using computer
Jihwan Park
GIST School of Life Sciences Associate Professor
Eyeoncell Co., Ltd. CEO
A 1987 American film
Director: Joe Dante, Executive producer: Steven Spielberg
Won an Oscar for Best Visual Effects
The sequencer is our inner space ship and new microscope
A large amount of information is in the individual cells of our body
2m long DNA in individual cells
X
37 trillion cells in our body
displayed at the Wellcome Collection, London
Understanding biology as a flow of information
DNA RNA Amino acids Protein
transcription
translation folding
DNA sequence of ASCL1
displayed at the Wellcome Collection, London
How to understand biology using computer
How to understand biology using computer
Cell types Phenotype
Brain diseases
Neurons
Genetic variants Muscle diseases
Muscle cells
Cardiac diseases
Cardiac cells
Obesity
Adipocytes
A neuron and a liver cell share the same genome
Chromatin structure and epigenetic regulation
Kidney Int. 2011 79(1):23-32.
1 GENOME Catalogue
> 200 EPIGENOME Operation manual
> 200 cell types
DNA sequencing- past
[Link]
breakthroughs-one-letter-at-a-time/
Credit: National Cancer Institute.
DNA sequencing- present
DNA sequencing- future?
Portable DNA sequencers are being
used to monitor epidemics, as shown
here during the Zika outbreak in Brazil.
University of Birmingham
What is the Human Genome Project?
Decoding the Human Genome
The DNA sequencing is used to determine the exact sequence of bases
in a DNA molecule (cost $3 billion, took 13 years)
Aims of the project:
• to identify all genes in the human DNA
• determine the sequences of the 3 billion bases that make up
human DNA
• to obtain a physical map of human genome
• to know the function of genes
• To find out which sequence of DNA is responsible for genetic
disorders
Surprise #1: most of the genome is non-coding
• only 1.5% of the DNA codes for proteins, tRNAs, or rRNAs
• remaining 98.5% of the DNA is noncoding DNA
transposable elements, telomeres, centromeres and
etc.
But non-coding does not mean non-functional
Development 2017 144: 2548-2559 Oncotarget. 2017; 8:48424-48435
Surprise #2: Human have a small number of protein-
coding genes
Organism # of Genome size # of protein-
chromosomes coding genes
Rice 12 389Mb ~37,000
Mouse 40 2.6Gb ~24,000
Human 46 3.2Gb ~21,000
Chimpanzee 48 2.7Gb ~21,000
Roundworm 12 97Mb ~19,000
Fruit Fly 8 137Mb ~15,000
Bacteria (E. coli) 1 4.6Mb ~3,200
The complexity of human genomes
• Alternative splicing
• Chemical modifications of proteins
• non-coding RNAs
• cis-regulatory sequences
Surprise #3: Human and chimp genomes differ only by 1.23%
Don’t worry, you share 99.9% of your DNA with them!
We are all essentially identical twins
but still there are many differences
0.1% of 3,2000,000,000 = 3.2 million/strand
x 2 strands
= 6.4 million differences!
How to understand biology using computer
Cell types Phenotype
Brain diseases
Neurons
Genetic variants Muscle diseases
Muscle cells
Cardiac diseases
Cardiac cells
Obesity
Adipocytes
RNA sequencing and data analysis
Process of quantifying RNA abundance by determining the precise order of nucleotides within a DNA strand
? ? ? ? ? ? ? ?? ? ? ? ? ? ? ? ? ? ? ?? ? ? ? ? ? ? ? ? ? ? ?? ? ? ? ? ? ? ? ? ? ? ?? ? ? ?
RNA ? ? ? ? ? ? ? ?? ? ? ? ? ? ? ? ? ? ? ?? ? ? ? ? ? ? ? ? ? ? ?? ? ? ?
? ? ? ? ? ? ? ?? ? ? ? ? ? ? ? ? ? ? ?? ? ? ?
? ? ? ? ? ? ? ?? ? ? ?
Seqeuncing
Sequencing
reads Alignment
Reference
genome
gene1 gene2 gene3 gene4 Gene
3 1 4 2 quantification
Gene expression patterns
Disease research using RNA sequencing
Normal Disease
Energy metabolism
Genes
Inflammation
Expression level
low high
Single cell RNA sequencing
G
Single cell data matrix
PT Cells 1 2 3 4 5
Gene 1 = 2 6 0 3 1
DCT Gene 2 = 0 0 4 5 3
Gene 3 = x
LOH
......
[Link] (Chapter 37) CD
G: glomerulus, PT: proximal tubule, DCT: distal convoluted tubule
LOH: loop of Henle, CD: collecting duct
Bulk RNA sequencing vs Single cell RNA sequencing
bulk RNA sequencing
Single cell RNA sequencing
How much RNA does a typical mammalian cell contain?
[Link]
• 360,000 mRNA molecules
• 10–30 pg total RNA
• 0.25-0.9 pg mRNA
Evolution of single cell RNAseq techniques
Svensson et al. 2018 Nature Protocol
Single cell RNA sequencing workflow
Single Cell Atlas of the Healthy Kidney (43,745 cells)
cluster nCells %
13 4 1 1,001 2.29 Endothelial cells
1
2 78 0.18 Podocytes
14
16 3 26,482 60.54 Proximal tubule epithelial cells
2 15 4 1,581 3.61 Loop of Henle
5 8,544 19.53 Distal tubule epithelial cells
12 6 6 870 1.99 Collecting duct/Principal cells
7 1,729 3.95 Collecting duct/Intercalated cells
3 8
8 110 0.25 Collecting duct
10 7 9 601 1.37 Novel cell type
9
10 549 1.26 Fibroblasts
11 228 0.52 Macrophages
12 74 0.17 Neutrophils
13 235 0.54 B lymphocytes
5 14 1,308 2.99 T lymphocytes
11 15 313 0.72 Lymphocytes/NK-cells
16 42 0.1 Novel cell type 2
The Human Cell Atlas project
The Scientist Magazine Domcke et al. 2020 Science
Single cell RNA sequencing for disease research
Normal Disease
Sandberg 2014 Nature Methods
Transcriptomics
Single cell long
(Science ‘18, Kidney Int. ‘19)
read sequencing
(under review) Technology Multi-omics
Cell lineage tracing Epigenomics
(Kidney Int. 2024) (Kidney Int. ’24)
Pipeline
development Spatial
(under revision) Single cell
transcriptomics
biology
Application
• Thyroid cancer (Sci. Adv. ’22, Adv. Sci. ‘24) • Kidney fibrosis (Cell Metabol. ‘20,
• Pan-cancer (Nat. Comm. ‘22) Kidney Int. ‘24)
• Acute myeloid leukemia (Mol. Cells ‘23) • Aging (JASN ‘24)
• Tumor microenvironment (Mil. Med. Res. ‘22) • Diabetes induced erectile dysfunction
(eLife ‘23)
Integration of single cell multi-omics data
• Biological and technical variation in genome/epigenome/transcriptome/proteome
• Identification of genetic or epigenetic effects on gene expression
• Pre-, post-transcriptional regulation
배서경, 김경대, 박지환 TiBMB 2020
The single-cell multi-omics studies
Single cell transcriptome profiling is highly scalable
Wen et al 2022 The Innovation
2세대 NGS – Short read sequencing
2세대 NGS (Illumina)
전체 NGS 시장의
약 70% 점유
최근까지 유전체 결정에 있어서 주된 data로 사용
진정한 인간유전체 결정 $1,000 시대
100~150 bp의 짧은(short-read) 서열을 제공
유전체를 조립하면 구조 적인 에러 발생 확률 높음
Third generation sequencing:
long-read, real-time, single molecule sequencing
Oxford Nanopore Technologies PacBio
Kraft and Kurth. Medizinische genetic 2019
High throughput sequencing core Single cell analysis
- Nanopore MinION, PromethION, PacBio Revio - 10x Genomics Single cell multi omics
(3rd generation sequencers) - MERSCOPE Spatial transcriptomics
- Elements AVITI NGS
High performance
computing servers
Short-read seq vs long read sequencing
Udine et al. Mol. Neurodegener. 2023
Solve more genetic diseases with long-read sequencing
L. Hickey [Link] 2020
The Functional Impact of Alternative Splicing in
Cancer and normal physiology
González et al 2017 Cell Reports Stevens and Oltean 2016 JASN
Long read sequencing for a single cell RNA sequencing libraries
• Cost-effective short-read sequencing can read only 3’ ends of mRNAs
• 3’ ends of mRNAs + cell barcodes can quantify the gene expression levels
polyA tail
economic,
Accurate,
Short-read
Expensive,
Low accuracy
Ultra long
Alternative splicing
Somatic mutations within gene body
[Link]
Transposable elements
Ouro-seq: Circularization-based targeted scRNAseq
11
Ouro-seq can efficiently cover full-length transcripts
cDNAs with cell
barcode sequence
Before Ouro-seq
After Ouro-seq
43
unpublished
Comparison with Previous Methods - Cell Barcode containing reads
Proportion of biological
3' or 5' end identifying
molecules (>3,000bp)
0.8
0.6
0.4
Ouro-Seq applied
0.2 Previous studies
(16 studies, 2017-2023)
0
0.6
molecules (>3,000bp)
biological full-length
Proportion of true
0.4
0.2
0
1,000bp 1,500bp 2,000bp 2,000bp 3,000bp
Median molecule length (N50) satisfying the criteria
(all molecule sizes)
unpublished
Ouro-tools, a comprehensive toolkit for QC and Ouro-seq data analysis
Ouro-tools provides accurate TSS and TES mapping with
unreferenced Gs and unreferenced As
External PolyA
Unreferenced Gs
46
unpublished
Ouro-tools provides accurate TSS and TES mapping with
unreferenced Gs and unreferenced As
47
unpublished
Constructing a mouse alternative splicing atlas
Throughput
Platform
per sample
HiSeq 50Gbp
(short-read) (175M reads)
PromethION 150Gbp Total 3.4Tbp (1,653M reads)
(long-read) (50M reads)
48
unpublished
Ouro-seq distinguishes transcript isoform expression (Hnrnpf)
49
unpublished
Cell type specific exon switching patterns
50
unpublished
Ouro-seq distinguishes transcript isoform expression (Tpm1)
Ouro-seq distinguishes transcript isoform expression (Tpm1)
Cell type-specific splicing isoforms
NKCC2-F: the lowest binding affinities for Na+, K-, and Cl-
NKCC2-B: the highest binding affinity for Na+, K-, and Cl-
NKCC2-A: intermediate binding affinity for Na+, K-, and Cl-
Discovery of novel promoters
Acknowledgement
GIST
Chungnam National University Hospital
HyunSu An PhD candidate
Sin Young Choi, PhD Yea Eun Kang, MD, PhD
So-I Shin, PhD Koo Bon Seok, MD, PhD
Gyeong Dae Kim, PhD candidate Jin Man Kim, MD, PhD
Jawoon Yi, MS
KyuMin Park Undergraduate Da Hyun Kang, MD, PhD
Minho Eun, MS candidate
Donggun Kim, PhD candidate
Kyung Hee University Hospital
Seo-Gyeong Bae, MS
Ju-Young Moon, MD, PhD
Seoul National University Hospital
Ha Jeong Lee, MD, PhD Inha University Hospital
Chung-Ang University Jun-Kyu Suh, MD, PhD
Kyoung-Dong Kim, PhD Ji-Kan Ryu, MD, PhD
Sichuan University Chonnam National University Hospital
Han Luo, MD, PhD Jae-Sook Ahn, MD
Jingqiang Zhu, MD Hyeoung-Joon Kim, MD, PhD
Heng Xu, PhD
KNIH
Byeong-Sun Choi, PhD