ChIP-Seq Data Analysis
Presentation by Group 7:
Maksim Danilchyk, Sofya Shorzhina,
Tilo Alves Radtke, Kristian Reinhart
01 02
INTRODUCTION & WHAT IS CHIP-
BIOLOGY SEQ
03 04 05
(PRE-)PROCESSING, PEAK CALLING & APPLICATIONS &
QUALITY CONTROL ANALYSIS RELATED
& CHALLENGES METHODS
2
01
INTRODUCTION &
BIOLOGY
3
WHAT IT
METHOD RESOLUTION OUTPUT BIOLOGICAL QUESTION
MEASURES
Gene expression Bulk cell Gene counts /
RNA-Seq Which genes are expressed?
levels population transcripts
Cell-specific
Gene expression Which cell types are present and
scRNA-Seq Single cell expression
levels what are they expressing?
profiles
Spatial gene
Spatial Gene expression Tissue region What are genes expressed within
expression
Transcriptomics leves + location / single cells tissues / regions?
maps
Single base /
DNA DNA methylation Methylation Which genomic regions are
genomic
Methylation levels percentages epigenetically silenced or active?
regions
Protein-DNA Where do transcription factors,
Binding sites Binding sites
interactions, DNA-binding proteins or
ChIP-Seq (~100-1000 across the
histone histone markers bind?
bp) genome
modifications
4
WHAT IS ChIP-seq?
● A technique combining two methods:
○ ChIP (Chromatin Immunoprecipitation) — uses an antibody to enrich DNA fragments
bound to a protein of interest
○ NGS (Next-Generation Sequencing) — reads out the captured DNA fragments genome-wide
● Together they answer: Where does a protein bind across the entire genome?
ChIP-seq Targets
Transcription
Transcription DNA-binding DNA Histone
CATEGORY factors & Chromatin remodelers
machinery proteins modifications marks
repressors
BRCA1, 5-methylcytosine H3K4me3,
EXAMPLE NANOG, CTCF RNA Pol II BMI1, EZH2
BRCA2 5-formylcytosin H3K27ac
5
Why Study DNA–Protein Interactions?
Understanding Gene Regulation
Genome-Wide Discovery
● Transcription factors bind specific DNA
● Identify where proteins bind across the
motifs (cis-regulatory elements)
entire genome
● These regulatory regions can control dozens
● Link transcription factors to their target
to hundreds of downstream genes
genes
● Protein binding determines whether genes
● Reveal interactions between regulatory
are activated or repressed
proteins
Relevance to Health & Disease
● Abnormal protein–DNA interactions
contribute to:
○ Cancer (e.g. breast cancer due to
estrogen receptor TF too active)
○ Developmental disorders
○ Immune and metabolic diseases 6
Gene Expression Regulation
Chromatin
Epigenetic control
Transcriptional control
Post-transcriptional control
Translational control
Post-translational control
ChIP-seq maps epigenetic and
transcriptional control genome-wide [Link]
you-are-10645 7
Transcription Factors
●Bind to specific DNA sequences near genes
○e.g. promoters, enhancers, silencers
●Recruit or block the transcription machinery
○e.g. RNA polymerase II, co-activators,
repressors
●Can activate or repress transcription
depending on context
Classic ChIP-seq targets → produce sharp, [Link]
narrow peaks
8
Chromatin proteins
●Histones package DNA into nucleosomes and
higher-order structures
●Chromatin-modifying enzymes add or remove
chemical marks
○HATs add acetyl groups (active); HDACs
remove them (silent)
○HMTs add methyl groups → can activate
or repress
●Chromatin remodelers open or close
chromatin access
Histone marks (e.g. H3K4me3, H3K27ac) are
ChIP-seq targets
→ produce broad chromatin domains [Link]
9
Regulatory mechanisms
Chromatin structure determines if a
gene is expressed or not
10
02
WHAT IS CHIP-SEQ
11
Video Animation Goes Here
12
Mapping ChIP-seq reads
Source: [Link] 13
Functional Workflow of ChIP-seq
Source: [Link] 14
WHAT QUESTIONS CAN CHIP-SEQ ANSWER?
● Where does a transcription factor bind across the genome?
○ e.g. CTCF binding sites, enhancer occupancy by NANOG
● Which genes are regulated by a specific protein or histone mark?
○ e.g. H3K4me3 at active promoters predicts transcribed genes
● What chromatin modifications mark a given genomic region?
○ e.g. H3K27me3 at Polycomb-repressed domain
● How does chromatin organization or protein-DNA binding change between conditions?
○ Healthy vs. diseased cells
○ Drug-treated vs. untreated cells
○ Different developmental stages
15
03
(PRE-)PROCESSING,
QUALITY CONTROL
& CHALLENGES
16
Noise Sources
Library Preparation Crosslink proteins to DNA Fragment (Sonication)
Protein specific antibody Immunoprecipitate Reverse crosslink and purify DNA
Source: [Link] 17
Noise Sources (2)
ChIP-Seq signal quality depends on:
● The number of active binding sites
● The number of starting genomes (number of cells)
○ ~2-10 million cells a good number
● Sequencing depth Globally
● IP efficiency (antibody quality, biological model used)
● GC rich content
○ Bias in fragment selection during amplification
● Open chromatin regions fragment more easily than closed regions
○ open region will generate more reads than closed one due to non-random
fragmentation
● Differential mappability of short reads to repeat-rich genomic regions
Locally
● Hyper-ChIPable regions
18
Read & Mapping Quality
● RAW reads quality control of FASTA / FASTQ files
○ Read quality (PHRED score, fragment sizes, fragment uniformity, …)
■ FastQC / MultiQC
○ Adapter removal, quality trimming, etc.
■ fastp / Cutadapt / Trimmomatic
● Alignment (BWA / Bowtie 2) & quality control (SAMtools / Qualimap / Picard):
○ Check control samples for even coverage over reference genome
■ Blacklist high / low coverage regions
■ Apply ENSEMBL blacklist, contains genomic regions with anomalous, unstructured, high signal or
read counts in NGS experiments, independent of cell type or experiment
○ Check for library contamination, sample mixup, etc.
○ Deduplicate reads
■ Remove copies of same fragment
○ Uniquely mapped reads
■ Among all reads generated, how many reads map to a unique location?
○ Uniquely mapped locations
■ Among the reads that map to a single location, how many different locations are there?
“Looking at a genomic positions, how many uniquely mapped reads are they supported by”
■ With little DNA may need to do many PCR cycles → overrepresentation of fragments
Source: [Link] 19
Library Quality Measures
Library Complexity (phantompeakqualtools):
● NRF (Non-Redundant Fraction)
○ How many distinct DNA fragments are present in your sequencing library
● PBC1 (ENCODE quality metric)
○ Fraction of occupied genomic positions that contain only a single read
○ High value means most reads come from unique fragments
● PBC2 (ENCODE quality metric)
○ Fraction of positions containing exactly one read
to positions containing exactly two reads
○ High value indicate fewer duplicated fragments
and therefore better library complexity
● FRiP score (Fraction of Reads in Peaks)
○ Higher values indicate stronger enrichment
○ Essentially signal-to-noise ratio
Sources: [Link] [Link] 20
04
PEAK CALLING &
ANALYSIS
21
Peak Shapes
ChIP-Seq for a transcription factor (Non-histone ChIP) ChIP-Seq for chromatin marks (Histone ChIP)
→ Protein-DNA-Interactions → Histone Modifications
→ Narrow peaks → Broad peaks
(Also for DNA-binding protein binding to many locations)
Source: [Link] 22
Signals in Practice
Transcription factors
Histone modifications
Enrichment over input
Source: [Link] 23
Peak Calling
● Peak calling for enriched regions
○ MACS3 (narrow & broad peaks) / epic2 (broad peaks) / SICER (broad peaks) / …
■ Returns a peak file with the fields:
chr start end length abs_summit pileup -log10(pvalue)
fold_enrichment -log10(qvalue) name
○ Peaks are relative to control
■ Reads do not follow Poisson or Normal distribution, some regions happen
to be sparser or more dense naturally → potential for false peaks
● Additional QC:
○ Biological replicates: Running an experiment and control in replicate allows for
checking of similar enrichment patterns in replicates (agreement >60% expected)
○ Expected binding motifs: Do binding motifs for top peaks look like expected for target protein?
○ What is the percentage of peaks in promoter / exon / intron / intergenic regions?
Sources: [Link] [Link] 24
Cross Correlation Analysis
● Reads are shifted in the direction of the strand they map to by an increasing number of base pairs
● Cross-correlation of strand-specific read positions (Watson & Cricks fragments) calculated
● Typically two peaks appear when cross-correlation is plotted against the shift value
○ Enrichment peak for the read length (“phantom” peak)
○ Enrichment peak for predominant fragment length (ChIP peak)
● Less dominant ChIP peak indicates poor quality ChIP-seq experiment
Sources: [Link] 25
Downstream Analysis & Normalization
The alignment and peak calling tools usually do not
require normalization of the data or handle it internally
as needed (mostly for scores to keep them comparable)
For downstream analysis, comparison between
samples, etc. normalization usually becomes necessary.
For simple analysis library-size normalization (e.g. using
DiffBind) to obtain CPM/RPM (counts / reads per million)
is sufficient.
More involved analysis may need more sophisticated
normalization approaches, e.g.:
● Differential binding
○ Statistical comparison
○ DiffBind, csaw, DESeq2, edgeR
● Global chromatin changes
○ Correct global shifts, e.g. true global
differences in TF binding between two
samples
○ ChIPseqSpike
Source: [Link] 26
05
APPLICATIONS &
RELATED METHODS
27
METHOD WHAT IT MAPS KEY DIFFERENCE FROM CHIP-SEQ INPUT NEEDED
No cross-linking, antibody-tethered
CUT&RUN TFs & Histone marks Low (≥10k cells)
MNase cuts in situ -> low background
Tagmentation in situ replaces
CUT&Tag TFs & Histone marks sonication -> very fast, single-cell Very low (single cell)
compatible
Open chromatin /
ATAC-Seq No antibodies, maps accessible regions Very low (500 cells)
accesibility
genome-wide regardless of protein
Exonuclease trims reads to ca. 1 bp of
ChIP-exo TF binding sites High
binding site -> near base-pair resolution
3D chromatin Combines ChIP enrichment with Hi-C
Moderate
HiChIP contacts + histone proximity ligation -> enhancer-
marks promotor loops
28
Additional Usecases
Predict transcription factor function
●Overlap peaks with active/repressive histone marks to infer whether a TF activates or represses
●Co-occurrence of peaks across multiple ChIP-seq datasets reveals co-regulatory complexes
Identify DNA binding motifs
●Scan peak sequences with HOMER or MEME to find DNA words a protein directly recognises
●Secondary motifs point to co-factors recruited alongside the primary protein
Link peaks to target genes & pathways
●Annotate peaks to nearest genes with ChIPseeker, then run GO/KEGG enrichment to map biology
●Combine with RNA-seq: peaks near differentially expressed genes identify direct regulatory
targets
29
Diseases Linked to Gene Regulation Defects
Cancer Developmental disorders Immune and metabolic diseases
●Mutated transcription ●Mutations in chromatin ●GWAS variants in enhancers
factors drive uncontrolled remodellers cause intellectual linked to type 2 diabetes and
proliferation (e.g. MYC, TP53) disability (e.g. ARID1B in autoimmune conditions
Coffin-Siris)
●Altered histone marks ●Aberrant NF-κB binding
silence tumour suppressors ●Disrupted HOX gene underlies chronic inflammation
(e.g. H3K27me3 at CDKN2A) regulation causes limb and
organ malformations
●ERα cistrome
redistribution underlies
tamoxifen resistance in
breast cancer
30
Tutorial
Please switch to your computers and open the prepared R notebook
31
Sources
● Given by course:
○ [Link]
○ [Link]
○ [Link]
○ [Link]
● Our additional sources:
○ [Link]
○ [Link]
○ [Link]
○ [Link]
○ [Link]
○ [Link]
○ [Link]
32