0% found this document useful (0 votes)
0 views32 pages

ChIPseq_presentation

The document is a presentation on ChIP-Seq data analysis, detailing its methodology, applications, and challenges. It explains the significance of ChIP-Seq in understanding gene regulation and protein-DNA interactions, as well as the technical aspects of data processing, peak calling, and quality control. Additionally, it discusses related methods and the implications of gene regulation defects in various diseases.

Uploaded by

arberlin0700288
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
0 views32 pages

ChIPseq_presentation

The document is a presentation on ChIP-Seq data analysis, detailing its methodology, applications, and challenges. It explains the significance of ChIP-Seq in understanding gene regulation and protein-DNA interactions, as well as the technical aspects of data processing, peak calling, and quality control. Additionally, it discusses related methods and the implications of gene regulation defects in various diseases.

Uploaded by

arberlin0700288
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

ChIP-Seq Data Analysis

Presentation by Group 7:
Maksim Danilchyk, Sofya Shorzhina,
Tilo Alves Radtke, Kristian Reinhart
01 02
INTRODUCTION & WHAT IS CHIP-
BIOLOGY SEQ

03 04 05
(PRE-)PROCESSING, PEAK CALLING & APPLICATIONS &
QUALITY CONTROL ANALYSIS RELATED
& CHALLENGES METHODS
2
01
INTRODUCTION &
BIOLOGY
3
WHAT IT
METHOD RESOLUTION OUTPUT BIOLOGICAL QUESTION
MEASURES

Gene expression Bulk cell Gene counts /


RNA-Seq Which genes are expressed?
levels population transcripts

Cell-specific
Gene expression Which cell types are present and
scRNA-Seq Single cell expression
levels what are they expressing?
profiles

Spatial gene
Spatial Gene expression Tissue region What are genes expressed within
expression
Transcriptomics leves + location / single cells tissues / regions?
maps

Single base /
DNA DNA methylation Methylation Which genomic regions are
genomic
Methylation levels percentages epigenetically silenced or active?
regions

Protein-DNA Where do transcription factors,


Binding sites Binding sites
interactions, DNA-binding proteins or
ChIP-Seq (~100-1000 across the
histone histone markers bind?
bp) genome
modifications
4
WHAT IS ChIP-seq?
● A technique combining two methods:
○ ChIP (Chromatin Immunoprecipitation) — uses an antibody to enrich DNA fragments
bound to a protein of interest
○ NGS (Next-Generation Sequencing) — reads out the captured DNA fragments genome-wide
● Together they answer: Where does a protein bind across the entire genome?

ChIP-seq Targets

Transcription
Transcription DNA-binding DNA Histone
CATEGORY factors & Chromatin remodelers
machinery proteins modifications marks
repressors

BRCA1, 5-methylcytosine H3K4me3,


EXAMPLE NANOG, CTCF RNA Pol II BMI1, EZH2
BRCA2 5-formylcytosin H3K27ac

5
Why Study DNA–Protein Interactions?

Understanding Gene Regulation


Genome-Wide Discovery
● Transcription factors bind specific DNA
● Identify where proteins bind across the
motifs (cis-regulatory elements)
entire genome
● These regulatory regions can control dozens
● Link transcription factors to their target
to hundreds of downstream genes
genes
● Protein binding determines whether genes
● Reveal interactions between regulatory
are activated or repressed
proteins
Relevance to Health & Disease

● Abnormal protein–DNA interactions


contribute to:
○ Cancer (e.g. breast cancer due to
estrogen receptor TF too active)
○ Developmental disorders
○ Immune and metabolic diseases 6
Gene Expression Regulation
Chromatin

Epigenetic control

Transcriptional control

Post-transcriptional control

Translational control

Post-translational control

ChIP-seq maps epigenetic and


transcriptional control genome-wide [Link]
you-are-10645 7
Transcription Factors

●Bind to specific DNA sequences near genes


○e.g. promoters, enhancers, silencers
●Recruit or block the transcription machinery
○e.g. RNA polymerase II, co-activators,
repressors
●Can activate or repress transcription
depending on context
Classic ChIP-seq targets → produce sharp, [Link]

narrow peaks

8
Chromatin proteins
●Histones package DNA into nucleosomes and
higher-order structures
●Chromatin-modifying enzymes add or remove
chemical marks
○HATs add acetyl groups (active); HDACs
remove them (silent)
○HMTs add methyl groups → can activate
or repress
●Chromatin remodelers open or close
chromatin access
Histone marks (e.g. H3K4me3, H3K27ac) are
ChIP-seq targets
→ produce broad chromatin domains [Link]

9
Regulatory mechanisms
Chromatin structure determines if a
gene is expressed or not

10
02
WHAT IS CHIP-SEQ

11
Video Animation Goes Here

12
Mapping ChIP-seq reads

Source: [Link] 13
Functional Workflow of ChIP-seq

Source: [Link] 14
WHAT QUESTIONS CAN CHIP-SEQ ANSWER?
● Where does a transcription factor bind across the genome?
○ e.g. CTCF binding sites, enhancer occupancy by NANOG
● Which genes are regulated by a specific protein or histone mark?
○ e.g. H3K4me3 at active promoters predicts transcribed genes
● What chromatin modifications mark a given genomic region?
○ e.g. H3K27me3 at Polycomb-repressed domain
● How does chromatin organization or protein-DNA binding change between conditions?
○ Healthy vs. diseased cells
○ Drug-treated vs. untreated cells
○ Different developmental stages

15
03
(PRE-)PROCESSING,
QUALITY CONTROL
& CHALLENGES
16
Noise Sources

Library Preparation Crosslink proteins to DNA Fragment (Sonication)

Protein specific antibody Immunoprecipitate Reverse crosslink and purify DNA

Source: [Link] 17
Noise Sources (2)
ChIP-Seq signal quality depends on:
● The number of active binding sites
● The number of starting genomes (number of cells)
○ ~2-10 million cells a good number
● Sequencing depth Globally
● IP efficiency (antibody quality, biological model used)
● GC rich content
○ Bias in fragment selection during amplification

● Open chromatin regions fragment more easily than closed regions


○ open region will generate more reads than closed one due to non-random
fragmentation
● Differential mappability of short reads to repeat-rich genomic regions
Locally
● Hyper-ChIPable regions

18
Read & Mapping Quality
● RAW reads quality control of FASTA / FASTQ files
○ Read quality (PHRED score, fragment sizes, fragment uniformity, …)
■ FastQC / MultiQC
○ Adapter removal, quality trimming, etc.
■ fastp / Cutadapt / Trimmomatic

● Alignment (BWA / Bowtie 2) & quality control (SAMtools / Qualimap / Picard):


○ Check control samples for even coverage over reference genome
■ Blacklist high / low coverage regions
■ Apply ENSEMBL blacklist, contains genomic regions with anomalous, unstructured, high signal or
read counts in NGS experiments, independent of cell type or experiment
○ Check for library contamination, sample mixup, etc.
○ Deduplicate reads
■ Remove copies of same fragment
○ Uniquely mapped reads
■ Among all reads generated, how many reads map to a unique location?
○ Uniquely mapped locations
■ Among the reads that map to a single location, how many different locations are there?
“Looking at a genomic positions, how many uniquely mapped reads are they supported by”
■ With little DNA may need to do many PCR cycles → overrepresentation of fragments

Source: [Link] 19
Library Quality Measures
Library Complexity (phantompeakqualtools):
● NRF (Non-Redundant Fraction)
○ How many distinct DNA fragments are present in your sequencing library
● PBC1 (ENCODE quality metric)
○ Fraction of occupied genomic positions that contain only a single read
○ High value means most reads come from unique fragments
● PBC2 (ENCODE quality metric)
○ Fraction of positions containing exactly one read
to positions containing exactly two reads
○ High value indicate fewer duplicated fragments
and therefore better library complexity
● FRiP score (Fraction of Reads in Peaks)
○ Higher values indicate stronger enrichment
○ Essentially signal-to-noise ratio

Sources: [Link] [Link] 20


04
PEAK CALLING &
ANALYSIS
21
Peak Shapes

ChIP-Seq for a transcription factor (Non-histone ChIP) ChIP-Seq for chromatin marks (Histone ChIP)
→ Protein-DNA-Interactions → Histone Modifications
→ Narrow peaks → Broad peaks
(Also for DNA-binding protein binding to many locations)

Source: [Link] 22
Signals in Practice

Transcription factors

Histone modifications

Enrichment over input

Source: [Link] 23
Peak Calling

● Peak calling for enriched regions


○ MACS3 (narrow & broad peaks) / epic2 (broad peaks) / SICER (broad peaks) / …
■ Returns a peak file with the fields:
chr start end length abs_summit pileup -log10(pvalue)
fold_enrichment -log10(qvalue) name
○ Peaks are relative to control
■ Reads do not follow Poisson or Normal distribution, some regions happen
to be sparser or more dense naturally → potential for false peaks
● Additional QC:
○ Biological replicates: Running an experiment and control in replicate allows for
checking of similar enrichment patterns in replicates (agreement >60% expected)
○ Expected binding motifs: Do binding motifs for top peaks look like expected for target protein?
○ What is the percentage of peaks in promoter / exon / intron / intergenic regions?
Sources: [Link] [Link] 24
Cross Correlation Analysis

● Reads are shifted in the direction of the strand they map to by an increasing number of base pairs
● Cross-correlation of strand-specific read positions (Watson & Cricks fragments) calculated
● Typically two peaks appear when cross-correlation is plotted against the shift value
○ Enrichment peak for the read length (“phantom” peak)
○ Enrichment peak for predominant fragment length (ChIP peak)
● Less dominant ChIP peak indicates poor quality ChIP-seq experiment

Sources: [Link] 25
Downstream Analysis & Normalization
The alignment and peak calling tools usually do not
require normalization of the data or handle it internally
as needed (mostly for scores to keep them comparable)

For downstream analysis, comparison between


samples, etc. normalization usually becomes necessary.

For simple analysis library-size normalization (e.g. using


DiffBind) to obtain CPM/RPM (counts / reads per million)
is sufficient.

More involved analysis may need more sophisticated


normalization approaches, e.g.:
● Differential binding
○ Statistical comparison
○ DiffBind, csaw, DESeq2, edgeR
● Global chromatin changes
○ Correct global shifts, e.g. true global
differences in TF binding between two
samples
○ ChIPseqSpike
Source: [Link] 26
05
APPLICATIONS &
RELATED METHODS
27
METHOD WHAT IT MAPS KEY DIFFERENCE FROM CHIP-SEQ INPUT NEEDED

No cross-linking, antibody-tethered
CUT&RUN TFs & Histone marks Low (≥10k cells)
MNase cuts in situ -> low background

Tagmentation in situ replaces


CUT&Tag TFs & Histone marks sonication -> very fast, single-cell Very low (single cell)
compatible

Open chromatin /
ATAC-Seq No antibodies, maps accessible regions Very low (500 cells)
accesibility
genome-wide regardless of protein

Exonuclease trims reads to ca. 1 bp of


ChIP-exo TF binding sites High
binding site -> near base-pair resolution

3D chromatin Combines ChIP enrichment with Hi-C


Moderate
HiChIP contacts + histone proximity ligation -> enhancer-
marks promotor loops

28
Additional Usecases
Predict transcription factor function
●Overlap peaks with active/repressive histone marks to infer whether a TF activates or represses
●Co-occurrence of peaks across multiple ChIP-seq datasets reveals co-regulatory complexes

Identify DNA binding motifs


●Scan peak sequences with HOMER or MEME to find DNA words a protein directly recognises
●Secondary motifs point to co-factors recruited alongside the primary protein

Link peaks to target genes & pathways


●Annotate peaks to nearest genes with ChIPseeker, then run GO/KEGG enrichment to map biology
●Combine with RNA-seq: peaks near differentially expressed genes identify direct regulatory
targets

29
Diseases Linked to Gene Regulation Defects
Cancer Developmental disorders Immune and metabolic diseases

●Mutated transcription ●Mutations in chromatin ●GWAS variants in enhancers


factors drive uncontrolled remodellers cause intellectual linked to type 2 diabetes and
proliferation (e.g. MYC, TP53) disability (e.g. ARID1B in autoimmune conditions
Coffin-Siris)
●Altered histone marks ●Aberrant NF-κB binding
silence tumour suppressors ●Disrupted HOX gene underlies chronic inflammation
(e.g. H3K27me3 at CDKN2A) regulation causes limb and
organ malformations
●ERα cistrome
redistribution underlies
tamoxifen resistance in
breast cancer

30
Tutorial

Please switch to your computers and open the prepared R notebook


31
Sources
● Given by course:
○ [Link]
○ [Link]
○ [Link]
○ [Link]

● Our additional sources:


○ [Link]
○ [Link]
○ [Link]
○ [Link]
○ [Link]
○ [Link]
○ [Link]

32

You might also like