Introduction to Dot Plots of Genetic
Sequences
Table of Contents
1. Introduction
2. What is a Dot Plot?
3. Basic Principles
4. Construction of Dot Plots
5. Interpreting Dot Plots
6. Types of Patterns
7. Filtering and Noise Reduction
8. Applications
9. Practice Examples
10. Software Tools
11. Summary
Introduction
Dot plots are fundamental visualization tools in bioinformatics used to compare genetic
sequences. They provide an intuitive graphical method to identify similarities, differences,
and structural features between DNA, RNA, or protein sequences. Originally developed by
Gibbs and McIntyre in 1970, dot plots remain one of the most straightforward approaches to
sequence analysis.
What is a Dot Plot?
A dot plot is a two-dimensional matrix where:
• X-axis: Represents one sequence (horizontal)
• Y-axis: Represents another sequence (vertical)
• Dots/Points: Placed at positions (i,j) where nucleotides or amino acids match
The resulting pattern reveals evolutionary relationships, structural similarities, and functional
domains between sequences.
Basic Principles
Simple Dot Plot Algorithm:
1. Create a matrix with sequence A on one axis and sequence B on the other
2. For each position (i,j), place a dot if A[i] = B[j]
3. Analyze the resulting pattern
Matrix Representation:
A T G C A T (Sequence 1)
+--+--+--+--+--+--+
A | •| | | | •| |
T | | •| | | | •|
C | | | | •| | |
G | | | •| | | |
A | •| | | | •| |
+--+--+--+--+--+--+
(Sequence 2)
Construction of Dot Plots
Step-by-Step Process:
1. Sequence Alignment: Place sequences on perpendicular axes
2. Comparison: Compare each nucleotide/amino acid pair
3. Dot Placement: Mark matches with dots
4. Pattern Analysis: Examine diagonal lines and patterns
Mathematical Representation:
For sequences A and B of lengths m and n:
• Matrix M[i][j] = 1 if A[i] = B[j], otherwise 0
• Visualize as dots where M[i][j] = 1
Interpreting Dot Plots
Key Features to Look For:
1. Diagonal Lines
• Main Diagonal: Perfect similarity (identical sequences)
• Parallel Diagonals: Similar regions or repeats
• Broken Diagonals: Gaps or mismatches
2. Pattern Orientations
• Positive Diagonal (/): Direct similarity
• Negative Diagonal (): Inverted similarity (reverse complement)
3. Density Patterns
• Dense regions: High similarity
• Sparse regions: Low similarity
• Empty regions: No similarity
Types of Patterns
1. Perfect Match
A T G C
A • . . .
T . • . .
G . . • .
C . . . •
Single diagonal line indicates identical sequences
2. Tandem Repeats
A T A T G C
A • . • . . .
T . • . • . .
A • . • . . .
T . • . • . .
G . . . . • .
C . . . . . •
Multiple parallel diagonals indicate repeated elements
3. Inverted Repeats
A T G C A T
A • . . . • .
T . • . . . •
G . . • . . .
C . . . • . .
A • . . . • .
T . • . . . •
Perpendicular diagonals indicate palindromic sequences
4. Insertions/Deletions
A T - G C
A • . . . .
T . • . . .
G . . . • .
C . . . . •
Broken diagonals indicate gaps
Filtering and Noise Reduction
Problems with Simple Dot Plots:
• Random matches: Create noise
• Low complexity regions: Produce dense, uninformative patterns
• Short matches: May not be biologically significant
Filtering Strategies:
1. Window Size Filtering
• Use sliding windows of size w
• Require minimum number of matches within window
• Reduces random noise
2. Scoring Matrices
• Apply substitution matrices (e.g., BLOSUM, PAM)
• Weight matches by biological significance
• Threshold-based filtering
3. Statistical Filtering
• Calculate expected number of random matches
• Filter based on statistical significance
• Use Z-scores or P-values
Applications
1. Sequence Similarity Detection
• Homology identification
• Evolutionary relationships
• Functional domain detection
2. Structural Analysis
• Repeat identification
• Palindrome detection
• Tandem repeat analysis
3. Genome Analysis
• Synteny mapping
• Chromosomal rearrangements
• Duplication events
4. Quality Control
• Sequence validation
• Contamination detection
• Assembly verification
Practice Examples
Example 1: Basic Interpretation
Sequences:
• Sequence A: ATGCATGC
• Sequence B: ATGCATGC
Expected Pattern: Perfect diagonal line from top-left to bottom-right
Biological Interpretation: Identical sequences, possibly from same gene or highly
conserved region
Example 2: Tandem Repeat Detection
Sequences:
• Sequence A: ATGCATGCATGC
• Sequence B: ATGC
Expected Pattern: Multiple short diagonal lines
Biological Interpretation: Sequence A contains tandem repeats of sequence B
Example 3: Inverted Repeat Analysis
Sequences:
• Sequence A: ATGCGCAT
• Sequence B: ATGCGCAT
Expected Pattern: Main diagonal plus perpendicular diagonal
Biological Interpretation: Palindromic sequence, possible hairpin structure
Example 4: Insertion/Deletion Detection
Sequences:
• Sequence A: ATGCATGC
• Sequence B: ATGC--GC
Expected Pattern: Broken diagonal with gap
Biological Interpretation: 2-nucleotide deletion in sequence B
Practice Exercise:
Given sequences:
• Seq1: ACGTACGTACGT
• Seq2: ACGTACGT
Questions:
1. What pattern would you expect?
2. What does this pattern indicate biologically?
3. How would you filter noise in this comparison?
Answers:
1. Multiple parallel diagonal lines
2. Tandem repeats of ACGT motif
3. Use window size of 4 nucleotides matching the repeat unit
Software Tools
Command-Line Tools:
• EMBOSS dottup/dotmatcher: Basic dot plot generation
• MUMmer: Genome-scale comparisons
• BLAST: Sequence similarity with graphical output
Graphical Tools:
• Gepard: Java-based dot plot generator
• JDotter: Interactive dot plot analysis
• FlexiDot: Flexible dot plot visualization
Web-Based Tools:
• NCBI BLAST: Online sequence comparison
• EBI Tools: Various dot plot utilities
• Galaxy: Workflow-based analysis
Summary
Key Takeaways:
1. Dot plots are intuitive: Visual representation of sequence similarities
2. Pattern recognition is crucial: Different patterns indicate different biological
features
3. Filtering is essential: Raw dot plots contain significant noise
4. Multiple applications: From simple comparisons to genome analysis
5. Tool selection matters: Choose appropriate software for your specific needs
Best Practices:
• Always consider biological context
• Use appropriate filtering parameters
• Validate findings with other methods
• Consider sequence length and complexity
• Document analysis parameters
Common Pitfalls:
• Over-interpreting random matches
• Ignoring statistical significance
• Inadequate filtering
• Misunderstanding pattern orientation
• Neglecting sequence quality
Further Reading
1. Gibbs, A.J. & McIntyre, G.A. (1970). The diagram, a method for comparing
sequences. European Journal of Biochemistry, 16(1), 1-11.
2. Maizel, J.V. & Lenk, R.P. (1981). Enhanced graphic matrix analysis of nucleic acid
and protein sequences. Proceedings of the National Academy of Sciences, 78(12),
7665-7669.
3. Sonnhammer, E.L. & Durbin, R. (1995). A dot-matrix program with dynamic
threshold control suited for genomic DNA and protein sequence analysis. Gene,
167(1-2), GC1-GC10.