0% found this document useful (0 votes)
8 views22 pages

E. coli Gene Expression Analysis with JactiveModules

The document outlines a practical session for network-based data analysis using the JactiveModules tool in Cytoscape, focusing on a dataset from Gene Expression Omnibus (GSE56133) that examines the effect of ampicillin on E. coli gene expression. It details the preprocessing of the dataset, the importation of differentially expressed genes, and the visualization of results, including node attributes and network styling. The session concludes with instructions for refining the analysis by adjusting p-values and performing Gene Ontology enrichment on selected networks.

Uploaded by

ali.mu9656
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views22 pages

E. coli Gene Expression Analysis with JactiveModules

The document outlines a practical session for network-based data analysis using the JactiveModules tool in Cytoscape, focusing on a dataset from Gene Expression Omnibus (GSE56133) that examines the effect of ampicillin on E. coli gene expression. It details the preprocessing of the dataset, the importation of differentially expressed genes, and the visualization of results, including node attributes and network styling. The session concludes with instructions for refining the analysis by adjusting p-values and performing Gene Ontology enrichment on selected networks.

Uploaded by

ali.mu9656
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

K.

Marchal JactiveModules practical session

NETWORK-BASED DATA ANALYSIS

INTRODUCTION

The described method can be imported in cytoscape.

DATASET

We will perform an example analysis on a publicly available data set (Gene Expression Omnibus, GSE56133)
measuring the effect of ampicillin on the expression behavior in E. coli.

FIND IN GENE EXPRESSION OMNIBUS THE DATASET. WHAT IS GENE OMNIBUS?

[Link]

“GEO is a public functional genomics data repository supporting MIAME-compliant data submissions. Array- and
sequence-based data are accepted. Tools are provided to help users query and download experiments and curated
gene expression profiles.”

1
K. Marchal JactiveModules practical session

Download the original article that describes the dataset.

Try to find what the dataset is about.

2
K. Marchal JactiveModules practical session

To be used as input for Jactivemodules/Phenetic the dataset was preprocessed as described in the original paper
(Dwyer et al. 2014) using standard microarray analysis procedures. All values were converted into differential gene
expression values. Each value in the processed dataset thus contains the average log ratio Test/control of the three
replicate experiments where the control is the condition with ampicillin added and the reference is the condition
without ampicillin.

“Samples for total RNA collection were taken immediately before treatment (time 0) and at 1 h posttreatment.
cDNA preparation and hybridization to Affymetrix GeneChip E. coli Genome 2.0 microarrays were performed as
previously described (7). CEL files for the resulting expression profiles were background adjusted and normalized
using RobustMultiarray Averaging (108). Statistical significance was computed using Welch’s t test. Triplicate
(technical replicate) measurements from treated MG1655 samples were compared with the triplicate (technical
replicate) untreated samples. For each set of comparisons, P values were corrected for falsediscovery rate (FDR)
(109)”

Data per experiment

#ID_REF =
#VALUE = Values in the data table represent RMA-normalized log2 expression
ID_REF VALUE
1759068_at 4.995971
1759069_at 11.155949
1759070_s_at 8.459495

3
K. Marchal JactiveModules practical session

1759071_s_at 7.511163
1759072_s_at 4.913314
1759073_at 6.765406
1759074_at 6.30017
1759075_at 8.648702
1759076_s_at 6.478148
1759077_s_at 3.15568
1759078_at 3.248215
1759079_at 5.906231
1759080_s_at 5.230258
1759081_s_at 6.274671
1759082_s_at 7.989912
1759083_at 11.506391
1759084_s_at 9.213291

INPUT DATA

E. coli interaction network

For all genes differentially expressed p-values

LOAD THE E. COLI NETWORK

4
K. Marchal JactiveModules practical session

5
K. Marchal JactiveModules practical session

6
K. Marchal JactiveModules practical session

IMPORT DIFFERENTIALLY EXPRESSED GENES

For all genes the level of differential expression is given and the p values.

These values are imported as node attributes and are visible in the table panel.

7
K. Marchal JactiveModules practical session

Import the gene names (synonyms of the b numbers) in the same way as node attributes.

The synonyms are added to the Table pane.

8
K. Marchal JactiveModules practical session

ADAPT THE STYLE OF THE NETWORK

9
K. Marchal JactiveModules practical session

Use the common names instead of b numbers to label the nodes.

Adapt the color of the nodes according to their level of differential expression.

10
K. Marchal JactiveModules practical session

Double click

Add a marker. Put the minimal value on green and the maximal on red. Zero levels (no differential expression) will
be white.

11
K. Marchal JactiveModules practical session

Use the p values to determine the size of the nodes. Zero values largest nodes and one values smallest nodes.

12
K. Marchal JactiveModules practical session

Use the edge attributes table to color the edges

13
K. Marchal JactiveModules practical session

RUN JACTIVEMODULES

Using the p values as input.

14
K. Marchal JactiveModules practical session

15
K. Marchal JactiveModules practical session

5 modules were detected. Put these subnetworks in the style we just created (these were obtained with the
greedy search)

16
K. Marchal JactiveModules practical session

Rather than the annealing option there is also the search option:
Instead of annealing, this will greedily search for local optima in the graph, starting from all nodes in the graph.
The depth specifies the number of moves required to consider an area a local maximum. The greater the depth,
the better the optima, and the longer the algorithm takes. Typical depths are 1 or 2. Searching (rather than
annealing) will return the set of all optima it found, in the order that it found them. They may overlap.

17
K. Marchal JactiveModules practical session

View module 5

Most selected nodes have low p values but relatively few are truly differentially expressed.

To zoom in and out use the network view .

To save the network figure

18
K. Marchal JactiveModules practical session

19
K. Marchal JactiveModules practical session

In general, Jactive modules selected large modules of which the nodes in general have a low p value (large nodes)
but many nodes that are barely differentially expressed (this indicates that the pvalue alone is not sufficient to
select the most important genes). How would you improve the result? We can tweek the input file of Jactive
modules. We need to have for every gene on the network a p-value. However if we put the p values for the genes
we are not interested in (the ones that should not end up in a module at a very high level (e.g. 1) the tool will use
less of these genes in its module. In a differential expression analysis often the genes with the lowest p values are
barely differentially expressed. We want to exclude these genes from our modules. So we will put the p values for
the genes that have a low fold change artificially at 1 and rerun the analysis.

data = [Link]("[Link]", sep = "\t", header = FALSE)


select_1 =which(abs(data[,2])<1)
data_select=data
data_select[select_1,3]=1
[Link](data_select,'[Link]',sep="\t", quote=FALSE, append=FALSE, [Link]=FALSE,
[Link]=FALSE)

20
K. Marchal JactiveModules practical session

Now you will notice that the networks become smaller and more enriched in genes that are significantly
differentially expressed and that exhibit a high log fold change.
Select one network and perform GO enrichment. Hereto you copy the b numbers of the genes in the network and
paste the in an online pathway enrichment tool e.g. [Link] Use common gene
names and match with Escherichia coli K12.

21
K. Marchal JactiveModules practical session

Compare the results also with those obtained with phenetic (CUT OFF ON GENE LIST P VALUE 0.05, FOLD CHANGE
1 AND EDGE COST 0.05, lower than default so you obtain a comparable number of genes in the network as
obtained with Jactive modules. .

22

Common questions

Powered by AI

Using p-values as the sole criterion for selecting differentially expressed genes can lead to the inclusion of genes with low fold changes, which might be statistically significant but not biologically meaningful. As JactiveModules analysis shows, genes with low p-values are often barely differentially expressed, which can skew module detection towards less relevant genes. Adjusting the p-values artificially for low fold change genes can mitigate this issue, ensuring that selected modules are enriched with truly relevant differentially expressed genes .

JactiveModules is used for detecting active modules within a network, based on differential expression data and statistical significance (p-values). Adjusting p-values can affect the selection of genes included in the modules by excluding genes with low fold changes that may not be biologically significant despite their statistical significance. By setting p-values for these genes artificially high, the algorithm reduces their inclusion in the modules, leading to a network more enriched with truly differentially expressed genes showing a high log fold change .

JactiveModules uses a greedy search strategy to detect active modules based on differential expression data and p-values, optimizing for local optima within a network. The depth of search can be adjusted to improve module detection. Phenetic, on the other hand, uses a cutoff-based approach with specified thresholds for p-values, fold changes, and edge cost. While JactiveModules uses an iterative approach for module expansion, Phenetic filters based on predefined criteria, enabling comparison of results under similar conditions but with potentially different emphases on gene inclusion based on statistical parameters .

Manipulating input data, particularly by artificially modifying the p-values of genes with low fold changes, affects both the size and composition of detected modules in JactiveModules. By raising the p-values of less interesting genes, the module detection algorithm is steered away from these genes, resulting in smaller, more focused modules enriched with significantly differentially expressed genes. This refined selection process ensures the biological relevance of detected modules .

To improve the detection of biologically meaningful modules, one can adjust the input data by setting high p-values for genes with low fold changes to exclude them, ensuring the inclusion of more biologically relevant genes. Adjusting the depth of search in module detection algorithms, such as those used in JactiveModules, can also enhance the capture of local optima. Moreover, integrating multi-omic datasets and employing cross-validation with known biological pathways can further refine module detection and ensure biological significance .

Performing GO enrichment analysis on a module helps identify the biological processes, cellular components, and molecular functions significantly associated with the genes in the network. It provides insights into the biological significance underlying the observed gene expression changes and can reveal the functional implications of certain genes and interactions within the network, thereby enhancing the biological interpretation of the data .

Gene expression data from microarray experiments is prepared by converting the raw expression values into differential expression values as average log ratios between test and control conditions. Robust Multi-array Averaging is used for background adjustment and normalization. Statistical significance tests, such as Welch’s t-test, are applied, and FDR correction is used to adjust p-values, enhancing the reliability of differential expression analysis .

Data for JactiveModules is preprocessed by converting gene expression levels into differential expression values, specifically the average log ratio of test/control conditions across replicates. The dataset from Dwyer et al. (2014) was prepared with standard microarray analysis procedures, including background adjustment and normalization using Robust Multi-array Averaging. Statistical significance was assessed through Welch’s t-test, and p-values were corrected for false discovery rate .

In network visualization, nodes can represent genes with colors indicating differential expression levels and sizes reflecting p-value significance, whereas edges can be colored based on interaction attributes. For instance, a node size could be largest for zero p-values and smallest for one p-values, emphasizing significant genes. Colors ranging from green (low expression) to red (high expression) provide intuitive visual cues that mirror underlying biological data, enhancing interpretability .

The visualization style of a network in Cytoscape can be adapted by labeling nodes with common gene names instead of identifiers, coloring nodes according to the level of differential expression (e.g., green for minimal values and red for maximal), and sizing nodes based on statistical significance (p-values) with zero indicating largest nodes and one smallest. Edges can also be colored based on attributes. This helps visually emphasize biologically relevant patterns in the data .

You might also like