Reference 3
Reference 3
DOI 10.1186/s12920-017-0293-y
Abstract
Background: Differential gene expression is important to understand the biological differences between healthy
and diseased states. Two common sources of differential gene expression data are microarray studies and the
biomedical literature.
Methods: With the aid of text mining and gene expression analysis we have examined the comparative properties
of these two sources of differential gene expression data.
Results: The literature shows a preference for reporting genes associated to higher fold changes in microarray data,
rather than genes that are simply significantly differentially expressed. Thus, the resemblance between the literature
and microarray data increases when the fold-change threshold for microarray data is increased. Moreover, the
literature has a reporting preference for differentially expressed genes that (1) are overexpressed rather than
underexpressed; (2) are overexpressed in multiple diseases; and (3) are popular in the biomedical literature at large.
Additionally, the degree to which diseases are similar depends on whether microarray data or the literature is used
to compare them. Finally, vaguely-qualified reports of differential expression magnitudes in the literature have only
small correlation with microarray fold-change data.
Conclusions: Reporting biases of differential gene expression in the literature can be affecting our appreciation of
disease biology and of the degree of similarity that actually exists between different diseases.
Background Our goal in this study was to compare two widely used
Investigating the differences between diseased and sources of DEG information, namely high-throughput
healthy state helps us understand the pathology of dis- microarray expression studies and the scientific litera-
eases and, eventually, treat them. One particular focus of ture. For that purpose, we mined the scientific literature
investigation is differentially-expressed genes (DEGs), and analyzed microarray datasets on a set of diseases to
which involves the identification of genes that are differ- study the similarities and differences of these two types
entially expressed in disease. In pharmaceutical and clin- of data within specific biological contexts.
ical research, DEGs can be valuable to pinpoint In the scientific literature, information about DEGs is
candidate biomarkers, therapeutic targets and gene sig- largely found in unstructured form and scattered across
natures for diagnostics. While particular gene expression publications. It can appear in the form of gradable state-
changes may not always translate into consequential bio- ments, which are statements that describe a measure-
logical activity, such data can nonetheless be pooled with ment with respect to a baseline, scale or norm [3]. For
other biological data in a high-throughput fashion to example, the sentence “The expression of protectin was
create integrated analyses, such as building the target found to be decreased in the epithelium of patients with
landscape of a disease [1, 2]. ulcerative colitis.” [4] compares the pathological expres-
sion of protectin to an implicit baseline, presumably the
* Correspondence: [Link]-esteban@[Link] expression level of protectin in healthy state. Such a sen-
1
Roche Pharmaceutical Research and Early Development, Roche Innovation tence describes a “negative regulation of gene expres-
Center Basel, Grenzacherstrasse 124, 4070 Basel, Switzerland
Full list of author information is available at the end of the article sion” as defined by the Gene Regulation Ontology [5].
© The Author(s). 2017 Open Access This article is distributed under the terms of the Creative Commons Attribution 4.0
International License ([Link] which permits unrestricted use, distribution, and
reproduction in any medium, provided you give appropriate credit to the original author(s) and the source, provide a link to
the Creative Commons license, and indicate if changes were made. The Creative Commons Public Domain Dedication waiver
([Link] applies to the data made available in this article, unless otherwise stated.
Rodriguez-Esteban and Jiang BMC Medical Genomics (2017) 10:59 Page 2 of 10
DEG information can also be found in non-gradable datasets not stored in standard repositories can be
statements in which a comparison is implicit. For hard to obtain.
example, in the sentence “Expression of the COX-2 The advent of gene expression measurement with
enzyme has been reported in animal models of inflam- RNA sequencing (RNA-seq) technology has affected the
matory bowel disease (IBD) as well as in patients af- number of microarray studies being undertaken. How-
fected by ulcerative colitis and Crohn's disease.” [6] it is ever, in 2016, GEO still released 4945 array expression
implied that there is lack of expression in wild-type ani- profiling series (“expression profiling by array”[DataSet
mals and healthy patient tissue. Type] AND “gse”[Entry Type]), or about the same
Statements about DEGs in the literature often lack de- quantity released for high-throughput sequencing series
tail or specificity, which is a challenge for human inter- (“expression profiling by high throughput sequencing,”
pretation and for their automatic extraction by n = 4894). Moreover, a large trove of microarray studies
computers. Thus, they can refer indistinctly to protein has been accumulating in GEO over time, with 49,026
or RNA [7], and use baselines that are not defined or array series available (search performed on 2017-2-14).
vague. Vagueness is a general feature of natural language While RNA-seq is increasingly favored for high-
and is a special problem with gradable statements [8]. throughput expression analysis, modern microarray and
For example, in the sentence “Involucrin [...] is markedly RNA-seq platforms produce expression values that are
increased in inflammatory skin diseases such as psoria- highly correlated and each possesses its own technical
sis.” [9], the magnitude implied by “markedly” is difficult advantages [18, 19].
to evaluate. Furthermore, the baseline of the statement
is implicit, although it is probably the expression level in Methods
healthy skin tissue. Finally, the source of supporting evi- For the text mining part of our study, there is no prior
dence and experimental details, such as the technique work focused specifically on DEGs in disease. The clos-
employed, is missing in the sentence and in the article in est work concerns the extraction of population percent-
which the sentence appears. Such a statement shows a ages of lymphoma tumors that show expression of a
low level of presented evidence as defined by [10]. gene in immunohistochemistry experiments [20]. In that
In contrast to DEG information found in scientific work, gene names were tagged using dictionary-
text, microarray expression data typically appear in matching and a set of rules was devised to identify sen-
structured form in numerical datasets that cover tences with potentially relevant information about gene
thousands of genes and can be stored in repositories expression. There have also been studies on identifying
such as the Gene Expression Omnibus (GEO) [11] and genes expressed in cell types [21] and anatomical loca-
ArrayExpress [12]. Such repositories allow the pooling of tions [22], or in both [23]. The identification of sen-
multiple datasets to create an aggregate view across dif- tences that describe gene expression, without any other
ferent experimental settings [13]. While better organized contextual details, has also been addressed as part of
than the literature, microarray expression datasets more-general event extraction tasks [24, 25].
present their own challenges. The raw expression data Our approach was to identify sentences from Medline
from these datasets require processing and quality as- abstracts that provide information of the type “X is
sessment [14], and resulting expression values convey a differentially regulated in Y” with respect to healthy
relative rather than an absolute measure. Thus, the ana- controls, with X being a gene and Y a disease. Such
lysis of microarray expression is usually restricted to information can be mapped to vectors of the type
identifying expression values with largest change (PMID, X, Y, Δ) where PMID is the corresponding
between samples (e.g., [15]) or that change beyond a PubMed ID of the abstract and Δ refers to the direction
certain statistically-significant threshold or a fixed fold- and magnitude of expression change between diseased
change threshold. and healthy states. The values that the vectors (PMID,
An important limitation of microarray expression X, Y, Δ) can take were based, in our case, on the content
studies is that they concern only mRNA and not of the sentences identified, thus coreferences or informa-
protein, and in particular only whole-cell mRNA tion from the rest of the document were not considered
[16]. Therefore, they lack the detail and granularity except in cases of ambiguity in the gene name or ana-
of experimental methods, such as immunohistochem- tomical location of the expression. In such cases this in-
istry, that can describe detailed spatial distributions. formation could come from the rest of the abstract if
Moreover, interpretation of microarray expression re- appearing therein. Redundant (PMID, X, Y, Δ) state-
sults is complicated by the natural variation that ex- ments were discarded.
ists across biological samples, as well as by Typically, qualifier keywords and phrases (such as
differences in technical settings across experiments “overexpressed,” “decreased expression” and “greatly
and laboratories [17]. Finally, microarray expression elevated”) helped determine the value of Δ. The set of
Rodriguez-Esteban and Jiang BMC Medical Genomics (2017) 10:59 Page 3 of 10
possible values for Δ were defined to be the following: the subset of genes (HUGO gene symbols) shared by
{high increase, increase, decrease, high decrease}. We both PubTator and GPL570 (n = 17,126).
tracked qualifiers that indicated a very large change in
expression to assign the Δ values high decrease and high Results
increase. For example, expression that was described as The focus of our work was on four diseases: Crohn’s dis-
“greatly elevated” in disease was mapped to high ease (CD), ulcerative colitis (UC), psoriasis (PS) and
increase, while “elevated” or “significantly overexpressed” atopic dermatitis (AD). Their choice stemmed partially
was mapped to increase. from their specificity to particular tissues: psoriasis and
Since DEG information can be conveyed through text atopic dermatitis to the skin, Crohn’s disease and ulcera-
in many ways, we devised a generally-inclusive method tive colitis to the gastrointestinal tract. Another reason
based on tri-occurrence. We searched first for abstracts for their selection was our interest in exploring similar
mentioning a disease Y and a gene X using disease and diseases that are often compared to each other, in our
human gene annotations from NCBI’s PubTator (down- case the pairs PS-AD and UC-CD. We collected DEG
load 2016-01-25) [26]. Those abstracts were then split statements from the literature and microarray datasets
into sentences with the aid of the JULIE Sentence concerning these four diseases (see Methods), focusing
Boundary Detector [27]. For each sentence we detected only on the main affected tissues (e.g., we discarded
whether there were mentions of gene X, disease Y (or serum measurements). We then compared the data re-
abbreviation) and trigger word (or substring). The trig- ported in the literature with the information contained
ger words were selected after [22] to be the following: in microarray datasets.
{express, production, produce, transcription, transcribe}. Through our text mining approach, we created a sam-
Finally, the resulting sentences were manually reviewed. ple of DEG statements coming from 200 Medline ab-
Our goal was to produce a sample of sentences that rep- stracts for AD, 308 for CD, 429 for PS and 273 for UC.
resented an unbiased view of the literature. Undefined These statements concerned 173 unique genes for AD,
expression changes were not considered (e.g. expression 240 for CD, 327 for PS and 285 for UC. (The text min-
described as altered/alteration, aberrant, abnormal, dys- ing results are available as supplementary information,
regulation, expressed differentially, modulated, discord- see Additional file 5.) The microarray datasets presented
ant). Moreover, names indicating protein complexes or different quantities of DEGs depending on fold change
families of proteins or genes that could not be mapped (FC) filtering. For example, for |FC| > 2, 110 unique
to at most three genes were not considered. genes were differentially expressed in AD, 92 in CD, 998
Human microarray expression series relative to each in PS and 2339 in UC.
disease were searched in GEO by using the correspond-
ing disease names as keywords. To maintain consistency, Overexpression is more reported than underexpression
all series selected were based on the same platform, As can be seen in Fig. 1 and Table 1, DEG reports
Affymetrix Human Genome U133 Plus 2.0 Array favor overexpressed genes 3-4 times more than
(GPL570), and included both diseased tissue samples underexpressed genes. Intriguingly, the magnitude of
and normal samples. Unaffected-tissue samples, samples this bias does not differ much between diseases.
after drug treatment or samples from non-primary dis- Microarray expression data shows no such systematic
ease tissue (such as blood peripheral samples) were not imbalance.
considered. From the series found following these cri- To simplify the discussion, the focus in the next sec-
teria, those with the largest sample size were prioritized. tions is on overexpressed genes, for which there exist
The series selected for in-depth analysis were: more data in the literature.
GSE36842 for atopic dermatitis, GSE36807 for Crohn’s
disease, GSE13355 for psoriasis and GSE38713 for ul- The reporting of high overexpression correlates with the
cerative colitis. DEGs were identified using the limma reporting of overexpression and only weakly with
Bioconductor R package using Benjamini & Hochberg microarray fold change
(false discovery rate) to correct for multiple testing and The more a gene is mentioned as overexpressed in a dis-
adjusted p-value <0.05. The fold change (FC) in expres- ease the more likely it will be mentioned as highly over-
sion was used as a variable filter (cutoff ) throughout the expressed (highly increased) in the same disease (Fig. 2).
study. We did not identify any covariates that required One potential explanation for this is that highly overex-
batch correction. Box plots and principal component pressed genes are the focus of more scrutiny due to their
analysis for each dataset are provided in the supplemen- presumed heightened biological relevance. There are ex-
tary information (see Additional files 1, 2, 3 and 4). For amples of this phenomenon that can be observed in the
calculating the positive likelihood ratio between micro- literature. Such is the case of gene S100A7 in psoriasis,
array data and the literature we took into account only which first raised interest as a highly expressed gene in
Rodriguez-Esteban and Jiang BMC Medical Genomics (2017) 10:59 Page 4 of 10
Table 1 Percentage of overexpressed vs. underexpressed unique DEGs in microarray data and the literature. |FC| > n indicates
microarray DEGs with absolute fold change above n
Microarray Literature Microarray Literature
|FC| > 0 |FC| > 2 |FC| > 0 |FC| > 2
AD PS
Overexpressed 44.6% 23.6% 79.4% Overexpressed 55.5% 58.5% 79.2%
Underexpressed 55.4% 76.4% 20.6% Underexpressed 44.5% 41.5% 20.8%
CD UC
Overexpressed 32.6% 30.4% 75.2% Overexpressed 50.4% 60.6% 77.4%
Underexpressed 67.4% 69.6% 24.8% Underexpressed 49.6% 39.4% 22.6%
Rodriguez-Esteban and Jiang BMC Medical Genomics (2017) 10:59 Page 5 of 10
Fig. 4 Number of overexpressed genes for each disease. Number of overexpressed genes for each disease (a) as reported in the literature and (b
and c) as appearing in microarray datasets (FC > 0 and FC > 2, respectively). The tables show the LR+ for genes overexpressed in one disease
(table headers) that are overexpressed in another disease (row names) based on (d) the literature or (e) microarray data with FC > 0 or (f) FC > 2
would be overstating the similarity of PS and AD the Thus, the information conveyed by these two sources
most, while the similarities between PS and CD would can be quite distinct when choosing a FC > 0 cutoff. In
be the least emphasized. Thus, it is possible that the the case showing highest LR+ (UC), the probability of a
similarities between PS and CD have received insuffi- gene being overexpressed in the microarray dataset goes
cient attention (see for example [29]) in comparison to up from 21 to 50% when the literature states that it is
the similarities between PS and AD, if microarray data is overexpressed. The probability of a gene being overex-
to be used as guidance. pressed in the UC literature goes up from 0.09 to 0.34%
when it is overexpressed in the microarray dataset.
As the microarray fold-change cutoff increases, micro-
array data and the literature increase in resemblance
The LR+ can also help us determine further the relation-
ship between overexpression in the literature and in mi-
croarrays. We can compute the LR+ of a gene being
overexpressed in the literature when it is overexpressed
in microarray data and vice versa. Our interest is in
knowing whether the odds of a gene being overex-
pressed in one of the sources change when it is known
to be overexpressed in the other source. Fig. 5 Positive likelihood ratio given microarray data and the
Our finding was that the LR+ depends on the FC cut- literature. Positive likelihood ratio (LR+) of (a) microarray data given
off chosen. For example, the LR+ of microarray overex- the literature and (b) the literature given microarray data for
pression for FC > 0 given the literature (and vice versa) different values of log2FC threshold and for each disease: AD
(diamonds), CD (squares), PS (triangles), UC (crosses). The higher the
is not significant for AD and CD. For PS and UC the LR
LR+ the more likely one data source can predict another one
+ is significant and ranges between 1.5 and 4 (see Fig. 5).
Rodriguez-Esteban and Jiang BMC Medical Genomics (2017) 10:59 Page 7 of 10
A different picture arises with increased FC thresholds, overrepresented in the microarray dataset and the 15
as can be seen in Fig. 5. The LR+ then increases sub- overrepresented in the literature.
stantially, which means that the literature becomes more For FC > 2, the similarities between microarray data
related to microarray data as the FC threshold increases. and the literature were greater. For UC there were 28
This is probably due to the fact that, as has been already shared functional classes between the 47 overrepre-
stated, the probability that a gene is mentioned as over- sented in the microarray dataset and the 38 overrep-
expressed in the literature increases with higher micro- resented in the literature. For PS, there were 11
array FC. shared functional classes between the 28 overrepre-
sented in the literature and the 18 overrepresented in
the microarray dataset.
Differences between microarray data and the literature
translate into alternative views of the underlying disease
biology Discussion
As could be expected, the differences that have been de- Our goal was to explore the relationship between micro-
scribed between microarray and literature data translate array expression data and the expression data reported
into different representations of the pathological pro- in the literature because in our daily work both of these
cesses that characterize each disease. To measure this data sources are used as complementary sources of in-
quantitatively, we looked at the level of enrichment of formation. From the therapeutic point of view, for ex-
Gene Ontology (GO) functional classes associated to the ample, every DEG in disease is a potential point of
genes overexpressed in microarray data and in the litera- intervention or target. Thus, the sole use of microarray
ture. Figure 6 shows the top 20 statistically overrepre- data or of the literature could lead to missing out on po-
sented GO functional classes in microarray and tential targets that appear in one source and not the
literature data for UC based on the PANTHER statistical other. For instance, EGFR does not appear upregulated
overrepresentation test with Bonferroni correction [30]. in the PS microarray dataset, while it is one of the most
For UC and FC > 0, 16 functional classes were shared frequently mentioned upregulated genes in the PS litera-
between the 38 overrepresented in the literature and the ture dataset. On the other hand, defensin beta 4B
36 overrepresented in the microarray dataset. For PS (DEFB4B) does not appear in the PS literature dataset
and FC > 0, on the other hand, only the “unclassified” despite showing the second-highest level of overexpres-
functional class was shared between the 17 sion in the PS microarray dataset.
Fig. 6 Statistically overrepresented Gene Ontology functional classes. Top-20 statistically overrepresented Gene Ontology functional classes based
on overexpressed genes in the UC literature (left) and in the UC microarray dataset (right)
Rodriguez-Esteban and Jiang BMC Medical Genomics (2017) 10:59 Page 8 of 10
Our strategy for gathering microarray data was to se- (e.g. mRNA vs. protein). On the other hand, even though
lect one dataset for each disease of interest, each dataset all microarray data in our study came from the same
created with the same platform to avoid variability platform from the same manufacturer, and each dataset
across manufacturers. For literature data, our approach was created within a single research study, microarray
was to gather a representative sample of the literature, data variability has been shown to be a challenge for re-
rather than to create an exhaustive representation. We, producibility [34–37].
moreover, focused on abstracts, rather than on full text Moreover, because experiments in the literature can
articles, due to limited full text availability. Thus, the be more fine-grained than microarray studies, it is pos-
true number of statements regarding differential expres- sible that a gene might be found to be upregulated in
sion in the literature is larger than what is reported here. some parts of a diseased tissue and downregulated in
The fact that more literature results were oriented to- others, confounding the simplified representation used
wards overexpression than underexpression, unlike in here and hampering comparisons with microarray data.
microarray data, indicates a scientific bias towards report- One additional aspect not considered in this study was
ing overexpression. This bias could be related to the fact the historical dimension. High-throughput techniques
that most drugs are inhibitors and therefore an overex- have been gaining in popularity only recently; therefore
pressed gene is more likely to represent a potential target. older publications would have been less affected by find-
Since, in principle, downregulation may have as much ings coming from high-throughput studies.
functional importance in disease as upregulation, this bias
could be distorting in our understanding of diseases.
We also noted that popular genes tend to be more Conclusion
often described in the literature as overexpressed in At the start of this study we had the expectation that
disease, an effect that is much milder or non- there would be certain biases in the literature in com-
existent for overexpressed genes from microarray parison to microarray data. The literature evidently has
data. This could explain partially why differential ex- a focus that is, at the very least, biased by past research
pression similarities between diseases are higher history, which does not affect microarray data. Our goal
within the literature in comparison to microarray was to quantify this bias, using microarray data as the
data. The quest for higher research impact could be unbiased “ground truth.” However, we did not expect
one of the drivers for the additional attention paid that the relationship between microarray data and the
to popular genes [31–33], leading to further amplifi- literature could be dependent on FC cutoff (which in
cation of their presumed biological importance be- retrospect appears to be naïve), and therefore that we
yond actual biological evidence. should not necessarily consider microarray data a
Our analysis also hints that our perception of the level of ground truth that the literature only partially represents.
similarity between certain diseases could be biased by gen- The use of an FC threshold does not in principle have
eral properties of the diseases that are not reflected in the a fixed biological meaning and its link to biological ac-
expression data. Thus, PS and AD, which share anatomical tivity can change from gene to gene. Moreover, different
location, appear more similar in the literature than UC and FC thresholds yield different outcomes from an expres-
AD, contrary to what is reflected in microarray data. sion study [38]. Based on our work, the literature has a
We also found that microarray data and the literature closer connection with microarray expression data fil-
can produce divergent views of the pathological mecha- tered with higher FC thresholds, which means that it
nisms driving diseases depending on the fold-change cut- may not track biological phenomena appropriately when
off. For FC > 0, the functional classes associated to the FC thresholds do not actually separate meaningful
overexpressed genes in the literature can be very different and non-meaningful expression changes.
from those associated to microarray data. As the threshold
for FC increases, the similarity between the literature and Additional files
microarray data increases, which is then reflected in
higher LR+ values and overlapping functional classes. Additional file 1: Boxplot and PCA for AD. Boxplot and principal
component analysis for the GSE36842 study. (TIFF 189 kb)
One explanation for the divergences between micro-
Additional file 2: Boxplot and PCA for CD. Boxplot and principal
array data and the literature comes obviously from the component analysis for the GSE36807 study. (TIFF 120 kb)
differences in experimental settings. Expression data Additional file 3: Boxplot and PCA for PS. Boxplot and principal
from the literature stem from a variety of sources involv- component analysis for the GSE13355 study. (TIFF 100 kb)
ing methods such as immunohistochemistry, flow cy- Additional file 4: Boxplot and PCA for UC. Boxplot and principal
tometry, in situ hybridization, RT-PCR, next-generation component analysis for the GSE38713 study. (TIFF 226 kb)
sequencing–and also microarrays. Each of these sources Additional file 5: Text mining results. Curated results produced by the
text mining algorithm. (XLSX 162 kb)
differs in level of granularity and molecule measured
Rodriguez-Esteban and Jiang BMC Medical Genomics (2017) 10:59 Page 9 of 10
Abbreviations 10. Wilbur WJ, Rzhetsky A, Shatkay H. New directions in biomedical text
AD: Atopic dermatitis; DEFB4B: Defensin beta 4B; DEG: Differentially- annotation: definitions, guidelines and corpus construction. BMC
expressed gene; FC: Fold change; GEO: Gene Expression Omnibus; GO: Gene Bioinformatics. 2006;7:356.
Ontology; IBD: Inflammatory bowel disease; LR: Likelihood ratio; 11. Edgar R, Domrachev M, Lash AE. Gene Expression Omnibus: NCBI gene
NCBI: National Center for Biotechnology Information; PCA: Principal expression and hybridization array data repository. Nucleic Acids Res. 2002;
component analysis; PMID: PubMed ID; PS: Psoriasis; RA: Rheumatoid arthritis; 30(1):207–10.
RNA-seq: RNA sequencing; UC: Ulcerative colitis 12. Parkinson H, Kapushesky M, Kolesnikov N, Rustici G, Shojatalab M,
Abeygunawardena N, Berube H, Dylag M, Emam I, Farne A, Holloway E,
Acknowledgements Lukk M, Malone J, Mani R, Pilicheva E, Rayner TF, Rezwan F, Sharma A,
Not applicable. Williams E, Bradley XZ, Adamusiak T, Brandizi M, Burdett T, Coulson R,
Krestyaninova M, Kurnosov P, Maguire E, Neogi SG, Rocca-Serra P, Sansone
Availability of data and material SA, Sklyar N, Zhao M, Sarkans U, Brazma A. ArrayExpress update–from an
The datasets supporting the conclusions of this article are available in the archive of functional genomics experiments to the atlas of gene expression.
Gene Expression Omnibus repository, [Link] or Nucleic Acids Res. 2009;37(Database issue):D868–72.
included within the article (and its additional files). 13. Kodama K, Horikoshi M, Toda K, Yamada S, Hara K, Irie J, Sirota M, Morgan
AA, Chen R, Ohtsu H, Maeda S, Kadowaki T, Butte AJ. Expression-based
Funding genome-wide association study links the receptor CD44 in adipose tissue
The authors received no specific funding for this work. with type 2 diabetes. Proc Natl Acad Sci U S A. 2012;109(18):7049–54.
14. Gusnanto A, Calza S, Pawitan Y. Identification of differentially expressed
Authors’ contributions genes and false discovery rate in microarray studies. Curr Opin Lipidol. 2007;
Conceived and designed the analysis: RR and XJ. Gathered the data: RR and 18(2):187–93.
XJ. Analyzed the data: RR. Wrote the paper: RR. All authors reviewed drafts of 15. Rivas MV, Jarvis ED, Morisaki S, Carbonaro H, Gottlieb AB, Krueger JG.
the manuscript and read and approved the final manuscript. Identification of aberrantly regulated genes in diseased skin using the cDNA
differential display technique. J Invest Dermatol. 1997;108(2):188–94.
Ethics approval and consent to participate 16. Trask HW, Cowper-Sal-lari R, Sartor MA, Gui J, Heath CV, Renuka J, Higgins
Not applicable. AJ, Andrews P, Korc M, Moore JH, Tomlinson CR. Microarray analysis of
cytoplasmic versus whole cell RNA reveals a considerable number of missed
Consent for publication and false positive mRNAs. RNA. 2009;15(10):1917–28.
Not applicable. 17. Bammler T, Beyer RP, Bhattacharya S, Boorman GA, Boyles A, et al.
Standardizing global gene expression analysis between laboratories and
Competing interests across platforms. Nat Methods. 2005;2(5):351–6.
The authors declare that they have no competing interests. 18. Fu X, Fu N, Guo S, Yan Z, Xu Y, Hu H, Menzel C, Chen W, Li Y, Zeng R,
Khaitovich P. Estimating accuracy of RNA-Seq and microarrays with
Publisher’s Note proteomics. BMC Genomics. 2009;10:161.
Springer Nature remains neutral with regard to jurisdictional claims in 19. Nazarov PV, Muller A, Kaoma T, Nicot N, Maximo C, Birembaut P, Tran NL,
published maps and institutional affiliations. Dittmar G, Vallar LRNA. sequencing and transcriptome arrays analyses show
opposing results for alternative splicing in patient derived samples. BMC
Author details Genomics. 2017;18(1):443.
1
Roche Pharmaceutical Research and Early Development, Roche Innovation 20. Chang JF, Popescu M, Arthur GL. Automated extraction of precise protein
Center Basel, Grenzacherstrasse 124, 4070 Basel, Switzerland. 2Biogen, expression patterns in lymphoma by text mining abstracts of
Cambridge, MA, USA. immunohistochemical studies. J Pathol Inform. 2013;4:20.
21. Hunter L, Lu Z, Firby J, Baumgartner WA Jr, Johnson HL, Ogren PV, Cohen
Received: 31 March 2017 Accepted: 2 October 2017 KB. OpenDMAP: an open source, ontology-driven concept analysis engine,
with applications to capturing knowledge regarding protein transport,
protein interactions and cell-type-specific gene expression. BMC
References Bioinformatics 2008;9:78.
1. Loging W, Harland L, Williams-Jones B. High-throughput electronic biology: 22. Gerner M, Nenadic G, Bergman CM. An exploration of mining gene
mining information for drug discovery. Nat Rev Drug Discov. 2007;6(3):220– expression mentions and their anatomical locations from biomedical text.
30. Proceedings of the 2010 Workshop on Biomedical Natural Language
2. Campbell SJ, Gaulton A, Marshall J, Bichko D, Martin S, Brouwer C, Harland Processing. 2010.
L. Visualizing the drug target landscape. Drug Discov Today. 2010;15(1-2):3– 23. Neves M, Damaschun A, Mah N, Lekschas F, Seltmann S, Stachelscheid H,
15. Fontaine JF, Kurtz A, Leser U. Preliminary evaluation of the CellFinder
3. Kennedy C. Comparatives, Semantics of. In: Brown K, editor. Encyclopedia of literature curation pipeline for gene expression in kidney cells and
Languages and Linguistics. 2nd ed. Oxford: Elsevier; 2006. anatomical parts. Database (Oxford). 2013;2013(0):bat020.
4. Scheinin T, Böhling T, Halme L, Kontiainen S, Bjørge L, Meri S. Decreased 24. Kim JD, Ohta T, Pyysalo S, Kano Y, Tsujii J. Overview of BioNLP’09 Shared
expression of protectin (CD59) in gut epithelium in ulcerative colitis and Task on Event Extraction. Proceedings of the Workshop on Current Trends
Crohn's disease. Hum Pathol. 1999;30(12):1427–30. in Biomedical Natural Language Processing: Shared Task. 2009.
5. Beisswanger E, Lee V, Kim JJ, Rebholz-Schuhmann D, Splendiani A, 25. Kim J, Pyysalo S, Ohta T, Bossy R, Nguyen N, Tsujii J. Overview of BioNLP
Dameron O, Schulz S, Hahn U. Gene Regulation Ontology (GRO): Design Shared Task 2011. Proceedings of the BioNLP Shared Task 2011 Workshop;
principles and use cases. Stud Health Technol Inform. 2008;136:9–14. 2011. p. 1–6.
6. Lesch CA, Kraus ER, Sanchez B, Gilbertsen R, Guglietta A. Lack of beneficial 26. Wei CH, Harris BR, Li D, Berardini TZ, Huala E, Kao HY, Lu Z. Accelerating
effect of COX-2 inhibitors in an experimental model of colitis. Methods Find literature curation with text-mining tools: a case study of using PubTator to
Exp Clin Pharmacol. 1999;21(2):99–104. curate genes in PubMed abstracts. Database (Oxford). 2012;2012:bas041.
7. Rodriguez-Esteban R, Roberts PM, Crawford ME. Identifying and classifying 27. Tomanek K, Wermter J, Hahn U. Sentence and token splitting based on
biomedical perturbations in text. Nucleic Acids Res. 2009;37(3):771–7. conditional random fields. Proceedings of the 10th Conference of the
8. Barker C. Vagueness. In: Brown K, editor. Encyclopedia of Languages and Pacific Association for Computational Linguistics; 2007. p. 49–57.
Linguistics. 2nd ed. Oxford: Elsevier; 2006. 28. Celis JE, Crüger D, Kiil J, Lauridsen JB, Ratz G, Basse B, Celis A. Identification
9. Takahashi H, Hashimoto Y, Ishida-Yamamoto A, Iizuka H. Roxithromycin of a group of proteins that are strongly up-regulated in total epidermal
suppresses involucrin expression by modulation of activator protein-1 and keratinocytes from psoriatic skin. FEBS Lett. 1990;262(2):159–64.
nuclear factor-kappaB activities of keratinocytes. J Dermatol Sci. 2005;39(3): 29. Najarian DJ, Gottlieb AB. Connections between psoriasis and Crohn's
175–82. disease. J Am Acad Dermatol. 2003;48(6):805–21.
Rodriguez-Esteban and Jiang BMC Medical Genomics (2017) 10:59 Page 10 of 10