Biomedical Ontology Development Insights
Biomedical Ontology Development Insights
Tool Dissemination
Doing It Right
NCBC UPDATE:
Shedding New Light On
Winter 2008/2009
contents
ContentsWinter 2008/2009 Winter 2008/2009
Volume 5, Issue 1
ISSN 1557-3192
13 NCBC Update:
Shedding New Light on Biological Complexity
Katharine Miller • Kristin Sainani, PhD
Cassandra Brooks • Emmanuel Romero
Hadley Leggett, MD • Kayvon Sharghi
Lisa Grossman • Lizzie Buchen
BY KATHARINE MILLER Michael M. Torrice • Michael Wall, PhD
Molly Davis • Stephanie Pappas
Community Contributors
Tool Dissemination—Doing It Right
21 BY KRISTIN SAINANI, PhD
Mark Musen, MD, PhD
Kartik Mani, PhD
Joy Ku, PhD
Layout and Design
DEPARTMENTS Wink Design Studio
1 GUEST EDITORIAL Printing
It Takes a Village: Building the Next Advanced Printing
Generation of Biomedical Ontologies BY Editorial Advisory Board
MARK A. MUSEN, MD, PhD Russ Altman, MD, PhD, Brian Athey, PhD,
Dr. Andrea Califano, Valerie Daggett, PhD,
Scott Delp, PhD, Eric Jakobsson, PhD,
3 SIMBIOS NEWS Ron Kikinis, MD, Isaac Kohane, MD, PhD,
Stop Wheel Reinvention, Mark Musen, MD, PhD, Tamar Schlick, PhD,
Share Your Simulations Jeanette Schmidt, PhD, Michael Sherman
Arthur Toga, PhD, Shoshana Wodak, PhD,
BY JOY P. KU, PhD John C. Wooley, PhD
For general inquiries,
5 NEWSBYTES subscriptions, or letters to the editor,
BY HADLEY LEGGETT, MD, LISA GROSSMAN, visit our website at
MICHAEL WALL, PhD, CASSANDRA BROOKS, [Link]
MICHAEL M. TORRICE, PhD, KAYVON SHARGHI,
Office
EMMANUEL ROMERO, LIZZIE BUCHEN, STEPHANIE Biomedical Computation Review
PAPPAS, MOLLY DAVIS Stanford University
• Modeling Cracks in Clogged Arteries 318 Campus Drive
• Modeling Muscles From the Inside Out Clark Center Room S231
• “Digital Embryo” Created Stanford, CA 94305-5444
• The Circuitry of Yeast Biomedical Computation Review
• Watching a Molecule Bind is published quarterly by:
• Identifying a Cell’s Weakest Link
• Diagnosing Cell Circuitry
• Cancer’s Signature—Written in Blood
• Blurring Data for Privacy and Usefulness
An NIH National
• Modular Modeling Center for Physics-
Based Simulation of
29 UNDER THE HOOD | Network-based Approaches to Biological Structures
Prediction of Disease Genes BY KARTIK MANI, PhD Publication is made possible through the NIH
Roadmap for Medical Research Grant U54
30 SEEING SCIENCE | Visualizing Ventricular Fibrillation GM072970. Information on the National Centers
for Biomedical Computing can be obtained from
BY KATHARINE MILLER [Link] The NIH
program and science officers for Simbios are:
COVER ART: Created by Rachel Jones of Wink Design Studio using a collection Peter Lyster, PhD (NIGMS)
Jennie Larkin, PhD (NHLBI)
of images from the NCBCs. Robotic arm is © Eraxion | [Link] Jennifer Couch, PhD (NCI)
Semahat Demir, PhD (NSF)
Peter Highnam, PhD (NCRR)
Jerry Li, MD, PhD (NIGMS)
Yuan Liu, PhD (NINDS)
Richard Morris, PhD (NIAID)
Joseph Pancrazio, PhD (NINDS)
Grace Peng, PhD (NIBIB)
Nancy Shinowara, PhD (NCMRR)
David Thomassen, PhD (DOE)
Ronald J. White, PhD (NASA/USRA)
Jane Ye, PhD (NLM)
It Takes a Village:
Building the Next Generation
of Biomedical Ontologies
A
lthough the notion of ontology has been around annotate the ontology via a Web-based wiki and
since Aristotle, the perceived need to develop suggest changes and extensions, although again, the
ontologies in biomedicine has accelerated in modification of the actual ontology content will be
recent years as investigators attempt to make sense of the channeled through a set of trained individuals who
terabytes of high-throughput data that are now finding understand principles of knowledge representation and
their way into public repositories. While the number of the use of knowledge-editing tools.
biomedical terminologies and ontologies continues to Probably the best exemplar of an open, nearly demo-
increase as new areas of biomedical content become for- cratic ontology-development initiative is the Open
malized, the creation and annotation of these resources Directory Project (ODP). Founded more than 10 years
can’t quite keep up. The flood of information may neces- ago, the ODP has enlisted more than 75,000 volunteers
sitate a new approach involving vastly more ontology to flesh out the extensive open-content ontology of Web
developers. It may, in fact, take a village. pages that has been adopted by Google, Yahoo!,
The construction of biomedical ontologies has long Netscape, and a host of other companies. The ODP has
been a cottage industry, with even vast systems such as generated an enormous ontology (commonly known as
SNOMED (the Systemized Nomenclature of Medicine) dmoz) that provides standard, categorized entrée to vir-
initially representing the handiwork of a very small group tually all the content on the Web. All of us use dmoz,
of dedicated individuals. Venerable ontologies such as the perhaps unknowingly, every time we browse the Web by
Foundational Model of Anatomy and the NCI Thesaurus categories in Google and Yahoo!, rather than searching
represent the work of a surprisingly small set of develop- the Web’s free text for particular terms. Embracing
ers. Nevertheless, as the demand for ever larger and more everything imaginable that a user could search for, dmoz
granular ontologies accelerates, and as large-scale systems is a remarkable demonstration of how scalable ontology
such as the International Classification of Disease are engineering can be, particularly when volunteers step
being reengineered, the scientific community has increas- forward to provide fine-grained descriptions of their par-
ingly raised concerns about whether ontology develop- ticular areas of personal interest.
ment ultimately can be a scalable enterprise. Practical
ontologies comprise tens of thousands of con-
cepts, and a handful of individuals can
never have personal knowledge
of everything that needs to
be represented in such
a system.
To address this
problem, workers
in biomedicine
are attempting to
democratize the
development of
large-scale ontolo-
gies. The engineering
of the Gene Ontology,
for example, has been char-
acterized by an open development
process to which nearly anyone can
contribute. The actual editing of the Gene
Ontology content, however, is still performed by only a At [Link], the user is presented with a list of pos-
handful of trusted curators. The National Cancer sible ontologies to explore and visualize. This screenshot of the
Institute is experimenting with an open process for Cellular Components ontology shows the relationship between major
extensions to the current content of the NCI Thesaurus components and some of the minor components. Any registered user
via the BiomedGT initiative. Here, nearly anyone can can comment on any ontology.
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 1
guest editorial
The dmoz ontology is very simple in its structure, and
lacks the rich semantics of ontologies developed in for-
It is clear that the
mal knowledge representation systems such as the Web
Ontology Language (OWL). When the developers of ontology-development
dmoz make modeling errors, the consequences are
unlikely ever to impede the advancement of science or community needs at least to
to threaten lives. Nevertheless, the dmoz ontology
stands as a stunning example of how legions of volun- experiment with new methods
teers can be mobilized to generate an enormous and
undeniably useful ontology. Imagine if the lessons of
dmoz could be applied to SNOMED or to BiomedGT!
of ontology engineering
At the National Center for Biomedical Ontology
(NCBO), we are experimenting with ways in which the
that can scale to future
biomedical community can take an active part in con-
tributing to the construction of scalable ontologies and biomedical requirements.
controlled terminologies. Our BioPortal system allows any
registered user to comment on any ontology in our dis- maintain the quality of ontologies if the development
tributed repository, to comment on the comments left by process is democratized. Organizations such as the Open
other users, and to demonstrate how the elements of one Biomedical Ontologies (OBO) Foundry have been
ontology may relate to those of another. We have used established under the assumption that there must always
this capability extensively in the engineering of the be central management of ontology development to
Biomedical Resource Ontology used to describe the ensure the quality of the content. And yet there contin-
online software and data resources developed by the ue to be too much data, too many medical records, and
National Centers for Biomedical Computing and by the too many experiments for the ontology-development
recipients of Clinical and Translational Science Awards. community to keep up with existing needs.
BioPortal, at present, does not play a role in completely I don’t know whether the dmoz approach will really
open ontology editing, however. be practical in biomedicine, but it is clear that the ontol-
There are very legitimate concerns about how we can ogy-development community needs at least to experi-
ment with new methods of ontology engineering that
can scale to future biomedical requirements. Surely there
DETAILS are ways to take advantage of the expertise distributed
BioPortal: [Link] among all biomedical investigators in a way that will
The Open Director Project: [Link] overcome many of the limitations of centralized ontology
curation. Workers at NCBO are extremely excited about
Mark A. Musen, M.D., Ph.D. is Professor of Medicine the possibilities that new technology might provide in
(Biomedical Informatics Research) and Computer enabling this more open approach to ontology engineer-
Science at Stanford University. He is Director of the
ing. Experimentation with community-based ontology
Stanford Center for Biomedical Informatics Research
development not only may accelerate the engineering of
and principal investigator of the National Center for
Biomedical Ontology (NCBO).
badly needed ontology content, but also can provide a
laboratory for the study of new mechanisms for collabo-
ration and interaction in biomedicine. ■
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 3
simbios news
Xuhui Huang, PhD, a research associate chief of the journal PLoS Computational A FIRST STEP TOWARD THE
in the bioengineering department, and Biology. “So by making the software PUBLICATION OF THE FUTURE
their colleagues developed Mapper, a available, researchers open up that possi- Bourne sees the Simbios publica-
tool that improves detection of low-den- bility to benefit from what other people tion repository as a very positive step:
sity states within a massive amount of do with the software as well.” “It actually speaks to the dream that I
data. After creating a project for Mapper Bourne acknowledges that the cur- have.” He envisions all aspects of
on [Link], they submitted a paper rent system does not always reward the research being accessible, with the
showing how Mapper could be used to work involved in preparing and support- paper being an access point to the
identify intermediate stable states during ing open-source software: answering experiment. From the paper, a
the RNA hairpin folding process, a diffi- questions from software users and provid- researcher could retrieve and manip-
cult task when those states represent ing documentation, examples, and tuto- ulate the associated data, and possibly
only two to three percent of the whole rials. All of that effort takes time that discover new links and relationships
data set. In order for someone else to could be spent doing research that would via the data and tools—not just the
replicate that research, Yao and Huang generate more publications—the metric paper citations—enhancing the
posted (on [Link]) not only the by which academics are primarily judged. research process.
Mapper software, but also the project’s To address this concern, Bourne says, Bourne observes that there are an
input data and instructions about how to PLoS is considering having a special sec- increasing number of efforts to cap-
use that data with Mapper. tion that only publishes articles report- ture this whole research work flow
“To reproduce the results from a ing on open-source compu- process. The goal of his BioLit
paper is not an easy task,” Huang says. tational biology software ([Link] project
“You need all the components togeth- that has been deposit- is to connect open access
er—the data, the program, your param- ed in an established articles with information in
eters, instructions—so that people can repository. existing biological databas-
easily reproduce the results. [Link] es, such as the Protein Data
provides such a platform, especially Bank (PDB). Another exam-
with this publication mode.” ple is the Insight Journal
While the information could have ([Link]
been posted on his own web- an open access on-line
site, Yao says that researchers publication focused on
from other fields would medical image pro-
not think to look there. cessing and visualiza-
For an interdisciplinary tion where authors are
field, a common platform encouraged to provide
like the [Link] publi- the data and software associ-
cation repository is particu- ated with their papers.
larly valuable. “Most people get togeth-
er because of content,”
REWARDS Bourne says. Efforts such as the
FOR SHARING [Link] publication repository
While some researchers think that provide the infrastructure to share
sharing their software or data means giv- different types of information,
ing up their competitive advantage, oth- enabling a dialog between the
ers believe that it is a great way to build Molecular dynamics were used to simulate the fold- people who are using and
a successful career. “Careers often come ing of the RNA hairpin structure (above), generating developing the content.
from the application of software to make hundreds of thousands of molecular structures. “My sense is that in the next
new discoveries in the life sciences,” says Mapper was then used to sort through all that data ten years, scientific discourse is
Philip Bourne, PhD, a professor in phar- to identify the relatively infrequent intermediate going to change very dramati-
macology at the University of California states that occur during the folding process. cally as a result of these kinds of
at San Diego, and founding editor-in- Courtesy of Yuan Yao. things.” Bourne says. ■
DETAILS
Learn more about the projects mentioned in this article.
Walking simulations: [Link]
Mapper: [Link] Simbios ([Link]
is a National Center for Biomedical Computing
Climber: [Link] located at Stanford University.
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 5
NewsBytes
model’s parameters (such as the num- work is still in progress,” he cautions. cells and a beating heart. To meet that
ber of contractions per second) until it The new model can simulate any challenge, Keller and his colleagues
matched the behavior. This led to some biological tissue that contracts, not just developed a new technique called digital
accurate but limited simulations. skeletal muscles, says Ellen Kuhl, PhD, scanned laser light sheet fluorescence
Böl’s work builds muscles from the professor of mechanical engineering at microscopy (DSLM).
inside out. He uses the finite element Stanford University. Kuhl was so DSLM generated a three-dimension-
method, originally developed by aero- impressed that she is now working with al image of the embryo by combining
space engineers to design planes, to Böl to model heart tissue, with the goal about 400 pictures taken along slightly
divide a muscle into discrete parts that of helping researchers develop a patch different planes. The team repeated this
each behave differently. Previous finite to replace dead tissue after a heart process every 60 to 90 seconds, tracking
element muscle models used a continu- attack. “I think the cardiac application changes as the zebrafish developed. In
um-based approach, which lumped all is even more sexy, because many more 24 hours, this amounted to about
muscle fibers together and treated them people could benefit from it,” she says. 400,000 images for each embryo.
as a single unit. But Böl gets into the —By Lisa Grossman To deal with this deluge of data—
nitty-gritty of each tiny fiber. In essence, three terabytes per embryo—the
his modeled muscles behave like a bunch “Digital Embryo” Created researchers developed a computational
of ropes of different thicknesses attached How does a humble zygote grow into pipeline. They wrote algorithms defin-
at the same point. Because the model a fully functioning animal, billions or ing the structure of cell nuclei, then
describes each rope independently, Böl trillions of cells strong? This question ran the microscopy data through a net-
can plug in any parameters he wants and has intrigued biologists for centuries. work of more than 1000 computers at
get realistic behavior back out. Now scientists have generated the first EMBL and the Karlsruhe Institute of
In his model, Böl splits the muscle complete developmental blueprint of a Technology in Germany.
into an active element (the contractile vertebrate—a “digital embryo” map- The computational analysis picked
muscle fibers) and a passive one (the ping the positions, divisions, and out every nucleus. Keller’s team then
incompressible tissue that surrounds movements of every cell during the first processed this information into compre-
them). Putting the “ropes” into the 24 hours of a zebrafish’s life. hensive databases of cell positions, divi-
realistic environment of soft tissue “Such reconstruction of a complex sions, and migratory tracks. In all, they
yields a more complete picture, he says. vertebrate embryo had not been achieved catalogued 55 million nucleus entries.
The model has both experimental before,” says Philipp Keller, a PhD candi- Digitizing the data was key.
and clinical value, Böl says. Scientists date at the European Molecular Biology “Microscopy tells you about phenom-
will use it to test the properties of living Laboratory (EMBL) in Heidelberg, ena from a qualitative point of view,”
muscle, or to help doctors design unique Germany. Keller is lead author of the Keller says. “But with digital
treatments for patients, he believes. He paper, which appeared in the October 9, embryos, we can count the number of
is now working with sports doctors to 2008 issue of Science. cells that are involved in a process
refine and implement his approach. “But Developmental biologists have long and see what they do.”
I have to say, these are first trials and coveted such a tool, but imaging a com- The digital embryo has many poten-
plex organism’s growth pres- tial uses. For example, the researchers
ents a serious hurdle. used it to determine that zebrafish germ
After just one day, for layers—which eventually give rise to
example, a zebrafish all of the fish’s tissues—form more syn-
already has 20,000 chronously across the embryo than pre-
viously thought.
Keller also envi-
7
NewsBytes
paper applied a similar engineering
approach to better understand how yeast
a means to tease out the structure and
function of the underlying system,” says
Until now, says
cells respond to fluctuations in nutrient James Collins, PhD, professor of bio-
levels. If yeast is deprived of its favorite medical engineering at Boston Emad Tajkhorshid,
sugar (glucose), it will consume an alter- University. “I am already beginning to
native and less nutritious sugar (galac- think about how these might be inter- “nobody has been
tose). The researchers created a sinu- esting tools to use to look at other sys-
soidal input by alternately feeding and tems, bacteria in particular.” able to capture and
starving yeast of glucose on different —By Cassandra Brooks
time scales while galactose was con- describe the full
stantly present in the environment. The Watching a Molecule Bind
cells responded to long-term changes in
glucose, but not to faster fluctuations.
Like a paper clip being pulled to a
magnet, a small molecule called ADP
process of ligand
The researchers then made a model gets pulled into its port in a new simu-
based on the well-known metabolism lation. Because of a simple case of
binding to a binding
of galactose. But the experimental opposites attract, it’s the first time com-
yeast was responding much faster to the putational biochemists have successful- site while permitting
glucose fluctuations than the model ly simulated a molecule—or ligand—
predicted. “This suggested something being drawn into its binding site in an natural motion of
was crucially missing from the model,” unbiased simulation.
says co-author Jeff Hasty, PhD, associ- “Nobody has been able to capture and the ligand.”
ate professor of bioengineering at the describe the full process of ligand binding
University of California, San Diego. to a binding site while permitting natural Tajkhorshid and graduate student Yi
Studying live yeast provided the motion of the ligand,” says Emad Wang describe their simulations in the
answer: The messenger RNA necessary Tajkhorshid, PhD, assistant professor of July 15, 2008 issue of the Proceedings of the
for the galactose metabolic pathway biochemistry, pharmacology and bio- National Academy of Sciences.
was degraded when glucose was pres- physics at the University of Illinois at Tajkhorshid and Wang simulated
ent. “The most exciting thing is that Urbana-Champaign. “We think we are the binding of adenosine disphosphate
without the model, none of this would getting the most faithful representation (ADP), a molecule involved in fueling
have happened,” said Hasty. of the binding site, because in our simula- the cell, to the ADP/ATP carrier pro-
“The broader contribution of each tions, the protein is dynamic and allowed tein (AAC) located in the membrane
of these pieces will be to point to the to freely react to and establish new inter- of mitochondria—the cell’s power gen-
value of using periodic input signals as actions with the ligand as it binds.” eration plants. For ADP to be shuttled
into the mitochondria, it must first
float into a cavity inside AAC and
bind to it—an event that lasted 100
nanoseconds in the simulations.
Previously, simulations of molecular
binding have required an active force to
produce the attachment. But placing
the ligand (in this case ADP) at the
mouth of the ligand binding site (here,
the AAC cavity) in molecular dynamics
simulations is more faithful to biological
reality. Initially, Tajkhorshid thought
that the ADP would just float away.
Instead it moved right into place. He
and his colleagues found that AAC uses
a special bait to lure ADP to its binding
site: Positively charged amino acids line
the sides and bottom of the AAC cavi-
Yeast grows in a microfluidic chamber designed at the University of California, San Diego. Regular ty, creating a surprisingly strong electro-
nutritional inputs, generated in a wave-like pattern, reveal aspects of how the cells regulate their static potential that attracts the nega-
metabolism and internal environments. The green background color signals that it is a galactose rich tively charged ADP. They called this
environment. Photo credit: UC San Diego Jacobs School of Engineering. process “electrostatic funneling.” And
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 9
NewsBytes
cate. “Signaling networks are so compli-
cated right now that common sense
doesn’t always hold true,” Yaffe says.
“The thing that makes me really stop
and pay attention is the methodology,
which I found of special note,” says
Raphael Levine, PhD, distinguished
professor of chemistry at the University
of California, Los Angeles. “Instead of
trying to see if the model can predict
something new, they tried to drive it to
say something which they know it
shouldn’t say. As a result, they were suc-
cessful in finding some new biology.”
—By Kayvon Sharghi
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 11
NewsBytes
“I’ve given Little b the power to
That works in theory, but the actual
risk depends on the type of data set and reason about biological objects,”
what an intruder wants from it. A prose-
cutor digging up dirt on a defendant Aneil Mallavarapu says.
would try to re-identify a specific person
in the database. A journalist trying to easy-to-use tool for biology labs. build entire virtual cells or virtual plants
discredit an organization’s data-security “I think that as an everyday tool, it collaboratively, increasing their ability
procedures would also only need to re- [Little b] is going to be kind of like the to study their projects in silico.
identify one person, but it wouldn’t mat- microscope,” says Aneil Mallavarapu, While the idea of breaking down bio-
ter who. El Emam set out to test whether PhD, lead developer of Little b and a logical systems into modular chunks
k-anonymity works in both circum- senior research scientist in systems biol- may seem logical, Little b may not arrive
stances. His findings: k-anonymity cor- ogy at Harvard Medical School. “We’re in the lab immediately, says Birgit
rectly predicts the risk of re-identifying essentially building a new kind of gel, a Schoeberl, PhD, a senior director of
one specific individual with minimal new type of microscope for the lab.” The research at Merrimack Pharmaceuticals,
harm to the value of the database (the work appears in the June 2008 issue of Inc, in Cambridge, Massachusetts. “I’m
prosecutor example). But using k- the Journal of the Royal Society Interface. excited about the concept and what I
anonymity to protect against re-identify- Biologists traditionally create models see, but in my own experience, it isn’t
ing an arbitrary person (the journalism to describe unique systems, such as the straightforward,” Schoeberl says. “I
example) is unnecessarily strict and com- development of fruit fly embryos or the think it’s not quite ready for non-
promises the research quality of the data. actions of a phosphorylation cascade on developers. I hope he keeps developing
Since researchers choose k based on gene transcription. Such computational it, or someone takes it on to keep
statistical theory, El Emam suggests data models are usually based on lists of the working on the idea.”
custodians run test cases to verify if the system’s properties, which detail every —By Molly Davis ■
k is sufficient, or if it’s overprotective, as molecular interaction in the
in the journalism example, before mak- system. This allows researchers
ing the data available to researchers. If to tailor models to the precise
needed, the number of groupings of k questions being asked, but it
identical data points could then be also constrains the model’s use-
adjusted to ensure that the actual risk fulness, because it can only
approximates the theoretical risk of 1/k probe into one area.
and, in this way, keep the risk accept- Little b strives to break down
ably low while preserving data. biological systems into modules
“What is needed are the steps to that can be used regardless of
turn this article into a practical tool the specific context, such as
that custodians can use in conjunction “nuclear export“ or “membrane
with researchers,” says Joan Roch, chief localization.“ It then defines
privacy officer for Canada Health those parts in a mathematical
Infoway in Montreal, Quebec. language. Researchers can use
El Emam says he plans to continue Little b to put together assorted
exploring actual risks in various data- modules to describe their sys-
security scenarios: “It’s a big problem, tem; Little b then uses those
and we’ve solved part of it.” symbolic modules to write out Little b is based on a core language, which includes
—By Stephanie Pappas executable code that a scientist the Lisp language it was created in (green) and the
could use in a simulation pro- knowledge base, symbolic mathematics and syntax
Modular Modeling gram like MATLAB. “I’ve modules that allow Little b to reason about biological
Biological models can quickly given Little b the power to rea- systems. It also includes modular libraries that
become as complex as the systems they son about biological objects,” describe specific biological interactions, and transla-
represent. And minor changes can Mallavarapu says. tors that can generate code used in simulations. Blue
necessitate a complete rewrite of the Mallvarapu is excited about areas exist within the current framework; yellow
model. But researchers may soon snap the possible use biologists might areas are currently under development or are envi-
their models together like LEGOs, using make of Little b. He would like sioned for future work. Reprinted with permission
a new programming language called to see the language help uncov- from Mallavarapu, A, et al., Programming with mod-
Little b, which uses modularity to sim- er the complex pathways els: modularity and abstraction provide powerful
plify biological modeling. Eventually, involved in diseases. He hopes capabilities for systems biology, Journal of the Royal
the authors hope to turn Little b into an that researchers will eventually Society Interface, online publication, July 23, 2008.
NCBC UPDATE:
Shedding New Light On
By Katharine Miller
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 13
B ut it’s the Centers’ wide-ranging impact on
biomedicine that takes center stage. From AIDS
to diabetes, prostate cancer or schizophrenia, the
NCBCs are changing the landscape of disease research
by shedding new light on biological complexity.
“The impact on biology and medicine hap- the NCBCs to successfully penetrate the
pened faster than anyone expected,” says Russ broader community with tools, techniques and
Altman, MD, PhD, co-principal investigator methodology,” he says. “We’ve shown what
for Simbios, the National Center for Physics- can be accomplished by applying these tools
based Simulation of Biological Structures, to biological problems.”
an NCBC grantee at Stanford University. And while the specific breakthroughs enabled
And that impact springs from the way the by NCBC tools varies with the tool being used or
NCBCs function, says Andrea Califano, PhD, the disease being studied, it is clear that they are
who heads the National Center for all helping researchers approach the complex sys-
Multiscale Analysis of Genomic and tem that is the human body. “Dealing with
Cellular Networks (MAGNet) at Columbia complexity is the essential challenge of this
University. “Developing new tools in the con- century in biology,” says Scott Delp, PhD,
text of solving specific scientific, biological or co-PI for Simbios. “And you can’t do it with-
medical problems is what I think has allowed out computers.”
NCBC
Y EAR
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 15
investigators can do research that will impact centered at the University of Michigan, agrees.
human health. Our goal is to create the kinds of While his center’s tools have contributed to a
tools that would be valuable to everybody.” better understanding of type 2 diabetes and
Ron Kikinis, PhD, head of NA-MIC, con- prostate cancer progression, the tools’ reach
curs. “We will not solve cancer but we will extends much farther: “We’re opening doors to
provide the people who are fighting cancer new research,” he says.
with better tools to fight their fight,” he says.
“And the DBPs will use these tools and pro- NCBC CHALLENGE:
mote those tools into their communities—so PUTTING IT ALL TOGETHER
that makes it possible for lots of different dis- For the last thirty years, biology has been
eases to be addressed.” about breaking things down into their funda-
Brian Athey, PhD, co-PI for the National mental parts to understand them. “But things
Center for Integrative Biomedical Informatics, don’t work as independent parts,” says Delp.
“Theoretical and computational biology let
you put things back together to understand
the whole system.”
Several of the NCBC PIs cite the re-
“The tools get developed assembling of biological pieces as a major
focus of their efforts. For example, literally
thousands of experiments have looked at how
because you couldn’t do elements of the neuromuscular system (mus-
cles, joints, connective tissue) operate inde-
something without them. pendently. But, Delp says, looking at those
elements separately doesn’t tell you how peo-
And vice versa, you get this ple move. OpenSim lets researchers put the
pieces together. “When you can code the
tool and you decide to pose details accurately in a computer framework,
then you can understand how the system
new questions. You end up works,” Delp says.
Likewise for the brain, says CCB’s Toga.
Brain researchers have typically focused on
pushing and pulling so only one variable at a time—for example,
electrical activity, blood flow, distribution of
that both are advanced,” receptors, gene expression patterns, or cortex
morphology. But, Toga says. “All of these
Art Toga says. brain changes are happening in concert.” To
understand the brain requires re-integration of
these events. CCB, Toga says, is providing the
tools, mechanisms, and strategies to put things
Toga: “Our hope is to continue to integrate what we know about the brain in a way
that allows us to ask questions such as: ‘How does the brain change throughout a per- NCBC
son’s life?’ These sorts of emerging questions are provocative. And we can only ask
them because of computation. So by the end of our ten years, GOALS
10 YEAR
Musen: “We are thinking about what it would mean to be able to move biomedical
knowledge from prose to machine-processable format. The long-term vision is to create the
infrastructure and tools so that biomedical literature could be intelligible to both people and
machines. Ultimately this could allow intelligent computer-based agents to read the litera-
ture, to make associations between scientific contributions, and to
synthesize ideas from the literature. That would obviously change the
NCBC way we do science in a very profound way. But there are lots of baby
YEAR
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 17
epigenetic, functional and structural data—and NCBCS:
getting an answer that can really dissect dis- MORE THAN THE SUM OF THEIR PARTS
ease,” he says. MAGNet’s goal is to establish The NCBCs are also working together in
such a framework and to show that the frame- various ways to ensure that they have a broad
work can integrate data in meaningful ways for impact. In some ways this is a surprise, say the
several diseases. “We already have proof of con- NCBC PIs, because the NIH cast such a wide
cept for glioblastoma multiforme—a cancer net—with centers that cover ontologies, simu-
that produces the worst possible prognosis in lations, clinical systems, systems biology and
patients,” Califano says. The results for that imaging. “Given the breadth of the needs and
work will be published in the next few months. the solutions to biomedical computing prob-
“This kind of proof of concept in a disease is of lems,” says Kohane, “it wouldn’t have been sur-
course important, but at the same time the prising if there had been no overlap and the syn-
methodology becomes universal.” ergies had been fewer.”
NCIBI is also integrating many different Yet the NCBCs have found overlap and
high-throughput data types to better understand have helped each other. For example, the i2b2
complexity. “We do not yet understood the full center collaborated with NCIBI around Type 2
complexity of the architecture of the human diabetes, Kohane says. And NA-MIC nicely
genome,” Athey says. “Only 2 percent of the complemented i2b2’s major depression DBP by
genome are ‘genes’ and we’re learning more and correlating patient imaging with what was
more that the other 98 percent are doing being seen genetically. Similarly, ontologies
things.” To tackle that problem, he says, com- from NCBO have been helpful to CCB in con-
putational biology is making huge strides. “What structing their brain atlas; and CCB and
bioinformatics was five years ago is frankly just a Simbios have used some of NA-MIC’s visuali-
glimmer of what it is today,” Athey says. “It’s zation tools.
exploding into something much more robust. Even though the NCBCs might be develop-
And that’s going to continue for a while.” ing different tools, Califano says, “when you
NCBC
Athey: “There’s much more work to do to figure out how to use systems biology more
1 0
YEAR
effectively to understand disease and its complications. The daunting complexity of biological
systems is becoming more and more clear. To gain an understanding of that complexity, we GOALS
need an integrative approach that’s iterative and that allows the integration of many different
kinds of data types around hypotheses and models. The abundance of high throughput data we’re
presented with from next generation sequencing, and what that’s revealing about the transcriptome
and alternative splicing, and all the components we haven’t yet annotated—
it’s just astounding. It’s literally changing our basic understanding of cells and
their complexity and function. And, frankly, it’s changing what our understand-
ing of a gene is. So there’s a lot of work to do. I think that’s
the theme. And each success brings on new challenges.”
Brian Athey, PhD, is the principal investigator for the National Center for
Integrative Biomedical Informatics (NCIBI), associate professor of biomed-
ical informatics at the University of Michigan, and director of the Michigan
Center for Biological Information.
NCBC
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 19
NCBCS: BENCH TO BEDSIDE time to build the tool, teach people how to use
Whether casting a wide net to enable it, get it adopted, make a discovery and then
research in lots of areas is enough to render the translate that into clinical care.” Currently, says
NCBCs successful remains to be seen. Curing a Delp, “OpenSim is only halfway down that
disease would be better. “If we actually success- pipeline and is just beginning to see the first
fully did a big population study and discovered examples where new discoveries will enhance
human health.”
Kikinis says NA-MIC’s tool kit is
similarly poised for bedside use. He’s
beginning to see the first signs—such
as questions at seminars, and email
inquiries—that companies are inter-
“Adoption by companies ested in it. “Adoption by companies is
one indication that what we’re doing
is one indication that what will eventually make a difference to
clinical practice,” he says. “We are not
we’re doing will eventually make yet at that point, but I have these
early indicators.”
a difference to clinical practice,” Migrating computational biology
from the bench to the bedside remains
a challenging goal for all the centers.
Ron Kikinis says. “We are But, as Toga sees it, “I think these com-
putational strategies, which are the
not yet at that point, but hallmark of this program, are having a
great effect on accelerating that.” CCB
I have these early indicators.” is modeling the effect that HIV and
Alzheimers have on the brain. These
are diseases that will strike people we
all know, Toga notes. “So our work
immediately transforms a mathematical
problem [shape modeling] into some-
something important or successfully calculated thing with obvious and immediate clinical
how to design a vaccine or predicted a new drug value,” he says. “And the time frame for doing
for a specific disease, then we’d be bringing our- that is getting shorter and shorter and shorter.”
selves to the next level,” says Kohane. “We’d be Kikinis summed it up succinctly: “What are
solving a biomedical problem of true health rel- the NCBCs doing for biology? Everything.
evance. In fairness, I think we’re all trying to get That’s by design, but now you can say that
there, but we’re not there yet.” they’re actually delivering, and there’s a sense of
“The challenge is,” says Delp, “that it takes excitement. It’s clear that things are moving.” ■
TooL
Dissemination DOING IT RIGHT
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 21
T
To overcome the cottage industry men- someone has to build “disseminability” ration with colleagues at the University
tality, the National Institutes of Health into the tool, with robust, flexible, and of California, San Diego; the program is
(NIH) is placing a greater emphasis on extensible code. Then, someone has to downloaded about 1000 times a month
dissemination as a piece of the National package the tool in a way that makes it ([Link]
Centers for Biomedical Computing accessible to a wide audience. Finally, When Baker realized that APBS
(NCBCs) as well as for other grantees. someone has to publicize the tool, build offered something new that might be
But what does it really take to turn a community of users, and support and widely useful, he says, “I took most of
an impressive algorithm into a widely maintain the tool. what I’d written at that point and just
disseminated, prolific computational In an ideal world, that “someone” deleted it and started over.” A tool that
tool? The transition might be harder would include a team of people with is going out to others has to be built
than you think. diverse skills—such as software engi- according to professional software
“Today, our software is very wide- neers, technical writers, and marketers. design principles, he says. The code
ly used, but it didn’t take off right But, in reality, it is often a scientist should be clean, bug-free, and robust;
away. It took years,” says Klaus moonlighting as all of the above. Tool and it should be built in a flexible, mod-
Schulten, PhD, speaking about the dissemination has traditionally been ular fashion so that others can add to the
molecular dynamics simulator NAMD underappreciated and underfunded, tool and adapt it to their own problems.
([Link] making it hard for researchers to dedi- “There’s a world of difference
and the molecular graphics viewer VMD cate resources to tools beyond what’s between developing code for yourself
([Link] needed for their science. Fortunately, and developing code that you want to
which together have more than this situation is changing—with initia- distribute,” Schulten agrees. Establishing
tives such as the NCBCs that recog- the proof of concept takes 10 percent of
nize the importance of tool develop- your time, whereas adhering to profes-
ment and dissemination—but there is sional design principles takes 90 percent,
“There’s a world of still a long way to go. he says. “And it is almost impossible to
So how do scientists manage to do it convince any normal scientist to spend
difference between right? Biomedical Computation Review that 90 percent.” Professional program-
spoke to a panel of individuals who have mers helped design VMD and NAMD,
developing code disseminated popular open source bio- and they were a key factor in the tools’
medical tools to find out what it takes to success, he says.
for yourself and succeed and how they pulled it off.
DRESSING YOUR TOOL FOR
developing code LAYING THE GROUND WORK
The ingredients for successful tool
SUCCESS: ACCESSIBLE, WELL
DOCUMENTED, WITH A GUI
that you want to dissemination have to be built into the
tool’s core from the start.
To become widely used, tools also
have to be accessible—which means
“You can’t assemble a software pack- open source, portable, well document-
distribute,” says age out of a bunch of code that your ed, and user-friendly.
graduate students wrote trying
Klaus Schulten. to get their theses done. It can’t
be an afterthought,” says
Nathan A. Baker, PhD, associ-
100,000 users. “We went through a ate professor of biochemistry
long initial phase where we were close and molecular biophysics at
to failure all the time.” Schulten is pro- Washington University in St.
fessor of physics at the University of Louis. “At some point in the
Illinois at Urbana-Champaign and design process you say, ‘oh,
director of the Theoretical and other people might want to use
Computational Biophysics Group at this.’” Baker wrote APBS—a
the university’s Beckman Institute. program that solves the Poisson-
For a tool to spread, it takes more Boltzmann equation for molec-
than a good algorithm. From the start, ular electrostatics—in collabo-
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 23
“Unless you have this great 10 million dollar idea that
will make you a fortune, the last thing you want to do
is to limit access to your work,” says Erik Lindahl.
reason scientists flock to commercial ages—and then good luck reading the deterred by the lack of a graphical
alternatives for open source software is documentation.” user interface (GUI). For example,
not because of superior performance To help make the documentation Baker says of APBS: “It’s no worse
(often the opposite is true), but because more user-friendly, several of our inter- than the other command-line compu-
of a great user interface and great docu- viewees advocate “learn by example” tational biology tools. But I would say
mentation, Lindahl says. Open source tutorials, which lead users step-by-step that maybe 80 percent of our audi-
tools often fall short on these aspects. through common research problems. ence would prefer to interact with it
“I’m a sucker for good documentation. Many potential users are also in some other way.”
If there are not clear
PDFs with graphics, I’m
extremely unlikely to
use it,” says Raymond
R. Balise, PhD, a bio-
statistical programmer
at Stanford University,
who uses the open
source statistical pack-
age R, which has hun-
dreds of thousands of
users ([Link]
[Link]/). But the
best programmers are
usually not the best
writers, he says. “So
you have brilliantly
designed elegant pack-
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 25
“That really was a godsend because it’s Cilk Arts, focused heavily on web out- questions specifically about FFTW, and
basically like we have a shop and our reach. They posted benchmarks com- it was especially important to respond
shopping window is the web.” he says. paring their software with other FFT to these—having a support presence
“It’s so easy to do and you reach so implementations; added FFTW links on public forums reassures people that
many people.” on websites that list FFT programs, as the software works and is actively
maintained,” Johnson says.
FFTW is now downloaded
about 10,000 times a month
Non-programming users of GenePattern send ([Link]
Active mailing lists and
effusive emails, says Jill Mesirov, “because they online forums help draw in new
users, support existing users,
say ‘Wow, this really lets me use all these and build a sense of communi-
ty. “I frequently get much bet-
sophisticated tools and I can do it on my own.’” ter support from open source
mailing lists than you get from
vendors,” Lindahl says.
Answering emails about the
To promote FFTW (“the Fastest well as on sites that catalog free-soft- tool also goes a long way: “We’ve
Fourier Transform in the West”)—a ware projects (such as [Link] received over 10,000 email messages
general-purpose tool that performs and [Link]); advertised on about FFTW over the past 10 years, and
Fourier transforms, which are often mailing lists; created their own mailing responded to a large fraction of them,”
used in molecular dynamics simula- list; and answered questions on online Johnson says.
tions—creators Steven G. Johnson, discussions about FFTs, including pro- Beyond the web, more “heavy-
PhD, assistant professor of applied viding links to FFTW and other free weight” outreach includes training ses-
mathematics at MIT, and Matteo Frigo, FFT software. sions, workshops, and conferences. For
PhD, chief scientist and founder of “Eventually, people began posting example, Simbios and NA-MIC as well
as other NCBCs hold training events at
conferences and stand-alone work-
shops for developers and general users.
Cytoscape developers run tutorials at
the major bioinformatics conferences
and some major disease conferences.
It’s hard to convince scientists to spend
time running training sessions rather
than improving the tool, Pieper says.
So, it’s important to involve people
who are specifically interested in and
passionate about teaching, he advises.
R, Bioconductor, and Cytoscape hold
their own annual conferences (funded
primarily by corporate sponsors and
paying participants), which help adver-
tise the tools as well as bring developers
together. “There’s definitely a commu-
nity, and the whole mentality of work-
ing as an international team is huge for
R,” Balise says.
High school teachers and college
professors also promote tools in their
classrooms. With VMD, “it became so
user friendly that it could actually
trickle down to college and high school
education,” Schulten says. “We were
very fortunate that these outreach
efforts were essentially ripped out of our
hands. So now there are many efforts,
Building a Pipeline. The GenePattern tool helps expert and non-expert users analyze genomic and pro- and we just happily receive the news.”
teomic data, while capturing the steps in a reproducible pipeline. The tool was built with non-expert Distributed computing efforts are all
users in mind, which has been a major factor in the popularity of the tool. Reproduced from Reich M, about outreach, since researchers must
GenePattern 2.0, Nature Genetics (2006) 38:500-501, supp. fig. 1. convince the general public to down-
load and run their tool. Coverage in
MAKING IT HAPPEN and in the States he’d “nudge” postdocs to turn code
they wrote for their research into for-
Successful tool dissemination can be mal GROMACS modules. Pande says
lengthy and costly, and it requires is that it’s hard he and his graduate students have to
diverse skills, such as programming, work 60 to 70-hour weeks to keep
writing, marketing, and teaching. So to get funded Folding@home going. “It’s just a lot of
how do scientists support these efforts? work to be running something like
“Up to now it’s frequently been the only for software this,” Pande says. Johnson says he and
case that you’re kind of moonlighting,” Frigo did most of the legwork for FFTW
Lindahl says. “One problem both in development,” themselves over the years, despite
Europe and in the States is that it’s many other time commitments.
hard to get funded only for software
development.” Many tools are support-
says Lindahl. Tool upkeep and dissemination are
also undervalued when it comes to
ed using bits and pieces of resources academic promotion—making it even
scrounged from science-driven grants harder to justify dedicating scarce time
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 27
and resources to these endeavors. n’t happen. It would be like, as with most into obscurity—but it is wasteful and
“Academic credit for maintaining previous funding, an afterthought in reflects poorly on the biomedical com-
software is not the same as producing some grant: ‘Oh, and by the way, I guess puting community. “There’s a huge
publications,” says BioPerl developer we’ll keep this tool limping along.’” amount of resource that goes into
Jason E. Stajich, PhD, Miller As part of the NCBCs, Simbios and making these things, and so much of it
Research Fellow in the department of NA-MIC have specific funding for tool is just lost.” Bourne says.
plant and microbial biology at the maintenance and dissemination. “One Fortunately, funding agencies and
University of California, Berkeley. of the things that’s great about the journals are beginning to acknowledge
BioPerl is a programming toolkit for NCBC program is that there’s funding the importance of tool upkeep and dis-
processing sequence data. It has been to do actual training events,” Pieper semination. In the past few years, the
cited more than 500 times says. Finally, Schulten has had long- National Science Foundation (NSF)
([Link] standing (two decades of) tool-specific and NIH have “come around to the
Stajich worked heavily on BioPerl funding through an NIH P41 grant— idea that software is not something to
before and during his graduate studies which specifically funds technology be dabbled with,” Pande says. Lindahl
but, as he transitions to a faculty posi- development. These funds allow him to has also noticed an increase in tool-
tion, he needs to focus more on his sci- hire professional programmers and run specific funding. Journals could also
ence; and many other developers are training events. help alter the reward system, Bourne
in the same situation. “We’d like to do says. PLoS is contemplating a software
more outreach, but it requires a criti- MEASURING SUCCESS section where papers will only be pub-
cal mass of people who actually have AND REFLECTING ON FAILURE lished if the software is deposited in an
time to do that,” he says. The final step in tool dissemination open source archive such as source-
To augment the piecemeal model of is evaluation—measuring how well the [Link] or [Link]. Online
tool dissemination, some groups have efforts are going. journal editors or readers could simply
formed non-profits. For example, “It is extremely difficult to measure add a comment to papers when the
Stajich and his colleagues formed the the popularity of a free software project software is no longer available, Bourne
Open Bioinformatics Foundation, like FFTW,” Johnson says. Citations says. “That would sort of be a black
which provides infrastructure for provide a rigorous measure of success, mark against the author, so I think that
BioPerl and related projects, such as but these take time to accumulate. So, might encourage the author to make
BioJava and BioPython. Similarly, the our interviewees also track softer meas- the software available longer.”
Cytoscape Consortium provides an ures including: registered users, down- Even with more incentives and
umbrella for the institutions involved loads, mailing list subscribers, mailing resources, tool dissemination will still be
in Cytoscape core development. The list activity, Web site visits, conference a challenge. Despite sufficient resources
non-profit model can help with logis- attendees, and the number of plugins and a proven track record in tool dis-
tics, including accepting donations and added to a tool. semination, Schulten says his latest
running conferences. This article focuses on tools that tool, BioCore ([Link]
Other tools in this article have succeeded. But, for every success story, Research/biocore/), is teetering on the
managed to obtain tool-specific fund- many more tools have failed. In a edge of failure. BioCore is a collabora-
ing, which was likely instrumental in recent editorial in PLoS Computational tive work environment for biomedical
their success. For example, APBS, Biology, founding editor-in-chief research, supporting tasks such as co-
GenePattern, and some members of Philip E. Bourne, PhD, a professor of authoring papers and sharing molecular
the Cytoscape Consortium have been pharmacology at the University of visualization results. The program hasn’t
funded through NIH’s R01 program for California, San Diego, and his col- taken off yet, in part because scientists
“software development and mainte- leagues describe their efforts to track are reluctant to try new technology, he
nance” (which has been available down 14 software programs (for parti- says. But Schulten is determined to
since 2002). GROMACS has also tioning proteins into domains) showcase the tool more and run more
obtained recent funding through the described in published papers. Eight training events. “We have to put more
European Union. The funding gives us programs were not even accessible in a energy into these efforts,” he says.
the ability to reply to user requests usable form, let alone widely used and Success requires persistence, Lindahl
within 24 to 48 hours and to develop popular. Given the difficulty of the agrees. “Don’t give up in the begin-
tutorials, Baker (of APBS) says. task and the lack of rewards, it’s not ning. It takes a while to build these
“Without that funding, that just would- surprising that so many tools languish communities.” ■
Network-based Approaches
to Prediction of Disease Genes
T
he recent surge of high-through- naling, or other), which control a large
put experimental data, such as set of genes differentially expressed in
gene expression microarrays, the disease state. The third, the focus one particular disease phenotype (P).
offers a profound opportunity to gain a here, relies on the fact that interaction Formulaically, this test is represented as
more detailed understanding of the networks are themselves dynamic and the difference (⌬I) between Iall(G1;G2)
genes involved in the progression of may change from a normal to disease and Iall-P(G1;G2), where Iall includes all
disease. While initial analyses of these state. Thus, if one identifies interactions sample points, and Iall-P excludes the phe-
data used statistical techniques to iden- that have actually changed between notype P. Biologically, a positive or nega-
tify genes capable of distinguishing dis- phenotypes, one might then work back- tive ⌬I implies that these two genes have
ease tissue from normal (biomarkers), wards to identify genes that could prove gained or lost an interaction in the phe-
researchers are now turning to the promising for further investigation. notype P respectively (e.g., an oncogene
analysis of gene interaction networks to We will detail two examples of the “loses” its ability to be regulated in can-
address this problem. third category, both of which inciden- cer). The genes participating in a statisti-
Gene interaction networks may be tally use an information-theoretic cally significant number of these interac-
developed from several sources includ- approach. The first defines a concept tions are then selected. When applied to
ing manual curation, high-throughput called synergy, which measures the coop- data from three primary B cell lym-
experiments (such as yeast 2-hybrid), erative effect of two variables on the phomas, IDEA correctly predicted the
literature mining and reverse engineer- state of a third. The two variables in this known oncogenes reported in the litera-
If one identifies interactions that have actually changed between phenotypes, one might
then work backwards to identify genes that could prove promising for further investigation.
ing algorithms. They can include many case are genes (G1 and G2), and the ture (e.g., MYC in Burkitt’s Lymphoma),
different types of interactions as well third is a binary state variable represent- as well as effector genes not identified by
(complexes, regulatory, signaling, etc). ing disease or normal (D). Formulaically, differential expression analysis.
Integrating and analyzing all of this this can be represented as the difference These network-based approaches,
information to discover genes relevant to between I(G1,G2;D) (the cooperative along with others, have shown promise
disease requires network-based algo- effect) and the sum I(G1;D) + I(G2;D) in more accurately delineating the mech-
rithms. Thus far, such algorithms fall (the individual effects), where I is mutu- anisms of disease progression. Like any
into three general (though not necessar- al information. Biologically, synergistic new class of methods, however, there are
ily mutually exclusive) categories. The interactions imply that the combined drawbacks. First and foremost, there is no
first predicts protein complexes, rather state of the two genes affects disease, “gold standard” of gene interactions that
than individual genes, associated with while individually the genes have a far can be used, although the knowledge
the disease phenotype. The second iden- lesser or no effect. This algorithm com- base is growing rapidly. They often
tifies key regulators (transcriptional, sig- putes this quantity across all gene pairs require large training sets or sample diver-
represented on the input microarray sity to be effective, which may not always
data, and a “synergy network” is generat- be available. Lastly, computational com-
DETAILS ed from the highest scoring interactions. plexity may limit their applicability.
When applied to publicly available Nevertheless, the application of net-
Kartik Mani received his PhD in
prostate cancer data, this approach works and these algorithms to the identi-
Biomedical informatics at Columbia
showed the RBP1I gene participating in fication of disease-causing genes remains
University, working in the Multi-Scale
Analysis of Genomic and Cellular
a large number of synergistic interac- an exciting new area of computational
Networks (MAGNet) Center under the tions. This finding along with others biology. Expect to see several new net-
direction of Dr. Andrea Califano. His indicated that the progression of prostate work-based approaches emerge as the
research focused on the application of cancer is linked with oxidative stress and body of high-throughput and interac-
interaction networks to gene-disease inhibition of the apoptosis pathway, con- tion-based data continues to grow.
association, and culminated in the sistent with previous hypotheses.
development of the IDEA algorithm The second algorithm, Interactome REFERENCES
described above. He is currently Dysregulation Enrichment Analysis 1. Watkinson, J., X. Wang, et al.
pursuing his MD at the Albert Einstein (IDEA), computes the mutual informa- (2008). BMC Syst Biol 2: 10.
College of Medicine in Bronx, NY. tion between two genes across a large, 2. Mani, K. M., C. Lefebvre, et
diverse dataset, including or excluding al. (2008). Mol Syst Biol 4: 169. ■
Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 29
Nonprofit Org.
U.S. Postage Paid
Permit No. 28
Palo Alto, CA
seeing science
SeeingScience
BY KATHARINE MILLER
Visualizing
Ventricular Fibrillation
U
nsynchronized twitching of the heart’s ventricles—known as ven-
tricular fibrillation—kills about 300,000 Americans yearly. Its
underlying cause: electrical spiral and scroll waves that
propagate through the heart. Simulation and visualization are
playing an important role in understanding that process.
In a novel approach to a review of the research, Flavio Fenton,
PhD, and Elizabeth Cherry, PhD, research associates in biomedical
sciences at Cornell University, simulated and visualized what’s cur-
rently known about how electrical spiral waves propagate through
the heart to cause tachycardia (rapid heart rate) and fibrillation.
The work was published in the December 2008 Visualization in
Physics focus issue of the New Journal of Physics. ■
Biomedical computational tools enable complex simulations and analyses that are crucial for understanding intricate biological systems. Their development and adoption ensure that researchers have access to state-of-the-art technologies, facilitate interdisciplinary collaboration, and promote discoveries that can significantly advance human health. These tools are crucial for addressing the growing complexity in biology and medicine .
Jerome Mettetal's team applied an engineering approach coupled with computer modeling and microfluidic arrays to study yeast cells. By confining yeast cells in chambers and feeding them in controlled, cyclical patterns, they investigated how yeast responds to varying osmotic pressures by adding bursts of salt. This method uncovered new roles for three different negative feedback loops in cells' equilibrium processes .
Studying yeast cell circuitry using information-processing perspectives and temporally varying inputs offers insights into key biological pathways and regulatory mechanisms, which can be extrapolated to more complex organisms. Such methods can reveal hidden dynamics in cellular regulation, thus contributing to broader applications in understanding human diseases and cellular responses .
The 'breaking point' methodology identifies critical signaling molecules by driving models beyond typical parameter ranges. This highlights previously unknown roles or interactions within signaling networks, facilitating hypothesis generation and verification. The approach can uncover novel insights into complex systems where traditional models fail, enhancing comprehensive understanding of cellular processes .
The digital embryo allows researchers to analyze developmental processes like zebrafish germ layer formation, which occur more synchronously than previously thought. Additionally, overlaying genomic data with the digital embryo can help identify genes that regulate vital processes, such as organ formation. This tool promotes advancements in tissue engineering and the study of tumor growth .
Computational models are used to simulate cellular responses by driving components beyond their observed experimental ranges to determine the weakest links, or proteins, causing computation failure. This technique highlights critical kinases leading to cytokine-induced apoptosis, challenging previously held models and enhancing understanding of protein roles in cell death mechanisms .
Disseminating biomedical tools involves ensuring the software is robust, well-documented, accessible, and supported, which often lacks funding and recognition. Overcoming these challenges requires building dissemination elements from the start, like including diverse skills in software development, documentation, and community building, as exemplified by the initiative under NCBCs .
NCBCs develop computational tools that facilitate research across various diseases by providing platforms for data analysis and integration. By enabling disease modeling, data sharing, and enhancing computational methods, these tools can transform biomedical research, improve clinical practices, and support personalized medicine approaches, showing their potential for widespread healthcare advances .
Fault diagnosis identifies critical molecules in cell pathways by drawing parallels with fault detection in electronic circuits. Highlighting crucial pathways can lead to more precise medicines by targeting safer molecules essential for cell function. This approach can reduce the toxicity of new drugs under trial by avoiding the targeting of molecules that are vital to cellular processes, paving the way for safer therapeutic interventions .
Integrating genomic data with clinical records allows for comprehensive large-scale studies that can identify reproducible patterns in genomics and clinical outcomes, potentially leading to the early detection of adverse drug events and comparisons of therapeutic efficacy across populations. This integration facilitates more cost-effective research and enhances personalized medicine by leveraging entire healthcare systems as study units .