0% found this document useful (0 votes)
11 views32 pages

Biomedical Ontology Development Insights

Uploaded by

kaksdennis
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views32 pages

Biomedical Ontology Development Insights

Uploaded by

kaksdennis
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

PLUS

Tool Dissemination
Doing It Right

NCBC UPDATE:
Shedding New Light On

Winter 2008/2009
contents
ContentsWinter 2008/2009 Winter 2008/2009
Volume 5, Issue 1
ISSN 1557-3192

Executive Editor David Paik, PhD


Managing Editor Katharine Miller
FEATURES Associate Editor Joy Ku, PhD
Science Writers

13 NCBC Update:
Shedding New Light on Biological Complexity
Katharine Miller • Kristin Sainani, PhD
Cassandra Brooks • Emmanuel Romero
Hadley Leggett, MD • Kayvon Sharghi
Lisa Grossman • Lizzie Buchen
BY KATHARINE MILLER Michael M. Torrice • Michael Wall, PhD
Molly Davis • Stephanie Pappas
Community Contributors
Tool Dissemination—Doing It Right
21 BY KRISTIN SAINANI, PhD
Mark Musen, MD, PhD
Kartik Mani, PhD
Joy Ku, PhD
Layout and Design
DEPARTMENTS Wink Design Studio
1 GUEST EDITORIAL Printing
It Takes a Village: Building the Next Advanced Printing
Generation of Biomedical Ontologies BY Editorial Advisory Board
MARK A. MUSEN, MD, PhD Russ Altman, MD, PhD, Brian Athey, PhD,
Dr. Andrea Califano, Valerie Daggett, PhD,
Scott Delp, PhD, Eric Jakobsson, PhD,
3 SIMBIOS NEWS Ron Kikinis, MD, Isaac Kohane, MD, PhD,
Stop Wheel Reinvention, Mark Musen, MD, PhD, Tamar Schlick, PhD,
Share Your Simulations Jeanette Schmidt, PhD, Michael Sherman
Arthur Toga, PhD, Shoshana Wodak, PhD,
BY JOY P. KU, PhD John C. Wooley, PhD
For general inquiries,
5 NEWSBYTES subscriptions, or letters to the editor,
BY HADLEY LEGGETT, MD, LISA GROSSMAN, visit our website at
MICHAEL WALL, PhD, CASSANDRA BROOKS, [Link]
MICHAEL M. TORRICE, PhD, KAYVON SHARGHI,
Office
EMMANUEL ROMERO, LIZZIE BUCHEN, STEPHANIE Biomedical Computation Review
PAPPAS, MOLLY DAVIS Stanford University
• Modeling Cracks in Clogged Arteries 318 Campus Drive
• Modeling Muscles From the Inside Out Clark Center Room S231
• “Digital Embryo” Created Stanford, CA 94305-5444
• The Circuitry of Yeast Biomedical Computation Review
• Watching a Molecule Bind is published quarterly by:
• Identifying a Cell’s Weakest Link
• Diagnosing Cell Circuitry
• Cancer’s Signature—Written in Blood
• Blurring Data for Privacy and Usefulness
An NIH National
• Modular Modeling Center for Physics-
Based Simulation of
29 UNDER THE HOOD | Network-based Approaches to Biological Structures
Prediction of Disease Genes BY KARTIK MANI, PhD Publication is made possible through the NIH
Roadmap for Medical Research Grant U54
30 SEEING SCIENCE | Visualizing Ventricular Fibrillation GM072970. Information on the National Centers
for Biomedical Computing can be obtained from
BY KATHARINE MILLER [Link] The NIH
program and science officers for Simbios are:
COVER ART: Created by Rachel Jones of Wink Design Studio using a collection Peter Lyster, PhD (NIGMS)
Jennie Larkin, PhD (NHLBI)
of images from the NCBCs. Robotic arm is © Eraxion | [Link] Jennifer Couch, PhD (NCI)
Semahat Demir, PhD (NSF)
Peter Highnam, PhD (NCRR)
Jerry Li, MD, PhD (NIGMS)
Yuan Liu, PhD (NINDS)
Richard Morris, PhD (NIAID)
Joseph Pancrazio, PhD (NINDS)
Grace Peng, PhD (NIBIB)
Nancy Shinowara, PhD (NCMRR)
David Thomassen, PhD (DOE)
Ronald J. White, PhD (NASA/USRA)
Jane Ye, PhD (NLM)

ii BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


guest editorial
GuestEditorial
BY MARK A. MUSEN, MD, PhD

It Takes a Village:
Building the Next Generation
of Biomedical Ontologies
A
lthough the notion of ontology has been around annotate the ontology via a Web-based wiki and
since Aristotle, the perceived need to develop suggest changes and extensions, although again, the
ontologies in biomedicine has accelerated in modification of the actual ontology content will be
recent years as investigators attempt to make sense of the channeled through a set of trained individuals who
terabytes of high-throughput data that are now finding understand principles of knowledge representation and
their way into public repositories. While the number of the use of knowledge-editing tools.
biomedical terminologies and ontologies continues to Probably the best exemplar of an open, nearly demo-
increase as new areas of biomedical content become for- cratic ontology-development initiative is the Open
malized, the creation and annotation of these resources Directory Project (ODP). Founded more than 10 years
can’t quite keep up. The flood of information may neces- ago, the ODP has enlisted more than 75,000 volunteers
sitate a new approach involving vastly more ontology to flesh out the extensive open-content ontology of Web
developers. It may, in fact, take a village. pages that has been adopted by Google, Yahoo!,
The construction of biomedical ontologies has long Netscape, and a host of other companies. The ODP has
been a cottage industry, with even vast systems such as generated an enormous ontology (commonly known as
SNOMED (the Systemized Nomenclature of Medicine) dmoz) that provides standard, categorized entrée to vir-
initially representing the handiwork of a very small group tually all the content on the Web. All of us use dmoz,
of dedicated individuals. Venerable ontologies such as the perhaps unknowingly, every time we browse the Web by
Foundational Model of Anatomy and the NCI Thesaurus categories in Google and Yahoo!, rather than searching
represent the work of a surprisingly small set of develop- the Web’s free text for particular terms. Embracing
ers. Nevertheless, as the demand for ever larger and more everything imaginable that a user could search for, dmoz
granular ontologies accelerates, and as large-scale systems is a remarkable demonstration of how scalable ontology
such as the International Classification of Disease are engineering can be, particularly when volunteers step
being reengineered, the scientific community has increas- forward to provide fine-grained descriptions of their par-
ingly raised concerns about whether ontology develop- ticular areas of personal interest.
ment ultimately can be a scalable enterprise. Practical
ontologies comprise tens of thousands of con-
cepts, and a handful of individuals can
never have personal knowledge
of everything that needs to
be represented in such
a system.
To address this
problem, workers
in biomedicine
are attempting to
democratize the
development of
large-scale ontolo-
gies. The engineering
of the Gene Ontology,
for example, has been char-
acterized by an open development
process to which nearly anyone can
contribute. The actual editing of the Gene
Ontology content, however, is still performed by only a At [Link], the user is presented with a list of pos-
handful of trusted curators. The National Cancer sible ontologies to explore and visualize. This screenshot of the
Institute is experimenting with an open process for Cellular Components ontology shows the relationship between major
extensions to the current content of the NCI Thesaurus components and some of the minor components. Any registered user
via the BiomedGT initiative. Here, nearly anyone can can comment on any ontology.

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 1
guest editorial
The dmoz ontology is very simple in its structure, and
lacks the rich semantics of ontologies developed in for-
It is clear that the
mal knowledge representation systems such as the Web
Ontology Language (OWL). When the developers of ontology-development
dmoz make modeling errors, the consequences are
unlikely ever to impede the advancement of science or community needs at least to
to threaten lives. Nevertheless, the dmoz ontology
stands as a stunning example of how legions of volun- experiment with new methods
teers can be mobilized to generate an enormous and
undeniably useful ontology. Imagine if the lessons of
dmoz could be applied to SNOMED or to BiomedGT!
of ontology engineering
At the National Center for Biomedical Ontology
(NCBO), we are experimenting with ways in which the
that can scale to future
biomedical community can take an active part in con-
tributing to the construction of scalable ontologies and biomedical requirements.
controlled terminologies. Our BioPortal system allows any
registered user to comment on any ontology in our dis- maintain the quality of ontologies if the development
tributed repository, to comment on the comments left by process is democratized. Organizations such as the Open
other users, and to demonstrate how the elements of one Biomedical Ontologies (OBO) Foundry have been
ontology may relate to those of another. We have used established under the assumption that there must always
this capability extensively in the engineering of the be central management of ontology development to
Biomedical Resource Ontology used to describe the ensure the quality of the content. And yet there contin-
online software and data resources developed by the ue to be too much data, too many medical records, and
National Centers for Biomedical Computing and by the too many experiments for the ontology-development
recipients of Clinical and Translational Science Awards. community to keep up with existing needs.
BioPortal, at present, does not play a role in completely I don’t know whether the dmoz approach will really
open ontology editing, however. be practical in biomedicine, but it is clear that the ontol-
There are very legitimate concerns about how we can ogy-development community needs at least to experi-
ment with new methods of ontology engineering that
can scale to future biomedical requirements. Surely there
DETAILS are ways to take advantage of the expertise distributed
BioPortal: [Link] among all biomedical investigators in a way that will
The Open Director Project: [Link] overcome many of the limitations of centralized ontology
curation. Workers at NCBO are extremely excited about
Mark A. Musen, M.D., Ph.D. is Professor of Medicine the possibilities that new technology might provide in
(Biomedical Informatics Research) and Computer enabling this more open approach to ontology engineer-
Science at Stanford University. He is Director of the
ing. Experimentation with community-based ontology
Stanford Center for Biomedical Informatics Research
development not only may accelerate the engineering of
and principal investigator of the National Center for
Biomedical Ontology (NCBO).
badly needed ontology content, but also can provide a
laboratory for the study of new mechanisms for collabo-
ration and interaction in biomedicine. ■

A Note from the Managing Editor:


Starting in our next issue, we will launch a new "Debate"
T hanks to all who participated in the BCR survey. Your
names were entered in a drawing for an iPod shuffle
which went to Alan Villalobos from DNA2.0. The survey
column, starting with the topic selected by the survey
respondents: "To Mine or Not to Mine: Are clinical data
results are helping us to plan for the future. repositories useful sources of untapped discoveries awaiting
data-mining algorithms or are they too noisy and messy."
If you didn't get a chance to answer the survey, you can
still give us feedback on the magazine by visiting Best,
[Link] and clicking
on the "Feedback" link.
Kathy Miller MANAGING EDITOR

2 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


simbios news
SimbiosNews
BY JOY P. KU, PhD

Stop Wheel Reinvention,


Share Your Simulations!
S
imbios has built a new publication University’s mechanical engineering
repository that links publications department, to create a publication proj-
to the research data and software ect on [Link]. She is sharing 32
behind them. The goal: to encourage walking simulations used to analyze how
and facilitate replication of published muscle functions change with walking
results and to foster use of what has speed in children. It’s the largest number
already been accomplished rather than of simulations ever included in a mus-
leaving others to reinvent the wheel. cle-driven simulation study, yet Liu sees
The repository is built upon [Link]— it as just the beginning of further
Simbios’ web-based infrastructure that research rather than the end of the line.
provides open access to simulation soft- “The simulations themselves could
ware tools and models—making it easy become the starting point for a number
to use and accessible to all. of other studies,” Liu says. “There’s no
“The publication repository is more reason why people should have to recre-
than just the collection of data, models, ate simulations that already exist.”
and software used in the publications,” Dahlia Weiss, a doctoral student in
says Jeanette Schmidt, PhD, Executive structural biology and chemistry at
Director of Simbios. “It provides the Stanford University, has a similar perspec-
means for others to reproduce and build tive. She established a publication project
upon the results of your publication.” for her article comparing her Climber
software tool against four other tools for
GIVING YOUR PUBLISHED interpolating between two molecular
RESEARCH A FUTURE structures as one morphs into the other.
Historically, when researchers have Climber, based on a non-linear interpola-
come across papers describing potentially tion method, turned out to be very good
useful software or data, their chances of at producing intermediate structures for
actually getting their hands on that soft- very large, complicated changes.
ware or data were hit or miss. The student “Knowing that we have a really
who did the research might have moved good tool and not making it publicly
on, or the software developer might want available just seems really pointless,”
to clean up the code first and take months says Weiss. She thinks Climber would
(or longer) to do so. The Simbios publica- be useful wherever high fidelity inter-
tion repository for physics-based simula- mediate structures are required, not just
tions of biological structures addresses this for looking at structural movement.
problem by providing a simple way to
share and access the software, data, and REPLICATING RESEARCH GOES
other materials that support a particular BEYOND SOFTWARE SHARING
research paper. It means that all the hard But the [Link] publication reposi-
work behind the paper—the hours of cod- tory is not just about software sharing. It
ing, the repetitive experiments to get use- supports and encourages sharing any-
able data—is captured and can easily thing needed to replicate research
enable future research. results. For example, Stanford University
That’s what motivated May Liu, researchers Yuan Yao, PhD, a post-doc-
PhD, a recent graduate from Stanford toral fellow in the math department, and
On [Link], May Liu created a publication project to share the three-dimensional sim-
ulation results from her latest publication analyzing eight subjects walking at four speeds
(very slow, slow, free, and fast). Shown here are still images from simulations of a rep-
resentative subject. The goal of the publication projects is to encourage and facilitate
replication of published results. Courtesy of May Liu. Reprinted with permission from Liu,
MQ, et al., Muscle contributions to support and progression over a range of walking
speeds, Journal of Biomechanics (2008) 41:3243–3252.

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 3
simbios news
Xuhui Huang, PhD, a research associate chief of the journal PLoS Computational A FIRST STEP TOWARD THE
in the bioengineering department, and Biology. “So by making the software PUBLICATION OF THE FUTURE
their colleagues developed Mapper, a available, researchers open up that possi- Bourne sees the Simbios publica-
tool that improves detection of low-den- bility to benefit from what other people tion repository as a very positive step:
sity states within a massive amount of do with the software as well.” “It actually speaks to the dream that I
data. After creating a project for Mapper Bourne acknowledges that the cur- have.” He envisions all aspects of
on [Link], they submitted a paper rent system does not always reward the research being accessible, with the
showing how Mapper could be used to work involved in preparing and support- paper being an access point to the
identify intermediate stable states during ing open-source software: answering experiment. From the paper, a
the RNA hairpin folding process, a diffi- questions from software users and provid- researcher could retrieve and manip-
cult task when those states represent ing documentation, examples, and tuto- ulate the associated data, and possibly
only two to three percent of the whole rials. All of that effort takes time that discover new links and relationships
data set. In order for someone else to could be spent doing research that would via the data and tools—not just the
replicate that research, Yao and Huang generate more publications—the metric paper citations—enhancing the
posted (on [Link]) not only the by which academics are primarily judged. research process.
Mapper software, but also the project’s To address this concern, Bourne says, Bourne observes that there are an
input data and instructions about how to PLoS is considering having a special sec- increasing number of efforts to cap-
use that data with Mapper. tion that only publishes articles report- ture this whole research work flow
“To reproduce the results from a ing on open-source compu- process. The goal of his BioLit
paper is not an easy task,” Huang says. tational biology software ([Link] project
“You need all the components togeth- that has been deposit- is to connect open access
er—the data, the program, your param- ed in an established articles with information in
eters, instructions—so that people can repository. existing biological databas-
easily reproduce the results. [Link] es, such as the Protein Data
provides such a platform, especially Bank (PDB). Another exam-
with this publication mode.” ple is the Insight Journal
While the information could have ([Link]
been posted on his own web- an open access on-line
site, Yao says that researchers publication focused on
from other fields would medical image pro-
not think to look there. cessing and visualiza-
For an interdisciplinary tion where authors are
field, a common platform encouraged to provide
like the [Link] publi- the data and software associ-
cation repository is particu- ated with their papers.
larly valuable. “Most people get togeth-
er because of content,”
REWARDS Bourne says. Efforts such as the
FOR SHARING [Link] publication repository
While some researchers think that provide the infrastructure to share
sharing their software or data means giv- different types of information,
ing up their competitive advantage, oth- enabling a dialog between the
ers believe that it is a great way to build Molecular dynamics were used to simulate the fold- people who are using and
a successful career. “Careers often come ing of the RNA hairpin structure (above), generating developing the content.
from the application of software to make hundreds of thousands of molecular structures. “My sense is that in the next
new discoveries in the life sciences,” says Mapper was then used to sort through all that data ten years, scientific discourse is
Philip Bourne, PhD, a professor in phar- to identify the relatively infrequent intermediate going to change very dramati-
macology at the University of California states that occur during the folding process. cally as a result of these kinds of
at San Diego, and founding editor-in- Courtesy of Yuan Yao. things.” Bourne says. ■

DETAILS
Learn more about the projects mentioned in this article.
Walking simulations: [Link]
Mapper: [Link] Simbios ([Link]
is a National Center for Biomedical Computing
Climber: [Link] located at Stanford University.

4 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


NewsBytes
Modeling Cracks ened, fat-lined arteries—sometimes can only describe an artery’s shape, not
with disastrous results. Now, structural its mechanical properties, such as resist-
in Clogged Arteries engineers have created the first fully ance. And these parameters vary from
Every year, doctors in the United three-dimensional model to predict how patient to patient, depending on the
States perform more than a million arteries fracture under such stress. extent of arterial disease. To get individ-
angioplasties: By inflating a tiny balloon “Once you have the true geometry [of ualized data, Pandolfi says, one must test
inside a clogged artery, cardiologists can the artery], this model applies pressure a piece of artery outside the body or do
compress fatty plaques and restore blood to simulate the presence of a balloon an in situ experiment—dangerous proce-
flow. But the balloon also applies high and evaluate the possibility of breaking dures in a patient with unstable arteries.
pressure that can crack the wall of hard- the plaque or rupturing the artery walls,” “The key thing is to get more data
says author Anna Pandolfi, and do more tests on human tissue,”
PhD, an associate professor of says Gerhard Holzapfel, PhD, profes-
structural mechanics at the sor of biomechanics at Graz University
Politecnico di Milano in Italy. in Austria who published his own
The research appears in the model of arterial fracture last year.
October 2008 issue of Computer “When we throw in more data,” he
Methods in Biomechanics and says, “I am very certain we can actual-
Biomedical Engineering. ly define a more optimal stent, on a
In lab experiments, arteries computer, for a specific lesion.”
tend to break when exposed to —By Hadley Leggett, MD
pressures of 0.3 megapascals or
more—about 20 times the Modeling Muscles
average human blood pressure. From the Inside Out
But angioplasty can easily gen- A new model of skeletal muscle
erate such forces, and some starts from the micro-mechanical prop-
areas of diseased arteries are erties of the smallest possible unit—the
particularly fragile. sarcomere—and builds up to the mus-
To better understand how cle fibers and then to the muscles
arteries fracture, Pandolfi and themselves. In addition, it places the
her colleague Anna Ferrara, fibers in their natural context—within
PhD, of the Politecnico di surrounding soft tissue. The effort
Milano, combined high-reso- brings a new degree of flexibility and
lution magnetic resonance realism to muscle simulation.
imaging (MRI) of a patient’s “The idea behind micromechanical
arteries with a model they pre- modeling is to imitate the behavior of
viously developed to describe the material as well as possible,” says
fracture in brittle solids, such lead researcher Markus Böl, PhD, pro-
as glass. Using a technique fessor of mechanics of polymers and
called finite element analysis, biomaterials at the Braunschweig
they divided the artery wall University of Technology in Germany.
into small volumes and “We’re trying to include all the micro-
assumed each chunk had a uni- parameters we can. In this way we do
form behavior. Then they sim- not have to fit the material behavior to
ulated several high-pressure the experimental data.” His work
scenarios and monitored the appears in the October 2008 issue of
Evolution of cracks in a clogged human artery depends on the evolution of arterial cracks. Computer Methods in Biomechanics and
geometry of the arterial wall and the pressure inside the artery. “What we got was an inter- Biomedical Engineering.
In the first simulation (left), a 40-percent-narrowed artery frac- esting correspondence with Scientists started making mathe-
tures at a blood pressure of 260 mmHg. In the second simulation the medical data,” Pandolfi matical models of muscles in the 1920s.
(right), an 80-percent-narrowed artery fractures at a blood pres- says: As others had seen in a Most attempts to date were one-dimen-
sure of 380 mmHg. Colors show the distribution of stress on the clinical setting, cracks usually sional, and they ignored the soft tissue
arterial wall, measured in megapascals. Courtesy of Anna began at the edge, or “shoul- surrounding muscle fibers, Böl says.
Pandolfi. Reprinted from Pandolfi A and Ferrara A, Numerical der,” of a fatty plaque. Also, they usually were built from the
modeling of fracture in human arteries, in Computer Methods in But, Pandolfi says, the model outside in: Scientists would look at the
Biomechanics and Biomedical Engineering (2008) 11(5):563. has limitations: An MRI scan way a muscle behaved and tweak their

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 5
NewsBytes
model’s parameters (such as the num- work is still in progress,” he cautions. cells and a beating heart. To meet that
ber of contractions per second) until it The new model can simulate any challenge, Keller and his colleagues
matched the behavior. This led to some biological tissue that contracts, not just developed a new technique called digital
accurate but limited simulations. skeletal muscles, says Ellen Kuhl, PhD, scanned laser light sheet fluorescence
Böl’s work builds muscles from the professor of mechanical engineering at microscopy (DSLM).
inside out. He uses the finite element Stanford University. Kuhl was so DSLM generated a three-dimension-
method, originally developed by aero- impressed that she is now working with al image of the embryo by combining
space engineers to design planes, to Böl to model heart tissue, with the goal about 400 pictures taken along slightly
divide a muscle into discrete parts that of helping researchers develop a patch different planes. The team repeated this
each behave differently. Previous finite to replace dead tissue after a heart process every 60 to 90 seconds, tracking
element muscle models used a continu- attack. “I think the cardiac application changes as the zebrafish developed. In
um-based approach, which lumped all is even more sexy, because many more 24 hours, this amounted to about
muscle fibers together and treated them people could benefit from it,” she says. 400,000 images for each embryo.
as a single unit. But Böl gets into the —By Lisa Grossman To deal with this deluge of data—
nitty-gritty of each tiny fiber. In essence, three terabytes per embryo—the
his modeled muscles behave like a bunch “Digital Embryo” Created researchers developed a computational
of ropes of different thicknesses attached How does a humble zygote grow into pipeline. They wrote algorithms defin-
at the same point. Because the model a fully functioning animal, billions or ing the structure of cell nuclei, then
describes each rope independently, Böl trillions of cells strong? This question ran the microscopy data through a net-
can plug in any parameters he wants and has intrigued biologists for centuries. work of more than 1000 computers at
get realistic behavior back out. Now scientists have generated the first EMBL and the Karlsruhe Institute of
In his model, Böl splits the muscle complete developmental blueprint of a Technology in Germany.
into an active element (the contractile vertebrate—a “digital embryo” map- The computational analysis picked
muscle fibers) and a passive one (the ping the positions, divisions, and out every nucleus. Keller’s team then
incompressible tissue that surrounds movements of every cell during the first processed this information into compre-
them). Putting the “ropes” into the 24 hours of a zebrafish’s life. hensive databases of cell positions, divi-
realistic environment of soft tissue “Such reconstruction of a complex sions, and migratory tracks. In all, they
yields a more complete picture, he says. vertebrate embryo had not been achieved catalogued 55 million nucleus entries.
The model has both experimental before,” says Philipp Keller, a PhD candi- Digitizing the data was key.
and clinical value, Böl says. Scientists date at the European Molecular Biology “Microscopy tells you about phenom-
will use it to test the properties of living Laboratory (EMBL) in Heidelberg, ena from a qualitative point of view,”
muscle, or to help doctors design unique Germany. Keller is lead author of the Keller says. “But with digital
treatments for patients, he believes. He paper, which appeared in the October 9, embryos, we can count the number of
is now working with sports doctors to 2008 issue of Science. cells that are involved in a process
refine and implement his approach. “But Developmental biologists have long and see what they do.”
I have to say, these are first trials and coveted such a tool, but imaging a com- The digital embryo has many poten-
plex organism’s growth pres- tial uses. For example, the researchers
ents a serious hurdle. used it to determine that zebrafish germ
After just one day, for layers—which eventually give rise to
example, a zebrafish all of the fish’s tissues—form more syn-
already has 20,000 chronously across the embryo than pre-
viously thought.
Keller also envi-

This simulation of muscle contraction shows the process from a sin-


gle twitch (at left) to continuous clenching (at right). Reprinted from
Böl M and Reese S, Micromechanical modeling of skeletal muscles
based on the finite element method, Computer Methods in
Biomechanics and Biomedical Engineering (2008) 11(5):489-504 with
permission from Taylor & Francis, publishers.

6 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


sions applications in tissue engineer- “This paper is groundbreaking,” said
ing and the study of tumor growth. Kees Weijer, PhD, professor of devel-
Overlaying the digital embryo with opmental physiology at the University
genomic data also could be powerful, of Dundee in Scotland. “And making
he adds. Researchers could learn all the data available is very helpful
which genes regulate vital develop- since these coordinates will be used to
mental processes, such as organ forma- compare the development of mutants.”
tion. To encourage such progress in —By Michael Wall, PhD
multiple fields, the researchers made
their data public. The Circuitry of Yeast
For centuries, yeast has helped scien-
tists understand how cells work. Now,
two inventive teams have applied an
engineering approach coupled with
computer modeling to reveal new
details about key biological pathways by
“Microscopy which yeast cells regulate themselves in
a changing environment, as reported in
tells you about the January 25, 2008 issue of Science and
the August 28, 2008 issue of Nature.
phenomena “What’s interesting to me was look-
ing at this biological system from an
information-processing perspective,”
from a qualitative says Jerome Mettetal, PhD, a physicist
at the Massachusetts Institute of
point of view,” Technology and lead author of the
Science paper. “By applying temporally
Philipp Keller says. varying inputs, you can find out a lot
about the system that you wouldn’t be
“But with digital able to see otherwise.”
Traditionally, biologists measure how
embryos, we can cells respond by adding or taking some-
thing away in a steady-state context. But
count the number in real cells, inputs from the environ-
ment vary constantly. To understand the
mechanisms by which cells respond to
of cells that are changes, the two teams created microflu-
idic arrays that confine yeast cells in a
involved in a chamber and feed them in regular cycles,
controlled by software. Based on the out-
process and see put, each team generated a model of the
inner workings of the cells.
what they do.” Mettetal’s team added bursts of salt to
the microfluidic array in order to tease
out how yeast responds to changes in
osmotic pressure—the salt level in the
surrounding medium. They then built a
model based on the response generated
by the yeast. When they compared their
DSLM images of a zebrafish embryo at four model to known cell responses to osmot-
different time periods, between 1.5 and 20 ic changes, they discovered new roles for
hours post-fertilization. Different colors three different negative feedback
indicate different densities of nuclei (blue loops—the processes by which a biolog-
and purple are least dense, while yellow is ical system reestablishes equilibrium.
most dense). Courtesy of Philipp Keller. The research team on the Nature

7
NewsBytes
paper applied a similar engineering
approach to better understand how yeast
a means to tease out the structure and
function of the underlying system,” says
Until now, says
cells respond to fluctuations in nutrient James Collins, PhD, professor of bio-
levels. If yeast is deprived of its favorite medical engineering at Boston Emad Tajkhorshid,
sugar (glucose), it will consume an alter- University. “I am already beginning to
native and less nutritious sugar (galac- think about how these might be inter- “nobody has been
tose). The researchers created a sinu- esting tools to use to look at other sys-
soidal input by alternately feeding and tems, bacteria in particular.” able to capture and
starving yeast of glucose on different —By Cassandra Brooks
time scales while galactose was con- describe the full
stantly present in the environment. The Watching a Molecule Bind
cells responded to long-term changes in
glucose, but not to faster fluctuations.
Like a paper clip being pulled to a
magnet, a small molecule called ADP
process of ligand
The researchers then made a model gets pulled into its port in a new simu-
based on the well-known metabolism lation. Because of a simple case of
binding to a binding
of galactose. But the experimental opposites attract, it’s the first time com-
yeast was responding much faster to the putational biochemists have successful- site while permitting
glucose fluctuations than the model ly simulated a molecule—or ligand—
predicted. “This suggested something being drawn into its binding site in an natural motion of
was crucially missing from the model,” unbiased simulation.
says co-author Jeff Hasty, PhD, associ- “Nobody has been able to capture and the ligand.”
ate professor of bioengineering at the describe the full process of ligand binding
University of California, San Diego. to a binding site while permitting natural Tajkhorshid and graduate student Yi
Studying live yeast provided the motion of the ligand,” says Emad Wang describe their simulations in the
answer: The messenger RNA necessary Tajkhorshid, PhD, assistant professor of July 15, 2008 issue of the Proceedings of the
for the galactose metabolic pathway biochemistry, pharmacology and bio- National Academy of Sciences.
was degraded when glucose was pres- physics at the University of Illinois at Tajkhorshid and Wang simulated
ent. “The most exciting thing is that Urbana-Champaign. “We think we are the binding of adenosine disphosphate
without the model, none of this would getting the most faithful representation (ADP), a molecule involved in fueling
have happened,” said Hasty. of the binding site, because in our simula- the cell, to the ADP/ATP carrier pro-
“The broader contribution of each tions, the protein is dynamic and allowed tein (AAC) located in the membrane
of these pieces will be to point to the to freely react to and establish new inter- of mitochondria—the cell’s power gen-
value of using periodic input signals as actions with the ligand as it binds.” eration plants. For ADP to be shuttled
into the mitochondria, it must first
float into a cavity inside AAC and
bind to it—an event that lasted 100
nanoseconds in the simulations.
Previously, simulations of molecular
binding have required an active force to
produce the attachment. But placing
the ligand (in this case ADP) at the
mouth of the ligand binding site (here,
the AAC cavity) in molecular dynamics
simulations is more faithful to biological
reality. Initially, Tajkhorshid thought
that the ADP would just float away.
Instead it moved right into place. He
and his colleagues found that AAC uses
a special bait to lure ADP to its binding
site: Positively charged amino acids line
the sides and bottom of the AAC cavi-
Yeast grows in a microfluidic chamber designed at the University of California, San Diego. Regular ty, creating a surprisingly strong electro-
nutritional inputs, generated in a wave-like pattern, reveal aspects of how the cells regulate their static potential that attracts the nega-
metabolism and internal environments. The green background color signals that it is a galactose rich tively charged ADP. They called this
environment. Photo credit: UC San Diego Jacobs School of Engineering. process “electrostatic funneling.” And

8 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


because of it, no additional forces are Tajkhorshid and Wang watched as ADP was
needed in the simulations of ADP pulled down into the cavity of the AAC protein. In
binding to AAC. this graphic, the AAC structure is outlined in black.
In addition, when the team scanned ADP molecules at different stages of the 100 ns simu-
the amino-acid sequences of other lation are shown in colors ranging from pink to red—
molecules that shuttle negatively pink represents ADP’s starting position and red
charged molecules across mitochon- denotes its final binding state. The strongest region of
drial membranes, they found large the protein’s positive electrostatic potential is shown
numbers of positively charged in blue mesh. Courtesy of Emad Tajkhorshid.
amino acids not present in other
membrane proteins, Tajkhorshid known as “survival stimuli.”
says. He suspects these other carri- The researchers then manipu-
ers also use electrostatic funneling lated the model to drive the activ-
to pull in their molecular quarries. ity levels of the proteins outside of
Alan Robinson, PhD, a their experimentally observed
researcher at the Medical Research ranges. When the model could no
Council Dunn Human Nutrition longer computationally fit one of the
Unit in Cambridge, U.K., says signal variables, it would stop making
Tajkhorshid has “published what looks cell death, are poorly understood. So predictions. This “breaking point” high-
like the most reasonable structure of Yaffe and colleagues Kevin Janes, PhD, a lighted the protein that caused the fail-
ADP bound to the carrier.” This struc- recent MIT graduate, and H. Christian ure. Thus, the technique acts as a sort of
ture may serve as the starting point for Reinhardt, PhD, a postdoctoral associ- high-throughput screen, revealing new
more detailed studies of how ADP ate at MIT, built a model of the cell hypotheses about proteins previously
binds to AAC and how it triggers the
protein to open, he says.
— By Michael M. Torrice, PhD “Signaling networks are so complicated
Identifying a right now that common sense doesn’t
Cell’s Weakest Link always hold true,” Michael Yaffe says.
To understand why bridges collapse
or computers fail, engineers might
create models of these systems and using carefully collected data. Included thought to have well-defined roles
push them beyond their limits. Now, in the model were nearly 8,000 meas- within the cell. The team then verified
computational biologists are using a urements of protein signals in response these hypotheses experimentally, lead-
similar approach to understand the to combinations of three cytokines ing to surprising new insights about
causes of cell death. By driving their that help dictate the fates of cells: how the signaling proteins communi-
model of the cell beyond experimen- tumor necrosis factor (TNF), known as
tally observed values of certain impor- the “death stimulus,” and epidermal
tant cellular ingredients, they push it growth factor (EGF) and insulin,
to the “breaking point”—uncovering
the weakest links. The process Breakpoint model analysis pushes cellular
revealed some new biological roles for ingredients beyond their normal ranges to
several key signaling molecules—the see which ones are critical to a particular
kinases ERK, Akt, and MK2. cellular process. Here we see fluorescent
“It showed us things that, in retro- proteins highlighting the subcellular loca-
spect, we couldn’t see looking by inspec- tion of several different key signaling mole-
tion of the original model,” says co-author cules (phosphoinositide-binding domains),
Michael Yaffe, PhD, associate professor of which function together with lipid and pro-
biology and biological engineering at the tein kinases and phosphoserine/threonine-
Massachusetts Institute of Technology binding domains, to control a wide variety
(MIT). The work was published in the of cellular events. These are the kinds of
October 17, 2008 issue of Cell. molecular interactions that could be studied
The mechanisms by which proteins using breakpoint model analysis. Courtesy
influence cytokine-induced apoptosis, or of Seth J. Field and Michael Yaffe.

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 9
NewsBytes
cate. “Signaling networks are so compli-
cated right now that common sense
doesn’t always hold true,” Yaffe says.
“The thing that makes me really stop
and pay attention is the methodology,
which I found of special note,” says
Raphael Levine, PhD, distinguished
professor of chemistry at the University
of California, Los Angeles. “Instead of
trying to see if the model can predict
something new, they tried to drive it to
say something which they know it
shouldn’t say. As a result, they were suc-
cessful in finding some new biology.”
—By Kayvon Sharghi

Diagnosing Cell Circuitry


To biologists, a computer’s mother-
board may just look like highways of cir-
cuitry connecting various chips. But if A simple model of the caspase3 network (top) shows the various regulatory molecules and
they focus harder, they might see a model their relationships to each other. Depending on which regulatory molecules are active or
for disease, according to new research. inactive, caspase3 will induce cell death. This network can be re-envisioned (below) as an
Just as a single corrupt circuit can electronic circuit after organizing previous knowledge of the molecules’ relationships using
foul a computer’s operation, a faulty Boolean logic. Algorithms applied to this circuit can predict molecules to which a pathway’s
molecule can upset a healthy body. “If signal is most vulnerable. Reprinted with permission from Abdi A, et al., Fault Diagnosis
your body is not functioning correctly, Engineering of Digital Circuits Can Identify Vulnerable Molecules in Complex Cellular
then the molecules inside your cells are Pathways, Science Signaling, (2008) 1(42):ra10.
causing the problem,” says Effat
Emamian, MD, president and CEO of death regulator caspase3, and a nerve- not a fundamental flaw,” Janes adds.
Advanced Technologies for Novel cell network called CREB. His recon- The team acknowledges these limi-
Therapeutics in New Jersey. structions used binary language to char- tations in its Science Signaling paper.
The parallels between signal trans- acterize a molecule’s state in its pathway The next step, Emamian says, is to
duction pathways in a cell and circuit as “active” or “inactive.” Relationships focus on larger networks, and not nec-
networking in a motherboard inspired between molecules were organized into essarily just signaling pathways. “We
Emamian’s team to identify defective decision-making operations using can analyze metabolic pathways, or
cell pathways in the same way that Boolean logic where each relationship pathways that also have several critical
engineers inspect faulty circuits. This contains only two possible values—on enzymes playing in the whole game.”
technique, known as fault diagnosis, or off. This allowed the researchers to —By Emmanuel Romero
can pinpoint the molecules that are write algorithms predicting which mole-
most critical to a cell’s function. cules were critical to a pathway’s smooth
Such an accurate assessment may functioning. The algorithms confirmed Cancer’s Signature—
lead to more precise medicines. Most what was known about p53 and cas- Written in Blood
new drugs in trial are toxic, Emamian pase3, but they also revealed new criti- When it comes to deciphering the
says, because they often target mole- cal molecules in the CREB network. health of the body, the blood carries a
cules essential for cell function. Fault The approach is a good start for quick- potential mother lode of protein clues.
diagnosis can reveal safer molecules to ly identifying essential points in cell net- Given the ease of extracting blood, such
target. The work appears in the October works, says Kevin Janes, PhD, assistant proteins could serve as efficient health
21, 2008 issue of Science Signaling. professor of biomedical engineering at barometers. But it’s tough to distinguish
Lead author Ali Abdi, PhD, associate the University of Virginia. But while between the multitude of proteins natu-
professor of electrical and computer Boolean logic can make good approxima- rally found in blood and those that are
engineering at the New Jersey Institute tions, it may oversimplify the relation- secreted into the blood—including those
of Technology, helped test Emamian’s ships for some networks, he says. For secreted by diseased tissue such as cancer.
theory. Abdi re-envisioned three previ- example, Emamian’s approach doesn’t Their signal may get swamped by the
ously studied cell pathways as electronic allow consideration for graded responses many other proteins present in blood,
circuits: tumor suppressor p53, cell between “active” and “inactive.” “But it’s thwarting efforts to discover useful infor-

10 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


“Figuring out which proteins are secreted
into the blood is like searching for a needle real-world data. The results, which show
promise for protecting privacy without
rendering the data set useless, appear in
in a big, big haystack,” says Ying Xu, PhD. the September/October 2008 issue of
the Journal of the American Medical
“This [algorithm] sorts through all that hay.” Informatics Association.
“It’s not a theoretical problem,” says
mation. Now, scientists have developed When the researchers applied the Khaled El Emam, PhD, associate pro-
an algorithm that sorts through the mul- classifier to other data sets, it could dis- fessor at the University of Ottawa and
titude, expediting the search for blood- tinguish proteins secreted into the Canada Research Chair in electronic
based cancer biomarkers. blood from all other proteins in the health information, who collaborated
“Figuring out which proteins are blood with more than 80 percent accu- with Fida Kamal Dankar, PhD, on the
secreted into the blood is like searching racy. The results appear in the October paper. “We’re trying to protect privacy,
for a needle in a big, big haystack,” says 2008 issue of Bioinformatics. but we need the tools.”
Ying Xu, PhD, professor of bioinfor- Xu and his colleagues are now using Just as the nightly news renders the
matics and computational biology at microarrays to identify differences in faces of anonymous sources unrecogniz-
the University of Georgia. “This [algo- gene expression levels between cancer- able, the approach known as k-anonymi-
rithm] sorts through all that hay.” ous and non-cancerous stomach tissue. ty blurs distinctive variables to reduce
To develop their algorithm, Xu and Using their classifier, they can then sift the risk that someone could trace
his colleagues began by scouring the liter- through the data to zero in on genes that patients with distinctive characteristics.
ature for all proteins known to be secret- produce proteins that are most likely to For example, the approach might cut
ed into the blood, regardless of their ori- be secreted into the blood, followed by birthdates down to birth years. And eas-
gins. They then analyzed the amino-acid validation with mass spectrometry. ily identifiable outliers—the octogenari-
sequences of these proteins to identify “We’ve already identified proteins an in a college town, the teenager in a
common features, such as signal peptides, that are elevated during different stages retirement community—are omitted.
transmembrane domains, solubility, and of stomach cancer,” Xu says. “Typically, The remaining information contains at
secondary structure. They discovered 18 in order to find out what stage it’s in, least k data points that look identical,
features that were powerful predictors of you’d have to actually cut the patients where 1/k is deemed
blood secretion, and used them to train a open and do a biopsy. Our markers an acceptable
computerized classifier. could be the first markers to provide level of risk.
information about cancer stage.”
By applying his biomarker discovery
pipeline to a range of cancers, Xu ulti-
mately hopes to identify general biomark-
ers that apply to any cancer. He envisions
doctors detecting various cancers at early
stages with a simple blood test.
Bo Huang, PhD, a post-doctoral fel-
low at Vanderbilt University, hopes to
use Xu’s classifier to find biomarkers for
breast cancer. “These results provide a
powerful method to discover potential
biomarkers, not only for cancers but also
for many other diseases,” Huang says.
—By Lizzie Buchen

Blurring Data for


Privacy and Usefulness
Hospitals with research agendas
This microarray shows genes that differ in share a common problem: how to use
regulation between cancerous and non- medical records for research while pro-
cancerous lung tissue. Ying Xu’s classifier can tecting patient privacy. One approach—
predict which of the proteins made by these the data-protection equivalent of blur-
genes may be useful as blood-based bio- ring the face of an anonymous source on
markers. Courtesy of Ying Xu. television—has now been tested using

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 11
NewsBytes
“I’ve given Little b the power to
That works in theory, but the actual
risk depends on the type of data set and reason about biological objects,”
what an intruder wants from it. A prose-
cutor digging up dirt on a defendant Aneil Mallavarapu says.
would try to re-identify a specific person
in the database. A journalist trying to easy-to-use tool for biology labs. build entire virtual cells or virtual plants
discredit an organization’s data-security “I think that as an everyday tool, it collaboratively, increasing their ability
procedures would also only need to re- [Little b] is going to be kind of like the to study their projects in silico.
identify one person, but it wouldn’t mat- microscope,” says Aneil Mallavarapu, While the idea of breaking down bio-
ter who. El Emam set out to test whether PhD, lead developer of Little b and a logical systems into modular chunks
k-anonymity works in both circum- senior research scientist in systems biol- may seem logical, Little b may not arrive
stances. His findings: k-anonymity cor- ogy at Harvard Medical School. “We’re in the lab immediately, says Birgit
rectly predicts the risk of re-identifying essentially building a new kind of gel, a Schoeberl, PhD, a senior director of
one specific individual with minimal new type of microscope for the lab.” The research at Merrimack Pharmaceuticals,
harm to the value of the database (the work appears in the June 2008 issue of Inc, in Cambridge, Massachusetts. “I’m
prosecutor example). But using k- the Journal of the Royal Society Interface. excited about the concept and what I
anonymity to protect against re-identify- Biologists traditionally create models see, but in my own experience, it isn’t
ing an arbitrary person (the journalism to describe unique systems, such as the straightforward,” Schoeberl says. “I
example) is unnecessarily strict and com- development of fruit fly embryos or the think it’s not quite ready for non-
promises the research quality of the data. actions of a phosphorylation cascade on developers. I hope he keeps developing
Since researchers choose k based on gene transcription. Such computational it, or someone takes it on to keep
statistical theory, El Emam suggests data models are usually based on lists of the working on the idea.”
custodians run test cases to verify if the system’s properties, which detail every —By Molly Davis ■
k is sufficient, or if it’s overprotective, as molecular interaction in the
in the journalism example, before mak- system. This allows researchers
ing the data available to researchers. If to tailor models to the precise
needed, the number of groupings of k questions being asked, but it
identical data points could then be also constrains the model’s use-
adjusted to ensure that the actual risk fulness, because it can only
approximates the theoretical risk of 1/k probe into one area.
and, in this way, keep the risk accept- Little b strives to break down
ably low while preserving data. biological systems into modules
“What is needed are the steps to that can be used regardless of
turn this article into a practical tool the specific context, such as
that custodians can use in conjunction “nuclear export“ or “membrane
with researchers,” says Joan Roch, chief localization.“ It then defines
privacy officer for Canada Health those parts in a mathematical
Infoway in Montreal, Quebec. language. Researchers can use
El Emam says he plans to continue Little b to put together assorted
exploring actual risks in various data- modules to describe their sys-
security scenarios: “It’s a big problem, tem; Little b then uses those
and we’ve solved part of it.” symbolic modules to write out Little b is based on a core language, which includes
—By Stephanie Pappas executable code that a scientist the Lisp language it was created in (green) and the
could use in a simulation pro- knowledge base, symbolic mathematics and syntax
Modular Modeling gram like MATLAB. “I’ve modules that allow Little b to reason about biological
Biological models can quickly given Little b the power to rea- systems. It also includes modular libraries that
become as complex as the systems they son about biological objects,” describe specific biological interactions, and transla-
represent. And minor changes can Mallavarapu says. tors that can generate code used in simulations. Blue
necessitate a complete rewrite of the Mallvarapu is excited about areas exist within the current framework; yellow
model. But researchers may soon snap the possible use biologists might areas are currently under development or are envi-
their models together like LEGOs, using make of Little b. He would like sioned for future work. Reprinted with permission
a new programming language called to see the language help uncov- from Mallavarapu, A, et al., Programming with mod-
Little b, which uses modularity to sim- er the complex pathways els: modularity and abstraction provide powerful
plify biological modeling. Eventually, involved in diseases. He hopes capabilities for systems biology, Journal of the Royal
the authors hope to turn Little b into an that researchers will eventually Society Interface, online publication, July 23, 2008.

12 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


A
fter four years, the seven National
Centers for Biomedical Computing
(NCBCs)—established largely to
build a national biocomputing
infrastructure—have, as one might
expect, produced an impressive
array of computer tools. >

NCBC UPDATE:
Shedding New Light On

By Katharine Miller

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 13
B ut it’s the Centers’ wide-ranging impact on
biomedicine that takes center stage. From AIDS
to diabetes, prostate cancer or schizophrenia, the
NCBCs are changing the landscape of disease research
by shedding new light on biological complexity.

“The impact on biology and medicine hap- the NCBCs to successfully penetrate the
pened faster than anyone expected,” says Russ broader community with tools, techniques and
Altman, MD, PhD, co-principal investigator methodology,” he says. “We’ve shown what
for Simbios, the National Center for Physics- can be accomplished by applying these tools
based Simulation of Biological Structures, to biological problems.”
an NCBC grantee at Stanford University. And while the specific breakthroughs enabled
And that impact springs from the way the by NCBC tools varies with the tool being used or
NCBCs function, says Andrea Califano, PhD, the disease being studied, it is clear that they are
who heads the National Center for all helping researchers approach the complex sys-
Multiscale Analysis of Genomic and tem that is the human body. “Dealing with
Cellular Networks (MAGNet) at Columbia complexity is the essential challenge of this
University. “Developing new tools in the con- century in biology,” says Scott Delp, PhD,
text of solving specific scientific, biological or co-PI for Simbios. “And you can’t do it with-
medical problems is what I think has allowed out computers.”

“Developing new tools in the context of solving


specific scientific, biological or medical problems is
what I think has allowed the NCBCs to successfully
penetrate the broader community with tools,
techniques and methodology.” says Andrea Califano.

NCBC
Y EAR

Califano: “One critical thing we hope to accomplish is to create a new


1GO0ALS breed of biologist trained both in computational and experimental sciences.
You already see evidence of this in some labs. Now, as never before, some of
the projects enabled by the NCBCs have computation and experimental biolo-
gy playing hand in hand rather than in a pipeline fashion. That is also reflected in the
tools that we generate. Unlike other platforms, geWorkbench was created for an experimental biol-
ogist who wants to learn enough computational biology to be able to analyze data. It’s easy and
intuitive to use, and the researcher doesn’t have to learn complex scripting languages. The empha-
sis has been on enabling experimental labs to use more and more computational tools. Across the
entire set of activities at MAGNet, the real aim is to fuse the two disciplines and to create a really
interdigitated boundary between the computational and
experimental life sciences.”
Andrea Califano, PhD, is the principal investigator for the National Center
for Multiscale Analysis of Genomic and Cellular Genomics (MAGNet) and
professor of biomedical informatics at Columbia University.

14 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


Here, following a few years of hard work, the cipal investigator Zak Kohane, MD, PhD.
NCBC PIs reflect on what they’ve accom- And imaging tools originally developed by
plished so far, how they’ve gained traction in CCB and the National Alliance for Medical
the research community, and what their goals Image Computing (NA-MIC) to study schiz-
are going forward. ophrenia in the brain are
now proving useful in study-
NCBC TOOLS: ENABLING DISCOVERY ing many other brain dis-
ACROSS THE DISEASE SPECTRUM eases, as well in prostate can-
From the start, each NCBC’s tool and cer (at NA-MIC) and cardio-
infrastructure development goals were driven vascular disease (in CCB’s
by a cluster of specific biological problems—
commonly referred to in NCBC parlance as
case). “You can begin to see
how the shape-modeling
“Dealing with
the “driving biological problems” or DBPs. approach [we’ve developed]
After a few years, these DBPs were replaced by is applicable to a whole range complexity is the
a new set of DBPs, ensuring that the tools of biological problems,” says
would be suitable for multiple purposes. That Toga of CCB. essential challenge
strategy has worked. Similarly, OpenSim, a soft-
“To a certain degree, the tools and biology ware program developed by of this century
are push-pull kinds of associations,” says Art Simbios to study human
Toga, PhD, principal investigator for the movement and movement dis- in biology,”
Center for Computational Biology (CCB) orders, was first used to con-
based at the University of California, Los
Angeles. “The tools get developed because you
duct research into one of the
Simbios DBPs, cerebral palsy,
says Scott Delp.
couldn’t do something without them. And vice but is now being used more
versa, you get this tool and you decide to pose broadly. Indeed, it has been
“And you can’t
new questions. You end up pushing and pulling adopted by more than one
so that both are advanced.” thousand individuals working do it without
Thus, NCBC tools that were developed to on any number of problems
address one biomedical problem have proven including osteoarthritis, computers.”
to be broadly useful. For example, at i2b2— Parkinson’s disease and stroke.
Informatics for Integrating Biology and the This is the vision of the
Bedside—an NCBC based at Harvard, tools NCBCs—to provide the under-
developed to allow the use of medical record lying computational tools that
systems for clinical research initially focused will advance the field of medi-
on diseases such as asthma, obesity and cine and biology, across a spec-
depression. Now, however, these tools have trum of diseases. As Mark Musen, PhD, of the
been adopted at 18 large academic health National Center for Biomedical Ontologies
centers with no apparent limit on the number (NCBO) at Stanford, says, “We’re enablers.
of diseases that can be studied, says i2b2 prin- We are providing the foundation by which

Kikinis: “We are developing algorithms and a plat-


form—the NAMIC kit—for analysis of diagnostic
images. I think that platform will be one of our major
NCBC
1 0
YEAR

accomplishments. It is free and open source with a


very liberal license, and it will continue to be devel- GOALS
oped. That will continue for a long time. So our goal is
to develop enabling technologies and make them accessi-
ble. That will be one of the legacies of the center.”
Ron Kikinis, PhD, is the principal investigator for the
National Center for Medical Image Computing (NA-MIC) as
well as director of the Surgical Planning Laboratory of the
Department of Radiology, Brigham and Women’s Hospital
and Harvard Medical School, and professor of radiology at
Harvard Medical School.

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 15
investigators can do research that will impact centered at the University of Michigan, agrees.
human health. Our goal is to create the kinds of While his center’s tools have contributed to a
tools that would be valuable to everybody.” better understanding of type 2 diabetes and
Ron Kikinis, PhD, head of NA-MIC, con- prostate cancer progression, the tools’ reach
curs. “We will not solve cancer but we will extends much farther: “We’re opening doors to
provide the people who are fighting cancer new research,” he says.
with better tools to fight their fight,” he says.
“And the DBPs will use these tools and pro- NCBC CHALLENGE:
mote those tools into their communities—so PUTTING IT ALL TOGETHER
that makes it possible for lots of different dis- For the last thirty years, biology has been
eases to be addressed.” about breaking things down into their funda-
Brian Athey, PhD, co-PI for the National mental parts to understand them. “But things
Center for Integrative Biomedical Informatics, don’t work as independent parts,” says Delp.
“Theoretical and computational biology let
you put things back together to understand
the whole system.”
Several of the NCBC PIs cite the re-
“The tools get developed assembling of biological pieces as a major
focus of their efforts. For example, literally
thousands of experiments have looked at how
because you couldn’t do elements of the neuromuscular system (mus-
cles, joints, connective tissue) operate inde-
something without them. pendently. But, Delp says, looking at those
elements separately doesn’t tell you how peo-
And vice versa, you get this ple move. OpenSim lets researchers put the
pieces together. “When you can code the
tool and you decide to pose details accurately in a computer framework,
then you can understand how the system
new questions. You end up works,” Delp says.
Likewise for the brain, says CCB’s Toga.
Brain researchers have typically focused on
pushing and pulling so only one variable at a time—for example,
electrical activity, blood flow, distribution of
that both are advanced,” receptors, gene expression patterns, or cortex
morphology. But, Toga says. “All of these
Art Toga says. brain changes are happening in concert.” To
understand the brain requires re-integration of
these events. CCB, Toga says, is providing the
tools, mechanisms, and strategies to put things

Toga: “Our hope is to continue to integrate what we know about the brain in a way
that allows us to ask questions such as: ‘How does the brain change throughout a per- NCBC
son’s life?’ These sorts of emerging questions are provocative. And we can only ask
them because of computation. So by the end of our ten years, GOALS
10 YEAR

we really hope that our center produces new research pro-


grams that can continue to evolve in accord with the basic
thrust of the NCBCs. Because you know, it doesn’t finish. They
haven’t finished mapping the earth yet and there’s only one of those!
How can anyone possibly suggest we will ever finish mapping the
human brain when there are billions of them? So we’ll continue to layer
on what we already know without throwing away our previous efforts.”
Arthur Toga, PhD, is the principal investigator for the Center for
Computational Biology (CCB), and a professor of Neurology and
Director of the Laboratory of Neuro Imaging at the University of
California, Los Angeles.

16 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


Working together NCBC researchers creat-
ed iTools—a way to manage the descrip-
tion of computational biology data, tools,
and services. Using the iTools hyperbolic
viewer a researcher can displays all of the
activities of the NCBCs organized by Center
(as shown here) or by activity. iTools also
lays the groundwork for interoperability
among diverse biomedical computing
tools. Reprinted from Dinov, ID, et al., 2008
iTools: A Framework for Classification,
Categorization and Integration of
Computational Biology Resources. PLoS
ONE (2008) 3:(5):e2265.

back together. “Observations from one project


in 2007 can be combined with other observa- “We’re enablers.” Mark Musen says.
tions in another laboratory using different
subjects and techniques in 2008,” Toga says. “We are providing the foundation by
“That transition in science is revolutionary,
and the computational strategies that enable
it are only now beginning to emerge.”
which investigators can do research that
MAGNet hopes to provide a similar service
at the genetic and cellular level. Very few dis-
will impact human health. Our goal
eases are caused by a single gene, Califano says.
Usually a complex interplay of genetic and epi- is to create the kinds of tools that
genetic factors is involved. “But what has been
lacking is a framework for integrating genetic, would be valuable to everybody.”

Musen: “We are thinking about what it would mean to be able to move biomedical
knowledge from prose to machine-processable format. The long-term vision is to create the
infrastructure and tools so that biomedical literature could be intelligible to both people and
machines. Ultimately this could allow intelligent computer-based agents to read the litera-
ture, to make associations between scientific contributions, and to
synthesize ideas from the literature. That would obviously change the
NCBC way we do science in a very profound way. But there are lots of baby
YEAR

1GO0ALS steps until we can do that.”


Mark Musen, MD, PhD,
is the principal investi-
gator for the National
Center for Biomedical Ontologies
(NCBO) and professor of medicine at Stanford
University School of Medicine.

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 17
epigenetic, functional and structural data—and NCBCS:
getting an answer that can really dissect dis- MORE THAN THE SUM OF THEIR PARTS
ease,” he says. MAGNet’s goal is to establish The NCBCs are also working together in
such a framework and to show that the frame- various ways to ensure that they have a broad
work can integrate data in meaningful ways for impact. In some ways this is a surprise, say the
several diseases. “We already have proof of con- NCBC PIs, because the NIH cast such a wide
cept for glioblastoma multiforme—a cancer net—with centers that cover ontologies, simu-
that produces the worst possible prognosis in lations, clinical systems, systems biology and
patients,” Califano says. The results for that imaging. “Given the breadth of the needs and
work will be published in the next few months. the solutions to biomedical computing prob-
“This kind of proof of concept in a disease is of lems,” says Kohane, “it wouldn’t have been sur-
course important, but at the same time the prising if there had been no overlap and the syn-
methodology becomes universal.” ergies had been fewer.”

“What bioinformatics was five years ago is frankly


just a glimmer of what it is today,” Brian Athey says.
“It’s exploding into something much more robust.
And that’s going to continue for a while.”

NCIBI is also integrating many different Yet the NCBCs have found overlap and
high-throughput data types to better understand have helped each other. For example, the i2b2
complexity. “We do not yet understood the full center collaborated with NCIBI around Type 2
complexity of the architecture of the human diabetes, Kohane says. And NA-MIC nicely
genome,” Athey says. “Only 2 percent of the complemented i2b2’s major depression DBP by
genome are ‘genes’ and we’re learning more and correlating patient imaging with what was
more that the other 98 percent are doing being seen genetically. Similarly, ontologies
things.” To tackle that problem, he says, com- from NCBO have been helpful to CCB in con-
putational biology is making huge strides. “What structing their brain atlas; and CCB and
bioinformatics was five years ago is frankly just a Simbios have used some of NA-MIC’s visuali-
glimmer of what it is today,” Athey says. “It’s zation tools.
exploding into something much more robust. Even though the NCBCs might be develop-
And that’s going to continue for a while.” ing different tools, Califano says, “when you

NCBC
Athey: “There’s much more work to do to figure out how to use systems biology more
1 0
YEAR
effectively to understand disease and its complications. The daunting complexity of biological
systems is becoming more and more clear. To gain an understanding of that complexity, we GOALS
need an integrative approach that’s iterative and that allows the integration of many different
kinds of data types around hypotheses and models. The abundance of high throughput data we’re
presented with from next generation sequencing, and what that’s revealing about the transcriptome
and alternative splicing, and all the components we haven’t yet annotated—
it’s just astounding. It’s literally changing our basic understanding of cells and
their complexity and function. And, frankly, it’s changing what our understand-
ing of a gene is. So there’s a lot of work to do. I think that’s
the theme. And each success brings on new challenges.”
Brian Athey, PhD, is the principal investigator for the National Center for
Integrative Biomedical Informatics (NCIBI), associate professor of biomed-
ical informatics at the University of Michigan, and director of the Michigan
Center for Biological Information.

18 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


tackle a biological problem you must tackle it
from several angles.” So for example, MAGNet
and NCIBI have several DBPs that focus on
analyzing genomic data as a way of studying
neurodegenerative diseases, diabetes or cancer.
But these same diseases also need to be studied
using data from large cohorts, which ties in to
what i2b2 does at Harvard to use medical “If we actually successfully did a big
records to study large populations. It also ties in
to the ontology work of Mark Musen, Califano
says, because ontologies provide an essential
population study and discovered
foundation for other work. And, he says, when
you look at the actual problem you’re trying to
something important or successfully
understand, all sorts of issues related to physical
modeling also come up. Indeed, according to calculated how to design a vaccine or
Altman, eventually cellular physics will
become an essential piece of systems biology. predicted a new drug for a specific
“The reality of why all the centers come
together is precisely around the biology, disease, then we’d be bringing
Califano says. “We develop all the different
techniques and infrastructure to tackle biology ourselves to the next level,” says
problems, but when you actually want to tack-
le one of these problems, you require all of
these approaches.”
Kohane. “We’d be solving a
And those multiple tools also need to be kept
organized. So one key activity that has united
biomedical problem of true health
all the centers, says Musen, is the creation of an
online tool that allows biomedical software relevance. In fairness, I think
resources to be easily identified and searched
online. Called Biositemaps, the tool, seeded we’re all trying to get there, but
with information about the NCBC tools, can
inform search engines about software available we’re not there yet.”
from any organization that creates a simple
Biositemap file as described on the site
([Link] NCBO is provid-
ing the ontology behind the tool but, Musen
says, “It’s a product of all the NCBCs that would
not have been possible without the cooperative
involvement of all the different centers.”

NCBC

Kohane: “Within the ten-year time frame, the goal would be to


10
GOALS
YEAR

establish a kind of scientific ecosystem around the country where we


can use entire healthcare systems as a unit of study. We’ll be able to look
at reproducibility across multiple academic health centers to see if we’re seeing, for example, the
same adverse drug events (so that we can push early warnings to prevent such events); or com-
pare efficacious therapies; or compare whether we have reproducible findings in genomics or pro-
teomics across populations. This approach will allow us to do research in a more cost-effective
way. And although it sounds venal to talk about cost, cost is a key rate-limiting factor in large pop-
ulation studies. So if we can do clinical research, including genomic measurements in populations
of 10,000 to 100,000, that’s really a game-changer.”
Isaac Kohane, MD, PhD, is the principal investigator for Informatics for
Integrating Biology and the Bedside (i2b2), as well as Lawrence J.
Henderson Associate Professor of Pediatrics and Health Sciences and
Technology at Harvard Medical School, and Chair of the Informatics
Program at Children’s Hospital, Boston.

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 19
NCBCS: BENCH TO BEDSIDE time to build the tool, teach people how to use
Whether casting a wide net to enable it, get it adopted, make a discovery and then
research in lots of areas is enough to render the translate that into clinical care.” Currently, says
NCBCs successful remains to be seen. Curing a Delp, “OpenSim is only halfway down that
disease would be better. “If we actually success- pipeline and is just beginning to see the first
fully did a big population study and discovered examples where new discoveries will enhance
human health.”
Kikinis says NA-MIC’s tool kit is
similarly poised for bedside use. He’s
beginning to see the first signs—such
as questions at seminars, and email
inquiries—that companies are inter-
“Adoption by companies ested in it. “Adoption by companies is
one indication that what we’re doing
is one indication that what will eventually make a difference to
clinical practice,” he says. “We are not
we’re doing will eventually make yet at that point, but I have these
early indicators.”
a difference to clinical practice,” Migrating computational biology
from the bench to the bedside remains
a challenging goal for all the centers.
Ron Kikinis says. “We are But, as Toga sees it, “I think these com-
putational strategies, which are the
not yet at that point, but hallmark of this program, are having a
great effect on accelerating that.” CCB
I have these early indicators.” is modeling the effect that HIV and
Alzheimers have on the brain. These
are diseases that will strike people we
all know, Toga notes. “So our work
immediately transforms a mathematical
problem [shape modeling] into some-
something important or successfully calculated thing with obvious and immediate clinical
how to design a vaccine or predicted a new drug value,” he says. “And the time frame for doing
for a specific disease, then we’d be bringing our- that is getting shorter and shorter and shorter.”
selves to the next level,” says Kohane. “We’d be Kikinis summed it up succinctly: “What are
solving a biomedical problem of true health rel- the NCBCs doing for biology? Everything.
evance. In fairness, I think we’re all trying to get That’s by design, but now you can say that
there, but we’re not there yet.” they’re actually delivering, and there’s a sense of
“The challenge is,” says Delp, “that it takes excitement. It’s clear that things are moving.” ■

Delp: “The goal is twofold, really. One, that


we’ll produce a set of tools that are ubiquitous NCBC
in biomedical research so that every investiga-
1 0
YEAR

tor who is interested in how physics affects bio- GOALS


logical function will have SimTK-based tools as
part of their laboratory. The second objective is
that we and others will use those tools to make new
discoveries that enhance human health.”
Scott Delp is co-principal
investigator for the National
Center for Physics Based
Simulation of Biological
Structures (Simbios) and a
professor of bioengineering
and mechanical engineering
at Stanford University.

20 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


By Kristin Sainani, PhD

TooL
Dissemination DOING IT RIGHT

B iomedical computing at academic


research centers has been compared to
a cottage industry. Lots of individuals
work away on their focused research projects,
generating useful algorithms. But quite often, the
knowledge gained is lost when researchers move
on to new projects. Yes, they might post their code
on Web sites. But is it useful to anyone else without
support and documentation? And how can people
find it in the first place?

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 21
T
To overcome the cottage industry men- someone has to build “disseminability” ration with colleagues at the University
tality, the National Institutes of Health into the tool, with robust, flexible, and of California, San Diego; the program is
(NIH) is placing a greater emphasis on extensible code. Then, someone has to downloaded about 1000 times a month
dissemination as a piece of the National package the tool in a way that makes it ([Link]
Centers for Biomedical Computing accessible to a wide audience. Finally, When Baker realized that APBS
(NCBCs) as well as for other grantees. someone has to publicize the tool, build offered something new that might be
But what does it really take to turn a community of users, and support and widely useful, he says, “I took most of
an impressive algorithm into a widely maintain the tool. what I’d written at that point and just
disseminated, prolific computational In an ideal world, that “someone” deleted it and started over.” A tool that
tool? The transition might be harder would include a team of people with is going out to others has to be built
than you think. diverse skills—such as software engi- according to professional software
“Today, our software is very wide- neers, technical writers, and marketers. design principles, he says. The code
ly used, but it didn’t take off right But, in reality, it is often a scientist should be clean, bug-free, and robust;
away. It took years,” says Klaus moonlighting as all of the above. Tool and it should be built in a flexible, mod-
Schulten, PhD, speaking about the dissemination has traditionally been ular fashion so that others can add to the
molecular dynamics simulator NAMD underappreciated and underfunded, tool and adapt it to their own problems.
([Link] making it hard for researchers to dedi- “There’s a world of difference
and the molecular graphics viewer VMD cate resources to tools beyond what’s between developing code for yourself
([Link] needed for their science. Fortunately, and developing code that you want to
which together have more than this situation is changing—with initia- distribute,” Schulten agrees. Establishing
tives such as the NCBCs that recog- the proof of concept takes 10 percent of
nize the importance of tool develop- your time, whereas adhering to profes-
ment and dissemination—but there is sional design principles takes 90 percent,
“There’s a world of still a long way to go. he says. “And it is almost impossible to
So how do scientists manage to do it convince any normal scientist to spend
difference between right? Biomedical Computation Review that 90 percent.” Professional program-
spoke to a panel of individuals who have mers helped design VMD and NAMD,
developing code disseminated popular open source bio- and they were a key factor in the tools’
medical tools to find out what it takes to success, he says.
for yourself and succeed and how they pulled it off.
DRESSING YOUR TOOL FOR
developing code LAYING THE GROUND WORK
The ingredients for successful tool
SUCCESS: ACCESSIBLE, WELL
DOCUMENTED, WITH A GUI
that you want to dissemination have to be built into the
tool’s core from the start.
To become widely used, tools also
have to be accessible—which means
“You can’t assemble a software pack- open source, portable, well document-
distribute,” says age out of a bunch of code that your ed, and user-friendly.
graduate students wrote trying
Klaus Schulten. to get their theses done. It can’t
be an afterthought,” says
Nathan A. Baker, PhD, associ-
100,000 users. “We went through a ate professor of biochemistry
long initial phase where we were close and molecular biophysics at
to failure all the time.” Schulten is pro- Washington University in St.
fessor of physics at the University of Louis. “At some point in the
Illinois at Urbana-Champaign and design process you say, ‘oh,
director of the Theoretical and other people might want to use
Computational Biophysics Group at this.’” Baker wrote APBS—a
the university’s Beckman Institute. program that solves the Poisson-
For a tool to spread, it takes more Boltzmann equation for molec-
than a good algorithm. From the start, ular electrostatics—in collabo-

VMD Visuals: (top) secY protein, (lower left) fibrino-


gen protein,(lower right) polio virus particle. Picture
made by the molecular graphics software VMD.
Despite initial challenges, VMD is now a clear dis-
semination success story. The software is even used
in high school classrooms. Courtesy of: the
Theoretical and Computational Biophysics Group,
NIH Resource for Macromolecular Modeling and
Bioinformatics, at the Beckman Institute, University
of Illinois at Urbana-Champaign.

22 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


Growing a Tool. The use of GROMACS software has spiked since 2000: There has
been growth every month in the number of citations to one or more of the three
GROMACS papers or the manual. Courtesy of Erik Lindahl.

ty also started voluntarily fix- Cytoscape is a software platform for


ing bugs and writing new mod- modeling molecular interaction net-
ules and patches. “Everybody works that gets about 3000 downloads
benefits from the openness. So per month ([Link]
I think overall it’s been an To be accessible, tools not only
incredibly positive experience have to be free but also have to work
for us,” Lindahl says. on the computers that biologists are
GROMACS follows the using, says Thomas L. Madden, PhD,
GPL-style open source license, a scientist at the National Center for
which requires those who Biotechnology Information at the
adapt the software to make U.S. National Library of Medicine.
their programs open source as Madden helped transform UNIX-
well. Other tools in this article based BLAST into a tool that runs on
“Science is about getting things out follow the less restrictive BSD-style multiple platforms, including
there,” says Erik Lindahl, PhD, associate license. “If I was starting from scratch, Windows and Mac OS. BLAST is a
professor in the Center for Biomembrane I’d seriously consider going with this sequence alignment tool and an undis-
Research and the department of bio- completely open license,” Lindahl says. puted tool success story—the original
chemistry & biophysics at Stockholm “BSD actually worked out quite well BLAST paper was the most highly
University in Sweden. “Unless you have for us,” says Steve Pieper, PhD, founder cited biomedical paper in the 1990s
this great 10 million dollar idea that will and CEO of Isomics, Inc., in ([Link]
make you a fortune, the last thing you Cambridge, MA, and the dissemina- “A lot of bioinformatics tools are

“You can’t assemble a software package out of a bunch of code


that your graduate students wrote trying to get their theses done.
It can’t be an afterthought,” says Nathan Baker.
want to do is to limit access to your tion core PI for the NCBC NA-MIC only made for Linux or Unix, but we’ve
work.” Lindahl is a primary developer of (National Alliance for Medical Image had just as many downloads of the PC
GROMACS, a molecular dynamics sim- Computing). The NA-MIC toolkit version of BLAST as the Linux ver-
ulation package developed at the includes visualization software: VTK, sion,” he says. “I think you can figure
University of Groningen, which has ITK, and Slicer ([Link] that just about every lab has a PC. So I
been cited more than 1000 times [Link]/Wiki/[Link]/NA-MIC- don’t think you can underestimate the
([Link] Kit). The BSD license has allowed importance of that.”
When GROMACS was released in medical imaging companies to incorpo- Once users have a tool in-hand, if it
the early 1990s, it was not open source— rate bits and pieces of the software into is technically difficult or poorly docu-
academic users had to sign a contract their equipment—which gets the tech- mented, they are likely to seek out
and industry users had to pay a fee. But nology out where it can directly benefit something easier to
the licenses were a hassle and Lindahl patients, Pieper says. use. The main
barely broke even paying for the secre- Cytoscape—which also follows the
tary to handle them, he says. “So, we BSD license—has similarly been incor-
realized this wasn’t really very smart.” porated into several commercial soft-
When they moved GROMACS to ware applications, says Trey G. Ideker,
open source, their user base quickly PhD, associate professor of bioengi-
jumped from 1000 to 5000 and contin- neering at the University of California,
ued to climb from there. The communi- San Diego, and on the Cytoscape board
of directors.

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 23
“Unless you have this great 10 million dollar idea that
will make you a fortune, the last thing you want to do
is to limit access to your work,” says Erik Lindahl.

reason scientists flock to commercial ages—and then good luck reading the deterred by the lack of a graphical
alternatives for open source software is documentation.” user interface (GUI). For example,
not because of superior performance To help make the documentation Baker says of APBS: “It’s no worse
(often the opposite is true), but because more user-friendly, several of our inter- than the other command-line compu-
of a great user interface and great docu- viewees advocate “learn by example” tational biology tools. But I would say
mentation, Lindahl says. Open source tutorials, which lead users step-by-step that maybe 80 percent of our audi-
tools often fall short on these aspects. through common research problems. ence would prefer to interact with it
“I’m a sucker for good documentation. Many potential users are also in some other way.”
If there are not clear
PDFs with graphics, I’m
extremely unlikely to
use it,” says Raymond
R. Balise, PhD, a bio-
statistical programmer
at Stanford University,
who uses the open
source statistical pack-
age R, which has hun-
dreds of thousands of
users ([Link]
[Link]/). But the
best programmers are
usually not the best
writers, he says. “So
you have brilliantly
designed elegant pack-

Cytoscape Pathways (Including background image on


page 21). Pictures generated from Cytoscape, software
for visualizing complex molecular interaction net-
works. Cytoscape follows a “non-viral” open source
license, which allows companies to incorporate the
software into their own commercial tools. Many
companies now rely on Cytoscape as a critical part of
their tools. Courtesy of: Vuk Pavlovic and Benjamin
Elliott, the University of Toronto.

24 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009


Similarly, R is a great tool for mathe-
maticians and statisticians who are used
to difficult programming languages, but
telling physicians or biologists to “learn
to program” just doesn’t fly, Balise says.
To make tools accessible to a wider audi-
ence, you need to wrap a nice GUI
around the package and build in checks
and balances to alert users if they’re
doing something wrong, he says.
Tool developers often resist these
steps for fear that they will have to sacri-
fice power and flexibility for usability.
An easy-to-use GUI-based interface is
too constraining for research-driven
tools, such as R and Bioconductor, that
need to keep up with the cutting edge of
science, says Martin Morgan, PhD, a
core developer for Bioconductor, an R-
based tool for analyzing high-throughput
genomic data that has tens of thousands
of users ([Link]
These tools may never be a satisfactory
solution for a general audience, says Power Surge. The number of active computers running Folding@home has surged since 2000. Courtesy
Morgan, who is also a staff scientist and of: Vijay Pande, Stanford University.
director of the Bioinformatics Shared
Resource at the Fred Hutchinson Cancer So, they developed the program to and how to use it, she says.
Research Center in Seattle, Washington. be modular and flexible for expert users, Outreach often starts with a publi-
But usability can evolve, even if the including allowing it to interface with cation that announces the tool. In the
tool was designed for expert users. For standard programming languages such early days, people discovered BLAST
example, community developers have as MATLAB, Java, and R; but they also primarily through the publication and
spontaneously added GUIs onto several provided a point-and-click GUI. word of mouth, Madden says. BLAST
programs—including R Commander for “I think it really is the non-pro- solved a key problem, so it was obvious
R, and PyMOL and VMD plugins for gramming community that has made how it was useful. Nowadays, “light-
APBS. Core developers may also revisit the package so popular,” she says. “We
usability as a tool matures. For example, get emails from both types of users,
BLAST’s core developers have become and we get really effusive ones from
more focused on ease of use in recent the non-programming users, because
years, particularly for the BLAST web-
page interface, Madden says.
they say ‘Wow, this really lets me use
all these sophisticated tools and I can
“I’m a sucker for good
In rarer instances, developers con-
sider usability from the start. This was
do it on my own,’” Mesirov says. documentation. If there
the case with GenePattern, says Jill CONNECTING TO
Mesirov, PhD, director of computa- YOUR AUDIENCE are not clear PDFs with
tional biology and bioinformatics and The next step in tool dissemination
chief informatics officer at the Broad is the actual dissemination—connect- graphics, I’m extremely
Institute of MIT and Harvard. ing the tool to users. This means not
GenePattern is an analysis program for only getting the word out about the unlikely to use it,”
genomic and proteomic data, which tool but also “selling” it.
also captures users’ steps in a repro- “There is a mentality that if the tool is says Raymond Balise.
ducible pipeline; the package, released good enough it will speak for itself,” says
in 2004, already has thousands of users Stanford University’s Joy Ku, PhD,
([Link] director of dissemination for Simbios and
ware/genepattern/). its tools, including SimTK Core, a toolk-
From the beginning, GenePattern’s it for physics-based biological simulations weight” outreach on the web can also
developers recognized that they were tar- ([Link] and go a long way, he says. You can reach
geting two audiences: “We have a num- OpenSim, a package for modeling mus- many potential users with little cost
ber of computational scientists who do a culoskeletal movement ([Link] through newsgroups, email lists, blog-
lot of their own coding. We also have a home/opensim). But, in many cases, gers, and even random web searches.
lot of bench biologists who want to do particularly for complex tools, you “One thing that worked very well
analyses but don’t want to write code. really need active outreach to show for us is the web,” Schulten agrees,
And why should they?” Mesirov says. people how the tool applies to them speaking about VMD and NAMD.

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 25
“That really was a godsend because it’s Cilk Arts, focused heavily on web out- questions specifically about FFTW, and
basically like we have a shop and our reach. They posted benchmarks com- it was especially important to respond
shopping window is the web.” he says. paring their software with other FFT to these—having a support presence
“It’s so easy to do and you reach so implementations; added FFTW links on public forums reassures people that
many people.” on websites that list FFT programs, as the software works and is actively
maintained,” Johnson says.
FFTW is now downloaded
about 10,000 times a month
Non-programming users of GenePattern send ([Link]
Active mailing lists and
effusive emails, says Jill Mesirov, “because they online forums help draw in new
users, support existing users,
say ‘Wow, this really lets me use all these and build a sense of communi-
ty. “I frequently get much bet-
sophisticated tools and I can do it on my own.’” ter support from open source
mailing lists than you get from
vendors,” Lindahl says.
Answering emails about the
To promote FFTW (“the Fastest well as on sites that catalog free-soft- tool also goes a long way: “We’ve
Fourier Transform in the West”)—a ware projects (such as [Link] received over 10,000 email messages
general-purpose tool that performs and [Link]); advertised on about FFTW over the past 10 years, and
Fourier transforms, which are often mailing lists; created their own mailing responded to a large fraction of them,”
used in molecular dynamics simula- list; and answered questions on online Johnson says.
tions—creators Steven G. Johnson, discussions about FFTs, including pro- Beyond the web, more “heavy-
PhD, assistant professor of applied viding links to FFTW and other free weight” outreach includes training ses-
mathematics at MIT, and Matteo Frigo, FFT software. sions, workshops, and conferences. For
PhD, chief scientist and founder of “Eventually, people began posting example, Simbios and NA-MIC as well
as other NCBCs hold training events at
conferences and stand-alone work-
shops for developers and general users.
Cytoscape developers run tutorials at
the major bioinformatics conferences
and some major disease conferences.
It’s hard to convince scientists to spend
time running training sessions rather
than improving the tool, Pieper says.
So, it’s important to involve people
who are specifically interested in and
passionate about teaching, he advises.
R, Bioconductor, and Cytoscape hold
their own annual conferences (funded
primarily by corporate sponsors and
paying participants), which help adver-
tise the tools as well as bring developers
together. “There’s definitely a commu-
nity, and the whole mentality of work-
ing as an international team is huge for
R,” Balise says.
High school teachers and college
professors also promote tools in their
classrooms. With VMD, “it became so
user friendly that it could actually
trickle down to college and high school
education,” Schulten says. “We were
very fortunate that these outreach
efforts were essentially ripped out of our
hands. So now there are many efforts,
Building a Pipeline. The GenePattern tool helps expert and non-expert users analyze genomic and pro- and we just happily receive the news.”
teomic data, while capturing the steps in a reproducible pipeline. The tool was built with non-expert Distributed computing efforts are all
users in mind, which has been a major factor in the popularity of the tool. Reproduced from Reich M, about outreach, since researchers must
GenePattern 2.0, Nature Genetics (2006) 38:500-501, supp. fig. 1. convince the general public to down-
load and run their tool. Coverage in

26 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


the popular press (Time, CNN, and the
New York Times, for example) helped
generate buzz for Folding@home
([Link] a distrib-
uted computing project at Stanford
University led by Vijay Pande, PhD,
associate professor of chemistry.
Distributed computing also uses com-
petition to stir up interest—partici-
pants collect points based on the
amount of computing power they con-
tribute. Capturing the high score is
reminiscent of holding the high score
on Asteroids at your local video arcade
back in the eighties, but this is on a
much grander scale, Pande says. “It’s
something on a very high profile site,
where you can be number one out of
hundreds of thousands.”
Competitions are something we’d
like to explore, Ku says. Already,
Simbios runs a traditional grant compe-
tition for seed projects, which gener-
ates interest in and awareness of their
center. “Ultimately you’re only going
to fund a small percentage of appli-
cants, but all the applicants have to
become familiar enough with what
you’re doing,” she says. A similar
approach could be used for software.
So, which of these outreach efforts is
most effective? Until this year, we’ve
just been going by an intuitive feel for
what works, Ku says. But, in an effort to
improve dissemination, they collected
eight months of data on how people R Gallery. Community developers have written so many graphical programs for data visuali-
find their software project repository zation in R that it’s hard to keep track of them; here the programs are cataloged visually for
Web sites, [Link]. The breakdown easier access. Contributions from the community have been critical to R’s growth and success.
is: 29% word of mouth; 25% publica- Screenshot from the R Graph Gallery, [Link]
tions and conferences; 24% web search;
13% mailing lists and newsgroups; 9% as well as many hours of volun-
other mechanisms (including use in the teerism—from professors, graduate stu-
classroom, Biomedical Computation dents, postdocs, and community mem-
Review, and links on other Web sites). “One problem bers. Under this piecemeal model,
Word of mouth leads the way, but it there’s no money to hire professional
accounts for less than one-third of
hits—so more active outreach is vital.
both in Europe programmers let alone technical writers
or outreach coordinators. Lindahl says

MAKING IT HAPPEN and in the States he’d “nudge” postdocs to turn code
they wrote for their research into for-
Successful tool dissemination can be mal GROMACS modules. Pande says
lengthy and costly, and it requires is that it’s hard he and his graduate students have to
diverse skills, such as programming, work 60 to 70-hour weeks to keep
writing, marketing, and teaching. So to get funded Folding@home going. “It’s just a lot of
how do scientists support these efforts? work to be running something like
“Up to now it’s frequently been the only for software this,” Pande says. Johnson says he and
case that you’re kind of moonlighting,” Frigo did most of the legwork for FFTW
Lindahl says. “One problem both in development,” themselves over the years, despite
Europe and in the States is that it’s many other time commitments.
hard to get funded only for software
development.” Many tools are support-
says Lindahl. Tool upkeep and dissemination are
also undervalued when it comes to
ed using bits and pieces of resources academic promotion—making it even
scrounged from science-driven grants harder to justify dedicating scarce time

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 27
and resources to these endeavors. n’t happen. It would be like, as with most into obscurity—but it is wasteful and
“Academic credit for maintaining previous funding, an afterthought in reflects poorly on the biomedical com-
software is not the same as producing some grant: ‘Oh, and by the way, I guess puting community. “There’s a huge
publications,” says BioPerl developer we’ll keep this tool limping along.’” amount of resource that goes into
Jason E. Stajich, PhD, Miller As part of the NCBCs, Simbios and making these things, and so much of it
Research Fellow in the department of NA-MIC have specific funding for tool is just lost.” Bourne says.
plant and microbial biology at the maintenance and dissemination. “One Fortunately, funding agencies and
University of California, Berkeley. of the things that’s great about the journals are beginning to acknowledge
BioPerl is a programming toolkit for NCBC program is that there’s funding the importance of tool upkeep and dis-
processing sequence data. It has been to do actual training events,” Pieper semination. In the past few years, the
cited more than 500 times says. Finally, Schulten has had long- National Science Foundation (NSF)
([Link] standing (two decades of) tool-specific and NIH have “come around to the
Stajich worked heavily on BioPerl funding through an NIH P41 grant— idea that software is not something to
before and during his graduate studies which specifically funds technology be dabbled with,” Pande says. Lindahl
but, as he transitions to a faculty posi- development. These funds allow him to has also noticed an increase in tool-
tion, he needs to focus more on his sci- hire professional programmers and run specific funding. Journals could also
ence; and many other developers are training events. help alter the reward system, Bourne
in the same situation. “We’d like to do says. PLoS is contemplating a software
more outreach, but it requires a criti- MEASURING SUCCESS section where papers will only be pub-
cal mass of people who actually have AND REFLECTING ON FAILURE lished if the software is deposited in an
time to do that,” he says. The final step in tool dissemination open source archive such as source-
To augment the piecemeal model of is evaluation—measuring how well the [Link] or [Link]. Online
tool dissemination, some groups have efforts are going. journal editors or readers could simply
formed non-profits. For example, “It is extremely difficult to measure add a comment to papers when the
Stajich and his colleagues formed the the popularity of a free software project software is no longer available, Bourne
Open Bioinformatics Foundation, like FFTW,” Johnson says. Citations says. “That would sort of be a black
which provides infrastructure for provide a rigorous measure of success, mark against the author, so I think that
BioPerl and related projects, such as but these take time to accumulate. So, might encourage the author to make
BioJava and BioPython. Similarly, the our interviewees also track softer meas- the software available longer.”
Cytoscape Consortium provides an ures including: registered users, down- Even with more incentives and

In the past few years, the National Science Foundation (NSF)


and NIH have “come around to the idea that software is not
something to be dabbled with,” Vijay Pande says.

umbrella for the institutions involved loads, mailing list subscribers, mailing resources, tool dissemination will still be
in Cytoscape core development. The list activity, Web site visits, conference a challenge. Despite sufficient resources
non-profit model can help with logis- attendees, and the number of plugins and a proven track record in tool dis-
tics, including accepting donations and added to a tool. semination, Schulten says his latest
running conferences. This article focuses on tools that tool, BioCore ([Link]
Other tools in this article have succeeded. But, for every success story, Research/biocore/), is teetering on the
managed to obtain tool-specific fund- many more tools have failed. In a edge of failure. BioCore is a collabora-
ing, which was likely instrumental in recent editorial in PLoS Computational tive work environment for biomedical
their success. For example, APBS, Biology, founding editor-in-chief research, supporting tasks such as co-
GenePattern, and some members of Philip E. Bourne, PhD, a professor of authoring papers and sharing molecular
the Cytoscape Consortium have been pharmacology at the University of visualization results. The program hasn’t
funded through NIH’s R01 program for California, San Diego, and his col- taken off yet, in part because scientists
“software development and mainte- leagues describe their efforts to track are reluctant to try new technology, he
nance” (which has been available down 14 software programs (for parti- says. But Schulten is determined to
since 2002). GROMACS has also tioning proteins into domains) showcase the tool more and run more
obtained recent funding through the described in published papers. Eight training events. “We have to put more
European Union. The funding gives us programs were not even accessible in a energy into these efforts,” he says.
the ability to reply to user requests usable form, let alone widely used and Success requires persistence, Lindahl
within 24 to 48 hours and to develop popular. Given the difficulty of the agrees. “Don’t give up in the begin-
tutorials, Baker (of APBS) says. task and the lack of rewards, it’s not ning. It takes a while to build these
“Without that funding, that just would- surprising that so many tools languish communities.” ■

28 BIOMEDICAL COMPUTATION REVIEW Winter 2008/2009 [Link]


under the hood
Under TheHood
BY KARTIK MANI, PhD

Network-based Approaches
to Prediction of Disease Genes

T
he recent surge of high-through- naling, or other), which control a large
put experimental data, such as set of genes differentially expressed in
gene expression microarrays, the disease state. The third, the focus one particular disease phenotype (P).
offers a profound opportunity to gain a here, relies on the fact that interaction Formulaically, this test is represented as
more detailed understanding of the networks are themselves dynamic and the difference (⌬I) between Iall(G1;G2)
genes involved in the progression of may change from a normal to disease and Iall-P(G1;G2), where Iall includes all
disease. While initial analyses of these state. Thus, if one identifies interactions sample points, and Iall-P excludes the phe-
data used statistical techniques to iden- that have actually changed between notype P. Biologically, a positive or nega-
tify genes capable of distinguishing dis- phenotypes, one might then work back- tive ⌬I implies that these two genes have
ease tissue from normal (biomarkers), wards to identify genes that could prove gained or lost an interaction in the phe-
researchers are now turning to the promising for further investigation. notype P respectively (e.g., an oncogene
analysis of gene interaction networks to We will detail two examples of the “loses” its ability to be regulated in can-
address this problem. third category, both of which inciden- cer). The genes participating in a statisti-
Gene interaction networks may be tally use an information-theoretic cally significant number of these interac-
developed from several sources includ- approach. The first defines a concept tions are then selected. When applied to
ing manual curation, high-throughput called synergy, which measures the coop- data from three primary B cell lym-
experiments (such as yeast 2-hybrid), erative effect of two variables on the phomas, IDEA correctly predicted the
literature mining and reverse engineer- state of a third. The two variables in this known oncogenes reported in the litera-

If one identifies interactions that have actually changed between phenotypes, one might
then work backwards to identify genes that could prove promising for further investigation.
ing algorithms. They can include many case are genes (G1 and G2), and the ture (e.g., MYC in Burkitt’s Lymphoma),
different types of interactions as well third is a binary state variable represent- as well as effector genes not identified by
(complexes, regulatory, signaling, etc). ing disease or normal (D). Formulaically, differential expression analysis.
Integrating and analyzing all of this this can be represented as the difference These network-based approaches,
information to discover genes relevant to between I(G1,G2;D) (the cooperative along with others, have shown promise
disease requires network-based algo- effect) and the sum I(G1;D) + I(G2;D) in more accurately delineating the mech-
rithms. Thus far, such algorithms fall (the individual effects), where I is mutu- anisms of disease progression. Like any
into three general (though not necessar- al information. Biologically, synergistic new class of methods, however, there are
ily mutually exclusive) categories. The interactions imply that the combined drawbacks. First and foremost, there is no
first predicts protein complexes, rather state of the two genes affects disease, “gold standard” of gene interactions that
than individual genes, associated with while individually the genes have a far can be used, although the knowledge
the disease phenotype. The second iden- lesser or no effect. This algorithm com- base is growing rapidly. They often
tifies key regulators (transcriptional, sig- putes this quantity across all gene pairs require large training sets or sample diver-
represented on the input microarray sity to be effective, which may not always
data, and a “synergy network” is generat- be available. Lastly, computational com-
DETAILS ed from the highest scoring interactions. plexity may limit their applicability.
When applied to publicly available Nevertheless, the application of net-
Kartik Mani received his PhD in
prostate cancer data, this approach works and these algorithms to the identi-
Biomedical informatics at Columbia
showed the RBP1I gene participating in fication of disease-causing genes remains
University, working in the Multi-Scale
Analysis of Genomic and Cellular
a large number of synergistic interac- an exciting new area of computational
Networks (MAGNet) Center under the tions. This finding along with others biology. Expect to see several new net-
direction of Dr. Andrea Califano. His indicated that the progression of prostate work-based approaches emerge as the
research focused on the application of cancer is linked with oxidative stress and body of high-throughput and interac-
interaction networks to gene-disease inhibition of the apoptosis pathway, con- tion-based data continues to grow.
association, and culminated in the sistent with previous hypotheses.
development of the IDEA algorithm The second algorithm, Interactome REFERENCES
described above. He is currently Dysregulation Enrichment Analysis 1. Watkinson, J., X. Wang, et al.
pursuing his MD at the Albert Einstein (IDEA), computes the mutual informa- (2008). BMC Syst Biol 2: 10.
College of Medicine in Bronx, NY. tion between two genes across a large, 2. Mani, K. M., C. Lefebvre, et
diverse dataset, including or excluding al. (2008). Mol Syst Biol 4: 169. ■

Published by Simbios, an NIH National Center for Physics-Based Simulation of Biological Structures 29
Nonprofit Org.
U.S. Postage Paid
Permit No. 28
Palo Alto, CA

Biomedical Computation Review


Simbios AN NIH NATIONAL CENTER FOR BIOMEDICAL COMPUTING
Stanford University
318 Campus Drive
Clark Center Room S231
Stanford, CA 94305-5444

seeing science
SeeingScience
BY KATHARINE MILLER

Visualizing
Ventricular Fibrillation

U
nsynchronized twitching of the heart’s ventricles—known as ven-
tricular fibrillation—kills about 300,000 Americans yearly. Its
underlying cause: electrical spiral and scroll waves that
propagate through the heart. Simulation and visualization are
playing an important role in understanding that process.
In a novel approach to a review of the research, Flavio Fenton,
PhD, and Elizabeth Cherry, PhD, research associates in biomedical
sciences at Cornell University, simulated and visualized what’s cur-
rently known about how electrical spiral waves propagate through
the heart to cause tachycardia (rapid heart rate) and fibrillation.
The work was published in the December 2008 Visualization in
Physics focus issue of the New Journal of Physics. ■

Cherry and Fenton simulated electrical spiral waves through


the three-dimensional heart. In the Java Applet of this 3-D
simulation, we see a so-called “mother rotor” spiral wave on the front
of the heart. Although this might suggest a single spiral wave that
would cause only tachycardia (rapid heart beat), the 3-D heart can be
rotated in the Java applet to show the breakup of the wave on the
back of the ventricles—a sign that this heart would begin to quiver or
twitch uncontrollably in fibrillation. When this happens, no blood gets
pumped to the body or lungs.

Images reprinted with permission from EM Cherry and FH Fenton,


Visualization of spiral and scroll waves in simulated and experimental
cardiac tissue, New Journal of Physics 10 (2008) 125016, Figure 33d
Java applet. Also visit [Link]

Common questions

Powered by AI

Biomedical computational tools enable complex simulations and analyses that are crucial for understanding intricate biological systems. Their development and adoption ensure that researchers have access to state-of-the-art technologies, facilitate interdisciplinary collaboration, and promote discoveries that can significantly advance human health. These tools are crucial for addressing the growing complexity in biology and medicine .

Jerome Mettetal's team applied an engineering approach coupled with computer modeling and microfluidic arrays to study yeast cells. By confining yeast cells in chambers and feeding them in controlled, cyclical patterns, they investigated how yeast responds to varying osmotic pressures by adding bursts of salt. This method uncovered new roles for three different negative feedback loops in cells' equilibrium processes .

Studying yeast cell circuitry using information-processing perspectives and temporally varying inputs offers insights into key biological pathways and regulatory mechanisms, which can be extrapolated to more complex organisms. Such methods can reveal hidden dynamics in cellular regulation, thus contributing to broader applications in understanding human diseases and cellular responses .

The 'breaking point' methodology identifies critical signaling molecules by driving models beyond typical parameter ranges. This highlights previously unknown roles or interactions within signaling networks, facilitating hypothesis generation and verification. The approach can uncover novel insights into complex systems where traditional models fail, enhancing comprehensive understanding of cellular processes .

The digital embryo allows researchers to analyze developmental processes like zebrafish germ layer formation, which occur more synchronously than previously thought. Additionally, overlaying genomic data with the digital embryo can help identify genes that regulate vital processes, such as organ formation. This tool promotes advancements in tissue engineering and the study of tumor growth .

Computational models are used to simulate cellular responses by driving components beyond their observed experimental ranges to determine the weakest links, or proteins, causing computation failure. This technique highlights critical kinases leading to cytokine-induced apoptosis, challenging previously held models and enhancing understanding of protein roles in cell death mechanisms .

Disseminating biomedical tools involves ensuring the software is robust, well-documented, accessible, and supported, which often lacks funding and recognition. Overcoming these challenges requires building dissemination elements from the start, like including diverse skills in software development, documentation, and community building, as exemplified by the initiative under NCBCs .

NCBCs develop computational tools that facilitate research across various diseases by providing platforms for data analysis and integration. By enabling disease modeling, data sharing, and enhancing computational methods, these tools can transform biomedical research, improve clinical practices, and support personalized medicine approaches, showing their potential for widespread healthcare advances .

Fault diagnosis identifies critical molecules in cell pathways by drawing parallels with fault detection in electronic circuits. Highlighting crucial pathways can lead to more precise medicines by targeting safer molecules essential for cell function. This approach can reduce the toxicity of new drugs under trial by avoiding the targeting of molecules that are vital to cellular processes, paving the way for safer therapeutic interventions .

Integrating genomic data with clinical records allows for comprehensive large-scale studies that can identify reproducible patterns in genomics and clinical outcomes, potentially leading to the early detection of adverse drug events and comparisons of therapeutic efficacy across populations. This integration facilitates more cost-effective research and enhances personalized medicine by leveraging entire healthcare systems as study units .

You might also like