0% found this document useful (0 votes)
10 views22 pages

In Nov Perspective

Uploaded by

AmitAgrawal
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views22 pages

In Nov Perspective

Uploaded by

AmitAgrawal
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Grid Computing: Past, Present and Future

An Innovation Perspective

June, 2006
Table of Contents

1. Introduction ............................................................................................................3

2. The Evolution of Grid Computing.........................................................................4


2.1 The Grid Value Proposition .................................................................................................4
2.2 Customer Adoption..............................................................................................................6

3. The Grid Spectrum.................................................................................................7

4. Grid Solutions ........................................................................................................9


4.1 IBM Grid and Grow™.........................................................................................................9
4.2 IBM Grid Offering for Engineering Design: Clash Analysis in Automotive and
Aerospace.................................................................................................................................10
4.3 IBM Optimized Analytic Infrastructure.............................................................................10
4.4 IBM Grid Medical Archive Solution .................................................................................11
4.5 IBM Economic Development Grid....................................................................................12
4.6 IBM Global Services Grid Offerings.................................................................................12

5. Information Grid: An Information Infrastructure for Grid Computing...............13


5.1 Information Infrastructure..................................................................................................13
5.2 Challenges and Solutions...................................................................................................14
5.3 The Optimal Information Infrastructure ............................................................................15

6. The Importance of Standards ...............................................................................16


6.1 Web Services: A Foundation for Grid Computing ............................................................16
6.2 Higher Level of Grid Specific Standards...........................................................................17

7. Technology Development and Integration ..........................................................18

8. Grid and Service Oriented Architectures.............................................................18

9. Grid Ecosystem ....................................................................................................19

10. Conclusion: Innovation that Matters! ................................................................20

Page 2
1. Introduction
Over the last few years we have seen grid computing evolve from a niche technology associated
with scientific and technical computing, into a business-innovating technology that is driving
increased commercial adoption. Grid deployments accelerate application performance, improve
productivity and collaboration, and optimize the resiliency of the IT infrastructure. By
accelerating application performance, companies can more quickly deliver business results;
achieving greater productivity, faster time to market, and increased customer satisfaction.

Grid technology also provides the ability to store, share and analyze large volumes of data,
ensuring that people have access to information at the right time, which can improve decision
making, employee productivity and collaboration. Grid technology improves resource utilization
and reduces costs, while maintaining a flexible infrastructure that can cope with changing
business demands, yet remain reliable, resilient and secure.

At its core, grid is about virtualization, of both information and workload. In non-grid
environments, existing infrastructures are very much “siloed;” resources are dedicated to
applications and information. Many such dedicated infrastructures exist for common applications
such as HR, Payroll, etc. and for data/information mining purposes. System response is limited
by server capacity – and access to the data stored. It is very difficult to dynamically respond to
new requirements, as a new infrastructure would be required, inefficiencies would predominate,
and full utilization across the many silos would be difficult to achieve.

In a grid environment, resources are virtualized to create a pool of assets. Workload is spread
across servers and data can be seamlessly retrieved. By separating applications and information
from the infrastructure they run on, and providing this abstract, “virtualized” view, a new level of
infrastructure flexibility can be achieved. Infrastructures can now dynamically adapt to business
requirements, instead of the other way around. Resources are more fully utilized, resulting in
decreased infrastructure costs, reduced processing time, increased responsiveness and faster
time-to-market.

IBM has been a strong advocate and practitioner in facilitating the commercial adoption of grid
computing, even before the topic became a focus of media hype and analyst attention. Over the
years, IBM has made investments in all aspects of the grid domain: standards definitions,
technical development, open-source contributions, deployment of innovative business solutions,
and nurturing of a robust ecosystem that extends to software developers and business partners.
We view grid as a game-changing technology that challenges basic assumptions of ownership,
access, usage, operating efficiency, utilization of assets and total operating costs. A technology
that fosters innovation and collaboration while helping customers establish a competitive
advantage in their market.

With particular emphasis on the future of the grid marketplace, this paper investigates the roots
and evolution of grid computing germane to customer needs for associated business solutions
that span the grid spectrum. Starting with a quick view into the origins of grid computing, we
proceed to analyze its value proposition with respect to innovation and collaboration.
Furthermore, we present a holistic view of technologies associated with the grid and
virtualization domain, and focus upon the importance of developing and adopting industry
standards. We also explore significant market trends and dynamics and investigate critical
Page 3
challenges, especially those that extend beyond the compute elements of grid into data and
information management. Throughout this paper, we define IBM’s vision for grid computing and
position grid relative to other IBM initiatives.

2. The Evolution of Grid Computing


The phrase grid computing implies different technologies, markets and solutions to different
people. Early on, much of the available literature focused on the compute-intensive problems
made tractable by grid, often associating it with cycle-scavenging or job scheduling technologies.
Although these are important and useful components of a grid, they do not by themselves deliver
the complete grid vision.

The real “innovation” in grid comes from the combination of technology domains that include
workload virtualization, information virtualization, system virtualization, storage virtualization,
provisioning, and orchestration. From this statement, one may already conclude that no single
technology constitutes a grid, but, instead, the method with which broad sets of resources are
accessed and combined. Grid computing is not about a specific hardware platform, a database or
a particular piece of job management software, but the way in which IT resources dynamically
interact to address changing business requirements.

The grid domain has developed over a relatively short time period, fueled by significant
technology advancements. Many grid technology roots can be traced back to the late 1980s in
areas related to distributed supercomputing for numerically intensive applications, with
particular emphasis on scheduling algorithms (e.g. Condor1, Load Sharing Facility2). By the late
1990s, a more generalized framework for accessing high-performance computing systems and
distributed data (e.g. Globus3) began to emerge, and then, at the turn of the millennium, the pace
of change quickened with the recognition of the potential synergies between grid and the
emerging Service Oriented Architectures (in particular through the creation of the Open Grid
Services Architecture – OGSA).

2.1 The Grid Value Proposition

The dynamic and flexible properties of the grid establish a more competitive and innovative
business by enabling a more responsive IT infrastructure. IBM believes that this is not best
achieved through the deployment of proprietary protocols or technologies. Invariably, these lead
to ‘brittle’ architectures and solutions. Businesses need to be able to address their business issues
of the day, leveraging assets they possess. Grid establishes a common vision and method for
managing, referencing, and accessing the valuable IT resources available in an enterprise. This
capability typically does not result from having applications locked to particular platforms, data
to particular databases and user-access unnecessarily to particular systems.

The inherent flexibility that users and administrators/operators can derive is, perhaps, the most
subtle and least quantifiable benefit availed by grid computing. Traditionally, IT infrastructure
1
[Link]
2
[Link]
3
[Link]

Page 4
has been procured and deployed with a single purpose or function in mind. Systems have been
established and grown in alignment with particular vertical businesses. Corporations have grown
accustomed to managing a set of relatively distinct business units.

As the global market encourages the lowering of barriers to international trade, deregulation
fosters competition. The cost of entry into existing markets drops dramatically, the pace of
market change increases, and the need to combine cross-organization resources, products and
services grows. With this context, the traditional or “stove-pipe” IT solution is inefficient – in
some cases we see it as a support for significant inertia which inhibits changes in business
design.

There are different implications and derived benefits to the different functions within an
organization. Understandably, there are distinctions between the implications to a Line of
Business (LOB) executive, a Chief Information or Operation Officer (CIO or COO) and the
Chief Operating Officer (CEO) of a corporation.

LOB executives suffer from time, performance and quality pressures when driving to deliver to
the business particular results or products. In this case grid provides a potential avenue for
improving time to market and/or enhancing product quality – the highest value properties of grid
are those linked to: process or application acceleration, improved access to underutilized IT
resources, and system scalability. Embracing grid can give a “quick win” for those applications
or problem sets which are easily subdivided to be run concurrently across the distributed IT
infrastructure within the LOB department – and this is a characteristic of much of the early
commercial interest in grid.

IBM discussions with customers and LOBs often begin with a business requirement to more
quickly process a job, transaction, or set of tasks to meet a particular business deadline. The
processing or analysis may be anything from wanting to know the level of risk associated with an
investment portfolio, which is needed during the trading day through to the execution of multiple
Computational Fluid Dynamics (CFD) simulations to meet a particular design revision.

With good reason, CIOs/COOs are more concerned with cross-organization integration and
utilization of resources, enabling better access to information, mitigating business operating risk,
more easily introducing new applications and systems across the company, and finally, managing
heterogeneous systems while “keeping the whole thing running” at a reasonable total cost of
ownership. The grid benefits that are most often aligned with their needs are:
a) improved sharing of all IT resources offered into the grid and greater opportunity for
cross-organizational collaboration;
b) greater scalability of infrastructure by removing limitations inherent in the artificial
IT boundaries existing between separate groups or departments;
c) increased ability to launch new projects or initiatives without being limited by what
systems are available to a single group or department;
d) reduction of overall IT costs.

But probably the biggest benefits of grids are derived from client ability to achieve new levels of
innovation that can differentiate their business by implementing new business processes and
applications that they would have been unable to accomplish using conventional information
technology. Grid provides a virtual, resilient, responsive, flexible and cost effective

Page 5
infrastructure that fosters innovation and collaboration. In that regard, grid addresses the higher
needs of organizations often associated with a CEO’s agenda to find innovative ways to grow the
business, improve the productivity of employees, and provide a sustainable competitive
advantage.

2.2 Customer Adoption

Along with our profile of the motivations for grid computing, it is helpful to understand where
the technology has gained greatest traction. Being intensely attuned to the commercial sectors
which have showed the greatest affinity for grid computing, IBM has, correspondingly, aligned
its efforts with five grid focus areas

1. Business Analytics Grid: Enabling faster and more comprehensive business


planning and analysis through the sharing of data and computing power;
2. Engineering and Design Grid: Sharing data and computing power, for
computing intensive engineering and scientific applications, to accelerate product
design;
3. Research and Development Grid: Accelerate and enhance the R&D process by
enabling the sharing data and computing power seamlessly for research intensive
applications;
4. Government Development Grid: Create large-scale IT infrastructures to drive
economic development and/or enable new government services;
5. Enterprise Optimization Grid: Optimize computing and data assets to improve
utilization, efficiency and business continuity.

Business Enterprise
Analytics Optimization
Research

Manufacturing
Financial Life
Services Sciences & Government
Health Care & Education
Energy Derivative Mechanical/ Telco &
Analysis Electric Media Collaborative
Design Research
Drug
Actuarial
Discovery
Analysis Economic
Engineering Process
Simulation Protein
Bandwidth Development
Government
Asset Liability Consumption
& Design Seismic
Analysis Management Finite Element
Folding
Weather Development
Digital Analysis
Analysis Medical Rendering
Portfolio Risk
Reservoir Analysis High Energy
Gaming Physics

Grid Infrastructure

Figure 1. Grid Computing Focus Areas & Industry Applications

Page 6
Each of these areas may span multiple sectors and industry applications (see Figure 1). Some of
the early commercial adopters can be found in industries that exhibit close affinity to the grid
paradigm such as Aerospace, Automotive, Agriculture, Petroleum, Electronics, Financial
Services, Government, Higher Education, and Life Sciences. Furthermore, grid deployments
have started to gain more traction in new industries such as Media & Entertainment for digital
rendering and gaming applications. To support the needs in each specific sector, IBM has
developed multiple targeted solution offerings accompanied by client engagements and
references that help customers understand how grid can benefit their business.

3. The Grid Spectrum


Grid is best understood as part of a continuum representing virtualization technologies and
solutions (See Figure 2).

Virtualize Outside the


Enterprise

Virtualize the
Enterprise
Suppliers, partners,
customers and external

Virtualize Econommic
Unlike Developm
ment Grid
Enterprise wide Grids,
Information Insight, and
Virtualize Like Global Fabrics Grid Medical
Resources Archive Solution
Heterogeneous systems, Optimized Analytic
storage, and networks;
Application-based Grids
Infrastructure
Cluster
Clash Analysis in Automotive and
Single System Aerospace
(Partitioning) Simple Sophisticated
(2-4) (4+)
Homogenous systems,
storage, and networks
IBM Grid & Grow Heterogeneous
Homogenous
Single Organization Multiple Organizations
Tightly Coupled Loosely Coupled

Figure 2. The Grid Spectrum

At their simplest form, grid technology and solutions can be deployed within a homogeneous
environment or within a single organization in a tightly coupled manner. Virtualizing “like”
resources often entails deploying grid and virtualization functionality on cluster environments or
multiple single systems with partitioning. As discussed earlier, a common entry point for grid
deployments is a single line of business or single department using grids or clusters to increase
business value; application acceleration, meeting service level agreements, or gaining insights to
Page 7
data. Such adoption has been fueled by the ongoing proliferation of clusters within enterprises,
which often allows for a smoother transition using parallel applications. Another driver in this
space is the continued acceptance of Linux and open source for multiple applications and
workloads within enterprises. Indeed, many grid implementations in this space do rely on Linux
and open source software functionality.

The next level of virtualization extends the grid concept – still within the domain of a single
department or organization – to “unlike” resources. As a matter of fact, most departments often
run applications and processes that are composed of unlike resources, whether these are servers,
storage or other operating systems and software. The same principles of seamless integration
using grid technology still apply, despite the loss of homogeneity.

Where we “cross the chasm” is when we move from virtualizing unlike resources within a
specific organization or line of business to virtualizing across multiple lines of business (or
departments) within an enterprise. That is the point where (real and perceived) issues regarding
security, ownership, and governance surface. Grid is a powerful technology that challenges long-
held assumptions over ownership, access and usage. Clients are increasingly confronted with the
prospect of allowing their databases, their storage devices, their system processors to be
leveraged and used by others who did not pay for those assets. For a challenged LOB, grid can
be mistakenly viewed as a method by which their department’s valued IT assets are wrestled
from their local control and made available to other users in the enterprise.

We should point out that while such political struggles are often viewed as barriers to “crossing
the chasm” into enterprise, adopting grid technology does not dictate that a department
administrator loses authority or control over the resource. To the contrary, the resources are
offered into the grid under a set of policies or conditions determined by the owner. For the grid
to be effective and scale across the enterprise, it is important that enterprise IT establishes, up
front, a base set of resource-allocating policies. As examples: in offering a computing cluster into
a grid, the owner may declare the hours during which grid workloads will be able to run or what
percentage of the systems will be available to the grid; a database owner is free to determine
what data services are to be allowed; a storage system administrator is able to insist on particular
security requirements users must meet to be able to access its files. To assist in this process, IBM
offers grid middleware solutions that provide provisioning, orchestration, billing and service
level agreement type solutions.

In short, there is no specific reason why allowing grid-based access to a resource dilutes the local
administrator/operator control. But rather, one can manage the grid, get the benefit of leveraging
the additional resources as needed, and transfer that benefit to others within the organization.
That is when organizations are able to get much better balanced resource utilizations across the
entire IT environment.

Once organizations are able to get past the political barriers to enterprise optimization, they
begin to attain previously unavailable business value from grid by starting to integrate
horizontally. In that journey – which often equates to becoming an on demand business –
organizations should consider implementing potential new governance models, new businesses
and processes, and financial/accounting systems that provide incentives to participating
organizations.

Page 8
Finally, at the top right side of the spectrum, organizations start to virtualize with others outside
the boundaries of the enterprise. A grid environment can leverage resources with suppliers and
business partners and integrate business processes with the rest of the value net. These types of
grids exist today primarily in the public sector where governments and academic institutions link
their resources and share information to support collaboration and relationships across
organizations, countries, states or local governments. There is incredible potential, however, in
other industries such as automotive and aerospace where the design time and quality of an
automobile or plane, for example, can be significantly improved by sharing design data across all
the suppliers associated with the product.

To summarize, as customers move from left to right on the spectrum, they are moving from a
homogeneous environment within a single organization that is very tightly coupled to one that is
heterogeneous, distributed and loosely-coupled. It is important to note that there are multiple
potential customer entry points, thus the spectrum does not necessarily imply a sequential
progression. What is very important, however, is for the customers to preserve their ability to
expand on their grid implementations and grow as their business requirements grow.

4. Grid Solutions
IBM has the expertise, technical capability and solutions, supported by relationships with
business partners, to help clients achieve immediate value with grid with solutions that are right
for them. In addition, IBM can provide customers a clear path for grid expansion along the
progression map. IBM’s solutions strategy offers a set of repeatable solutions that can reduce the
time and risk of implementation. These solutions are based on repeatable patterns found across
numerous client engagements. Although there is a plethora or grid offerings in targeted market
segments and industries, in the following sections, we just highlight a few key solutions.

4.1 IBM Grid and Grow™

The IBM Grid and Grow™ offering provides an easy to deploy, integrated solution for customers
interested in beginning the grid journey. It includes industry-leading IBM and business partner
technology along with a "get started" services package to help first time grid customers
maximize the benefits of grid computing. Grid and Grow leverages a customer’s existing
investment and lays the foundation to expand to larger, more robust grid implementations,
further optimizing their infrastructure as their needs grow. This solution can also be leveraged to
expand capacity or build redundancy or to existing resources. The offering is also part of the
IBM “Express” solutions portfolio.

The offering features the IBM BladeCenter with seven blades and a choice of three server
architectures: Intel HS20, AMD LS20, or POWER JS20. These blades can be mixed and
matched in a single chassis and run one or more of RedHat or Novell SUSE Linux OS, AIX 5L,
or Windows. Additional blades can be easily added to fill out the existing chassis, or expanded to
multiple chassis for optimal scalability.

Customer choice continues with grid scheduler options: Altair's PBS Professional™, Platform
Computing's LSF, DataSynapse’s GridServer or IBM’s LoadLeveler. These enable application
scheduling, efficient dynamic resource allocation and resource sharing. Rounding out the
Page 9
offering is a services package that includes: a site readiness assessment, hardware and software
installation and tailoring, grid application readiness assessment, testing, documentation and
client training.

Customers can expand the initial Grid and Grow implementation to a larger, more robust
deployment to address additional business needs. Tools that could add significant value are
included in products from the IBM Tivoli and WebSphere portfolios. Some examples include
dynamic provisioning, dynamic software license tracking, and a standards-based secure portal.
IBM has designed the Grid and Grow platform as a solution building block that application
vendors can built upon with key applications that may be important to client business.

4.2 IBM Grid Offering for Engineering Design: Clash Analysis in Automotive
and Aerospace

In an intensely competitive marketplace, automotive and aerospace companies must achieve


faster time to market by decreasing the turnaround time for product design. By significantly
decreasing the amount of time required to assess whether new designs affect, or clash, with
existing product structures, designers can more quickly evaluate alternatives. In addition,
automobile and aircraft companies face great pressure to reduce IT investments and increase
return on investment (ROI), requiring them to maximize the use of their existing infrastructures.
Meanwhile, their heterogeneous environments – spanning many departments and
partners – are inherently complex. In this case, optimizing existing compute resources holds the
key to balancing market needs and costs.

The IBM Grid Offering for Engineering Design: Clash Analysis in Automotive and Aerospace
helps automotive and aerospace design engineers use grid technology for more rapid evaluation
of design alternatives during sub-assembly clash analysis. Developed in cooperation with
Platform Computing, the offering includes CATIA® and ENOVIA® application software. It helps
reduce the time required to capture, compile and analyze clash research data and can accelerate
product development and time to market.

The offering also includes a Grid Innovation Workshop for assessing and planning a grid
network, a pilot design and implementation services and comprehensive portfolio of IBM Global
Services product lifecycle management (PLM) for implementing and tuning product design, data
management and clash analysis software.

4.3 IBM Optimized Analytic Infrastructure

To be competitive in today's financial services marketplace requires performance of complex


analytics and computations in near real-time for a broad portfolio of products and activities such
as derivatives, structured and fixed-income products, risk management, program trading,
actuarial analysis, hedging and portfolio rebalancing.

From a business perspective, firms need to accelerate development of complex financial products
with shorter life cycles. A repeatable and consistent process for global deployment of
applications is also critical. This enables the pursuit of higher margins and revenue growth while
meeting client demands for innovation and specialization.

Page 10
From a technical perspective, financial services firms demand a dynamic IT infrastructure that
rapidly responds to changing business needs and requirements. This requires an operationally
efficient analytics infrastructure scalable for increasing data volumes, complexity, and a
spectrum of computational profiles. As well, this infrastructure needs to be inherently low
latency, extremely fast, standards-based and highly flexible, and address data center constraints
of space, power and cooling.

IBM Optimized Analytic Infrastructure (OAI) responds to these requirements and has been
designed to help financial services firms address their business and technical concerns so they
can compete more successfully in today's environment. The IBM OAI addresses a primary
requirement for financial market firms, to make rapid and extremely accurate decisions in a
stable, scalable and robust environment to support trading, analytics and risk-management. The
IBM OAI supports a broad spectrum of numerically intensive business processes and
applications by leveraging grid computing, HPC, Linux and blade technologies from both IBM
and Business Partners.

The IBM OAI includes products such as IBM’s GPFS for data management, IBM LoadLeveler®
for workload management, Cluster System Management for centralized administration, and the
IBM ApplicationWeb. These solutions have been used by some of the largest supercomputing
labs for over a decade to support a range of applications, such as high-energy physics, search
analytics, weather modeling, and electronic chip design on geographically distributed systems in
very large and dispersed user communities.

The IBM Optimized Analytic Infrastructure solution complements IBM technologies with
products from ISVs such as Altair PBS Professional (highly scalable scheduling environment),
Scali MPI Connect™ (MPI programming model) and GemStone Systems (message board/virtual
shared memory application environment) and the Linux operating systems (Red Hat and Novell
SUSE). All the IBM and ISV products have been tested using representative workloads to help
ensure full interoperability. Solutions like the IBM OAI will enable businesses to significantly
improve the speed and accuracy of decisions through the use of grid and HPC technologies.

4.4 IBM Grid Medical Archive Solution

The IBM Grid Medical Archive Solution (GMAS) allows multi-campus hospitals and imaging
centers to link geographically disparate sites – and modalities – together, helping optimize
storage utilization and eliminate redundancy. IBM GMAS brings together shared pools of
storage via deployment of intelligent grid software. Grid middleware create a “virtual” medical
imaging archive that pools distributed storage, yet does not require the consolidation of physical
images. By applying configurable business rules, storage grids can keep multiple copies of
images in geographically distributed sites, eliminating the need for a physical disaster recovery
site.

IBM GMAS is designed to be self-healing and self optimizing. It can also be deployed in support
of multiple applications and across heterogeneous storage hardware. Healthcare providers can
potentially gain significant ROI through increased storage utilization and simplified hardware
administration.

Based on Bycast StorageGRID software, IBM GMAS is designed to cost effectively deliver

Page 11
enterprise-wide medical image access, regardless of the image’s physical location or sourcing
system in an environment rich with security features. IBM GMAS delivers a unified storage
system that can support multiple picture archiving and communication systems (PACS), enabling
clinicians to view and share patient images at any time, from any location, using familiar PACS
interface.

4.5 IBM Economic Development Grid

IBM has launched an initiative to enable communities worldwide to stimulate economic growth
through the use of grid computing and other open standard technologies, such as Linux.
Cleveland is the first region to benefit from this Economic Development Grid initiative, which is
part of IBM's government development focus area to allow state and local governments, higher
education establishments, and local businesses to share information by leveraging computing
power and resources that benefit communities.

State and local governments continue to be challenged to find ways to collaborate and attract
new businesses to their communities to support economic growth, deliver and improve essential
services such as education and health care for their citizens, and create a climate for innovation
and growth to address future needs. Grid computing allows organizations to dynamically share
information and computer resources, because it provides benefits that are not available with
traditional IT infrastructure strategies. Grid computing applications in healthcare, life sciences,
software development, digital media, manufacturing and petroleum can all enable economic
benefits.

There are many types of grid computing implementations that can help communities drive
economic growth. A couple of examples include compute-intensive grids for software
development and medical research, as well as more data intensive grids that deliver collaboration
benefits for healthcare and education.

4.6 IBM Global Services Grid Offerings

IBM Global Services offers a variety of offerings aimed at moving customers beyond the
concept phase or helping customers expand their existing implementations. Some key offering
include:
• Accelerated Design Services for Grid: Developed by the IBM Grid Integration
Center in Austin, Texas - which integrates best-of-breed technology from IBM and
IBM Business Partners - the Grid Accelerated Design Services enable clients to build
grid solutions faster and more efficiently based on the experiences and results of other
clients that have successfully deployed grid solutions. The offering supports multiple
applications in a heterogeneous environment integrating grid middleware packages,
workload virtualization, storage virtualization, orchestration and provisioning, and
license management. In addition, the offering can help define and enforce policies
and priorities to control resource sharing across organizations.

• Grid Innovation Workshops: These two- and three-day workshops introduce grid
computing concepts, benefits and adoption frameworks, along with industry-specific
value propositions. Initial opportunities to leverage grid computing technologies are
identified, including business process considerations, top-line economics, technology
Page 12
architecture and potential risks.
• Grid Value At Work Tool: Developed by financial optimization experts from IBM
Research and IBM Global Services, this tool provides detailed and quantifiable
business value output for grid computing. Using industry templates, it examines
multiple grid and non-grid scenarios prior to implementation to predict application
performance and return on investment.

• Grid Computing Application Enablement: A service that enables applications to be


adapted to operate in a heterogeneous environment and to exploit grid computing
resources. The service also includes porting efforts to one or more of the platforms
running the grid for improved processing performance.

• Grid Solution Deployment Services: A service that includes the implementation of


infrastructure, application software, middleware, management tools and management
processes as needed for a successful grid deployment.

5. Information Grid: An Information Infrastructure for Grid Computing


Every year, data volumes increase 800 MB per user. The sheer quantity of information can
overwhelm the IT systems that must collect, store, retrieve, manage and protect it. Nearly one-
third of an IT staff’s time is spent searching for relevant data. Although the data may be timely,
is it easy to use? Is it integrated? Is it tailored to business needs, and is it cost-effective to
manage and retrieve? With a dynamic infrastructure, enabled by IBM grid computing
technology, customers can deploy capabilities that allow them to capture and analyze customer
information to increase the speed and accuracy of business decisions. IBM calls this capability
“Information on Demand”. IBM can help turn disparate data into true business insights. As a
result, more useful business information can be gained from the raw data that is stored.

5.1 Information Infrastructure

Establishing an information infrastructure for grid is a core component of the grid computing
model. It allows end users and applications secure access to any information source – regardless
of where it exists – over a local and/or distributed network in intranet, internet or even extranet
environments. It provides access to heterogeneous files, databases, storage systems and supports
data sharing for processing and/or large-scale collaboration.

The initial focus of a grid implementation may be on shortening the processing time of a single
application. However, as more applications and system resources become associated with a grid
environment, the need to consider how data is accessed and managed must be taken into
consideration. While a grid may be optimally constructed to intelligently schedule and manage
workloads, a poor information-access and storage-system deployment scheme could significantly
reduce the benefits that can be derived from a grid implementation. In a grid environment,
applications are not necessarily dedicated to running on specific processors (or nodes).

Applications can be provisioned onto different processors within the grid at any given time,
depending on their business priority. Moreover, nodes in the grid can be geographically

Page 13
dispersed. The challenge is making sure that data is easily accessible and doesn’t create network
bandwidth problems as a result of transporting it to computing locations in a distributed
environment. Customers would like to be able to have an easier method of pulling together data
and information from multiple “business” areas within and beyond the enterprise, without
disturbing the original format of the data (or how the data is managed at its source). Therefore, it
is necessary to ensure that any node in the grid can access the data ubiquitously, without having
to build a new path to access it. Physically consolidating data can be incredibly expensive, time-
consuming and negatively impact performance of applications.

Information-grid solutions may also address the following customer challenges:


• Fragmentation of data resources and assets due to a heterogeneous environment or
underutilized compute and storage resources
• Cumbersome data access and poor integration
• Data security and protection
• Complex management of decentralized systems and resources

5.2 Challenges and Solutions

The information grid solves the problem of managing information, which may include databases,
files, storage spanning across heterogeneous resources and software and hardware. The
following are some common computing challenges and their solutions.

• Challenge No. 1: Accessing “heterogeneous data” stored in different formats across


multiple “business” areas. The application must perform multiple I/O requests to retrieve
the data that slows down the execution of the job. Programmers that build and maintain
these applications must be aware of different formats and determine how to transform
and join the data within their applications.

Solution: Data-access virtualization technology across diverse data formats is


instrumental in helping solve the challenge of reading data stored in different formats.
Programmers can simplify access to data that are stored in mixed formats (e.g.,
multivendor rational databases, flat files) by enabling these data to be accessed with a
single structured query language instruction. Such access also helps reduce the need to
move remote files.

The data’s virtual view is also known as federated access to the data. That is, making the
data appear as one source even though the data are distributed and stored in mixed
formats. In the event that large volumes of data need to be tailored for an application,
specific extraction, transformation and loading functions can be preformed on the grid.
Once the data has been prepared into a proper format for an application, the data can be
temporally “transported” to and cashed at a location where the processing will take place.

The ability to cleanse, transform, federate and analyze data across “heterogeneous data”
sources is crucial to gaining information insight and making more effective business
decisions. WebSphere DataStage, ProfileStage, QualityStage and WebSphere Information
Integrator address this "heterogeneous data" challenge. These products are being
integrated into the “IBM WebSphere Information Server”.

Page 14
• Challenge No. 2: Data discovery and information delivery in a grid environment with
mixed file system types. It is difficult for application developers and end-users to locate
and access data because data are stored under multiple directories that are associated with
each file system type.

Solution: Poor storage-resource utilization can be resolved through the use of SAN
technology. Optimal solutions would include SAN software that enables system
administrators to create a virtual view of all of the SAN storage, making them appear to
be one homogeneous set of files with a common name space. Also, it is necessary to
move large data volumes across a network to facilitate remote processing.

A software solution that addresses this challenge should enable data to be cached close to
where distributed processing occurs. An ideal solution would include global naming,
secure wide-area access to consistent, current data and distributed data access including
POSIX/NFS interface, access control and remote-data caching. Similarly, virtualization
of heterogeneous file systems can help manage a complex SAN environment. Creating a
single name space for the file system helps programmers and administrators locate and
access data more easily versus having to identify files individually and determine what
access path is required to reference the data. IBM’s General Parallel File System or SAN
File System can be used as foundations for this solution

• Challenge No. 3: Mixed vendor storage is common within any given enterprise. It is
costly for administrators to manually manage data placement across heterogeneous sets
of storage devices. In many cases, bottlenecks occur in retrieving data from these devices
due to congestion of over utilized devices even though there may be space available on
underutilized storage resources.

Solution: Frequently, customers have heterogeneous (multi-vendor) storage devices


installed. Each vendor’s storage device comes with its own management console, which
makes it difficult to efficiently manage data placement across the various devices and
ensure that there isn’t uneven data loading. Uneven data distribution could cause some of
the devices to be over utilized while others remain underutilized. This unbalanced
condition could lead to bottlenecks when attempting to retrieve data, thus slowing the
application processing.

A virtualization portal that consolidates the view across all of the SAN devices allows a
single administrator to see how data is being loaded on these devices. Administrators are
then able to shift data from over utilized devices to those that are underutilized without
disrupting how applications access the data. IBM has the ability to deliver the building
blocks for such solution having industry-leading storage virtualization products such as
SAN Volume Controller and SAN File System. Other considerations in a SAN
environment include error detection and data resiliency. It is important that data are
protected and secure while providing the right data to applications.

5.3 The Optimal Information Infrastructure

Once the aforementioned solutions have been applied, the information grid has been established.
Page 15
These solutions address many potential problems of accessing data, managing heterogeneous file
and storage systems and removing the network effect of supplying data for remote processing.
By applying each of these solutions to create a virtual environment, distributed computing in a
grid environment can achieve its maximum benefits. The enterprise has the flexibility to utilize
all of its computing assets by enabling information to be virtually managed and presented.

6. The Importance of Standards


IBM is a strong advocate for open industry standardization in information technology. Not only
do open standards enhance interoperability, integration and customer choice, but also create an
important opportunity for communities of vendors, governments, universities and researchers to
innovate in the development of new computing paradigms. The very nature of grid computing,
which tries to take broadly distributed, heterogeneous computing and data resources and
aggregate them into an “abstracted” set of capabilities, almost demands open standards for
integration and interoperability.

Most organizations that “build” grids do not purposefully acquire new hardware, operating
systems and software to create their solutions, but rather integrate existing heterogeneous
systems into a fabric over which they can deploy applications. While many vendors have
successfully implemented distributed systems with proprietary interfaces, it can be argued that
only “true grid” systems based on standards are capable of achieving the “scale out” promised by
the “grid vision” – where an application can exploit any processing capability required, access
any data it needs, and not be concerned with the specifics of configuration, management, or
infrastructure.

6.1 Web Services: A Foundation for Grid Computing

Our fundamental approach over the past three years has been to base grid computing architecture
and standards on the emerging foundation of a “services oriented” model that is increasingly
being adopted by the IT industry. Web services provide a component model for the composition
of functions that make up and support distributed systems and grids. The loosely coupled model
of constructing applications, management functions or infrastructure using a collection of web
services is ideal for building scalable, flexible, dynamic systems like grids.

IBM has provided significant technical leadership to develop the WS-Resource Framework (WS-
RF) and WS-Notification Framework under the auspices of the Organization for the
Advancement of Structured Information Standards (OASIS). WS-RF and WS-Notification
provide the needed “stateful” web services environment on top of which other grid specific
standards can be implemented. More recently, IBM in partnership with Microsoft, Intel, and
Hewlett Packard has been working on a “convergence” plan for these and a number of other
overlapping specifications.

In addition to these very fundamental web services standards, there are additional standards
being worked on that add important functional capabilities like security (WS-Security), service
level management (WS-Agreement) , policy expression (WS-Policy), etc.

Page 16
6.2 Higher Level of Grid Specific Standards

Grid Standards

Beyond these fundamental and very general purpose web services specifications, is another
group of standards that builds on top of them, to define more functional protocols and operations
(e.g. scheduling and workload management, application deployment, resource provisioning, data
movement, and data access). The Global Grid Forum (GGF) is involved in the definition of a
number of important standards at this level. Most notable is a broad architectural focus called the
Open Grid Services Architecture (OGSA) which defines a very rich vision of execution,
management and data services to enable the creation of grids. OGSA is not a single specification,
but rather a set of related standards in several areas. Some important OGSA-related
specifications include:
• OGSA Basic Profile
• OGSA Security Profile
• Basic Execution Services (OGSA-BES)
• Job Submission Description Language (JSDL)
• Data Access and Integration Services (DAIS)
• Configuration Description, Deployment, and Lifecycle Management (CDDLM)
• OGSA Byte I/O (ByteIO)

Information Model

While several existing standards have embedded within them some form of resource
representation, it has become apparent that a truly common framework for an abstract
information model is needed. To this end, the Global Grid Forum (GGF) and other standards
bodies have increasingly turned to the Distributed Management Task Force (DMTF) Common
Information Model (CIM) as a touchstone. CIM has been developed over a number of years to
describe all kind of IT resources, from very high level conceptual capabilities to very specific
low level components. While the GGF and OGSA working groups have not yet formally
identified DMTF CIM as the information model that they will use for grid computing, they are
working towards that direction.

Management Standards

Along with the push to develop a common modeling approach for web services based standards,
there has been a strong motivation to establish common high level protocols and operations for
managing resources that are exposed as web services. OASIS’ Web Services Distributed
Management (WS-DM) is an industry-wide standard for management both using web services
and managing web services. WS-DM attempts to exploit web services technology to create a
universal and consistent abstraction for management and manageability interfaces that leverage
key features of web services protocols. The specific types of management capabilities exposed
by WS-DM include:
• Monitoring the quality of a service associate with a service
• Enforcing a service level agreement (quality of service)
• Querying or controlling the basic operational state of a resource
• Managing a resources lifecycle (create / destroy)
Page 17
As in the case of the foundational web services standards, there is an ongoing effort to develop a
“convergence” plan for overlapping management standards such as WS-DM and WS-
Management.

In summary, there is a growing collection of related standards and architecture being developed
in open standards bodies like IETF, W3C, OASIS, DMTF and GGF that are all based on web
services and can be composed to help develop interoperable grid middleware and infrastructure.
IBM has been a leader in the definition and development of these open standards and is driving
its important implementations along this standards roadmap.

7. Technology Development and Integration


IBM is integrating grid functionality (provisioning, workload scheduling, resource management
and information virtualization) across the software portfolio, and taking a leadership position in
accelerating grid standards development and adoption across the industry. IBM's grid technical
strategy extends the company's leadership in software, systems, and storage virtualization by
focusing on three main areas: workload virtualization, information virtualization, and grid
management. In addition, IBM is actively working on new programming models, tools, and
techniques for developing and enabling distributed grid applications. For all of these areas, IBM
is actively involved and is a leader in creating and developing the appropriate grid and Web
services standards.

IBM's workload virtualization strategy is to create a single, logical view of workload scheduling,
differentiated with IBM automation and management technologies. This will enable clients to
dramatically accelerate performance of multiple large application workloads across their
enterprise, leveraging and orchestrating IT resources in a more flexible and dynamic fashion than
ever before. Interoperability, driven by standards, is important in this space to provide a cross-
organization workload management capability spanning different types of scheduling
environments and domains.

IBM’s information virtualization strategy is to create an integrated view of storage, file systems,
and databases driven by standards, interoperability and advanced technologies like data
transformation, security, caching and replication. The end result of this strategy will be the
ability for clients to gain insight from disparate information federated systems in ways never
before envisioned. The integration of these core grid technologies enables HPC solutions to
deliver accelerated application performance, increased asset utilization, and information insight
across disparate IT resources, data centers and geographies.

8. Grid and Service Oriented Architectures


Customers are implementing Service Oriented Architectures (SOAs) to reduce complexity, to
enable flexibility, and to streamline business processes. A SOA is a framework that lets
customers build, deploy and integrate services for IT resources, applications and business
process flows. It facilitates integration, offers modularization of applications and provides a
coherent view of a business process as a set of coordinated services. Therefore, customers can
Page 18
build and integrate applications and business processes across their enterprise, as well as with
partners, suppliers and customers, in a much more on demand fashion.

In order to realize these benefits, companies need to establish a flexible infrastructure capable of
supporting dynamic operations with flexibility inherent throughout. As companies continue to
implement SOAs throughout their enterprise, there are going to be a lot of services, some very
fine grained and some larger, that will require a very responsive, dynamic and scalable
infrastructure. These services can be mobile. They need to be resilient. And the whole process
needs to be accomplished in a simple manner. The underlying infrastructure needs to be
adaptable and autonomic in providing the right resources to the right application services and
business policies to meet service level agreements and business performance needs.

The necessary infrastructure consists of our leading capabilities across system and resource
virtualization, and grid computing. IBM views grid computing as critical to the ongoing
development of a dynamic and flexible infrastructure that enables SOA:
• Just as an SOA allows customers to separate applications from services, grids allow
customers to separate both applications and services from the infrastructure and systems
resources. Scheduling and workload management are key capabilities here supporting the
placement and mobility of services and composite applications to the appropriate
resources. Grids provide an underlying foundation to support the dynamic nature of SOA.
• Companies can pool resources for services, improve availability and reliability, and
rapidly deploy and scale this new class of composite applications.
• With a virtualized infrastructure, customers can much more easily scale their support for
SOAs. They can harness all of their resources to accelerate time to results and to better
align their infrastructure performance to their business goals.

None of this will be possible at the SOA level without a "commitment to openness". It is the only
way to link virtualization and grids with SOAs in a consistent and uniform manner. The
infrastructure has to thoroughly support the ability to be managed in real-time, with sophisticated
monitoring capabilities.

Finally, the SOA and virtualization landscapes require a significant ecosystem of partners and
capabilities. We need to see cross industry collaboration to drive the innovation required for
SOA. In simple terms, grid computing, based on open standards and supported by collaborative
communities, delivers the underlying dynamic infrastructure that enables competitive advantage
for customers implementing SOA-based solutions.

9. Grid Ecosystem
The primary basis of grid is that it can accommodate highly-heterogeneous technologies.
Therefore, the viability of such technology requires the emergence of a strong, and viable,
ecosystem that includes software vendors and business partners. By abstracting the physical
resources, as well as the virtual resources, software developers can widen their target customer
base. Specifically, by preparing their applications to run across infrastructures, as opposed to
being the infrastructure, they make a better case for utilizing equipment that their application
process requires.

Page 19
To ensure that software developers and partners are able to adapt their products and services for
use in the grid, IBM is nurturing an ecosystem to assist the grid solution creation process. Our
efforts range from educating, enabling and assisting partners and independent software vendors
(ISVs) on when and where grid technologies apply, through to helping create reference
implementations for particular grid solutions. Some key examples include:
• developerWorks GridZone: The developerWorks GridZone provides software developers
with tools, online training, IBM Redbooks, articles, emerging technologies from IBM
Research, and more, to help them develop grid computing applications.
• IBM Innovation Centers: At the centers, the technical consultants work with the ISVs to
help them implement their application topologies on a grid infrastructure. The
infrastructure can consist of any hardware platforms or any of the supported operations
systems.
• The Solutions Enablement Virtual Loaner Program (VLP): The VLP uses grid computing
and other on demand technologies, such as the IBM Tivoli Provisioning Manager, to
provide a rich and flexible software development environment for remote-access use by
ISVs. ISVs are able to reserve, in advance, resources on the grid to satisfy their need for
low-cost access to current IBM hardware and middleware to develop, port, test and
validate their applications.
• IBM Ready for Grid Program: IBM's Ready for Grid computing program validates that
an application is capable of executing and realizing benefits from running in a grid
computing environment. The new program also includes "The Ready for IBM GRID
Computing" mark, which is a critical component of IBM's strategy to create a robust
ecosystem with our partners around open grid standards.
• Value Network Initiative: The Value Network Initiative builds networks of partners who
can effectively deliver grid solutions. This program offers select partners access to
enhanced PartnerWorld Industry Network (PWIN) co-marketing benefits.

10. Conclusion: Innovation that Matters!


Throughout the last decade, grid computing has emerged as a promising virtualization paradigm
capable of providing a flexible, dynamic, resilient and cost effective infrastructure that can
promote both collaboration and innovation within enterprises. Advancements in grid computing
are driving increased adoption into commercial lines of business spearheaded by customer desire
to create a competitive advantage through more responsive SOAs.

IBM is poised to extend its leadership in grid computing in order to advance innovation that
matters across a much broader commercial marketplace, extending its benefits to community and
society. Life sciences organizations, for example, already are leveraging the new efficiency
afforded by grid technology to advance human health. In one example, scientists at the
University of Pennsylvania have found a way to use grid computing to promote early detection
of a disease that affects one out of every eight women today: breast cancer. The technology
solution is known as the National Digital Mammography Archive (NDMA). The product uses
grid technology to facilitate the complicated data access and analysis required for timely and
accurate breast cancer screening.

In another example, the LA Grid program will link faculty, students and researchers from the

Page 20
world renowned IBM T.J. Watson Research Centers across the United States, Latin America and
Spain to collaborate on innovative industry projects for applications in areas such as health care,
life sciences and nanotechnology; and in regionally-specific concerns like hurricane mitigation.

Furthermore, IBM announced a new research effort to help battle AIDS using the massive
computational power of World Community Grid, a global community of computer users who
have joined the philanthropic technology initiative by simply donating unused time on their
personal computers. The new World Community Grid initiative will deploy massive computer
power to develop novel chemical strategies effective in the treatment of HIV-infected individuals
in the face of evolving drug resistance in the virus. Developing new, more robust therapies to
prevent the onset of AIDS in individuals infected with HIV will be the focus of this innovative
project.

Grid computing continues IBM's history of IT innovation for business. IBM has the skills,
experience and resources to attack the most challenging customer problems in the industry – on a
global scale, with a deep understanding of clients’ environments and requirements.

Page 21
Author:
Elias Kourpas, IBM

© IBM Corporation 2006

IBM Corporation
Systems and Technology Group
Route 100
Somers, New York 10589

Produced in the United States of America


June 2006; All Rights Reserved

This document was developed for products and/or


services offered in the United States. IBM may not
offer the products, features, or services discussed in
this document in other countries.

The information may be subject to change without


notice. Consult your local IBM business contact for
information on the products, features and services
available in your area.

All statements regarding IBM future directions and


intent are subject to change or withdrawal without
notice and represent goals and objectives only.

Linux is a trademark of Linus Torvalds in the United


States, other countries or both.

Other company, product, and service names may be


trademarks or service marks of others.

IBM hardware products are manufactured from new


parts, or new and used parts. Regardless, our
warranty terms apply.

Photographs show engineering and design models.


Changes may be incorporated in production models.

Copying or downloading the images contained in this


document is expressly prohibited without the written
consent of IBM

This equipment is subject to FCC rules. It will comply


with the appropriate FCC rules before final delivery to
the buyer.
Information concerning non-IBM products was
obtained from the suppliers of these products or other
public sources. Questions on the capabilities of the
non-IBM products should be addressed with those
suppliers.

The IBM home page on the Internet can be found at:


[Link]

The IBM Grid Computing home page on the Internet


can be found at: [Link]

Page 22

You might also like