Design Notes
Design Notes
Centers
CSIWZG522
BITS Pilani Prepared by
[Link]
BITS Pilani
• Students are strongly advised to make notes from RLs & CSs for examination
purpose right from the beginning.
• Unavailability of the textbook will not be an excuse to allow print copies of the e-
book during examinations.
RECAP OF PREVIOUS SESSION
Availability & Reliability (Uptime): Defined by tier levels (Tier 1-4), requiring
redundant UPS, generators, and cooling to prevent downtime.
Site Selection & Connectivity: Proximity to fiber providers, power grid reliability,
and low latency for users.
Source: H5Datacenter
Data Center Risk Assessment
metrics…
Data Center Risk Assessment
metrics…
Assurance Strategies
Security and Reliability Benchmarks: Adherence to ISO/IEC 27001 for information security
and SOC 2 Type II for operational integrity is the "gold standard" for building client trust and
ensuring reliable service delivery.
Evolving Sustainability Reporting: New 2026 mandates, such as the EU Energy Efficiency
Directive (EED), now require data centers to report precise performance metrics
like PUE (Power Usage Effectiveness) and WUE (Water Usage Effectiveness).
Uptime Institute Tiering: While a certification rather than a law, this defines the global
standard for infrastructure resilience and redundancy from Tier I to Tier IV.
ISO/IEC 22237 (Facility Design): A specific holistic standard for the physical
infrastructure of data centers, covering power, cooling, and cabling systems.
Environmental Mandates (EED & LEED): Regulations like the EU’s Energy Efficiency
Directive force facilities to report energy use, while LEED certifications track
sustainability.
FedRAMP (Government Services): Essential for data centers in the U.S.
providing cloud services to federal agencies, requiring high-level security authorization.
ISO 50001 (Energy Management): An international framework that helps data centers
improve energy performance and reduce carbon footprints through structured
management.
Avoiding Lock In
What is Lock-In? Risks of Lock-In
• Dependence on a single vendor or technology, Financial Risks
making it difficult to switch providers or systems. Increased costs over time, lack of competitive pricing.
Operational Risks
Importance of Avoiding Lock-In Difficulty adapting to new technologies, inability to
scale.
• Flexibility, cost control, and future-proofing.
Strategic Risks
Reduced innovation, limited access to new features.
Types of Lock-In
Vendor Lock-In Strategies to Avoid Lock-In
– Dependency on specific hardware, software, or Open Standards and Protocols
services. Use open-source technologies and standards to ensure
Technology Lock-In compatibility.
Modular Architecture
– Relying on proprietary technologies that limit
Design systems that can be easily modified or replaced.
integration.
Multi-Cloud and Hybrid Solutions
Contractual Lock-In Leverage multiple cloud providers to prevent
– Long-term contracts that restrict mobility and dependency on one.
negotiation.
Changing the perception of IT Cost
Changing the perspective of IT cost in data center design means shifting from CapEx-
heavy infrastructure (building for peak capacity) to OpEx-focused
efficiency (building for actual utilization).
• Total Cost of Ownership (TCO) vs. Initial Capital: Modern design prioritizes TCO over 10–15 years rather
than the initial build cost.
• Power Usage Effectiveness (PUE) as a Cost Lever: Efficiency is no longer just "green"—it's a direct cost
saving. Reducing PUE from 2.0 to 1.2 can slash millions in annual utility bills for large-scale facilities.
• Modular & Scalable "Pay-as-you-grow" Design: Instead of over-provisioning (which wastes capital), modular
data center architecture allows IT to add power and cooling units only when demand rises, deferring capital
expenditure.
• Software-Defined Infrastructure (SDI): Shifting complexity from hardware to software allows for the use of
"commodity" hardware, reducing vendor lock-in costs and making hardware refresh cycles less expensive.
• Data Gravity & Egress Optimization: Design around "data gravity" by placing infrastructure near cloud on-
ramps. This slashes the massive monthly costs of moving data between private servers and public clouds.
The Changing Budgetary Landscape
Budgets are moving from monolithic capital investments toward flexible, high-density, and speed-oriented
spending models.
• Infrastructure Investment Supercycle: Global data center expenditures are projected to reach $3
trillion by 2030. Construction costs for shell and core alone are rising at a 7% CAGR, averaging
over $11 million per MW in 2026.
• Massive AI "Tech Fit-Out" Costs: While shell construction is expensive, the IT equipment for AI
(GPUs, high-speed networking) can cost an additional $25 million per MW.
• CapEx to OpEx Transition: To maintain liquidity, enterprises are shifting away from owning physical
hardware. Instead, they are adopting "as-a-Service" (DaaS, IaaS) models, converting large upfront
costs into predictable monthly operational expenses.
• Premium for Liquid Cooling: AI workloads require high-density racks (up to 100kW), making liquid
cooling a necessity. These facilities carry a 10% cost premium over traditional air-cooled sites.
• Modular & Standardized Spending: To combat long construction timelines (often 4+ years for grid
connections), budgets are prioritizing modular construction and on-site power generation (gas
The Changing Career Landscape
The "traditional" data center technician is being replaced by hybrid roles that blend physical engineering with
software automation.
The Rise of the "Super-Specialist": High demand exists for Electrical and Power Infrastructure Engineers
capable of designing microgrids and high-voltage systems to handle the doubling of global data center
power use by 2030.
Sustainability & ESG Analysts: Specialists who can optimize Power Usage Effectiveness (PUE) and
integrate renewable energy are now "must-have" roles due to rising regulatory and carbon-tracking
pressures.
Shift from "Break/Fix" to "Supervise/Optimize": AI tools now handle routine thermal audits and hardware
troubleshooting. Human roles are evolving into AI Infrastructure Operations Engineers who validate AI
decisions and manage digital twins.
Infrastructure as Code (IaC) Mastery: Knowledge of software-defined everything (SDx) is no longer just for
DevOps. Physical infrastructure managers must now be comfortable with automation and scripting to
manage zero-touch provisioning.
Commissioning & Project Engineers: With 10 gigawatts of new capacity breaking ground annually, there is
a fierce talent war for engineers with hands-on build experience who can coordinate complex, multi-site
construction.
Complexity, Agility and Performance
AI-Driven Density (Performance): The shift from 10kW to 100kW per rack is the new performance baseline. To
achieve this, designs are moving from air cooling to Direct-to-Chip Liquid Cooling, which is essential for the
thermal demands of next-gen GPUs.
Modular "Block" Architecture (Agility): Forget 24-month builds. Using pre-fabricated modular units allows
operators to deploy capacity in "just-in-time" increments, slashing time-to-market and keeping capital fluid.
Digital Twin Management (Complexity): To manage the sheer density of components, designers now use Digital
Twins. These virtual replicas simulate power and thermal changes before physical deployment, preventing
"stranded capacity" and catastrophic failures.
Software-Defined Everything (Agility/Complexity): By abstracting hardware into Software-Defined Data Centers
(SDDC), networking and storage become programmable. This reduces manual configuration errors (Complexity)
and enables instant resource reallocation (Agility).
Microgrid Integration (Performance/Agility): With utility grids facing 5–10 year wait times, modern designs
integrate on-site microgrids (battery storage, hydrogen fuel cells). This provides the power performance needed
for AI without being held hostage by city infrastructure.
Edge-to-Core Continuity (Performance): Performance is now measured by latency. Placing high-
performance Edge Data Centers closer to users reduces data "travel time," while high-speed fiber backhauls link
them to the massive "training" core.
Data Center Automation
Data center automation is the process by which routine workflows and processes of a data
center—scheduling, monitoring, maintenance, application delivery, and so on—are
managed and executed without human administration.
The massive growth in data and the speed at which businesses operate today mean that
manual monitoring, troubleshooting, and remediation is too slow to be effective and can
put businesses at risk.
Automation can make day-two operations almost autonomous. Ideally, the data center
provider would have API access to the infrastructure, enabling it to inter-operate with
public clouds so that customers could migrate data or workloads from cloud to cloud.
Data center automation is predominantly delivered through software solutions that grant
centralized access to all or most data center resources.
What can you automate in your data
center?
Key Element Description
Automate host setups (both physical and virtual) to minimize manual errors and
Host automation
enhance the speed of deployments.
Automate patching at scale, lifecycle management, auditing, monitoring, security
Day-2 Automation
management, and cost management.
Storage Automate data assignment to different types of storage media for increased
automation performance, reduced costs, and efficient data management.
Network Automate network functions, including configurations and management to reduce
automation manual errors and increase the deployment velocity of network services.
Security Automate repetitive and low-level tasks to enhance the identification and
automation remediation of security and compliance risks.
Environment Utilize software-defined Infrastructure to manage and control data center
control automation utilization and energy consumption of all IT-related equipment.
Infrastructure and
Main mechanism to achieve data center automation.
policy as code
Orchestration for self-service data
center
This enables automated, on-demand provisioning of IT infrastructure via
portals, eliminating manual, error-prone configuration tasks.
By using Infrastructure as Code (IaC) and Multi-Domain Service
Orchestration (MDSO), organizations can rapidly deploy virtual/physical
resources, ensuring consistency and compliance while enabling developers
to request, manage, and deploy services independently.
– Infrastructure as Code (IaC): Use code to define, manage, and provision infrastructure, which ensures
consistency and allows for version control.
For modern data centers, downtime is not an option, particularly not unplanned
downtime. It is therefore vital that all data centers achieve the highest
possible standards of data center resilience.
In the context of data centers, the term “resilience” refers to the ability of an
infrastructure to withstand and recover from disruptions. The more resilience
a data center has, the more likely it is to be able to maintain uninterrupted
services at all times.
In other words, higher resilience generally translates into lower downtime.
Strategies for achieving high availability
Here are five key strategies for achieving and maintaining high availability in data
centers.
Redundancy and replication:
• By duplicating critical components and services, organizations create backups
that can seamlessly take over in case of failures. This applies not only to hardware
but also to data.
• Replicating data across different servers or even geographical locations ensures
that if one location experiences an issue, another can seamlessly continue
operations.
Load balancing:
• Load balancing is a key practice to prevent any single server from being
overwhelmed by incoming requests. By distributing traffic across multiple servers, the
load on each server is optimized, and the risk of overload-induced failures is
significantly reduced.
Strategies for achieving high
availability….
Failover mechanisms: Failover mechanisms ensure a swift transition from a failing
component to a backup. They often take the form of an active-passive setup where a
standby system takes over when the primary system experiences issues. Automated
failover systems are designed to detect problems and initiate the switch to maintain
continuous service availability.
Scalability and elasticity: Systems must be able to adapt to changing workloads by
automatically allocating or deallocating resources. This ensures continuous performance
even during sudden traffic spikes or increased demand. It therefore enhances overall
system robustness and minimizes the risk of downtime due to resource constraints.
Data center resilience best practices: These include comprehensive disaster recovery
planning, designing for loose coupling among components, and prioritizing security
measures. Resilient architectures consider the entire system’s lifecycle, including startup
dependencies and graceful degradation under stress.
Case studies
Here are three real-world case studies of data centers that have been designed for high resilience.
Google: Google’s data center resilience is based on redundancy and automatic failover mechanisms.
By replicating data across different locations, Google minimizes the risk of a single point of failure.
Load balancing ensures even distribution of traffic among multiple servers, preventing overload on
any single server. Automatic failover systems swiftly transition from primary to standby systems,
ensuring uninterrupted service.
Amazon Web Services (AWS): AWS prioritizes resilience through diverse availability zones, each
with its own infrastructure and facilities. This geographical separation minimizes the impact of
localized failures. Load balancing distributes traffic across multiple servers, maintaining optimal
performance and mitigating the risk of overload. AWS’s auto-scaling feature dynamically adjusts
resources based on demand, ensuring scalability and consistent availability.
Microsoft Azure: Azure focuses on resilience through distributed architecture, redundancy, and
failover strategies. Availability zones enhance fault tolerance, and Azure’s scalability ensures
adaptability to varying workloads. Data replication across regions boosts Azure’s disaster recovery
capabilities.
Design and Operation of Data
Centers
CSIWZG522
BITS Pilani
Pilani Campus
IMP Note to Self
IMP Note to Students
• It is important to know that just login to the session does not guarantee the
attendance.
• Once you join the session, continue till the end to consider you as present in the
class.
• IMPORTANTLY, you need to make the class more interactive by responding to
Professors queries in the session.
• Whenever Professor calls your number / name ,you need to respond, otherwise it
will be considered as ABSENT.
• Students are strongly advised to make notes from RLs & CSs for
examination purpose right from the beginning.
• Compute virtualization
• Software-defined storage
• Virtual networking
• In traditional infrastructures, compute, storage, and networking are delivered as separate, standalone
solutions
• Each layer has its own hardware, management tools, and operational lifecycle
• In hyperconverged infrastructure (HCI), these previously disparate services are consolidated into a
single software-defined platform
• Compute, storage, and networking functions run on the same set of nodes
• Integration occurs at the software layer, not through factory cabling or hardware bundling.
• Results in - single logical system instead of multiple silos, Unified deployment, monitoring, and
upgrades and Simplified operations
• SDS focuses on decoupling storage intelligence from physical hardware and implementing it
in software.
• Because of this philosophical nature, the boundaries of what qualifies as SDS are not always
rigid, and vendor interpretations may vary.
• Policy-Driven Services
Applies protection and data services based on policy
• Programmability
Exposes standard APIs for automation and integration
• Elastic Scalability
Scales capacity and performance as demand grows
Hyperconvergence is one common deployment model, but not a requirement for SDS
2. Kernel-Level Model
Policies govern: Data placement, Protection levels, Performance behavior, Service availability
• Retain snapshots:
• 3 days onsite
• Modify configurations
• Orchestration requires each subsystem to be: Discoverable, Controllable and Automatable via APIs
• Compute
• Storage (SDS)
• Networking
• Management platforms
• Scalability is enabled by hardware abstraction, which hides physical changes from workloads
• SDS allows:
Automation enables:
• Compute virtualization
• Virtual networking
• Hyperconverged Infrastructure (HCI) is an architecture that integrates compute, storage, and networking
into a single software-defined system
• Compute virtualization
• Virtual networking
Overall, HCI transforms multiple infrastructure silos into one integrated, software-defined
platform.
BITS Pilani, Pilani Campus
Hyperconverged Infrastructure
Architectural Components of HCI
Architecture Characteristics:
• Node-based, scale-out design
• No external shared storage arrays
• Tight coupling of compute and storage
BITS Pilani, Pilani Campus
Hyperconverged Infrastructure
• Linear scalability
Commodity hardware
• Enterprise virtualization
Limitations:
The Comparison
• This platform delivers multiple infrastructure services from the same system
• Compute virtualization
• This dual nature often creates confusion for administrators new to HCI
• Common question:
Is HCI a special hardware appliance or a software platform?
• Provides:
• Higher stability
• Better performance
• Software-Defined Storage (SDS) and Hyperconverged Infrastructure (HCI) are closely related but not
interchangeable
• Without SDS, HCI would require external shared storage, defeating its purpose
• HCI uses SDS to integrate compute and storage into a single platform
• Together, they form a stepping stone toward the Software-Defined Data Center (SDDC)
• Flexibility vs performance
Understanding the relationship between SDS and HCI is critical for designing scalable, efficient,
and operable modern data centers.
• Hyperconverged Infrastructure consolidates compute and storage workloads on the same nodes
• Low latency
• High IOPS
• Consistent performance
• Flash improves:
• VM responsiveness
• Storage efficiency
All-Flash HCI
Enables:
• Higher VM density per node
• Faster rebuilds and resynchronization
• Improved resilience during failures
• Higher workload demand ⇒ more disks required, even if capacity wasn’t needed
Challenge in Hyperconvergence
STOP RECORDING
50
Design and Operation of Data Centers
CSIWZG522
BITS Pilani Contributed by
Pilani Campus
Dr. Suhaas K P
IMP Note to Self
IMP Note to Students
• It is important to know that just login to the session does not guarantee the
attendance.
• Once you join the session, continue till the end to consider you as present in the
class.
• IMPORTANTLY, you need to make the class more interactive by responding to
Professors queries in the session.
• Whenever Professor calls your number / name ,you need to respond, otherwise it
will be considered as ABSENT.
• Students are strongly advised to make notes from RLs & CSs for
examination purpose right from the beginning.
• Virtualization
• Benefits of Virtualization
• Types of Virtualization
• VM Hardware Components
• Hyperconvergence
• The Cloud OpenStack
• The Role of Cloud
• Cloud Types
• Cloud Drivers
• Case Study : Cisco Hyperconvergence Strategy
• An SDDC virtualizes and delivers all infrastructure — compute, storage, networking, and
security — as software services.
• Management, provisioning and policy enforcement are performed in software rather than by
manual hardware configuration.
• Goal: deliver IT-as-a-Service (ITaaS) with policy-driven automation and pooled resources
• Growth of dynamic workloads and cloud-style consumption models (need for rapid
provisioning).
• Hardware heterogeneity and hardware-centric silos made manual ops slow and error-prone.
• Need for consistent operations across on-prem, edge and public cloud (hybrid/multi-cloud).
Private & hybrid cloud — consistent environment for workloads across on-prem and cloud.
Multi-tenant service providers / Telco clouds — isolation and rapid tenant provisioning.
Security & policy correctness: Software misconfiguration risks; need for strong testing.
Legacy application constraints: Not all apps are cloud-native or easily virtualizable.
Solution: VMware Cloud Foundation (VCF) deployed on Equinix Metal (bare-metal), forming
a managed SDDC endpoint that can run VMware SDDC Manager and provide local
compute/storage with global interconnection.
Outcomes / Benefits: faster rollout of distributed SDDC endpoints, improved latency for
edge apps, consistent management and security policies across locations, and simpler
hybrid operational model.
BITS Pilani, Pilani Campus
VMware Cloud Foundation on Equinix
Metal
• Traditional data centers relied on tightly coupled hardware and operating systems
Operational Impact:
Design Implication:
Data center layouts now optimize for scalability, density, and uniform hardware profiles rather
than application-specific servers.
Operational Outcome:
Business Value:
Higher infrastructure efficiency with lower operational overhead and improved service agility.
• Traditional data centers relied on dedicated storage arrays and proprietary controllers
• Storage capacity and performance were tightly bound to specific hardware platforms
Operational Impact:
Design Implication:
Storage design shifts from monolithic arrays to distributed, software-managed architectures
Business Value:
Higher storage efficiency with reduced cost and improved agility.
• Traditional data center networks relied on tightly coupled routers and switches
Operational Impact:
Design Implication:
Network design shifts from distributed device intelligence to centralized software control.
Operational Outcome:
Operations shift from manual device management to software-driven network services
Business Value:
Higher network agility with improved reliability and lower operational cost.
Operational Impact:
Design Implication:
Security design shifts from device-based protection to software-enforced, pervasive protection.
STOP RECORDING
48
CSIWZG522 :
Design and Operations of Data Center
CS 04
Click to edit Session title
2
Important Note to Students
⮚ It is important to know that just login to the session does not guarantee the
attendance.
⮚ Once you join the session, continue till the end to consider you as present in the
class.
⮚ Answering to Polls & Quiz questions during / after the class are mandatory
⮚ Students are advised join within 10 minutes once the class starts; Professor
will lock the entry to the meeting. 3
Disclaimer
• The slides presented here are obtained from the authors of the books and
from various other contributors. I hereby acknowledge all the contributors
for their material and inputs.
• I have added and modified a few slides to suit the requirements of the course.
4
Textbook & Reference
T1 Building a Modern Data Center:
Principles and Strategies of Design
by
Scott D. Lowe, David M. Davis, James Green
Reference Book
R1 Data Center for Beginners:
A beginner's guide towards understanding Data
Center
Design
by5
Advisory
• Students are strongly advised to make notes from RLs & CSs for examination
purpose right from the beginning.
• Unavailability of the textbook will not be an excuse to allow print copies of the e-
book during examinations.
6
Recap of Last Class
• Data Centre Evolution
• Data Center’s Computing Components
• Requirements for a modern data center networking platform
• Case Study - Google data centers’ Networking
• Data Center Storage
• Direct Attached Storage (DAS)
• Network Attached Storage (NAS)
• Storage Area Network (SAN)
• Case Study : AWS Data Center for Storage of AWS IoT and S3
• The Rise of the Monolithic Storage Array
• Block vs. File Storage
• The Virtualization of Compute — Software Defined Servers
• Case Study : Microsoft Underwater Data Center 7
Virtualization
Virtualization
Transforming a Classic Data Center (CDC) into a Virtualized Software Defined Data
Center
8
Data Center Evolution & Key Concepts
REAL-LIFE SIMPLE
TOPIC DESCRIPTION
EXAMPLE UNDERSTANDING
Servers: processors,
memory, storage, Intel Xeon + DDR5 RAM + Building blocks of any
Computing Components
network as core NVMe SSDs DC
infrastructure
Hardware Coupling Tightly coupled s/w & h/w OS/app h/w independent Agility
Hardware-assisted
virtualization (AMD-V,
Modern Standards CPU-level support Native efficiency
Intel VT) is industry
Hypervisor Types & Architecture
TYPE/COMPONENT DESCRIPTION CHARACTERISTICS USE CASE
Manages physical
Kernel Component hardware resources and Privileged operations System stability
core OS functions
Stores VM creation
choices: CPU count,
Configuration File .vmx file (VMware) VM blueprint
memory, network
adapters, disk types
Stores VM disk
contents; appears as
Virtual Disk File VMDK/VHDX files Data persistence
physical disk; VM can
have multiple disks
BIOS (virtual state),
Swap (paging when
Support Files Associated files Operational support
running), Log
(troubleshooting)
Enables VM connection to
Multi-homed: internal +
vNIC (Virtual NIC) other physical and virtual 1-4 NICs typical
external
machines on network
DVD/CD-ROM, Floppy
drive, SCSI, USB
Virtual Peripherals As needed CD-ROM for OS install
controllers, Graphic card,
IDE controllers
Serial/Com, Parallel ports
(legacy), Keyboard,
Interface Ports Legacy support Backward compatibility
Mouse interfaces for full
compatibility
Computing Infrastructure - Server Types & Evolution
MODERN
COMPONENT DEFINITION TYPE/CATEGORY
IMPLEMENTATION
Rack-mount (pizza-box),
Blade (compact in 1U/2U rack servers
Server Types Form factors
chassis), Mainframe primary
(multiple processors)
Provide processing,
memory, local storage, Compute nodes in
Server Functions Core functions
and network connectivity clusters
for applications
Integration of switches,
routing, load balancing, Software-defined
DC Networking Network infrastructure
analytics for networking
storage/processing
From physical to
virtualized: VMs,
Network Evolution containers, bare metal Technology shift Microservices architecture
with centralized
management
From hardware-centric
Virtualization Impact silos to integrated, flexible Architectural change Cloud-native design
infrastructure
Evolution: Traditional →
Virtualized → Software-
SDDC Journey Maturity progression Industry direction 2026+
Defined Data Center
(SDDC)
Storage Evolution - Magnetic to Flash Era
ERA/TECHNOLOGY DEFINITION CHARACTERISTICS CURRENT STATUS
Scale-out architectures
Scalability Solutions enabling better utilization Horizontal scaling Hyperscale model
and flexibility
Integrates compute,
storage, networking,
HCI Definition virtualization into single Nutanix platform All-in-one solution
software-defined platform
on x86
HCI eliminates traditional
storage and controller Performance
Bottleneck Elimination Removes SAN complexity
bottlenecks via improvement
distributed SDS
Centralized software-
based control with faster
Simplified Management Single console for all OpEx reduction
deployment and reduced
complexity
Converged: Proprietary,
scale-up; HCI:
Converged vs HCI HCI wins on flexibility Cost efficiency
Commodity, scale-out,
symmetric
Nutanix, SimpliVity,
Atlantis, Pivot3, Maxta;
HCI Vendors Market leaders Enterprise adoption
Use: VDI, virtualization,
private cloud
Cloud Types & Deployment Models
CLOUD TYPE DEFINITION CHARACTERISTICS BEST FOR
Flexibility, scalability,
Value Proposition automation, reduced IT Business outcomes Strategic enabler
ops complexity & capex
Real-World Case Studies & Implementation
SIMPLE
CASE STUDY COMPANY TECHNOLOGY
UNDERSTANDING
Global distributed
Smart routing across
Google Data Centers Google architecture with custom
world-wide servers
networking
Desktop-as-a-service
Enterprise VDI Multiple HCI for virtual desktops
delivery
Complete transformation to
Software-Defined Data Center - Review VMware SDDC
Emergence of SDDC all infrastructure abstracted and architecture diagram + watch
controlled through software VMware SDDC overview video
policies
STOP RECORDING
CSIWZG522 :
Design and Operations of Data Center
CS 03
BITS Pilani
Pilani Campus
⮚ Once you join the session, continue till the end to consider you as present in the class.
⮚ IMPORTANTLY, you need to make the class more interactive by responding to Professors
queries in the session.
⮚ Answering to Polls & Quiz questions during / after the the class are mandatory
⮚ Whenever Professor calls your number / name ,you need to respond, otherwise it will
be considered as ABSENT and you will be removed from the meeting
⮚ Students are advised join within 10 minutes once the class starts, Professor will lock
the entry to the meeting.
Declaimer
• The slides presented here are obtained from the authors of the books and
from various other contributors. I hereby acknowledge all the contributors
for their material and inputs.
• I have added and modified a few slides to suit the requirements of the course.
Text Book & Reference
A book cover of a data center
Reference Book
R1 Data Center for Beginners: A beginner's guide
towards understanding Data Center Design , B.A.
Ayomaya (Author)
Advisory
• Students are strongly advised to make notes from RLs & CSs for
examination purpose right from the beginning.
Industry
Case Study: NVIDIA (AI Infrastructure)
Ref.
Data Centre Evolution
A History of the Modern Data Center (1)
A History of the Modern Data Center (2)
• Tape-based data storage technology began
to be displaced when IBM released the
first disk-based storage unit in 1956
(Figure 2-1).
• It was capable of storing a whopping 3.
• 75 megabytes — paltry by today’s terabyte
standards.
• It weighed over a ton, was moved by
forklift, and was delivered by cargo plane.
• Magnetic, spinning disks continue to
increase in capacity to this day, although
the form factor and rotation speed have
been fairly static in recent years.
Figure 2-1: The spinning disk system
for the IBM RAMAC 305
A History of the Modern Data Center (3)
• The last time a new rotational speed was introduced was in
2000 when Seagate introduced the 15,000 RPM Cheetah
drive.
• CPU clock speed and density has increased many times
over since then.
• These two constantly developing technologies — the
microprocessor/x86 architecture and disk-based storage
medium — form the foundation for the modern data center.
• In the 1990s, the prevailing data center design had each
application running on a server, or a set of servers, with
locally attached storage media.
• As the quantity and criticality of line-of-business
applications supported by the data center grew, this
architecture began to show some dramatic inefficiency
when deployed at scale.
• Plus, the process of addressing that inefficiency has
characterized the modern data center for the past two
decades.
Data Center Components (Computing)
• Computing Components
⮚Servers
⮚Networking
⮚Storage
Computing Components: Servers
Computing Components : Networking
Requirements for a modern data center networking
platform:
1 2 3
Types of Data
Center
Storage
• Just as its name shows, DAS is attached directly to a host server, instead
of connecting through a network, like Ethernet.
Direct Attached Storage(Cont..)
• Advantages of DAS:
• Cost-saving: DAS is much cheaper than other storage technologies, such
as NAS and SAN. And the price per GB for these types of storage devices
is very low, which continues to trend downward.
• Better performance: Compared with other networked storage solutions,
DAS cannot be affected by network bottlenecks, such as network
congestion.
• Disadvantages of DAS:
• Limited scalability
• Not shareable enough: Since data on DAS cannot be connected through
the internet, data sharing can be a big problem. If
Network Attached Storage (NAS)
• Network Attached Storage (NAS) is a file-level
data center storage device that supports multiple
users to retrieve data from centralized disk
capacity over a TCP/IP network.
• It usually has its node on the local area network
(LAN), without the intervention of the application
server, allowing users to access data on the
network.
Storage Area Network (SAN)
• Storage Area Network (SAN) is a dedicated and high-speed network
established for storage that is independent of the TCP/IP network. It
connects servers to their logical disk units (LUNs) and provides block-level
network access to data center storage.
Storage Area Network (SAN)
How to Choose Suitable Data Center Storage?
• Choose a suitable based on the scalability, performance, IT staff, and
usage case.
• Scalability:
• Performance
• IT staff
• Usage case:
Click to edit Session title
CASE STUDY: Data Center Storage – Impact of Cloud & IoT
• File level access means just what it sounds like, “the granularity of access
is a full file.”
Block vs. File Storage (1)
• Data Services
• Most storage platforms come with a variety of different data services that allow the
administrator to manipulate and protect the stored data.
• These are a few of the most common.
• Snapshots
• A storage snapshot is a storage feature that allows an administrator to capture the
state and contents of a volume or object at a certain point in time.
• A snapshot can be used later to revert to the previous state.
• Snapshots are also sometimes copied off site to help with recovery from site-level
disasters.
• Replication
• Replication is a storage feature that allows an administrator to copy a duplicate of a
data set to another system.
• Replication is most commonly a data protection method; copies of data a replicated
off site and available for restore in the event of a disaster.
• Replication can also have other uses, however, like replicating production data to a
testing environment.
Block vs. File Storage (2)
Data Reduction
• Especially in enterprise environments, there is generally a large amount of
duplicate data.
• Virtualization compounds this issue by allowing administrators to very
simply deploy tens to thousands of identical operating systems.
• Many storage platforms are capable of compression and deduplication,
which both involve removing duplicate bits of data.
• The difference between the two is scope.
• Compression happens to a single file or object, whereas deduplication
happens across an entire data set.
• By removing duplicate data, often only a fraction of the initial data must
be stored.
Virtualization of Computers : Software Defined Servers
The Virtualization of Computers: Software Defined Servers
STOP RECORDING
CSIWZG522 :
Design and Operations of Data Center
CS 02
BITS Pilani
Pilani Campus
CS 02
Click to edit Session title
Important Note to Students
⮚It is important to know that just login to the session does not guarantee the
attendance.
⮚Once you join the session, continue till the end to consider you as present in the
class.
⮚Answering to Polls & Quiz questions during / after the class are mandatory
⮚Students are advised join within 10 minutes once the class starts,
Professor will lock the entry to the meeting.
Declaimer
• The slides presented here are obtained from the authors of the books and
from various other contributors. I hereby acknowledge all the contributors
for their material and inputs.
• I have added and modified a few slides to suit the requirements of the course.
Text Book & Reference
• Students are strongly advised to make notes from RLs & CSs for examination
purpose right from the beginning.
• Unavailability of the textbook will not be an excuse to allow print copies of the e-
book or PPTs during examinations.
Lecture Plan
Faster time-to-
Express lane at the
Efficiency market for AI Deploy AI fast
airport.
products.
NVIDIA - Focus on Cost Reduction
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Consolidating
One bus replacing
Hardware many CPUs into Do more, smaller
50 cars.
few GPUs.
Higher Using LED bulbs
Energy performance per instead Save power bills
watt of power. incandescent.
Dense computing
Fitting bunk beds
Space requires less floor Save floor space
in room.
space.
Automatic lawn
Lowering "Staff
mower (System
Management Overhead" via Lower human cost
Management)
automation.
saves time.
Handles 1M new
Scalability Capacity Size Big growth ready
users.
Saves money at
Elasticity Cost & Flex Flex with traffic
night.
Fast training of
Case Study NVIDIA AI AI power leader
GPT.
Elegant design
Minimalist home
Simplicity reduces human Simple is better
with no clutter.
error.
Future Trends in Design
SIMPLE
REAL-LIFE
TREND DESCRIPTION UNDERSTANDIN
EXAMPLE
G
Next-gen
Teleporting data
Quantum computing for Future fast math
across world.
complex math.
Adapt or Die & IT Flattening
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Organizations
Blockbuster failing
must evolve or
Adaptation to adopt Evolve or fail
face market
streaming.
extinction.
Removing silos
A small startup
between server,
Flattening where everyone Merge tech teams
storage, and
helps.
network teams.
Speed of delivery
becomes the Fast-food prep
Agility Speed is king
primary versus fine dining.
competitive edge.
Generalists (Full-
Handyman fixing
stack) replacing Broaden your
Skill Shift plumbing and
narrow niche skills
electric.
specialists.
Reducing
management
Direct chat with
IT as an Operational Expense (OpEx)
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Moving from
buying hardware Leasing a car
CapEx Shift Rent don't buy
to renting instead buying.
services.
Paying only for the
Paying the
Consumption resources actually Pay per use
monthly water bill.
used monthly.
Keeps capital free Using cash for
Cash Flow for core business marketing, not Keep cash free
investments. desks.
Expanding in
small, predictable Buying food for
Scalability Grow in bits
chunks based on one week.
need.
Ability to cancel or Ending a Netflix
Flexibility reduce services subscription Cancel any time
instantly. anytime.
Next Session:
Data Center Design- Components
STOP RECORDING
CSIWZG522 :
Design and Operations of Data Center
CS 01
BITS Pilani
Pilani Campus
CS 01
Click to edit Session title
Important Note to Students
It is important to know that just login to the session does not guarantee the
attendance.
Once you join the session, continue till the end to consider you as present in the
class.
Answering to Polls & Quiz questions during / after the class are mandatory
Students are advised join within 10 minutes once the class starts,
Professor will lock the entry to the meeting.
Declaimer
• The slides presented here are obtained from the authors of the books and
from various other contributors. I hereby acknowledge all the contributors
for their material and inputs.
• I have added and modified a few slides to suit the requirements of the course.
Text Book & Reference
• Students are strongly advised to make notes from RLs & CSs for examination
purpose right from the beginning.
• Unavailability of the textbook will not be an excuse to allow print copies of the e-
book or PPTs during examinations.
Lecture Plan
Data Centers are critical assets for everyday operations (Mission Infra).
01 02 03 04
• Security & Access (BlueField DPU (NVIDIA’s Brand Name of Data Processing Unit
called BlueField) security - Isolated zero‑trust infrastructure offload)
- Protecting data and systems with physical security, firewalls, and access controls, while
allowing authorized users access.
• SDS [Software Defined Storage] offers IT organizations some of the greatest storage
flexibility and ease of use (Virtual Storage).
IT Is Changing . . . and it Must
• With exponential growth, there are more opportunities for entrepreneurs, for organizations
to multiply their revenues, for top-tier graduate students to create projects that once
seemed like pure science fiction (Digital Boom).
• While bleeding edge technology trickles down into the mainstream IT organization, the
changing nature of the IT business itself rivals the pace of technological change
(Innovation Diffusion).
• Over the next couple of years, IT organizations and business leaders will have the
opportunity to focus on dramatic transformations in the way they conduct themselves.
Technological change will create completely new ways of doing business (Operating
Overhaul).
• The exponential growth in technology in the next decade will generate the rise of entirely
new industries and cause everything in our day-to-day experience to be different, from
the way we shop to the way we drive our cars (Emerging Sectors).
Simplification Is the New Black
• Thanks to the ever-growing number of devices interconnected on private and public networks, and
thanks to the Internet, the scope of IT’s responsibility continues to expand (Device Explosion).
• The computer is responsible for operating the press at maximum efficiency, and sometimes it does
so well that the print shop doubles its profits, and the computer becomes critical to the business
(Automation Gains).
• The boom was dubbed the “Internet of Things” (IoT) way back in 1999! (IOT Boom)
• Today, we’re still right at the beginning of this paradigm shift (Early Stage).
• The estimations of IoT proliferation in the next few years are staggering (Massive Scale).
• A 2014 Gartner report shows an estimated 25 trillion connected devices by 2020 (Device
Tsunami).
• Because this trend means that IT departments become responsible for more systems and
endpoints, one of the primary goals for these departments in the near future is to simplify
administration of their systems (Management Simplification).
Simplification Is the New Black
• Despite the fact that the number of
integrations and the amount of data IT has
to manage because of these devices is
increasing, in many industries budgets are
not (Budget Squeeze).
• This means that, in order to succeed, CIOs
will need to find ways to boost efficiency in
major ways (Efficiency Drive).
• Five years from now, the same team of administrators may be managing 10 times the
number of resources that they’re managing today, and expectations for stability,
performance and availability will continue to increase (Uptime Demands).
• There are two main venues where the push for simplicity must come from: the
manufacturers and the IT executives (Vendor Simplification).
Simplification Is the New Black
• First of all, manufacturers must begin to design and redesign their products with
administrative simplicity in mind (Simpler Consoles).
• To be clear, this doesn’t mean the product must be simple. In fact, the product will almost
assuredly be even more complex. However, the products must have the intelligence to
self-configure, solve problems, and make decisions so that the administrator doesn’t have
to (Autonomous Systems).
• Secondly, IT Directors and CIOs must survey their entire organization and ruthlessly
eliminate complexity (Complex Cleanup).
• In order to scale to meet the needs of the future, all productivity-sapping complexity must
be replaced with elegant simplicity such that the IT staff can spend time on valuable
work, rather than on putting out fires and troubleshooting mysterious failures that take
hours to resolve (Friction Reduction).
• One of the primary ways this might be done is to eliminate the organizational siloes that
prevent communication and collaboration (Silo busting).
Focus on Cost Reduction
• As shown in Figure 1-2, a recent
Gartner report shows that IT
budgets are predicted to remain
nearly flat as far out as 2020
(Budget Plateau).
• This is despite 56% of respondents
reporting that overall revenues are
expected to increase (Revenue
Growth).
• The reality is that many of those IT
organizations will be required to do
more with less, or at the very least,
do more without any increase in
budget (Resource Squeeze; Cost
Focus on Cost Reduction
• As this trend isn’t likely to change, the IT department of the future will be focused on
reducing expenses where possible and increasing control over the absolutely necessary
expenses (Cost Discipline).
• Thanks again to Moore’s Law, it will be possible to complete projects in 2016 for a fraction
of the cost that the same project would have cost in 2013 (Hardware Deflation).
• One of the most common examples of this is the cost of storage for a server virtualization
project (Virtual Storage).
• Due to the challenges of performance and workload characteristics, a properly built server
virtualization storage platform has historically been expensive and quite complex
(Enterprise Arrays).
• Thanks to denser processors though, cheaper and larger RAM configurations, the falling
cost of flash storage (solid state drives [SSDs]), and innovation toward solving the storage
problem, a server virtualization project can be completed successfully today for a much
lower cost, relatively speaking, and with much more simplicity than ever before
(Commodity Hardware; Simpler Stacks).
Focus on Cost Reduction
• As mentioned in the previous section, the IT department of
the future will be looking to manufacturers and consultants
to help them build systems that are so simple to manage
that less IT staff is required. This is because hiring
additional IT staff to manage complex solutions is costly
(Simpler Platforms; Staff Overhead).
• Of course, this doesn’t necessarily mean that IT jobs are at
risk; it means that since there’s no budget to grow the IT
staff, the current workforce must accomplish more duties
as the responsibility of IT grows. This additional
responsibility means that IT jobs are only safe assuming
that each IT practitioner is growing and adapting with the
technology (Role Expansion; Continuous Upskilling).
• Those who do not learn to handle the additional
responsibility will be of little use to the organization moving
forward (Skill Stagnation).
Focus on Customer Service and the Business
• IT departments face pressure to do more with less and seek top-
notch vendor support (Vendor Reliance).
• IT professionals lack time to troubleshoot solutions extensively
(Time Crunch).
• White-glove service from technical resources is becoming the
norm (Premium Support).
• Leading manufacturers use connectivity and big data to proactively
address customer issues (Proactive Analytics).
• The ideal support call is one that never needs to happen (Zero
Tickets).
• CIOs and IT Directors plan to shift infrastructure maintenance to
partners and vendors (Managed Services).
• IT staff will focus more on innovation and adding business value
(Value Creation). (Value = Utility + Warranty)
Focus on Customer Service and the Business
• IT departments must focus on providing value to their internal customers as well
(Employee enablement).
• As new technology enables business agility on the manufacturer side, IT will have to
continue providing the services users need and want, or users will assuredly find them
elsewhere (Service Competitiveness)
• This means that shadow IT poses significant security, compliance, and control risk to the
entire business, and the only way to really stop it is to serve the internal customers so
well that they don’t need to look elsewhere for their technical needs (Better
Alternatives).
• Shadow IT is a term used to describe business units (or individuals) other than IT who
deploy technical solutions — not sanctioned or controlled by IT — to solve their business
problems (Unsanctioned Tools).
• A simple example of this phenomenon is a few individuals in the finance department
using personal Dropbox folders to share files while the business’ chosen direction is to
share company files in SharePoint (Rogue Deployments; Unauthorized Sharing).
Adapt or Die
• The world of IT operates in the same way as the rest of the world: things change over time, and
those companies, technologies, or individuals who choose not to keep up with the current state of
the industry get left behind. It’s unfortunate, but it is reality (Constant Change).
• The movie rental giant Blockbuster was dominating the market until Netflix and Redbox innovated
and Blockbuster failed to adapt. Blockbuster eventually went bankrupt and is now an afterthought
in the movie consumption industry while Netflix’s fortunes are at an all-time high (Disrupted
Incumbent; Streaming Success).
• This lifecycle is exactly the same in IT; there are household names in the industry that are more or
less irrelevant (or quickly becoming so) at this point because of their failure to adapt to the
changing market (market obsolescence).
• Unfortunate as it may be, this is also happening at the individual level (Career Risk).
• As IT administrators and architects adapt or do not adapt over the course of time, they either
advance in their organization and career, or they become irrelevant (Skill Evolution).
IT as an Operational Expense
• Especially in the enterprise environment, getting budgetary approval for operational
expenses can prove to be easier than getting approval for large capital expenditures
(Operational Expenses Preference).
• As such, the operating model of many IT departments is shifting away from capital
expenses (Capex) when possible and toward a primarily operational expense-funded
(Opex-funded) model for completing projects (Budget Shift).
• A large component of recent success in these areas is due to a shift of on-premises,
corporately managed resources to public cloud infrastructure and Software-as-a-Service
(SaaS) platforms (Cloud Migration).
• Since cloud resources can be billed just like a monthly phone bill, shifting IT resources to
the cloud also shifts the way the budget for those resources is allocated (Subscription
Billing).
• While buying a pile of servers and network equipment to complete projects over the next
year is budgeted as a capital expenditure, the organization’s “cloud bill” will be slotted as
an operational expenditure (Hardware Capex; Service Opex).
IT as an Operational Expense
• This is because, with a capital purchase, once the equipment is purchased, it is owned
and depreciating in value from the moment it hits the loading dock. This removes an
element of control from IT executives as compared to an operational expense (Asset
lock-in).
• If a SaaS application is billed per user on a monthly basis, there’s no need to pay for
licenses now to accommodate growth in headcount six months down the road (License
Elasticity).
• It can also be in use this month and cancelled next month (Easy Off-boarding).
• This is in contrast to the IT director who can’t just “cancel” the stack of servers purchased
six months ago because the project got cancelled (Stranded Hardware) (Unusable
sunk Infrastructure)
• Due to these advantages from a budgeting and control standpoint, products and services
offering a model that will require little or no capital expense and allow budgeting as an
operational expense will be preferred (Opex Advantage).
IT as an Operational Expense
• What this means for manufacturers is that transparency, granular control, and offering a
OpEx-based model like renting or leasing, billing based on monthly usage, and
expanding in small, predictable chunks based on need, will position them for adoption by
the IT department of the future (Consumption Pricing).
• Also, offering insight and helping the customer to increase efficiency and accuracy over
the course of the billing cycle will create lasting loyalty (Value Analytics).
• This shifting focus from CapEx to OpEx is also giving rise to new kinds of Data Center
architectures that allow organizations to keep Data Centers on premises and private, but
that enable some economic aspects similar to cloud (Hybrid Models).
• Rather than having to overbuy storage, for example, companies can begin to adopt
software defined storage (SDS) or hyperconverged infrastructure (HCI) solutions that
enable pay-as-you grow adoption methodologies (Scalable Storage).
Class Summary
• What is Data Center?
• Core components of Data Centers
• Types of Data Centers
• Modern Data Center
• IT Challenges
• Simplification Is the New Black
• Focus on Cost Reduction
• IT as an Operational Expense
Next Session:
Distributed programming