0% found this document useful (0 votes)
22 views270 pages

Design Notes

The document outlines the design and operation of data centers, focusing on various infrastructure types including traditional, converged, and hyperconverged systems. It discusses business requirements, risk factors, regulatory compliance, and the evolving landscape of data center management and operations. Key topics include automation, disaster recovery, and the changing perceptions of IT costs and budgets in the context of modern data center challenges.

Uploaded by

rthe3010
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
22 views270 pages

Design Notes

The document outlines the design and operation of data centers, focusing on various infrastructure types including traditional, converged, and hyperconverged systems. It discusses business requirements, risk factors, regulatory compliance, and the evolving landscape of data center management and operations. Key topics include automation, disaster recovery, and the changing perceptions of IT costs and budgets in the context of modern data center challenges.

Uploaded by

rthe3010
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Design and Operation of Data

Centers
CSIWZG522
BITS Pilani Prepared by
[Link]
BITS Pilani

CSIWZG522: Design and Operation of Data Centers


CS 7 : Data Center Design- Infrastructure types

BITS Pilani, Deemed to be University under Section 3 of UGC Act, 1956


IMP Note to Self

BITS Pilani, Deemed to be University under Section 3 of UGC Act, 1956


TEXTBOOK AND REFERENCES

T1 Building a Modern Data Center: Principles


and Strategies of Design
by
Scott D. Lowe, David M. Davis, James Green

R1 Data Center for Beginners: A beginner's guide


towards understanding Data Center Design
by
B.A. Ayomaya

BITS Pilani, Deemed to be University under Section 3 of UGC Act, 1956


Advisory
BITS Pilani

• The printed book is available in the market.

• The e-book is available on the Internet.

• Students are strongly advised to make notes from RLs & CSs for examination
purpose right from the beginning.

• Unavailability of the textbook will not be an excuse to allow print copies of the e-
book during examinations.
RECAP OF PREVIOUS SESSION

•Deployment and consolidation of servers


•Performance of server.
• The Parallel Paths of SDS and Hyperconvergence
• The Details of SDS
• What Is Hyperconverged Infrastructure?
• The Relationship Between SDS and HCI
• The Role of Flash in Hyperconvergence
• Where Are We Now?

BITS Pilani, Deemed to be University under Section 3 of UGC Act, 1956


SESSION 7 Topics
Time Type Description Content
Reference
Pre CH RL4.1 Infrastructure Types- Traditional
4.2 Converged Infrastructure
4.3 Hyper Converged Infrastructure
During CH CH The Business Requirement Challenge T1: ch-4
Risk
Assurance, business continuity, and disaster recovery
The regulatory landscape
Avoiding lock-in
Changing the perception of IT Cost
Changing budgetary landscape
The changing career landscape
Complexity
Agility and performance
Automation and orchestration
Self service
The data growth challenge
Resiliency and availability

BITS Pilani, Deemed to be University under Section 3 of UGC Act, 1956


Infrastructure Types- Traditional Vs
Converged Infrastructure
Traditional data center:
infrastructure (legacy) uses separate,
individually managed silos of compute,
storage, and networking hardware,
offering high customization but
increased complexity and management
overhead.

Converged infrastructure (CI) packages


these components into a single pre-
configured, validated unit to simplify
deployment, improve performance, and
lower operational costs.

BITS Pilani, Deemed to be University under Section 3 of UGC Act, 1956


Infrastructure Types-
Hyperconverged Infrastructure (HCI)
Hyperconverged Infrastructure (HCI) is a software-defined, unified data center
architecture that integrates compute, storage, networking, and virtualization into a
single, x86-based appliance
Advantages for Data Centers:
• Reduced Costs: Lower capital expenditure
(CapEx) by avoiding expensive specialized storage
hardware and lower operational expenditure
(OpEx).
• Simplified Operations: Streamlines management
and accelerates deployment times.
• Agility & Scalability: Enables quick,,, easy,
scalability to meet changing business needs with
minimal disruption.
• Modernization: Ideal for building software-defined
data centers (SDDC) and supporting hybrid cloud
strategies.
The Business Requirement Challenge

Availability & Reliability (Uptime): Defined by tier levels (Tier 1-4), requiring
redundant UPS, generators, and cooling to prevent downtime.

Scalability & Flexibility: Modular, flexible design allows for expanding IT


hardware, power, and cooling without major structural disruptions.

Security & Compliance: Physical (biometrics, security guards) and logical


security, combined with compliance certifications (e.g., ISO, HIPAA, PCI-DSS).
The Business Requirement Challenge…

Power & Cooling Efficiency: Focus on reducing Power Usage Effectiveness


(PUE) through advanced cooling, hot/cold aisle containment, and sustainable
energy.

Cost Management (TCO): Balancing capital expenditure (CAPEX) for


construction with operational expenditure (OPEX) for energy and maintenance.

Site Selection & Connectivity: Proximity to fiber providers, power grid reliability,
and low latency for users.

Disaster Recovery (DR): Geographic redundancy and contingency planning for


environmental or technical disasters.
Risk Factors in Data Center
Infrastructure
Power & Energy: Insufficient power, lack of redundancy (generators/UPS), and
grid instability can lead to catastrophic outages.
Cooling & Thermal Management: Inadequate cooling leads to high-density heat,
requiring advanced liquid cooling to prevent equipment failure.
Environmental & Location Risks: Natural disasters (floods, earthquakes), rising
temperatures, and water scarcity for cooling systems.
Supply Chain & Technical Obsolescence: Shortages of components and the
rapid shift in AI hardware can make new designs outdated quickly.
Cybersecurity & Physical Security: Cyberattacks during construction and
physical security lapses.

BITS Pilani, Deemed to be University under Section 3 of UGC Act, 1956


Data Center Risk Assessment Table

Source: H5Datacenter
Data Center Risk Assessment
metrics…
Data Center Risk Assessment
metrics…
Assurance Strategies
Security and Reliability Benchmarks: Adherence to ISO/IEC 27001 for information security
and SOC 2 Type II for operational integrity is the "gold standard" for building client trust and
ensuring reliable service delivery.

Infrastructure Design Standards: Compliance with the TIA-942 or Uptime Institute


Tier classifications is crucial for verifying physical infrastructure resilience, including power
redundancy and environmental controls.

Mandatory Data Privacy: Facilities must strictly follow regional laws


like GDPR (EU), CCPA (USA), or the DPDP Act (India), which mandate specific protocols for
data sovereignty and immediate breach notification.

Evolving Sustainability Reporting: New 2026 mandates, such as the EU Energy Efficiency
Directive (EED), now require data centers to report precise performance metrics
like PUE (Power Usage Effectiveness) and WUE (Water Usage Effectiveness).

Industry-Specific Mandates: Providers serving specialized sectors must maintain niche


certifications, such as PCI DSS for financial transactions or HIPAA for healthcare data, to
Business Continuity Planning

Business Impact Analysis


(BIA): Categorise
applications into "Continuity
Tiers" (Mission-Critical vs.
Deferrable) to allocate
resources effectively based
on specific Recovery Time
Objectives
(RTO) and Recovery Point
Objectives (RPO).
Business Continuity Planning …

Concurrent Maintainability: Aim for Tier III or Tier IV standards, ensuring


that any component (power, cooling, network) can be removed for
maintenance without impacting the load.
Self-Healing Infrastructure: Implement AI that can autonomously detect
hardware anomalies and reroute traffic to healthy nodes before a failure
occurs.
Predictive Maintenance: Use DCIM (Data Center Infrastructure
Management) platforms with machine learning to forecast equipment failures
and optimize thermal efficiency.
N+1 or 2N Redundancy: Deploy extra UPS systems, standby generators, and
dual-utility feeds to ensure power stability during grid failures.
BITS Pilani, Deemed to be University under Section 3 of UGC Act, 1956
Disaster Recovery –Use case
Business Continuity and Disaster
Recovery Plan
Regulatory Compliances

ISO/IEC 27001 (Security Management): The global benchmark for establishing


an Information Security Management System, focusing on rigorous risk management
and data confidentiality.
SOC 1 & SOC 2 (Operational Audits): Developed by the AICPA, SOC 1 audits financial
reporting controls, while SOC 2 evaluates security, availability, and privacy "Trust
Services Criteria."
PCI DSS (Payment Security): A mandatory requirement for any facility that stores,
processes, or transmits credit card information to prevent fraud and data theft.
GDPR & Regional Privacy Laws: Strict mandates like the EU’s GDPR or India’s DPDP
Act govern how personal data is handled and require immediate reporting of data
breaches.
HIPAA (Healthcare): U.S. federal law requiring physical and technical safeguards for
data centers hosting electronic protected health information (ePHI).
Regulatory Compliances ….

Uptime Institute Tiering: While a certification rather than a law, this defines the global
standard for infrastructure resilience and redundancy from Tier I to Tier IV.
ISO/IEC 22237 (Facility Design): A specific holistic standard for the physical
infrastructure of data centers, covering power, cooling, and cabling systems.
Environmental Mandates (EED & LEED): Regulations like the EU’s Energy Efficiency
Directive force facilities to report energy use, while LEED certifications track
sustainability.
FedRAMP (Government Services): Essential for data centers in the U.S.
providing cloud services to federal agencies, requiring high-level security authorization.
ISO 50001 (Energy Management): An international framework that helps data centers
improve energy performance and reduce carbon footprints through structured
management.
Avoiding Lock In
What is Lock-In? Risks of Lock-In
• Dependence on a single vendor or technology, Financial Risks
making it difficult to switch providers or systems. Increased costs over time, lack of competitive pricing.
Operational Risks
Importance of Avoiding Lock-In Difficulty adapting to new technologies, inability to
scale.
• Flexibility, cost control, and future-proofing.
Strategic Risks
Reduced innovation, limited access to new features.
Types of Lock-In
Vendor Lock-In Strategies to Avoid Lock-In
– Dependency on specific hardware, software, or Open Standards and Protocols
services. Use open-source technologies and standards to ensure
Technology Lock-In compatibility.
Modular Architecture
– Relying on proprietary technologies that limit
Design systems that can be easily modified or replaced.
integration.
Multi-Cloud and Hybrid Solutions
Contractual Lock-In Leverage multiple cloud providers to prevent
– Long-term contracts that restrict mobility and dependency on one.
negotiation.
Changing the perception of IT Cost

Changing the perspective of IT cost in data center design means shifting from CapEx-
heavy infrastructure (building for peak capacity) to OpEx-focused
efficiency (building for actual utilization).
• Total Cost of Ownership (TCO) vs. Initial Capital: Modern design prioritizes TCO over 10–15 years rather
than the initial build cost.

• Power Usage Effectiveness (PUE) as a Cost Lever: Efficiency is no longer just "green"—it's a direct cost
saving. Reducing PUE from 2.0 to 1.2 can slash millions in annual utility bills for large-scale facilities.

• Modular & Scalable "Pay-as-you-grow" Design: Instead of over-provisioning (which wastes capital), modular
data center architecture allows IT to add power and cooling units only when demand rises, deferring capital
expenditure.

• Software-Defined Infrastructure (SDI): Shifting complexity from hardware to software allows for the use of
"commodity" hardware, reducing vendor lock-in costs and making hardware refresh cycles less expensive.

• Data Gravity & Egress Optimization: Design around "data gravity" by placing infrastructure near cloud on-
ramps. This slashes the massive monthly costs of moving data between private servers and public clouds.
The Changing Budgetary Landscape
Budgets are moving from monolithic capital investments toward flexible, high-density, and speed-oriented
spending models.

• Infrastructure Investment Supercycle: Global data center expenditures are projected to reach $3
trillion by 2030. Construction costs for shell and core alone are rising at a 7% CAGR, averaging
over $11 million per MW in 2026.

• Massive AI "Tech Fit-Out" Costs: While shell construction is expensive, the IT equipment for AI
(GPUs, high-speed networking) can cost an additional $25 million per MW.

• CapEx to OpEx Transition: To maintain liquidity, enterprises are shifting away from owning physical
hardware. Instead, they are adopting "as-a-Service" (DaaS, IaaS) models, converting large upfront
costs into predictable monthly operational expenses.

• Premium for Liquid Cooling: AI workloads require high-density racks (up to 100kW), making liquid
cooling a necessity. These facilities carry a 10% cost premium over traditional air-cooled sites.

• Modular & Standardized Spending: To combat long construction timelines (often 4+ years for grid
connections), budgets are prioritizing modular construction and on-site power generation (gas
The Changing Career Landscape
The "traditional" data center technician is being replaced by hybrid roles that blend physical engineering with
software automation.
The Rise of the "Super-Specialist": High demand exists for Electrical and Power Infrastructure Engineers
capable of designing microgrids and high-voltage systems to handle the doubling of global data center
power use by 2030.
Sustainability & ESG Analysts: Specialists who can optimize Power Usage Effectiveness (PUE) and
integrate renewable energy are now "must-have" roles due to rising regulatory and carbon-tracking
pressures.
Shift from "Break/Fix" to "Supervise/Optimize": AI tools now handle routine thermal audits and hardware
troubleshooting. Human roles are evolving into AI Infrastructure Operations Engineers who validate AI
decisions and manage digital twins.
Infrastructure as Code (IaC) Mastery: Knowledge of software-defined everything (SDx) is no longer just for
DevOps. Physical infrastructure managers must now be comfortable with automation and scripting to
manage zero-touch provisioning.
Commissioning & Project Engineers: With 10 gigawatts of new capacity breaking ground annually, there is
a fierce talent war for engineers with hands-on build experience who can coordinate complex, multi-site
construction.
Complexity, Agility and Performance
AI-Driven Density (Performance): The shift from 10kW to 100kW per rack is the new performance baseline. To
achieve this, designs are moving from air cooling to Direct-to-Chip Liquid Cooling, which is essential for the
thermal demands of next-gen GPUs.
Modular "Block" Architecture (Agility): Forget 24-month builds. Using pre-fabricated modular units allows
operators to deploy capacity in "just-in-time" increments, slashing time-to-market and keeping capital fluid.
Digital Twin Management (Complexity): To manage the sheer density of components, designers now use Digital
Twins. These virtual replicas simulate power and thermal changes before physical deployment, preventing
"stranded capacity" and catastrophic failures.
Software-Defined Everything (Agility/Complexity): By abstracting hardware into Software-Defined Data Centers
(SDDC), networking and storage become programmable. This reduces manual configuration errors (Complexity)
and enables instant resource reallocation (Agility).
Microgrid Integration (Performance/Agility): With utility grids facing 5–10 year wait times, modern designs
integrate on-site microgrids (battery storage, hydrogen fuel cells). This provides the power performance needed
for AI without being held hostage by city infrastructure.
Edge-to-Core Continuity (Performance): Performance is now measured by latency. Placing high-
performance Edge Data Centers closer to users reduces data "travel time," while high-speed fiber backhauls link
them to the massive "training" core.
Data Center Automation

Data center automation is the process by which routine workflows and processes of a data
center—scheduling, monitoring, maintenance, application delivery, and so on—are
managed and executed without human administration.

Data center automation increases agility and operational efficiency.

The massive growth in data and the speed at which businesses operate today mean that
manual monitoring, troubleshooting, and remediation is too slow to be effective and can
put businesses at risk.

Automation can make day-two operations almost autonomous. Ideally, the data center
provider would have API access to the infrastructure, enabling it to inter-operate with
public clouds so that customers could migrate data or workloads from cloud to cloud.

Data center automation is predominantly delivered through software solutions that grant
centralized access to all or most data center resources.
What can you automate in your data
center?
Key Element Description
Automate host setups (both physical and virtual) to minimize manual errors and
Host automation
enhance the speed of deployments.
Automate patching at scale, lifecycle management, auditing, monitoring, security
Day-2 Automation
management, and cost management.
Storage Automate data assignment to different types of storage media for increased
automation performance, reduced costs, and efficient data management.
Network Automate network functions, including configurations and management to reduce
automation manual errors and increase the deployment velocity of network services.
Security Automate repetitive and low-level tasks to enhance the identification and
automation remediation of security and compliance risks.
Environment Utilize software-defined Infrastructure to manage and control data center
control automation utilization and energy consumption of all IT-related equipment.
Infrastructure and
Main mechanism to achieve data center automation.
policy as code
Orchestration for self-service data
center
This enables automated, on-demand provisioning of IT infrastructure via
portals, eliminating manual, error-prone configuration tasks.
By using Infrastructure as Code (IaC) and Multi-Domain Service
Orchestration (MDSO), organizations can rapidly deploy virtual/physical
resources, ensuring consistency and compliance while enabling developers
to request, manage, and deploy services independently.
– Infrastructure as Code (IaC): Use code to define, manage, and provision infrastructure, which ensures
consistency and allows for version control.

– Multi-Domain Orchestration (MDSO): Manages end-to-end services across hybrid, multi-vendor


environments by interfacing with SDN controllers, NFV orchestrators, and data center management tools.
AI & Non AI
workloads
for Data
Center
Importance of resilience

For modern data centers, downtime is not an option, particularly not unplanned
downtime. It is therefore vital that all data centers achieve the highest
possible standards of data center resilience.

In the context of data centers, the term “resilience” refers to the ability of an
infrastructure to withstand and recover from disruptions. The more resilience
a data center has, the more likely it is to be able to maintain uninterrupted
services at all times.
In other words, higher resilience generally translates into lower downtime.
Strategies for achieving high availability

Here are five key strategies for achieving and maintaining high availability in data
centers.
Redundancy and replication:
• By duplicating critical components and services, organizations create backups
that can seamlessly take over in case of failures. This applies not only to hardware
but also to data.
• Replicating data across different servers or even geographical locations ensures
that if one location experiences an issue, another can seamlessly continue
operations.
Load balancing:
• Load balancing is a key practice to prevent any single server from being
overwhelmed by incoming requests. By distributing traffic across multiple servers, the
load on each server is optimized, and the risk of overload-induced failures is
significantly reduced.
Strategies for achieving high
availability….
Failover mechanisms: Failover mechanisms ensure a swift transition from a failing
component to a backup. They often take the form of an active-passive setup where a
standby system takes over when the primary system experiences issues. Automated
failover systems are designed to detect problems and initiate the switch to maintain
continuous service availability.
Scalability and elasticity: Systems must be able to adapt to changing workloads by
automatically allocating or deallocating resources. This ensures continuous performance
even during sudden traffic spikes or increased demand. It therefore enhances overall
system robustness and minimizes the risk of downtime due to resource constraints.
Data center resilience best practices: These include comprehensive disaster recovery
planning, designing for loose coupling among components, and prioritizing security
measures. Resilient architectures consider the entire system’s lifecycle, including startup
dependencies and graceful degradation under stress.
Case studies
Here are three real-world case studies of data centers that have been designed for high resilience.
Google: Google’s data center resilience is based on redundancy and automatic failover mechanisms.
By replicating data across different locations, Google minimizes the risk of a single point of failure.
Load balancing ensures even distribution of traffic among multiple servers, preventing overload on
any single server. Automatic failover systems swiftly transition from primary to standby systems,
ensuring uninterrupted service.
Amazon Web Services (AWS): AWS prioritizes resilience through diverse availability zones, each
with its own infrastructure and facilities. This geographical separation minimizes the impact of
localized failures. Load balancing distributes traffic across multiple servers, maintaining optimal
performance and mitigating the risk of overload. AWS’s auto-scaling feature dynamically adjusts
resources based on demand, ensuring scalability and consistent availability.
Microsoft Azure: Azure focuses on resilience through distributed architecture, redundancy, and
failover strategies. Availability zones enhance fault tolerance, and Azure’s scalability ensures
adaptability to varying workloads. Data replication across regions boosts Azure’s disaster recovery
capabilities.
Design and Operation of Data
Centers
CSIWZG522
BITS Pilani
Pilani Campus
IMP Note to Self
IMP Note to Students

• It is important to know that just login to the session does not guarantee the
attendance.
• Once you join the session, continue till the end to consider you as present in the
class.
• IMPORTANTLY, you need to make the class more interactive by responding to
Professors queries in the session.
• Whenever Professor calls your number / name ,you need to respond, otherwise it
will be considered as ABSENT.

BITS Pilani, Pilani Campus


BITS Pilani
Pilani Campus

CSIWZG522: Design and Operation of Data Centers


CS 6 : Data Center Design-Understanding Servers
TEXTBOOK AND REFERENCES A book cover of a data center

T1 Building a Modern Data Center: Principles and


Strategies of Design
by
Scott D. Lowe, David M. Davis, James Green
Click to edit Session title

R1 Data Center for Beginners: A beginner's guide towards


understanding Data Center Design
by
B.A. Ayomaya
BITS Pilani, Pilani Campus
Advisory

• The printed book is available in the market.

• The e-book is available on the Internet.

• Students are strongly advised to make notes from RLs & CSs for
examination purpose right from the beginning.

• Unavailability of the textbook will not be an excuse to allow print copies


of the e-book during examinations.

BITS Pilani, Pilani Campus


RECAP OF PREVIOUS SESSION

• The Emergence of SDDC


• Commoditization of Hardware
• Shift to Software Defined Compute
• Shift to Software Defined Storage
• Shift to Software Defined Networking
• Shift to Software Defined Security

BITS Pilani, Pilani Campus


CONTACT SESSION 06

Data Center Design-Understanding Servers – Continued..

BITS Pilani, Pilani Campus


The Parallel Paths of SDS and
Hyperconvergence
Software-Defined Storage and Hardware Commoditization

• Software-Defined Storage (SDS) is a primary enabler of hardware commoditization in the


data center

• Separates storage intelligence from proprietary storage hardware

• Enables the use of low-cost, industry-standard x86 servers

• Reduces dependence on vendor-specific storage appliances

• Shifts value from hardware to software

BITS Pilani, Pilani Campus


The Parallel Paths of SDS and
Hyperconvergence
Enterprise Storage Services on Commodity Hardware
– SDS delivers advanced storage functionality in software:
• Replication and fault tolerance
• Snapshots and cloning
• Thin provisioning
• Erasure coding and data distribution
– Provides high availability and data resilience
– Matches or exceeds traditional enterprise storage capabilities

BITS Pilani, Pilani Campus


The Parallel Paths of SDS and
Hyperconvergence
From Converged Infrastructure to Hyperconvergence

• Hyperconvergence is the evolution of converged infrastructure

• Integration shifts from hardware to software-defined integration

• Storage is no longer external; it runs inside the compute nodes

• Tight coupling of:

• Compute virtualization

• Software-defined storage

• Virtual networking

BITS Pilani, Pilani Campus


The Parallel Paths of SDS and
Hyperconvergence
Hyperconvergence: From Disparate Solutions to a Single System

• In traditional infrastructures, compute, storage, and networking are delivered as separate, standalone
solutions

• Each layer has its own hardware, management tools, and operational lifecycle

• In hyperconverged infrastructure (HCI), these previously disparate services are consolidated into a
single software-defined platform

• Compute, storage, and networking functions run on the same set of nodes

• Integration occurs at the software layer, not through factory cabling or hardware bundling.

• Results in - single logical system instead of multiple silos, Unified deployment, monitoring, and
upgrades and Simplified operations

BITS Pilani, Pilani Campus


Traditional Defined Storage

BITS Pilani, Pilani Campus


Software Defined Storage

BITS Pilani, Pilani Campus


The Comparison of Traditional Storage
and SDS

BITS Pilani, Pilani Campus


The Details of SDS

• Software-Defined Storage (SDS) is not a single product or technology, but a design


philosophy and architectural approach to storage.

• Similar to Cloud computing and DevOps, software-defined represents a way of designing


and operating systems, rather than a specific implementation.

• SDS focuses on decoupling storage intelligence from physical hardware and implementing it
in software.

• Because of this philosophical nature, the boundaries of what qualifies as SDS are not always
rigid, and vendor interpretations may vary.

BITS Pilani, Pilani Campus


The Details of SDS

Fundamental Characteristics of SDS


A solution can be considered SDS only if it satisfies the following core properties:
• Hardware Abstraction
Decouples storage software from physical devices

• Policy-Driven Services
Applies protection and data services based on policy

• Programmability
Exposes standard APIs for automation and integration

• Elastic Scalability
Scales capacity and performance as demand grows

• Automation and Orchestration


Integrates with cloud and infrastructure orchestration tools
BITS Pilani, Pilani Campus
The Details of SDS - Abstraction

Hardware Abstraction in Software-Defined Storage


• SDS is fundamentally an abstraction layer over physical storage
• Similar to compute virtualization, it decouples logical storage from hardware devices
• This abstraction is the primary source of SDS flexibility
• Enables portability, automation, and hardware independence
• SDS can abstract:
• Commodity disks in servers
• Or storage from traditional monolithic arrays

Hyperconvergence is one common deployment model, but not a requirement for SDS

BITS Pilani, Pilani Campus


The Details of SDS - Abstraction

SDS Implementation Models for Hardware Abstraction

1. Virtual Appliance Model

• SDS runs as a virtual machine in the infrastructure

• Abstracts backend storage from front-end workloads

• Manages data services through a centralized software layer

2. Kernel-Level Model

• SDS runs inside the hypervisor kernel

• Provides storage services directly from the host

• Enables tighter integration and lower overhead

BITS Pilani, Pilani Campus


The Details of SDS – Policy Driven
Services
Policy-Driven Management in Software-Defined Storage

Policies reduce: Administrative effort, Configuration errors and Long-term inconsistency

Policies govern: Data placement, Protection levels, Performance behavior, Service availability

Example Policy (VM-Centric):

• Stripe data across a defined number of disks or nodes

• Snapshot every 6 hours

• Retain snapshots:

• 3 days onsite

• 7 days offsite (replicated)


BITS Pilani, Pilani Campus
The Details of SDS – Policy Driven
Services
Policy-Driven Management in Software-Defined Storage

BITS Pilani, Pilani Campus


The Details of SDS – Programmability
Programmability as a Core Characteristic of SDS
• Automation is a foundational requirement of the Software-Defined Data Center (SDDC)

• For automation to work, system functions must be exposed through APIs

• APIs provide a standard, programmatic interface to:

• Query system state

• Modify configurations

• Trigger operational actions

• Common API models:

• SOAP – legacy, declining usage

• REST – dominant modern standard

BITS Pilani, Pilani Campus


The Details of SDS – Programmability

APIs as the Foundation of Orchestration


• SDDC relies on an orchestration engine to coordinate all components

• Orchestration requires each subsystem to be: Discoverable, Controllable and Automatable via APIs

• APIs provide the integration point between:

• Compute

• Storage (SDS)

• Networking

• Management platforms

• The Programmable Data Center vision:

• Every infrastructure function is accessible and controllable via API

BITS Pilani, Pilani Campus


The Details of SDS – Scalability

Elastic Scalability as a Core Property of SDS


• SDS is inherently designed for elastic, non-disruptive scaling

• Scalability is enabled by hardware abstraction, which hides physical changes from workloads

• SDS allows:

• Addition of new storage nodes

• Removal or replacement of hardware

• Migration to new platforms without impacting running workloads

• This provides a major advantage over traditional scaling, which required:

• Purchasing larger monolithic arrays

• Lengthy and risky data migrations

BITS Pilani, Pilani Campus


The Details of SDS – Automation and
Orchestration
SDS is designed to operate in an automated, orchestration-driven environment

Manual, device-level administration is replaced by workflow-driven operations

Automation enables:

• Automatic provisioning of storage

• Policy-based placement of workloads

• Automated failure recovery and rebalancing

Orchestration coordinates SDS with:

• Compute virtualization

• Virtual networking

• Cloud and container platforms

BITS Pilani, Pilani Campus


The Details of SDS – Summary

BITS Pilani, Pilani Campus


Hyperconverged Infrastructure

Definition of Hyperconverged Infrastructure

• Hyperconverged Infrastructure (HCI) is an architecture that integrates compute, storage, and networking
into a single software-defined system

• All core services run on the same set of x86 nodes

• Built on three foundations:

• Compute virtualization

• Software-Defined Storage (SDS)

• Virtual networking

• Managed through a single, unified management plane

Overall, HCI transforms multiple infrastructure silos into one integrated, software-defined
platform.
BITS Pilani, Pilani Campus
Hyperconverged Infrastructure
Architectural Components of HCI

• Each HCI node provides:


• CPU and memory (compute)
• Local disks (storage)
• Virtual networking

• Nodes are clustered to form a distributed system

• Storage is provided by an embedded SDS layer

• Management is centralized and policy-driven

Architecture Characteristics:
• Node-based, scale-out design
• No external shared storage arrays
• Tight coupling of compute and storage
BITS Pilani, Pilani Campus
Hyperconverged Infrastructure

Why Organizations Adopt HCI


• Simplified deployment and operations

Single system instead of multiple platforms

• Linear scalability

Add nodes to scale compute and storage together

• Reduced CapEx and OpEx

Commodity hardware

Fewer specialized devices

• Unified lifecycle management

Automated provisioning, upgrades, and patching

BITS Pilani, Pilani Campus


Hyperconverged Infrastructure

Where HCI Fits the Best


Common Use Cases:

• Enterprise virtualization

• Virtual Desktop Infrastructure (VDI)

• Remote Office / Branch Office (ROBO)

• Private cloud deployments

Limitations:

• Inefficient for storage-heavy workloads

• Fixed compute–storage scaling ratios

• Strong dependence on vendor ecosystem

BITS Pilani, Pilani Campus


Hyperconverged Infrastructure

BITS Pilani, Pilani Campus


Hyperconverged Infrastructure

The Comparison

BITS Pilani, Pilani Campus


One Platform, Many Services - Overview
• Modern data centers are built around a single, software-defined platform

• This platform delivers multiple infrastructure services from the same system

• Services are implemented in software, not in separate hardware appliances

• A single SDS / HCI platform can provide:

• Compute virtualization

• Block, file, and object storage

• Data protection (snapshots, replication, backup)

• High availability and fault tolerance

• Performance optimization and caching

• Monitoring, analytics, and reporting

BITS Pilani, Pilani Campus


One Platform, Many Services - Overview
How One Platform Enables Many Services

• Hardware abstraction separates services from physical devices

• Policy-driven control defines how each service behaves

• APIs and automation expose services on demand

• New services can be added or updated through software

Benefits of the One-Platform Model

• Reduced infrastructure complexity

• Fewer management tools and interfaces

• Faster service provisioning

• Lower CapEx and OpEx

• Consistent policy enforcement across services

BITS Pilani, Pilani Campus


One Platform, Many Services -
Comparison

BITS Pilani, Pilani Campus


The Evolution of IT Infrastructure

BITS Pilani, Pilani Campus


Is Hyperconvergence Hardware or Software?

Understanding the Nature of Hyperconverged Infrastructure


• Hyperconvergence combines software intelligence with physical compute, storage, and network
resources

• This dual nature often creates confusion for administrators new to HCI

• Common question:
Is HCI a special hardware appliance or a software platform?

• Short answer: It is both software and hardware

• HCI is fundamentally software-defined, but it always runs on physical infrastructure

BITS Pilani, Pilani Campus


Is Hyperconvergence Hardware or Software?

HCI with Specialized Hardware:

• Limits hardware choice to vendor-certified platforms

• Provides:

• Higher stability

• Better performance

• Increased usable capacity per node (all else being equal)

• Reduced risk due to tightly integrated hardware–software stack

HCI with Software-Only VSA (Virtual Storage Appliance):

• Runs as software on the hypervisor

• Enables broad hardware flexibility

• Avoids vendor lock-in at the hardware level

BITS Pilani, Pilani Campus


The Relationship Between SDS and HCI

BITS Pilani, Pilani Campus


The Relationship Between SDS and HCI

SDS and HCI in Modern Data Center Design:

• Software-Defined Storage (SDS) and Hyperconverged Infrastructure (HCI) are closely related but not
interchangeable

• SDS is a foundational technology, while HCI is an architectural deployment model

• SDS can exist independently of HCI

• HCI depends on SDS to provide distributed, software-defined storage

• In HCI, SDS is embedded within each node

• Local disks are pooled and presented as shared, resilient storage

• Without SDS, HCI would require external shared storage, defeating its purpose

BITS Pilani, Pilani Campus


The Relationship Between SDS and HCI

SDS and HCI in the SDDC Evolution Path:

• SDS enables hardware abstraction and storage virtualization

• HCI uses SDS to integrate compute and storage into a single platform

• Together, they form a stepping stone toward the Software-Defined Data Center (SDDC)

• Design tradeoffs include:

• Flexibility vs performance

• Hardware choice vs optimization

Understanding the relationship between SDS and HCI is critical for designing scalable, efficient,
and operable modern data centers.

BITS Pilani, Pilani Campus


The Role of Flash in Hyperconvergence

Importance of Flash in HCI

• Hyperconverged Infrastructure consolidates compute and storage workloads on the same nodes

• This consolidation significantly increases I/O demand and latency sensitivity

• Traditional HDD-based storage struggles to meet the performance requirements of HCI

• Flash storage (SSDs/NVMe) addresses these challenges by delivering:

• Low latency

• High IOPS

• Consistent performance

BITS Pilani, Pilani Campus


The Role of Flash in Hyperconvergence

Flash as a Performance Layer in SDS

• In HCI, flash is managed by the Software-Defined Storage (SDS) layer

• SDS uses flash to provide:

• Read and write caching

• Tiering between flash and capacity drives

• Log-structured or distributed metadata services

• Flash improves:

• VM responsiveness

• Storage efficiency

• Predictable workload performance

BITS Pilani, Pilani Campus


The Role of Flash in Hyperconvergence

How Flash Is Used in Hyperconverged Nodes

All-Flash HCI

• All storage media are SSDs or NVMe

• Maximum performance and lowest latency


Choice depends on:
• Higher cost per node
• Application I/O profile
Hybrid HCI • Budget constraints
• Flash for cache + HDD for capacity • Performance SLAs
• Balanced cost and performance

• Common in mid-sized data centers

BITS Pilani, Pilani Campus


The Role of Flash in Hyperconvergence

Flash in Data Center Design and Operations

Flash reduces the need for:


• Complex storage tuning
• Overprovisioning of hardware

Enables:
• Higher VM density per node
• Faster rebuilds and resynchronization
• Improved resilience during failures

Design considerations include:


• Endurance and write amplification
• Cost vs performance tradeoffs
• Network bandwidth alignment

BITS Pilani, Pilani Campus


The Role of Flash in Hyperconvergence

BITS Pilani, Pilani Campus


The Role of Flash in Hyperconvergence

Why Flash Storage Is Essential in Hyperconverged Infrastructure

Legacy Monolithic Storage

• Performance scaled by adding more disks

• Spinning HDDs capped at ~160–180 IOPS per disk

• Higher workload demand ⇒ more disks required, even if capacity wasn’t needed

• Large monolithic arrays could scale easily by adding disk shelves

Challenge in Hyperconvergence

• Storage is node-limited (e.g., 2U server ≈ 24 disks max)

• Maximum HDD-based performance per node ≈ 4,000 IOPS

• Cannot scale performance endlessly by adding disks

BITS Pilani, Pilani Campus


The Role of Flash in Hyperconvergence

Why Flash Storage Is Essential in Hyperconverged Infrastructure

Flash Storage Solution

• SSDs deliver thousands to tens of thousands of IOPS

• 1 SSD ≈ performance of 24 HDDs

• Eliminates performance bottlenecks without increasing node count

Flash storage is not optional—it is fundamental to achieving performance scalability


in hyperconverged systems

BITS Pilani, Pilani Campus


Class Summary

• The Parallel Paths of SDS and Hyperconvergence


• The Details of SDS
• What Is Hyperconverged Infrastructure?
• The Relationship Between SDS and HCI
• The Role of Flash in Hyperconvergence

BITS Pilani, Pilani Campus


Next Session:
Data Center Design- Infrastructure Types

STOP RECORDING
50
Design and Operation of Data Centers
CSIWZG522
BITS Pilani Contributed by
Pilani Campus
Dr. Suhaas K P
IMP Note to Self
IMP Note to Students

• It is important to know that just login to the session does not guarantee the
attendance.
• Once you join the session, continue till the end to consider you as present in the
class.
• IMPORTANTLY, you need to make the class more interactive by responding to
Professors queries in the session.
• Whenever Professor calls your number / name ,you need to respond, otherwise it
will be considered as ABSENT.

BITS Pilani, Pilani Campus


BITS Pilani
Pilani Campus

CSIWZG522: Design and Operation of Data Centers


CS 5 : Data Center Design-Understanding Servers
TEXTBOOK AND REFERENCES A book cover of a data center

T1 Building a Modern Data Center: Principles and


Strategies of Design
by
Scott D. Lowe, David M. Davis, James Green
Click to edit Session title

R1 Data Center for Beginners: A beginner's guide towards


understanding Data Center Design
by
B.A. Ayomaya
BITS Pilani, Pilani Campus
Advisory

• The printed book is available in the market.

• The e-book is available on the Internet.

• Students are strongly advised to make notes from RLs & CSs for
examination purpose right from the beginning.

• Unavailability of the textbook will not be an excuse to allow print copies


of the e-book during examinations.

BITS Pilani, Pilani Campus


RECAP OF PREVIOUS SESSION

• Virtualization
• Benefits of Virtualization
• Types of Virtualization
• VM Hardware Components
• Hyperconvergence
• The Cloud OpenStack
• The Role of Cloud
• Cloud Types
• Cloud Drivers
• Case Study : Cisco Hyperconvergence Strategy

BITS Pilani, Pilani Campus


CONTACT SESSION 05

Data Center Design-Understanding Servers

BITS Pilani, Pilani Campus


What is an SDDC?

Data Center Design-Understanding Servers

• An SDDC virtualizes and delivers all infrastructure — compute, storage, networking, and
security — as software services.

• Management, provisioning and policy enforcement are performed in software rather than by
manual hardware configuration.

• Goal: deliver IT-as-a-Service (ITaaS) with policy-driven automation and pooled resources

BITS Pilani, Pilani Campus


Why did SDDC emerge?

• Growth of dynamic workloads and cloud-style consumption models (need for rapid
provisioning).

• Hardware heterogeneity and hardware-centric silos made manual ops slow and error-prone.

• Need for consistent operations across on-prem, edge and public cloud (hybrid/multi-cloud).

• Rise of software primitives: SDN (software-defined networking), SDS (software-defined


storage), and mature hypervisors & orchestration tools.

BITS Pilani, Pilani Campus


Typical SDDC components / architecture

• Compute virtualization (hypervisors, container runtimes).


• Software-defined storage (SDS): virtual storage pools, policy-driven provisioning.
• Software-defined networking (SDN): programmable control plane, overlays (VXLAN, etc.).
Clicklayer:
• Management & orchestration to edit Session
SDDC manager, title automation/orchestration
policy engine,
(Terraform, Ansible).

• Security & policy: micro-segmentation, centralized policy enforcement.

BITS Pilani, Pilani Campus


Enabling technologies

• SDN controllers and virtual


switches (Open vSwitch, vendor
controllers).
• SDS solutions (block/object
abstraction, erasure coding, thin
provisioning).
• Orchestration & IaaC
(Infrastructure-as-a-Code tools).
• Telemetry/observability: software
agents, metrics and logging.

BITS Pilani, Pilani Campus


Use cases (where SDDC adds value)

Private & hybrid cloud — consistent environment for workloads across on-prem and cloud.

Disaster recovery & DRaaS — automated failover and replication.

Multi-tenant service providers / Telco clouds — isolation and rapid tenant provisioning.

Edge/Distributed data centers — centralized management for many small footprints.

Dev/Test and CI/CD — on-demand environments and fast tear-down.

BITS Pilani, Pilani Campus


Advantages

• Agility: fast provisioning — minutes vs days/weeks.


• Operational efficiency: policy automation reduces manual configuration and errors.
• Resource utilization & scalability: pooled resources and elastic allocation.
• Consistency & portability: identical management planes across sites/clouds.
• Lower TCO (over time) through automation, reduced manual labor and hardware consolidation.

BITS Pilani, Pilani Campus


Risks & Challenges

Skills & organizational change: Need for software/DevOps capabilities.

Security & policy correctness: Software misconfiguration risks; need for strong testing.

Vendor lock-in and interoperability: Proprietary SDDC stacks vs open APIs.

Legacy application constraints: Not all apps are cloud-native or easily virtualizable.

Operational complexity at scale: Telemetry, debugging overlays and overlays-of-overlays.

BITS Pilani, Pilani Campus


VMware Cloud Foundation on Equinix
Metal
Case Study:
Context / Goal: Deliver a hybrid/distributed SDDC to support latency-sensitive, distributed
workloads and provide consistent operations across on-prem and distributed endpoints.

Solution: VMware Cloud Foundation (VCF) deployed on Equinix Metal (bare-metal), forming
a managed SDDC endpoint that can run VMware SDDC Manager and provide local
compute/storage with global interconnection.

Outcomes / Benefits: faster rollout of distributed SDDC endpoints, improved latency for
edge apps, consistent management and security policies across locations, and simpler
hybrid operational model.
BITS Pilani, Pilani Campus
VMware Cloud Foundation on Equinix
Metal

BITS Pilani, Pilani Campus


Commoditization of Hardware in Data
Centers

Transformation of servers, storage, and networking devices into standardized,


low-cost, interchangeable components

Hardware becomes less differentiated; intelligence shifts to software layers


Click to
Focus moves from proprietary edittoSession
systems title platforms
x86-based commodity

Enables large-scale, vendor-neutral data center architectures

BITS Pilani, Pilani Campus


Why Did Hardware Become a Commodity?

• Cloud-scale demand for cost efficiency and rapid scalability

• Standardization of x86 architecture and Ethernet networking

• Rise of hyperscalers (Google, Amazon, Meta) designing in-house hardware


Click(Open
• Open hardware initiatives to edit Session
Compute title networking)
Project, white-box

Impact on Design & Operations:


• Disaggregated architectures (compute, storage, network as separate pools)

• Higher hardware refresh rates

• Vendor lock-in reduced; interoperability increased

• Strong dependency on automation and software management

BITS Pilani, Pilani Campus


Commoditized Hardware and Its Role in
SDDC

BITS Pilani, Pilani Campus


Software Defined Compute

Abstraction of Compute Resources


Decouples applications and virtual machines from
physical servers, enabling flexible placement and
migration.
Virtualization and Container Support
Supports multiple execution models including virtual
machines, containers, and serverless workloads.
Automated Provisioning and Scaling
Enables dynamic allocation, scaling, and
deprovisioning of compute resources using policies.
Centralized Management and Control
Provides policy-based monitoring, and lifecycle
management from a unified control plane.
BITS Pilani, Pilani Campus
Shift to Software-Defined Compute

From Hardware-Centric to Software-Defined Compute

• Traditional data centers relied on tightly coupled hardware and operating systems

• Compute capacity was fixed, statically provisioned, and manually managed

• Growing workload diversity and cloud adoption demanded greater flexibility

• Result: Emergence of software-defined compute as a core design principle

BITS Pilani, Pilani Campus


Shift to Software-Defined Compute

Why Did Data Centers Move to Software-Defined Compute?


• Rapid growth of virtualization and cloud-native applications

• Need for elastic scaling and on-demand provisioning

• Increasing hardware heterogeneity and large-scale deployments

• Business demand for agility, faster time-to-service, and cost optimization

Operational Impact:

• Shift from static capacity planning to dynamic resource allocation

• Increased reliance on automation and orchestration

BITS Pilani, Pilani Campus


Shift to Software-Defined Compute

Changes in Data Center Compute Architecture

• Decoupling of applications from physical hardware through hypervisors and containers


• Introduction of a virtualization layer above commodity servers
• Centralized management plane controlling distributed compute resources
• Formation of large, shared compute pools across clusters

Design Implication:

Data center layouts now optimize for scalability, density, and uniform hardware profiles rather
than application-specific servers.

BITS Pilani, Pilani Campus


Shift to Software-Defined Compute

Operational Changes Enabled by Software-Defined Compute

• Automated provisioning and lifecycle management of compute instances

• Policy-driven scheduling, placement, and resource allocation

• Centralized monitoring, telemetry, and capacity optimization

• Faster recovery, live migration, and workload mobility

Operational Outcome:

Operations shift from manual server administration to software-driven service management

BITS Pilani, Pilani Campus


Shift to Software-Defined Compute

Advantages in Modern Data Centers

• Improved resource utilization and consolidation ratios

• Rapid scalability and workload elasticity

• Reduced provisioning time and operational errors

• Enhanced availability through live migration and fault tolerance

Business Value:

Higher infrastructure efficiency with lower operational overhead and improved service agility.

BITS Pilani, Pilani Campus


Shift to Software-Defined Compute

CASE STUDY: VMware vSphere — Enterprise Software-Defined Compute


• Hypervisor (ESXi) abstracts physical servers
• vCenter provides centralized management and clustering
Features:
• VM live migration (vMotion)
• Distributed Resource Scheduler (DRS)
• High Availability (HA)
Operational Impact:
• Dynamic workload placement across clusters
• Automated failover and load balancing
• Significant reduction in manual server management

BITS Pilani, Pilani Campus


Shift to Software-Defined Compute

BITS Pilani, Pilani Campus


Shift to Software-Defined Storage

From Hardware-Centric Storage to Software-Defined Storage

• Traditional data centers relied on dedicated storage arrays and proprietary controllers

• Storage capacity and performance were tightly bound to specific hardware platforms

• Scaling required purchasing and integrating new storage appliances

• Result: Emergence of software-defined storage as a key architectural shift

BITS Pilani, Pilani Campus


Shift to Software-Defined Storage

Why Did Data Centers Move to Software-Defined Storage?

• Rapid growth of data volumes and unstructured data

• Need for elastic capacity scaling and tiered storage

• High cost and vendor lock-in of proprietary storage systems

• Demand for unified storage across on-premises and cloud environments

Operational Impact:

• Shift from static capacity planning to dynamic storage provisioning

• Increased reliance on automation and policy-based management

BITS Pilani, Pilani Campus


Shift to Software-Defined Storage

Changes in Storage Architecture

• Decoupling of storage services from dedicated storage hardware

• Introduction of a storage virtualization layer above commodity disks

• Formation of distributed storage pools across multiple servers

• Centralized control plane managing data placement and replication

Design Implication:
Storage design shifts from monolithic arrays to distributed, software-managed architectures

BITS Pilani, Pilani Campus


Shift to Software-Defined Storage

Advantages in Modern Data Centers

• Lower capital cost using commodity hardware

• Elastic scalability and pay-as-you-grow expansion

• Improved utilization through pooling and thin provisioning

• Unified management across block, file, and object storage

Business Value:
Higher storage efficiency with reduced cost and improved agility.

BITS Pilani, Pilani Campus


Shift to Software-Defined Storage

BITS Pilani, Pilani Campus


Shift to Software-Defined Storage

BITS Pilani, Pilani Campus


Shift to Software-Defined Networking

From Hardware-Centric Networking to Software-Defined Networking

• Traditional data center networks relied on tightly coupled routers and switches

• Control logic and forwarding functions were embedded in proprietary hardware

• Network configuration was manual, device-by-device, and error-prone

• Result: Emergence of software-defined networking as a major architectural shift

BITS Pilani, Pilani Campus


Shift to Software-Defined Networking

Why Did Data Centers Move to SDN?

• Rapid growth of virtualized and cloud-native workloads

• Explosion of east–west traffic inside data centers

• Need for rapid provisioning and network automation

• Limitations of static VLAN-based and box-by-box configurations

Operational Impact:

• Shift from static network design to programmable network fabrics

• Increased reliance on centralized control and APIs

BITS Pilani, Pilani Campus


Shift to Software-Defined Networking

Changes in Network Architecture

• Decoupling of control plane from data plane

• Introduction of centralized SDN controllers

• Programmable forwarding devices using open or vendor APIs

• Logical overlay networks built on top of physical fabrics

Design Implication:
Network design shifts from distributed device intelligence to centralized software control.

BITS Pilani, Pilani Campus


Shift to Software-Defined Networking

Operational Changes Enabled by SDN

• Centralized configuration and policy-based networking

• Automated network provisioning and service chaining

• Faster troubleshooting through global network visibility

• Dynamic traffic engineering and load balancing

Operational Outcome:
Operations shift from manual device management to software-driven network services

BITS Pilani, Pilani Campus


Shift to Software-Defined Networking

Advantages in Modern Data Centers

• Rapid service provisioning and reduced configuration errors

• Improved network utilization and traffic optimization

• Enhanced security through micro-segmentation and policy enforcement

• Better support for multi-tenant and hybrid cloud environments

Business Value:
Higher network agility with improved reliability and lower operational cost.

BITS Pilani, Pilani Campus


Shift to Software-Defined Networking

BITS Pilani, Pilani Campus


Shift to Software-Defined Networking

BITS Pilani, Pilani Campus


Shift to Software-Defined Security

From Perimeter-Based Security to Software-Defined Security

• Traditional security relied on perimeter firewalls and hardware appliances

• Security controls were static, device-centric, and manually configured

• Virtualization and cloud broke the notion of a fixed network perimeter

• Result: Emergence of software-defined security as a core architectural shift

BITS Pilani, Pilani Campus


Shift to Software-Defined Security

Why Did Data Centers Move to Software-Defined Security?

• Rapid adoption of virtualization, cloud, and microservices

• Growth of east–west traffic inside data centers

• Inadequacy of perimeter-only security models

• Need for fine-grained, workload-level security controls

Operational Impact:

• Shift from perimeter defense to distributed, workload-centric security

• Increased reliance on automation and centralized policy management

BITS Pilani, Pilani Campus


Shift to Software-Defined Security

Changes in Security Architecture

• Decoupling of security functions from physical appliances

• Introduction of centralized security control planes

• Distributed enforcement points embedded in hypervisors and virtual switches

• Policy-based security applied at VM, container, and application level

Design Implication:
Security design shifts from device-based protection to software-enforced, pervasive protection.

BITS Pilani, Pilani Campus


Shift to Software-Defined Security

BITS Pilani, Pilani Campus


Overall SDDC

BITS Pilani, Pilani Campus


Class Summary

• The Emergence of SDDC


• Commoditization of Hardware
• Shift to Software Defined Compute
• Shift to Software Defined Storage
• Shift to Software Defined Networking

• Shift to Software Defined Security

BITS Pilani, Pilani Campus


Next Session:
Data Center Design-Understanding Servers

STOP RECORDING
48
CSIWZG522 :
Design and Operations of Data Center

Dr. SRINIVASA KOSIGANTI


srinikosi@[Link]

CS 04
Click to edit Session title

2
Important Note to Students
⮚ It is important to know that just login to the session does not guarantee the
attendance.

⮚ Once you join the session, continue till the end to consider you as present in the
class.

⮚ IMPORTANTLY, you need to make the class more interactive by responding to


Professors queries in the session.

⮚ Answering to Polls & Quiz questions during / after the class are mandatory

⮚ Whenever Professor calls your number / name ,you need to respond,


otherwise it will be considered as ABSENT, and you will be removed from
the meeting

⮚ Students are advised join within 10 minutes once the class starts; Professor
will lock the entry to the meeting. 3
Disclaimer

• The slides presented here are obtained from the authors of the books and
from various other contributors. I hereby acknowledge all the contributors
for their material and inputs.

• I have added and modified a few slides to suit the requirements of the course.

4
Textbook & Reference
T1 Building a Modern Data Center:
Principles and Strategies of Design
by
Scott D. Lowe, David M. Davis, James Green

Reference Book
R1 Data Center for Beginners:
A beginner's guide towards understanding Data
Center
Design
by5
Advisory

• The printed book is available in the market.

• The e-book is available on the Internet.

• Students are strongly advised to make notes from RLs & CSs for examination
purpose right from the beginning.

• Unavailability of the textbook will not be an excuse to allow print copies of the e-
book during examinations.

6
Recap of Last Class
• Data Centre Evolution
• Data Center’s Computing Components
• Requirements for a modern data center networking platform
• Case Study - Google data centers’ Networking
• Data Center Storage
• Direct Attached Storage (DAS)
• Network Attached Storage (NAS)
• Storage Area Network (SAN)
• Case Study : AWS Data Center for Storage of AWS IoT and S3
• The Rise of the Monolithic Storage Array
• Block vs. File Storage
• The Virtualization of Compute — Software Defined Servers
• Case Study : Microsoft Underwater Data Center 7
Virtualization

Virtualization

Transforming a Classic Data Center (CDC) into a Virtualized Software Defined Data
Center

8
Data Center Evolution & Key Concepts
REAL-LIFE SIMPLE
TOPIC DESCRIPTION
EXAMPLE UNDERSTANDING

Historical progression DCs changed from basic


Google → AWS →
Data Centre Evolution from traditional to server rooms to
Microsoft cloud DCs
modern data centers intelligent cloud facilities

Servers: processors,
memory, storage, Intel Xeon + DDR5 RAM + Building blocks of any
Computing Components
network as core NVMe SSDs DC
infrastructure

Case Study: Google Data


Google B4 network Smart network routing
Networking Platform Centers' Networking
connecting global DCs for efficiency
Architecture

Case Study: AWS Data


S3 object storage, EC2 Multiple storage options
Storage (DAS, NAS, SAN) Center for IoT & S3
instances for different needs
Storage implementation

Traditional storage Old approach:


Monolithic Storage NetApp FAS series, EMC
systems; advantages and centralized controller
Arrays VMAX
limitations bottleneck

Case Study: Microsoft Innovation: cooling


Project Natick -
Software Defined Servers Underwater Data Center efficiency, real estate
underwater servers
using virtualization optimization
Virtualization Fundamentals - Core
Concepts
Virtualized
REAL-WORLD
Data CONCEPT DETAILS FEATURE
Center IMPACT
(VDC)
Transforming
Virtua Traditional DC into
lize Hardware abstraction Run multiple OS on
Virtualization Virtualized DC through
layer single physical machine
Netw core element
Virtua ork virtualization
lize Software allowing
Stora multiple OSs on
ge VMM + Kernel Foundation of virtual
Hypervisor physical machine,
components infrastructure
Virtual interacting with
ize hardware
Comp Kernel (OS functions) +
ute Virtual Machine Monitor Efficient resource
Components Two-layer architecture
Classic (VMM) managing VMs allocation
and resources
Data
Center Logical entity with
(CDC) vCPU, vRAM, vDisk,
Virtual Machine (VM) vNIC - looks like Standardized resources Portability across hosts
physical machine to
guest OS
Resides between
hardware (CPU,
Complete resource
Virtualization Layer Memory, NIC, HDD) and Hardware abstraction
independence
VMs; abstracts physical
resources
Smoother transition:
Benefits of Virtualization - Before vs After
BEFORE AFTER
ASPECT BUSINESS IMPACT
VIRTUALIZATION VIRTUALIZATION

OS Model Single OS Multiple OS Flexibility

Hardware Coupling Tightly coupled s/w & h/w OS/app h/w independent Agility

Issues Conflicts, resource waste VM isolation Stability

Resource Utilization Underutilized (20-30%) Improved (70-80%) Cost efficiency

Infrastructure Inflexible, expensive Flexible, low cost OpEx reduction

Scalability Limited Dynamic scaling Growth enabler

Management Manual, complex Automated, centralized Operational efficiency


x86 Virtualization Requirements & Architecture
TECHNICAL
REQUIREMENT DETAILS IMPLEMENTATION
ASPECT
Ring 0 (OS/privileged),
Ring 1, Ring 2, Ring 3
x86 Privilege Rings CPU privilege levels Protection mechanism
(User Apps); OS in Ring
0
Placing virtualization
layer below OS;
Virtualization Challenge capturing & translating VMM in Ring -1 Hardware trick
privileged instructions
at runtime
Full virtualization,
Virtualization Paravirtualization,
Three approaches Different trade-offs
Techniques Hardware-assisted
virtualization
Real-time capture and
translation of privileged Full virtualization
Binary Translation Overhead reduction
OS instructions at mechanism
runtime

Improved utilization and


Resource Optimization flexibility through CPU, Memory sharing Better TCO
hardware abstraction

Hardware-assisted
virtualization (AMD-V,
Modern Standards CPU-level support Native efficiency
Intel VT) is industry
Hypervisor Types & Architecture
TYPE/COMPONENT DESCRIPTION CHARACTERISTICS USE CASE

Operating system installs


on x86 bare-metal
VMware ESXi, Hyper-V,
Type 1: Bare-Metal hardware directly; Enterprise data centers
KVM
requires certified
hardware
Installs as application on
existing OS; relies on OS VMware Workstation,
Type 2: Hosted Development, testing
for device support & VirtualBox
resource management

Manages physical
Kernel Component hardware resources and Privileged operations System stability
core OS functions

Virtual Machine Monitor:


manages VM isolation,
VMM Component resource allocation, Resource scheduler VM independence
virtual hardware
abstraction

Each VM provided with


Virtual Components standardized resources: Standardization Portability
vCPU, vRAM, vDisk, vNIC

x86 Architecture support:


Architecture Support CPU, Memory, NIC Card, Complete virtualization Full compatibility
Virtualization Techniques Comparison
TECHNIQUE HOW IT WORKS PROS CONS MODERN USE
VMM in Ring 0;
Binary Translation
(BT) of non- Performance
Full Virtualization Works with any OS Legacy systems
virtualizable overhead
instructions; Guest
OS unaware
Guest OS aware of
virtualization;
Modified kernel Requires OS Linux-heavy
Paravirtualization Better performance
(Linux, OpenBSD); modification environments
runs in Ring 0; No
Windows
Hypervisor-aware
CPU handles
privileged
Hardware-Assisted Optimal performance CPU dependency All modern systems
instructions;
Reduced overhead;
AMD-V & Intel VT
Real-time capture and
translation of
Full virtualization
Binary Translation privileged OS Transparent to guest High overhead
only
instructions at
runtime
Hardware-assisted
Performance reduces virtualization 10-15% overhead
Cost of CPU support Enterprise preference
Advantage overhead vs. reduction
full/paravirtualization
Virtual Machine Architecture & Files
PURPOSE & STORAGE
FILE TYPE IMPORTANCE
DESCRIPTION LOCATION
Logical compute system
running OS and
VM User View Virtual environment Transparency to apps
applications like
physical machine

Discrete set of files


VM Hypervisor View (configuration, disk, Datastore/NAS/SAN Management abstraction
BIOS, swap, log)

Stores VM creation
choices: CPU count,
Configuration File .vmx file (VMware) VM blueprint
memory, network
adapters, disk types
Stores VM disk
contents; appears as
Virtual Disk File VMDK/VHDX files Data persistence
physical disk; VM can
have multiple disks
BIOS (virtual state),
Swap (paging when
Support Files Associated files Operational support
running), Log
(troubleshooting)

VMFS: cluster filesystem


File Systems (FC, iSCSI) | NFS: Storage backend Data accessibility
remote NAS storage
VM Hardware Components - Virtual Resources
REAL-WORLD
COMPONENT DEFINITION CONFIGURATION
EXAMPLE
One or more virtual CPUs;
configurable and Enterprise app: 8 vCPU
vCPU (Virtual CPU) 1-16+ vCPUs per VM
changeable per VM VM
requirement
Memory amount
presented to guest OS; Database server: 128GB
vRAM (Virtual Memory) 512MB to 512GB+
size changeable based on vRAM
requirements
Stores VM's OS and
application data; OS disk: 50GB, Data disk:
Virtual Disk 10GB-10TB+
minimum one virtual disk 500GB
required per VM

Enables VM connection to
Multi-homed: internal +
vNIC (Virtual NIC) other physical and virtual 1-4 NICs typical
external
machines on network

DVD/CD-ROM, Floppy
drive, SCSI, USB
Virtual Peripherals As needed CD-ROM for OS install
controllers, Graphic card,
IDE controllers
Serial/Com, Parallel ports
(legacy), Keyboard,
Interface Ports Legacy support Backward compatibility
Mouse interfaces for full
compatibility
Computing Infrastructure - Server Types & Evolution
MODERN
COMPONENT DEFINITION TYPE/CATEGORY
IMPLEMENTATION
Rack-mount (pizza-box),
Blade (compact in 1U/2U rack servers
Server Types Form factors
chassis), Mainframe primary
(multiple processors)
Provide processing,
memory, local storage, Compute nodes in
Server Functions Core functions
and network connectivity clusters
for applications
Integration of switches,
routing, load balancing, Software-defined
DC Networking Network infrastructure
analytics for networking
storage/processing
From physical to
virtualized: VMs,
Network Evolution containers, bare metal Technology shift Microservices architecture
with centralized
management

From hardware-centric
Virtualization Impact silos to integrated, flexible Architectural change Cloud-native design
infrastructure

Evolution: Traditional →
Virtualized → Software-
SDDC Journey Maturity progression Industry direction 2026+
Defined Data Center
(SDDC)
Storage Evolution - Magnetic to Flash Era
ERA/TECHNOLOGY DEFINITION CHARACTERISTICS CURRENT STATUS

Hard disk drives (HDD);


High latency, high
Magnetic Storage dominant in data center Legacy systems only
capacity, low cost
history

VMs retrieve hot blocks


from local DRAM/SSD
VDI Boot Storms Performance optimization Solved with tiering
cache before network
traversal
Solid State Drives;
improved latency &
Flash/SSD Advantages Low latency, high IOPS Standard in new DCs
throughput; DRAM/SSD
caching in front of arrays
Radical architecture
replacing monolithic
Software-Defined Storage Distributed storage Industry standard
arrays for general-
purpose workloads

Scale-out architectures
Scalability Solutions enabling better utilization Horizontal scaling Hyperscale model
and flexibility

AI-ready storage with


cloud integration,
Next-Gen Approach Intelligent storage Future direction (2026+)
containers, consumption-
based models
Storage Array Evolution Timeline - 1990 to AI-Ready
ERA STORAGE TYPE ARCHITECTURE TIMELINE USE CASE
Disk-based,
Traditional Disk controller-centric, Legacy enterprise
1990-2010 35 years dominant
Arrays scale-up, very low apps
latency
Disk+Flash,
moderate
2007-2015 Hybrid Arrays performance, scale- Transition period Mixed workloads
up, mixed enterprise
workloads
Flash-based, high
consistent
2009-2018 All-Flash Arrays Flash era Performance-critical
performance, scale-
up, databases & VMs
Software-defined,
node-based, scale-
2012-2020 HyperConverged out, high Modern approach Virtualized DCs
performance for
virtualization, VDI
Disaggregated,
NVMe/Object,
2018-Present Cloud-Native Current standard Public cloud
hyperscale, cloud &
microservices
Composable,
intelligent,
2024-2026+ AI-Ready Next-Gen NVMe/SCM, web- Emerging AI/ML workloads
scale, AI/ML & big
data
Hyperconvergence (HCI) - Definition & Benefits
CONCEPT DEFINITION REAL-LIFE EXAMPLE BUSINESS VALUE

Integrates compute,
storage, networking,
HCI Definition virtualization into single Nutanix platform All-in-one solution
software-defined platform
on x86
HCI eliminates traditional
storage and controller Performance
Bottleneck Elimination Removes SAN complexity
bottlenecks via improvement
distributed SDS

Capacity and performance


Add 1 node = linear
Scale-Out Model increase by adding more Flexibility
growth
nodes horizontally

Centralized software-
based control with faster
Simplified Management Single console for all OpEx reduction
deployment and reduced
complexity
Converged: Proprietary,
scale-up; HCI:
Converged vs HCI HCI wins on flexibility Cost efficiency
Commodity, scale-out,
symmetric
Nutanix, SimpliVity,
Atlantis, Pivot3, Maxta;
HCI Vendors Market leaders Enterprise adoption
Use: VDI, virtualization,
private cloud
Cloud Types & Deployment Models
CLOUD TYPE DEFINITION CHARACTERISTICS BEST FOR

Third-party DCs (AWS,


Azure, GCP, VMware);
Public Cloud External hosting Startups, rapid scaling
light IT footprint; very
affordable; less overhead
Corporate data center as
cloud; on-premises
Private Cloud resource control; Internal control Regulated industries
compliance & data
sovereignty
On-premises + public
cloud combination;
Hybrid Cloud Flexible deployment Enterprise balance
workload flexibility;
bursting during peak
Retailer: DB on-prem,
website on cloud |
Hybrid Examples Developers: production Real scenarios Practical models
on-prem, dev/test on
cloud
On-demand services via
OpenStack; consumption-
Cloud Services SaaS/PaaS/IaaS Agile operations
based model; flexible
scaling
Workload portability
between on-premises and
Key Benefit Strategic advantage Business flexibility
public cloud based on
Key Drivers of DC Transformation
TOPIC DEFINITION INDUSTRY IMPACT FUTURE TREND

Move from magnetic


(HDD) storage to Complete HDD exit by
The No-Spin Zone Storage revolution
Flash/SSD; improved I/O 2026
performance & latency
Centralized controller
bottleneck limitations;
Fall of Monolithic Architectural shift All systems distributed
transition to distributed
architectures
From traditional silos to
converged, then
Convergence Emergence Integration trend HCI + cloud standard
hyperconverged
infrastructure models
Virtualization benefits
extended; performance,
Cloud Role Market expansion Cloud-first strategy
availability, cost
optimization drivers
Cisco Hyperconvergence
Strategy demonstrating
Case Study: Cisco HCI Enterprise adoption Industry validation
modern data center
architecture

Flexibility, scalability,
Value Proposition automation, reduced IT Business outcomes Strategic enabler
ops complexity & capex
Real-World Case Studies & Implementation
SIMPLE
CASE STUDY COMPANY TECHNOLOGY
UNDERSTANDING

Global distributed
Smart routing across
Google Data Centers Google architecture with custom
world-wide servers
networking

Cloud object storage with Flexible storage for IoT


AWS IoT & S3 Amazon
massive scale data

Project Natick - Innovation in energy


Microsoft Underwater DC Microsoft
underwater cooling efficiency

Converged infrastructure Integrated solutions for


Cisco HCI Strategy Cisco
approach enterprises

Desktop-as-a-service
Enterprise VDI Multiple HCI for virtual desktops
delivery

GPU-optimized, high- Next-gen compute for


AI-Ready DCs Google, Meta, AWS
speed interconnects AI/ML
Next Session
TOPIC DETAILS STUDY STRATEGY

Complete transformation to
Software-Defined Data Center - Review VMware SDDC
Emergence of SDDC all infrastructure abstracted and architecture diagram + watch
controlled through software VMware SDDC overview video
policies

Shift from proprietary hardware


Compare Dell/HP server pricing
to commodity x86 servers
Commoditization of Hardware vs traditional enterprise servers
running software-defined
+ study Nutanix TCO calculator
solutions

Virtualization + containers + Practice Kubernetes deployment


Software Defined Compute orchestration replacing physical on Minikube + study VMware
servers Tanzu vs OpenShift comparison

SDS replacing SAN/NAS with Analyze vSAN vs Ceph


Software Defined Storage distributed storage across performance benchmarks +
commodity nodes deploy simple Ceph cluster

SDN replacing VLANs with virtual Study NSX micro-segmentation


Software Defined Networking overlays and policy-based demos + compare Cisco ACI vs
networking VMware NSX
Next Session:
Emergence of SDDC

STOP RECORDING
CSIWZG522 :
Design and Operations of Data Center

Dr. Rama Satish K V


satishkvr@[Link]

CS 03
BITS Pilani
Pilani Campus

Click to edit Session title

CSIWZG522 - Design and operations of Data


center
Click to edit Session title
Important Note to Students
⮚ It is important to know that just login to the session does not guarantee the attendance.

⮚ Once you join the session, continue till the end to consider you as present in the class.

⮚ IMPORTANTLY, you need to make the class more interactive by responding to Professors
queries in the session.

⮚ Answering to Polls & Quiz questions during / after the the class are mandatory

⮚ Whenever Professor calls your number / name ,you need to respond, otherwise it will
be considered as ABSENT and you will be removed from the meeting

⮚ Students are advised join within 10 minutes once the class starts, Professor will lock
the entry to the meeting.
Declaimer

• The slides presented here are obtained from the authors of the books and
from various other contributors. I hereby acknowledge all the contributors
for their material and inputs.

• I have added and modified a few slides to suit the requirements of the course.
Text Book & Reference
A book cover of a data center

AI-generated content may be incorrect.

T1 Building a Modern Data Center: Principles and


Strategies of Design by Scott D. Lowe (Author), David
M. Davis (Author), James Green (Author), Seth Knox
(Editor), Stuart Miniman (Foreword)

Reference Book
R1 Data Center for Beginners: A beginner's guide
towards understanding Data Center Design , B.A.
Ayomaya (Author)
Advisory

• The printed book is available in the market.

• The e-book is available on the Internet.

• Students are strongly advised to make notes from RLs & CSs for
examination purpose right from the beginning.

• Unavailability of the textbook will not be an excuse to allow print copies


of the e-book during examinations.
Recap of Last Class
• What is Data Centre
• Data Centre and its Values
• Data Centre Principles
• Reliability
• Scalability
• Elasticity
• Common Sense Principles Improve Data Center Operations
• Adapt or Die
• Flattening of the IT Organization
• IT as an Operational Expense
• Capital Expenses (CapEx) Vs Operational Expense (OpEx)
Lecture Plan
Time Type Description Content
Reference
Pre CH RL2.1, Data Center Design- Components

During CH2 A History of the Modern Data Center T1: Ch-2


CH The Rise of the Monolithic Storage Array
The Virtualization of Compute — Software Defined servers.

Industry
Case Study: NVIDIA (AI Infrastructure)
Ref.
Data Centre Evolution
A History of the Modern Data Center (1)
A History of the Modern Data Center (2)
• Tape-based data storage technology began
to be displaced when IBM released the
first disk-based storage unit in 1956
(Figure 2-1).
• It was capable of storing a whopping 3.
• 75 megabytes — paltry by today’s terabyte
standards.
• It weighed over a ton, was moved by
forklift, and was delivered by cargo plane.
• Magnetic, spinning disks continue to
increase in capacity to this day, although
the form factor and rotation speed have
been fairly static in recent years.
Figure 2-1: The spinning disk system
for the IBM RAMAC 305
A History of the Modern Data Center (3)
• The last time a new rotational speed was introduced was in
2000 when Seagate introduced the 15,000 RPM Cheetah
drive.
• CPU clock speed and density has increased many times
over since then.
• These two constantly developing technologies — the
microprocessor/x86 architecture and disk-based storage
medium — form the foundation for the modern data center.
• In the 1990s, the prevailing data center design had each
application running on a server, or a set of servers, with
locally attached storage media.
• As the quantity and criticality of line-of-business
applications supported by the data center grew, this
architecture began to show some dramatic inefficiency
when deployed at scale.
• Plus, the process of addressing that inefficiency has
characterized the modern data center for the past two
decades.
Data Center Components (Computing)
• Computing Components
⮚Servers
⮚Networking
⮚Storage
Computing Components: Servers
Computing Components : Networking
Requirements for a modern data center networking
platform:

Requirements for a modern data center


networking platform:

Automation. Achieving speed and agility in


modern data centers depends greatly on
automated provisioning of networking services
for applications.

Far faster and more reliable than a human


administrator, modern networking platforms
not only find the most efficient way to program
a network, balance workloads, and automate
time-consuming tasks, they also respond
dynamically to changes in usage.
Important Areas That Need To Be Monitored On A
Regular Basis In Data Centre Network

1 2 3

Some of the Real-time Bandwidth


Network
important areas that configuration
availability monitoring
management
need to be monitored
on a regular basis:
Click to edit Session title
Computing Components : Storage
• Data center storage also includes
the policies and procedures that
govern data storage and retrieval
such as data collection and
distribution, access control, storage
security, data availability, storage
quotas, backup schedules, data
retention schedules, and so on.

• In financial, medical and other


highly regulated industries, data
center storage must comply with
government and industry
regulations for data storage,
information privacy and data
security.
Types of Data Center Storage

Types of Data
Center
Storage

Direct Network Storage Area


Attached Attached Network
Storage (DAS) Storage (NAS) (SAN)
Direct Attached Storage (DAS)

• (DAS) typically refers to hard disk drives (HDDs) or solid-state drives


(SSDs) and is the most common type of data center storage.

• Just as its name shows, DAS is attached directly to a host server, instead
of connecting through a network, like Ethernet.
Direct Attached Storage(Cont..)
• Advantages of DAS:
• Cost-saving: DAS is much cheaper than other storage technologies, such
as NAS and SAN. And the price per GB for these types of storage devices
is very low, which continues to trend downward.
• Better performance: Compared with other networked storage solutions,
DAS cannot be affected by network bottlenecks, such as network
congestion.
• Disadvantages of DAS:
• Limited scalability
• Not shareable enough: Since data on DAS cannot be connected through
the internet, data sharing can be a big problem. If
Network Attached Storage (NAS)
• Network Attached Storage (NAS) is a file-level
data center storage device that supports multiple
users to retrieve data from centralized disk
capacity over a TCP/IP network.
• It usually has its node on the local area network
(LAN), without the intervention of the application
server, allowing users to access data on the
network.
Storage Area Network (SAN)
• Storage Area Network (SAN) is a dedicated and high-speed network
established for storage that is independent of the TCP/IP network. It
connects servers to their logical disk units (LUNs) and provides block-level
network access to data center storage.
Storage Area Network (SAN)
How to Choose Suitable Data Center Storage?
• Choose a suitable based on the scalability, performance, IT staff, and
usage case.
• Scalability:
• Performance
• IT staff
• Usage case:
Click to edit Session title
CASE STUDY: Data Center Storage – Impact of Cloud & IoT

• With hundreds of massive data centers


spread across the planet, cloud vendors like
Amazon Web Services, Microsoft Azure,
Google Cloud, IBM Cloud, and Alibaba have
the scale and global reach to satisfy the
capacity and data locality demands of
enterprises all over the world.
• Hence the Cloud, IoT, and Data Center
storage trends begin.
The Rise of the Monolithic Storage Array
• The inefficiency at scale actually had two components.
The first is that servers very commonly only used a
fraction of the computing power they had available.
• It would have been totally normal at this time to see a
server that regularly ran at 10% CPU utilization, thus
wasting massive amounts of resources.
• (The solution to this problem will be discussed in the
next section.) The second problem was that data
storage utilization had the same utilization issue.
• With the many, many islands of storage created by
placing direct attached storage with every server,
there came a great inefficiency caused by the need to
allow room for growth.
• As an example, imagine that an enterprise had 800
servers in their data [Link] each of those servers
had 60 GB of unused storage capacity to allow for
growth. That would mean there was 48 TB of unused
capacity across the organization.
The Rise of the Monolithic Storage Array
• Using the lens of today’s data center to look at
this problem, paying for 48 TB of capacity to
just sit on the shelf seems absurd, but until
this problem could be solved, that was the
accepted design (Figure 2-2).
• This problem was relatively easily solved,
however.
• Rather than provision direct-attached storage
for each server, disks were pooled and made
accessible via the network.
• This allowed many devices to draw from one capacity pool and increase utilization
across the enterprise dramatically.
• It also decreased the management overhead of storage systems, because it meant
that rather than managing 800 storage silos, perhaps there were only 5 or 10.
The Rise of the Monolithic Storage Array
• These arrays of disks (“storage arrays”)
were connected on a network segregated
from the local area network.
• This network is referred to as a storage
area network, or SAN, as shown in
Figure 2-3.
• The network made use of a different
network protocol more suited for storage
networking called Fibre Channel
Protocol.
• It was more suited for delivering storage because of its “lossless” and high-
speed nature. The purpose of the SAN is to direct and store data, and
therefore the loss of transmissions is unacceptable. This is why the use of
something like TCP/IP networking was not used for the first SANs.
Block vs. File Storage

• Data stored on a shared storage device is typically accessed in one of two


ways: at the block level or at the file level.

• File level access means just what it sounds like, “the granularity of access
is a full file.”
Block vs. File Storage (1)
• Data Services
• Most storage platforms come with a variety of different data services that allow the
administrator to manipulate and protect the stored data.
• These are a few of the most common.
• Snapshots
• A storage snapshot is a storage feature that allows an administrator to capture the
state and contents of a volume or object at a certain point in time.
• A snapshot can be used later to revert to the previous state.
• Snapshots are also sometimes copied off site to help with recovery from site-level
disasters.
• Replication
• Replication is a storage feature that allows an administrator to copy a duplicate of a
data set to another system.
• Replication is most commonly a data protection method; copies of data a replicated
off site and available for restore in the event of a disaster.
• Replication can also have other uses, however, like replicating production data to a
testing environment.
Block vs. File Storage (2)
Data Reduction
• Especially in enterprise environments, there is generally a large amount of
duplicate data.
• Virtualization compounds this issue by allowing administrators to very
simply deploy tens to thousands of identical operating systems.
• Many storage platforms are capable of compression and deduplication,
which both involve removing duplicate bits of data.
• The difference between the two is scope.
• Compression happens to a single file or object, whereas deduplication
happens across an entire data set.
• By removing duplicate data, often only a fraction of the initial data must
be stored.
Virtualization of Computers : Software Defined Servers
The Virtualization of Computers: Software Defined Servers

• The impact of virtualization changed


networking as well, all the way down to the
physical connections.
• Where there may have once been two data
cables for every server in the environment, in
a post-virtualization data center there are
perhaps two data cables per hypervisor with
many virtual machines utilizing those
physical links.
• This creates a favorable level
oversubscription of the network links as
compared to the waste from the legacy
model.
• See Figure 2-4 for an example of this
consolidation.
Consolidation Ratios
• The consolidation ratio is a way of referring to the effectiveness of some
sort of reduction technique.
• One example of a consolidation ratio would be the amount physical
servers that can be consolidated to virtual machines on one physical host.
• Another example is the amount of copies of duplicate data that can be
represented by only one copy of the data.
• In both cases, the ratio will be expressed as [consolidated amount]:1.
• For example, 4:1 vCPUs to physical cores would indicate in a P2V project
that for every 1 physical CPU core available, 4 vCPUs are allocated and
performing to expectations.
Click to edit Session title
Why Microsoft Has Underwater Data Centers
Aspect Details
Project Name Microsoft Project Natick
Core Idea Placing data centers underwater instead of on land
Primary Improve energy efficiency, sustainability, and
Objective reliability of data centers
• Natural seawater cooling reduces energy used for air
Why
conditioning
Underwater?
• Stable temperature environment year-round
Near coastal regions where most of the world’s
Location Choice
population lives
Deployment
About 35 meters (117 feet) below sea level
Depth
Test Location Orkney Islands, Scotland
Infrastructure Sealed steel container connected to land via power and
Setup fiber-optic cables
Why Microsoft Has Underwater Data Centers
Aspect Details
• Servers showed much lower failure rates than land-
Key Findings based data centers• Minimal human interference
improved reliability
Cooling Mechanism Passive cooling using surrounding seawater
Energy Source Powered by renewable energy (wind and solar)
• Reduced cooling costs• Faster deployment using
Major Benefits prefabricated units• Lower carbon footprint• Reduced
latency for coastal users

Operational Model “Lights-out” operation – no on-site human maintenance

• Difficult physical maintenance• Corrosion and


Challenges
pressure resistance• Environmental impact assessment
Environmental
Found to have minimal disturbance to marine life
Impact
Active experimentation completed; learnings applied to
Project Status
future data-center designs
Why Microsoft Has Underwater Data Centers
Class Summary
• Data Centre Evolution
• Data Center Components (Computing)
• Requirements for a modern data center networking platform
• Case Study - Google data centers’ Netwtorking
• Data Center Storage
Click(DAS)
• Direct Attached Storage to edit Session title
• Network Attached Storage (NAS)
• Storage Area Network (SAN)
• Case Study : AWS Data Center for Storage of AWS IoT and S3
• The Rise of the Monolithic Storage Array
• Block vs. File Storage
• The Virtualization of Compute — Software Defined Servers
• Case Study : Microsoft Underwater Data Center
Interesting Videos
• Why Microsoft Has Underwater Data Centers

• Networking across Google’s data centers?

* Amazon Operates 900 Data Centers as It Tries to Meet AI Demand


Next Session:
Data Center Design - Components

STOP RECORDING
CSIWZG522 :
Design and Operations of Data Center

CS 02
BITS Pilani
Pilani Campus

Click to edit Session title

CSIWZG522 - Data Center Design- Introduction contd..:

CS 02
Click to edit Session title
Important Note to Students
⮚It is important to know that just login to the session does not guarantee the
attendance.

⮚Once you join the session, continue till the end to consider you as present in the
class.

⮚IMPORTANTLY, you need to make the class more interactive by responding to


Professors queries in the session.

⮚Answering to Polls & Quiz questions during / after the class are mandatory

⮚Whenever Professor calls your number / name ,you need to respond,


otherwise it will be considered as ABSENT and you will be removed from
the meeting

⮚Students are advised join within 10 minutes once the class starts,
Professor will lock the entry to the meeting.
Declaimer

• The slides presented here are obtained from the authors of the books and
from various other contributors. I hereby acknowledge all the contributors
for their material and inputs.

• I have added and modified a few slides to suit the requirements of the course.
Text Book & Reference

No Author(s), Title, Edition, Publishing House


T1 Building a Modern Data Center: Principles and Strategies
of Design by Scott D. Lowe (Author), David M. Davis
(Author), James Green (Author), Seth Knox (Editor), Stuart
Miniman (Foreword)
Reference Book(s) & other resources

R1 Data Center for Beginners: A beginner's guide towards


understanding Data Center Design , B.A. Ayomaya
(Author)
Advisory

• The printed book is available in the market.

• The e-book is available on the Internet.

• Students are strongly advised to make notes from RLs & CSs for examination
purpose right from the beginning.

• Unavailability of the textbook will not be an excuse to allow print copies of the e-
book or PPTs during examinations.
Lecture Plan

NO TOPIC DESCRIPTION CONTENT REF

Principle: Scalability (Horizontal


1 T1: Ch-2
vs. Vertical)

Principle: Reliability (Redundancy


2 T1: Ch-2
& Uptime)

Principle: Elasticity (On-demand


3 T1: Ch-2
Resources)

Case Study: NVIDIA (AI


4 Industry Ref
Infrastructure)

5 Summary & Quiz Class Activity

RL: RECORDED LECTURE


Principle 1 - Scalability (Definition)
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Ability to handle Expanding a
Concept growing work by warehouse to Grow on demand
adding resources. store more goods.
Maintain
Adding checkout
performance levels
Goal lanes during Handle more load
during traffic
holiday sales.
spikes.
Prevents system Electricity grid
Benefit crashes during handling summer Stay always up
high usage. peak loads.
Measured by
Highway lanes
requests Speed remains
Metric added for rush
processed per fast
hour traffic.
second.
Long-term growth
without Building extra
Outcome Future Proof build
redesigning the floors on an office.
system.
Vertical Scalability (Scaling Up)
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Adding power Upgrading a car
Stronger single
Action (CPU/RAM) to an with a faster
unit
existing server. engine.
Restricted by the
hardware's A suitcase can Physical ceiling
Limit
maximum only hold so much. reached
capacity.
Simple to manage
Replacing a small
Complexity as it's one Easy to manage
bulb with brighter.
machine.
Usually requires a Changing tires
Downtime restart to upgrade while car is Pause for upgrade
parts. parked.
High cost for
Buying a luxury Pricey power
Cost specialized high-
sports car engine. boost
end hardware.
Horizontal Scalability (Scaling Out)
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Adding more Hiring more
Action servers to the workers to finish Power in numbers
existing pool. tasks.
Virtually limitless
Adding more
Limit growth by adding No upper limit
buses to a fleet.
more nodes.
If one fails, others
One bulb breaks, Fault tolerant
Redundancy keep the system
others stay lit. design
running.
Uses cheaper,
Buying multiple
commodity Cheaper mass
Cost standard delivery
hardware in large growth
bikes.
groups.
Best for cloud and Building many
Cloud ready
Benefit hyperscale data small houses in
growth
centers. colony.
Principle 2 - Reliability (Definition)
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Probability that a
A clock that never
Concept system functions Trust the system
stops ticking.
without failure.
Consistent A car starting
Focus performance under every cold Works every time
stated conditions. morning.
Mean Time
Lightbulb rated for
Measure Between Failures Long life span
10,000 hours.
(MTBF).
Provides Airplanes having
Assurance confidence to the rigorous safety Build user trust
end users. checks.
Critical for
mission-essential Hospital power Mission critical
Impact
business never going out. uptime
operations.
Reliability - Redundancy (N+1)
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Having extra
Carrying a spare Extra backup
Concept components as
tire in car. ready
backups (N+1).
Dual power feeds
Flashlight with
Power to prevent Dual power source
extra battery pack.
blackouts.

Multiple paths for Multiple roads to


Network Never lose path
data to travel. reach the airport.

RAID or mirroring Making a


Copy prevents
Data to prevent data photocopy of a
loss
loss. contract.
System stays
Twin-engine plane Keep moving
Outcome running even if
flying on one. forward
parts fail.
Reliability - Tier Classifications
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Basic capacity
with no A small home
Tier 1 Basic simple setup
redundancy office setup.
(99.67%).
Redundant
capacity Office with a Better partial
Tier 2
components backup generator. backup
(99.74% uptime).
Concurrently
maintainable; Hospital with dual
Tier 3 High uptime goal
multiple paths power grids.
(99.98%).
Fault tolerant; no
National security Zero downtime
Tier 4 single point
data center. allowed
failure.
Established by
Star ratings for Gold standard
Standard Uptime Institute
hotel quality. uptime
for benchmarking.
Principle 3 - Elasticity (Definition)
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Automatically A rubber band
Concept scaling resources stretching/shrinkin Flex with need
up and down. g.
Happens in real-
Turning on more Instant resource
Speed time or near real-
taps for water. change
time.
Handled by
Motion sensor
software without
Automation lights turning Smart auto scaling
human
on/off.
intervention.
Minimizes waste Paying for
Efficiency by releasing electricity only Pay for usage
unused resources. when used.
Key differentiator
Renting more Match demand
Value for Cloud Data
seats for event. exactly
Centers.
Scalability vs. Elasticity
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Long-term growth Short-term
Nature and capacity fluctuations and Long vs Short
planning. spikes.
Manual or Automatic and
Control scheduled dynamic resource Manual vs Auto
resource addition. adjustment.
Handling
Handling a 1-hour
Purpose permanent Growth vs Spike
"Flash Sale."
increase in users.
Resources are
Resources are
Release returned once Keep vs Return
kept for future use.
done.
Building a bigger
Optimizing current
Focus infrastructure Build vs Optimize
operational costs.
foundation.
Case Study Introduction – NVIDIA with AI
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G

A global leader in The "brain" maker


Entity AI Chip Giant
AI computing. for robots.

Huge data needs Teaching a child


Massive data
Problem for training AI million words
hunger
models. daily.
Building
specialized "AI Building a super-
Solution Build AI Factory
Factories" (Data fast car factory.
Centers).
Enables ChatGPT
Making computers Powering modern
Result and autonomous
talk like humans. AI
driving.
Moving from
A toy maker
Scope gaming to data Core AI Provider
building rockets.
center hub.
NVIDIA - Scalability in Action with AI
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
GPU Clusters
(Connecting Linking 1000
Component Group for power
thousands of laptops for power.
GPUs).
InfiniBand high-
A 10-lane super-
Network speed Super fast links
fast highway.
interconnect.
Massive data
A library with
Storage storage for AI Huge data vault
billion books.
training.
Modular pods that
Adding Lego Easy modular
Method can be added
blocks to build. growth
easily.
Scales from one
Small shop to
Growth server to data Tiny to Huge
global mall.
center.
NVIDIA - Reliability Features with AI
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
BlueField DPU
(Data Processing Security guard at
Security Secure chip level
Unit) for isolated every door.
security.
AI-driven health
Smart car warning Self healing
Monitoring checks for
of low oil. systems
hardware.
Liquid cooling to
Radiator keeping
Redundancy prevent Cooling stays on
engine cool.
overheating.
Specialized
software to Flight control for
Stability Steady AI work
manage GPU smooth flying.
health.
Omniverse
Nucleus (Shared
3D Collaboration Auto-saving your
Backup Data always safe
Database) game progress.
Simplification in NVIDIA Design
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Reducing
One remote for all
Goal management Make IT easy
devices.
complexity.
NVIDIA AI
An "Easy" button One software
Tool Enterprise (unified
for tech. platform
software).
Self-configuring
Self-parking car
Automation networking for Auto setup ready
technology.
clusters.

Fewer staff needed Self-checkout in a


Outcome Less staff effort
to manage more. supermarket.

Faster time-to-
Express lane at the
Efficiency market for AI Deploy AI fast
airport.
products.
NVIDIA - Focus on Cost Reduction
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Consolidating
One bus replacing
Hardware many CPUs into Do more, smaller
50 cars.
few GPUs.
Higher Using LED bulbs
Energy performance per instead Save power bills
watt of power. incandescent.
Dense computing
Fitting bunk beds
Space requires less floor Save floor space
in room.
space.
Automatic lawn
Lowering "Staff
mower (System
Management Overhead" via Lower human cost
Management)
automation.
saves time.

Shift to Opex via Leasing a car


Capex No big upfront
Cloud partners. monthly.
Summary - Principle Comparison
SIMPLE
REAL-LIFE
PRINCIPLE FOCUS AREA UNDERSTANDIN
EXAMPLE
G

Handles 1M new
Scalability Capacity Size Big growth ready
users.

99.999% "Five Never stops


Reliability System Uptime
Nines" availability. working

Saves money at
Elasticity Cost & Flex Flex with traffic
night.

Fast training of
Case Study NVIDIA AI AI power leader
GPT.

Modern Data Modern design


Integration All Three
Center Success. core
Critical Success Factors
SIMPLE
REAL-LIFE
FACTOR IMPORTANCE UNDERSTANDIN
EXAMPLE
G
Detect issues
Health checkup Watch system
Monitoring before they cause
once a year. health
downtime.
Zero-trust
ID check at every
Security architecture for Trust no one
door.
data safety.

Adapting to AI and Upgrading old


Innovation Stay tech current
Deep Learning. phone to new.

Staff must learn Learning to drive


Skillset Learn new tech
new AI tools. electric.

Elegant design
Minimalist home
Simplicity reduces human Simple is better
with no clutter.
error.
Future Trends in Design
SIMPLE
REAL-LIFE
TREND DESCRIPTION UNDERSTANDIN
EXAMPLE
G

Computing closer Local grocery


Edge DC Compute near user
to the user. store delivery.

Focus on 100% Solar panels on a


Green DC Save the planet
renewable energy. house.

Using water/fluids Water-cooled


Liquid Cool Water cools best
to cool chips. racing car engine.

AI managing the Robot butler


AI Ops AI runs DC
data center itself. cleaning house.

Next-gen
Teleporting data
Quantum computing for Future fast math
across world.
complex math.
Adapt or Die & IT Flattening
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Organizations
Blockbuster failing
must evolve or
Adaptation to adopt Evolve or fail
face market
streaming.
extinction.
Removing silos
A small startup
between server,
Flattening where everyone Merge tech teams
storage, and
helps.
network teams.
Speed of delivery
becomes the Fast-food prep
Agility Speed is king
primary versus fine dining.
competitive edge.
Generalists (Full-
Handyman fixing
stack) replacing Broaden your
Skill Shift plumbing and
narrow niche skills
electric.
specialists.
Reducing
management
Direct chat with
IT as an Operational Expense (OpEx)
SIMPLE
REAL-LIFE
FEATURE DEFINITION UNDERSTANDIN
EXAMPLE
G
Moving from
buying hardware Leasing a car
CapEx Shift Rent don't buy
to renting instead buying.
services.
Paying only for the
Paying the
Consumption resources actually Pay per use
monthly water bill.
used monthly.
Keeps capital free Using cash for
Cash Flow for core business marketing, not Keep cash free
investments. desks.
Expanding in
small, predictable Buying food for
Scalability Grow in bits
chunks based on one week.
need.
Ability to cancel or Ending a Netflix
Flexibility reduce services subscription Cancel any time
instantly. anytime.
Next Session:
Data Center Design- Components

STOP RECORDING
CSIWZG522 :
Design and Operations of Data Center

Dr. Rama Satish K V


satishkvr@[Link]

CS 01
BITS Pilani
Pilani Campus

Click to edit Session title

CSIWZG522 - Design and operations of Data center

CS 01
Click to edit Session title
Important Note to Students
It is important to know that just login to the session does not guarantee the
attendance.

Once you join the session, continue till the end to consider you as present in the
class.

IMPORTANTLY, you need to make the class more interactive by responding to


Professors queries in the session.

Answering to Polls & Quiz questions during / after the class are mandatory

Whenever Professor calls your number / name ,you need to respond,


otherwise it will be considered as ABSENT and you will be removed from
the meeting

Students are advised join within 10 minutes once the class starts,
Professor will lock the entry to the meeting.
Declaimer

• The slides presented here are obtained from the authors of the books and
from various other contributors. I hereby acknowledge all the contributors
for their material and inputs.

• I have added and modified a few slides to suit the requirements of the course.
Text Book & Reference

No Author(s), Title, Edition, Publishing House


T1 Building a Modern Data Center: Principles and Strategies of
Design by Scott D. Lowe (Author), David M. Davis (Author), James
Green (Author), Seth Knox (Editor), Stuart Miniman (Foreword)

Reference Book(s) & other resources

R1 Data Center for Beginners: A beginner's guide towards


understanding Data Center Design , B.A. Ayomaya (Author)
Advisory

• The printed book is available in the market.

• The e-book is available on the Internet.

• Students are strongly advised to make notes from RLs & CSs for examination
purpose right from the beginning.

• Unavailability of the textbook will not be an excuse to allow print copies of the e-
book or PPTs during examinations.
Lecture Plan

Time Type Description Content


Reference

Pre CH R.L -1.1, What is Data Centre


1.2. Data Center and its values

During CH1 T1: Ch-1


CH IT Is Changing . . . and It Must
Simplification Is the New Black
Focus on Cost Reduction
Focus on Customer Service and the Business
Overview of RL* : What is Data Center?
• A Data Center consists of networked
computers, storage systems, and computing
infrastructure (server farm).​

• It is used by organizations to assemble,


process, store, and disseminate large
amounts of data (enterprise hub).​

• Businesses rely heavily on applications,


services, and data within the Data Center
(business backbone).​

Data Centers are critical assets for everyday operations (Mission Infra).

RL: RECORDED LECTURE


Core components of Data Centers

01 02 03 04

Facility Data Support Operational


Storage infrastructure staff
Key Functions of a Data Center (1)
• Hardware (Compute Stack)
• Mainframes, servers (computers), storage drives, routers, switches, and firewalls.
• Infrastructure (Facility Backbone)
- Critical support systems for power (redundancy/backup),
- Cooling (to prevent overheating),
- Fire suppression, and
- Physical security.
• Purpose (Data Hub): Collects, processes, stores, and distributes vast amounts of
digital data for organizations.
• Connectivity (Network Hub): Provides high-speed network connections for
reliable data flow.
Key Functions of a Data Center (2)
• Data Storage (Storage Vault)
- Securely keeping massive amounts of data (files, databases, media) for
organizations and users.

• Data Processing (Compute Engine)


- Running applications, performing calculations, and enabling complex tasks like
AI model training and big data analytics.

• Application Hosting (App Platform)


- Housing the servers that run web applications, e-commerce sites, email,
productivity tools, and more.
Key Functions of a Data Center (3)
• Network Connectivity (NVIDIA InfiniBand fabric - High‑speed GPU cluster
interconnect)
- Providing high-speed, redundant network connections for reliable data transfer.

• Security & Access (BlueField DPU (NVIDIA’s Brand Name of Data Processing Unit
called BlueField) security - Isolated zero‑trust infrastructure offload)
- Protecting data and systems with physical security, firewalls, and access controls, while
allowing authorized users access.

• Backup & Recovery (Omniverse Nucleus (Shared3D Collaboration Database)


backups - Automated 3D collaboration data‑protection)
- Implementing strategies for data redundancy, backups, and disaster recovery to prevent
data loss and ensure business continuity.
Types of Data Centers

Enterprise data centres Colocation data centres

Services data centres Edge data centres

Cloud data centres Hyperscale data centres


Introduction to Modern Data Center
• The modern Data Center looks nothing like the
Data Center of only 10 years past (Cloud Native).
• The processing power still doubles every 2.5 to 3
years (Moore Trend – Power doubles but cost
falls).
• In the Modern Data Center, tasks that were
accepted as unattainable only a few years ago are
commonplace today (AI Workloads).
• Models that would have likely been cost-prohibitive
and really slow just a few years back are not the
same (Deep Learning).

• SDS [Software Defined Storage] offers IT organizations some of the greatest storage
flexibility and ease of use (Virtual Storage).
IT Is Changing . . . and it Must
• With exponential growth, there are more opportunities for entrepreneurs, for organizations
to multiply their revenues, for top-tier graduate students to create projects that once
seemed like pure science fiction (Digital Boom).
• While bleeding edge technology trickles down into the mainstream IT organization, the
changing nature of the IT business itself rivals the pace of technological change
(Innovation Diffusion).
• Over the next couple of years, IT organizations and business leaders will have the
opportunity to focus on dramatic transformations in the way they conduct themselves.
Technological change will create completely new ways of doing business (Operating
Overhaul).
• The exponential growth in technology in the next decade will generate the rise of entirely
new industries and cause everything in our day-to-day experience to be different, from
the way we shop to the way we drive our cars (Emerging Sectors).
Simplification Is the New Black
• Thanks to the ever-growing number of devices interconnected on private and public networks, and
thanks to the Internet, the scope of IT’s responsibility continues to expand (Device Explosion).
• The computer is responsible for operating the press at maximum efficiency, and sometimes it does
so well that the print shop doubles its profits, and the computer becomes critical to the business
(Automation Gains).
• The boom was dubbed the “Internet of Things” (IoT) way back in 1999! (IOT Boom)
• Today, we’re still right at the beginning of this paradigm shift (Early Stage).
• The estimations of IoT proliferation in the next few years are staggering (Massive Scale).
• A 2014 Gartner report shows an estimated 25 trillion connected devices by 2020 (Device
Tsunami).
• Because this trend means that IT departments become responsible for more systems and
endpoints, one of the primary goals for these departments in the near future is to simplify
administration of their systems (Management Simplification).
Simplification Is the New Black
• Despite the fact that the number of
integrations and the amount of data IT has
to manage because of these devices is
increasing, in many industries budgets are
not (Budget Squeeze).
• This means that, in order to succeed, CIOs
will need to find ways to boost efficiency in
major ways (Efficiency Drive).
• Five years from now, the same team of administrators may be managing 10 times the
number of resources that they’re managing today, and expectations for stability,
performance and availability will continue to increase (Uptime Demands).
• There are two main venues where the push for simplicity must come from: the
manufacturers and the IT executives (Vendor Simplification).
Simplification Is the New Black
• First of all, manufacturers must begin to design and redesign their products with
administrative simplicity in mind (Simpler Consoles).
• To be clear, this doesn’t mean the product must be simple. In fact, the product will almost
assuredly be even more complex. However, the products must have the intelligence to
self-configure, solve problems, and make decisions so that the administrator doesn’t have
to (Autonomous Systems).
• Secondly, IT Directors and CIOs must survey their entire organization and ruthlessly
eliminate complexity (Complex Cleanup).
• In order to scale to meet the needs of the future, all productivity-sapping complexity must
be replaced with elegant simplicity such that the IT staff can spend time on valuable
work, rather than on putting out fires and troubleshooting mysterious failures that take
hours to resolve (Friction Reduction).
• One of the primary ways this might be done is to eliminate the organizational siloes that
prevent communication and collaboration (Silo busting).
Focus on Cost Reduction
• As shown in Figure 1-2, a recent
Gartner report shows that IT
budgets are predicted to remain
nearly flat as far out as 2020
(Budget Plateau).
• This is despite 56% of respondents
reporting that overall revenues are
expected to increase (Revenue
Growth).
• The reality is that many of those IT
organizations will be required to do
more with less, or at the very least,
do more without any increase in
budget (Resource Squeeze; Cost
Focus on Cost Reduction
• As this trend isn’t likely to change, the IT department of the future will be focused on
reducing expenses where possible and increasing control over the absolutely necessary
expenses (Cost Discipline).
• Thanks again to Moore’s Law, it will be possible to complete projects in 2016 for a fraction
of the cost that the same project would have cost in 2013 (Hardware Deflation).
• One of the most common examples of this is the cost of storage for a server virtualization
project (Virtual Storage).
• Due to the challenges of performance and workload characteristics, a properly built server
virtualization storage platform has historically been expensive and quite complex
(Enterprise Arrays).
• Thanks to denser processors though, cheaper and larger RAM configurations, the falling
cost of flash storage (solid state drives [SSDs]), and innovation toward solving the storage
problem, a server virtualization project can be completed successfully today for a much
lower cost, relatively speaking, and with much more simplicity than ever before
(Commodity Hardware; Simpler Stacks).
Focus on Cost Reduction
• As mentioned in the previous section, the IT department of
the future will be looking to manufacturers and consultants
to help them build systems that are so simple to manage
that less IT staff is required. This is because hiring
additional IT staff to manage complex solutions is costly
(Simpler Platforms; Staff Overhead).
• Of course, this doesn’t necessarily mean that IT jobs are at
risk; it means that since there’s no budget to grow the IT
staff, the current workforce must accomplish more duties
as the responsibility of IT grows. This additional
responsibility means that IT jobs are only safe assuming
that each IT practitioner is growing and adapting with the
technology (Role Expansion; Continuous Upskilling).
• Those who do not learn to handle the additional
responsibility will be of little use to the organization moving
forward (Skill Stagnation).
Focus on Customer Service and the Business
• IT departments face pressure to do more with less and seek top-
notch vendor support (Vendor Reliance).
• IT professionals lack time to troubleshoot solutions extensively
(Time Crunch).
• White-glove service from technical resources is becoming the
norm (Premium Support).
• Leading manufacturers use connectivity and big data to proactively
address customer issues (Proactive Analytics).
• The ideal support call is one that never needs to happen (Zero
Tickets).
• CIOs and IT Directors plan to shift infrastructure maintenance to
partners and vendors (Managed Services).
• IT staff will focus more on innovation and adding business value
(Value Creation). (Value = Utility + Warranty)
Focus on Customer Service and the Business
• IT departments must focus on providing value to their internal customers as well
(Employee enablement).
• As new technology enables business agility on the manufacturer side, IT will have to
continue providing the services users need and want, or users will assuredly find them
elsewhere (Service Competitiveness)
• This means that shadow IT poses significant security, compliance, and control risk to the
entire business, and the only way to really stop it is to serve the internal customers so
well that they don’t need to look elsewhere for their technical needs (Better
Alternatives).
• Shadow IT is a term used to describe business units (or individuals) other than IT who
deploy technical solutions — not sanctioned or controlled by IT — to solve their business
problems (Unsanctioned Tools).
• A simple example of this phenomenon is a few individuals in the finance department
using personal Dropbox folders to share files while the business’ chosen direction is to
share company files in SharePoint (Rogue Deployments; Unauthorized Sharing).
Adapt or Die
• The world of IT operates in the same way as the rest of the world: things change over time, and
those companies, technologies, or individuals who choose not to keep up with the current state of
the industry get left behind. It’s unfortunate, but it is reality (Constant Change).
• The movie rental giant Blockbuster was dominating the market until Netflix and Redbox innovated
and Blockbuster failed to adapt. Blockbuster eventually went bankrupt and is now an afterthought
in the movie consumption industry while Netflix’s fortunes are at an all-time high (Disrupted
Incumbent; Streaming Success).
• This lifecycle is exactly the same in IT; there are household names in the industry that are more or
less irrelevant (or quickly becoming so) at this point because of their failure to adapt to the
changing market (market obsolescence).
• Unfortunate as it may be, this is also happening at the individual level (Career Risk).
• As IT administrators and architects adapt or do not adapt over the course of time, they either
advance in their organization and career, or they become irrelevant (Skill Evolution).
IT as an Operational Expense
• Especially in the enterprise environment, getting budgetary approval for operational
expenses can prove to be easier than getting approval for large capital expenditures
(Operational Expenses Preference).
• As such, the operating model of many IT departments is shifting away from capital
expenses (Capex) when possible and toward a primarily operational expense-funded
(Opex-funded) model for completing projects (Budget Shift).
• A large component of recent success in these areas is due to a shift of on-premises,
corporately managed resources to public cloud infrastructure and Software-as-a-Service
(SaaS) platforms (Cloud Migration).
• Since cloud resources can be billed just like a monthly phone bill, shifting IT resources to
the cloud also shifts the way the budget for those resources is allocated (Subscription
Billing).
• While buying a pile of servers and network equipment to complete projects over the next
year is budgeted as a capital expenditure, the organization’s “cloud bill” will be slotted as
an operational expenditure (Hardware Capex; Service Opex).
IT as an Operational Expense
• This is because, with a capital purchase, once the equipment is purchased, it is owned
and depreciating in value from the moment it hits the loading dock. This removes an
element of control from IT executives as compared to an operational expense (Asset
lock-in).
• If a SaaS application is billed per user on a monthly basis, there’s no need to pay for
licenses now to accommodate growth in headcount six months down the road (License
Elasticity).
• It can also be in use this month and cancelled next month (Easy Off-boarding).
• This is in contrast to the IT director who can’t just “cancel” the stack of servers purchased
six months ago because the project got cancelled (Stranded Hardware) (Unusable
sunk Infrastructure)
• Due to these advantages from a budgeting and control standpoint, products and services
offering a model that will require little or no capital expense and allow budgeting as an
operational expense will be preferred (Opex Advantage).
IT as an Operational Expense
• What this means for manufacturers is that transparency, granular control, and offering a
OpEx-based model like renting or leasing, billing based on monthly usage, and
expanding in small, predictable chunks based on need, will position them for adoption by
the IT department of the future (Consumption Pricing).
• Also, offering insight and helping the customer to increase efficiency and accuracy over
the course of the billing cycle will create lasting loyalty (Value Analytics).
• This shifting focus from CapEx to OpEx is also giving rise to new kinds of Data Center
architectures that allow organizations to keep Data Centers on premises and private, but
that enable some economic aspects similar to cloud (Hybrid Models).
• Rather than having to overbuy storage, for example, companies can begin to adopt
software defined storage (SDS) or hyperconverged infrastructure (HCI) solutions that
enable pay-as-you grow adoption methodologies (Scalable Storage).
Class Summary
• What is Data Center?
• Core components of Data Centers
• Types of Data Centers
• Modern Data Center
• IT Challenges
• Simplification Is the New Black
• Focus on Cost Reduction
• IT as an Operational Expense
Next Session:
Distributed programming

You might also like