5/13/2026
MADDA WALABU UNIVERSITY
COMPUTING COLLEGE
NAME : Ararso Jaleto ID : Ugrr/50198/15
Dep’t : Information Technology
Course : Distributed System
INSTR: Muluken .W
Submission Date : 6/14/2026
BALE ROBE , OROMIA , ETHIOPIA
[Link] Effective Resource Sharing ?
The primary goal of a distributed system is to make it easy for users and
applications to access remote resources and share them in a controlled and
efficient way .
❖ Economic Efficiency: Sharing is often driven by economics.
It is more cost-effective to allow multiple users in an office to share a single high-
end printer or expensive storage system than to purchase individual units for
everyone.
❖ Collaboration: Connecting users and resources facilitates the
exchange of information and collaborative work. This is exemplified by the
Internet's success in allowing geographically dispersed groups to work
together using "groupware" for editing, teleconferencing, and electronic
commerce.
❖ Concurrency Management: Effective sharing requires
concurrency transparency, ensuring that multiple independent users
accessing a shared resource (like a database) do not notice each other's
presence and that the resource remains in a consistent state.
6. Define Global States ?
In a distributed system, a consistent global state is also referred
to as a distributed snapshot.
➢ Definition: A global state is considered consistent if, for every
message recorded as received by a process, there is a corresponding
record of that message being sent by another process
1|P age
➢ Purpose: Recording these states is essential for backward
error recovery, allowing a system to "checkpoint" its status onto
stable storage so it can return to a known correct state following a
failure
8. Explain Distributed Termination Detection
?
While the sources do not provide a general definition for "distributed
termination detection" as a standalone algorithmic topic, they explain
termination protocols in the context of distributed commit (like Two-
Phase Commit).
❖ Mechanism: If a participant in a distributed transaction
times out while waiting for a final decision from the
coordinator, it executes a termination protocol.
❖ Process: The participant multicasts a DECISION_REQUEST to other
participants. If any other participant has already reached a final decision
(Commit or Abort), it informs the requesting process, allowing the protocol
to terminate correctly despite the coordinator's failure.
12. Explain TCP Stream Communication ?
TCP (Transmission Control Protocol) provides a
reliable, connection-oriented communication
service.
2|P age
The "Big Tube" Model: It maintains the illusion
that a connection is like a large tube: messages are put in
at one end
and come out at the other undamaged and in the exact same
order in which they were sent.
Stream Nature: A TCP connection is an example of a
discrete data stream, which is a sequence of data units where
the timing of delivery is generally not fundamental to the
correctness of the data interpretation.
Transmission Mode: It typically operates in an
asynchronous transmission mode, meaning data items are
transmitted one after another without strict timing constraints
on when each transfer must complete.
Session Support: In server clusters, TCP is used
to set up sessions where application-level requests are
sent over a single connection that is torn down only
when the session ends.
15. Explain File Service Architecture in
Detail ?
Distributed file systems are typically organized
into one of three architectural styles:
3|P age
A. Client-Server Architectures (e.g., NFS):
Remote Access Model:
This is the most common organization. A server provides
a standardized view of its local file system, and clients are
offered transparent access through an interface.
Layering:
In UNIX-based systems like NFS, a Virtual File System (VFS) layer hides
the differences between various file systems. The VFS passes requests
to an NFS client, which uses RPCs to communicate with the remote
NFS server.
[Link]-Based Architectures (e.g., Google
File System):
➢ File Striping:
Large files may be "striped" across multiple servers to allow for
parallel access, which is useful for compute-intensive applications.
➢ Master/Chunk Server Model:
The Google File System (GFS) uses a single master to handle metadata
(file names and chunk locations) while data itself is stored in 64MB
chunks across many "chunk servers." Clients communicate directly with
chunk servers for data transfer to avoid bottlenecks at the master.
4|P age
[Link]/Peer-to-Peer Architectures:
➢ DHT-Based Storage:
These systems use Distributed Hash Tables (DHTs) for decentralized
data lookup.
➢ Implementation: Some systems build a file system on top of
a block-oriented storage layer (like Ivy), while others distribute whole files
across nodes based on directory levels.
18. Explain Remote Procedure Call in Detail ?
The Remote Procedure Call (RPC) is a model designed to make
communication between processes on different machines appear as if
it were a local procedure call.
Basic Operation Steps:
1. Client Call: The client application calls a local client stub.
2. Marshaling: The stub packs the parameters into a
message (marshaling) and requests the local OS to send it.
3. Transmission: The message is sent across the network
to the server's OS.
4. Server Stub: The server's OS passes the message to a server
stub, which unpacks the parameters.
5. Execution: The server stub calls the actual local procedure on
the server.
5|P age
6. Return: The result follows the reverse path back to the
client.
Key Components:
i. Transparency: Ideally, neither the client nor the
server is aware that the call is remote or that a network
exists.
ii. IDL (Interface Definition Language): To ensure
both sides agree on the message format and data types,
interfaces are specified in an IDL, which is then used to
automatically generate the stubs.
iii. Data Representation: Marshaling must
handle differences in machine architectures, such as
byte ordering (endianness).
19. Explain Fault Tolerance Mechanism
in Detail ?
Fault tolerance is the ability of a system to mask the occurrence and
recovery from failures, appearing to continue operating normally to
the user. Fault tolerance is the characteristic that allows a distributed
system to continue operating normally even in the presence of
failures by masking their occurrence.
6|P age
Key Mechanisms in Fault Tolerance.
1. Architectural & Structural Mechanisms.
➢ Failure Masking through Redundancy: Replicating
data, time, or physical hardware (like Triple Modular Redundancy) to hide errors
from the user.
➢ Process Resilience and Groups: Organizing
components into groups to survive up to k failures and using consensus
algorithms to maintain agreement.
2. Operational & Monitoring Mechanisms.
➢ Failure Detection: Using active heartbeats or timeouts to identify
when a node or process has crashed.
➢ Reliable Communication: Ensuring message delivery guarantees
(at-least-once, at-most-once, or atomic multicast) between distributed components.
3. Transactional & Consistency Mechanisms.
➢ Distributed Commit Protocols: Using atomic protocols
like 2PC or non-blocking 3PC to ensure all nodes execute or abort an
operation together.
4. State & System Recovery Mechanisms.
➢ Recovery Mechanisms: Restoring system health using
backward recovery (checkpoints), forward recovery (erasure codes),
message logging, or microreboots.
THANK YOU
END
7|P age