OS Module - 5
OS Module - 5
CPU
Module V
5.1 File System
The file system consists of two distinct parts: a collection of files, each storing related data, and a
directory structure, which organizes and provides information about all the files in the system.
File Operations
The operating system provides system calls to create, write, read, reposition, delete,
and truncate files.
[Link].1 Creating a file - Two steps are necessary to create a file
Find space in the file system for the file.
Make an entry for the new file in the directory.
Writing a file - To write a file, the system call consists of both the name of the file and the information
to be written to the file. Given the name of the file, the system searches the directory to find the file's
location. The system must keep a write pointer to the location in the file where the next write is to take
place. The write pointer must be updated whenever a write occurs.
Reading a file - To read from a file, the system call that specifies the name of the file and where the
next block of the file should be put. The directory is searched for the file, and the system needs to keep
a read pointer to the location in the file where the next read is to take place. Once the read has taken
place, the read pointer is updated.
Repositioning within a file - The directory is searched for the file, and the file pointer is repositioned
to a given value. This file operation is also known as a file seek.
Deleting a file – To delete a file, search the directory for the file. Release all file space, so that it can be
reused by other files, and erase the directory entry.
Truncating a file - The user may want to erase the contents of a file but keep its attributes. Rather than
forcing the user to delete the file and then recreate it, this function allows all attributes to remain
unchanged –except for file length. The file size is reset to zero.
Information about currently open files is stored in an open file table. It contains
information‘s like:
o File pointer - records the current position in the file, for the next read or write
access.
o File-open count - How many times has the current file been opened by different
processes, at the same time and not yet closed? When this counter reaches zero
the file can be removed from the table.
o Disk location of the file – The information needed to locate the file on disk is
kept in memory so that the system does not have to read it from disk for each
operation.
o Access rights – The file access permissions are stored on the per-process table so
that the operating system can allow or deny subsequent I/O requests.
Some systems provide support for file locking.
o A shared lock is for reading only.
o An exclusive lock is for writing as well as reading.
o An advisory lock, it is up to the software developers to ensure that locks are acquired
or released.
o A mandatory lock, prevents any other process from accessing the locked file. (A
truly locked door.)
For instance, only a file with a ―. corn", ".exe", or ".bat‖, can be executed.
c) File Structure
The study of different ways of storing files in secondary memory such that they can be
easily accessed.
File types can be used to indicate the internal structure of the file. Certain files must be in
a particular structure that is understood by the operating system.
For example, the operating system requires that an executable file have a specific
structure so that it can determine where in memory to load the file and the location of the
first instruction.
UNIX treats all files as sequences of bytes, with no further consideration of the internal
structure.
Macintosh files have two forks - a resource fork, and a data fork. The resource fork
contains information relating to the UI, such as icons and button images. The data fork
contains the traditional file contents-program code or data.
a) Sequential Access
Here information in the file is processed in order, one record after the other.
This mode of access is a common method; for example, editors and compilers
usually access files in this fashion.
A sequential access file emulates magnetic tape operation, and generally supports a
few operations:
o read next - read a record and advance the file pointer to the next position.
o write next - write a record to the end of file and advance the file pointer to
the next position as shown in Fig. 5.23.
o skip n records - May or may not be supported. ‗n‘ may be limited to
positive numbers, or may be limited to +/- 1.
III Semester OPERATING SYSTEMS(BCS303) 5
RV Institute of Technology & Management®
b) Direct Access
A file is made up of fixed-length logical records that allow programs to read and write
records randomly. The records can be rapidly accessed in any order.
Direct access is of great use for immediate access to large amount of information.
Eg: Database file. When a query occurs, the query is computed and only the selected rows are
access directly to provide the desired information.
Operations supported include:
read n - read record number n. (position the cursor to n and then read the record)
write n - write record number n. (position the cursor to n and then write the record)
jump to record n – move to nth record (n- could be 0 or the end of file)
If the record length is L, there is a request for record ‗N‘. Then the direct access
to the starting byte of record ‗N‘ is at L*(N-1)
Eg: if 3rd record is required and length of each record(L) is 50, then the starting
position of 3rd record is L*(N-1)
Address = 50*(3-1) = 100.
These methods generally involve the construction of an index for the file called index file.
The index file is like an index page of a book, which contains key and address. To find a
record in the file, we first search the index and then use the pointer to access the record
directly and find the desired record.
An indexed access scheme can be easily built on top of a direct access system.
For very large files, the index file itself is very large. The solution index for to this is to create an
index file. i.e. multi-level indexing as shown in Fig. 5.25.
5.3Directory Structure
Directory is a structure which contains filenames and information about the files like
location, size, type etc. The files are put in different directories. Partitioning is useful for limiting
the sizes of individual file systems, putting multiple file-system types on the same device, or
leaving part of the device available for other uses.
Partitions are also known as slices or minidisks as shown in Fig. 5.25. A file system can
be created on each of these parts of the disk. Any entity containing a file system is generally
known as a volume.
a) Single-Level Directory
It is the simplest directory structure.
All files are contained in the same directory, which is easy to support and understand.
b) Two-Level Directory
Each user gets their own directory space - user file directory (UFD)
File names only need to be unique within a given user's directory.
A master file directory (MFD) is used to keep track of each user‘s directory,
and must be maintained when users are added to or removed from the system.
When a user refers to a particular file, only his own UFD is searched.
All the files within each UFD are unique.
To create a file for a user, the operating system searches only that user's UFD to
ascertain whether another file of that name exists.
To delete a file, the operating system confines its search to the local UFD; thus, it
cannot accidentally delete another user's file that has the same name. The user
directories themselves must be created and deleted as necessary.
This structure isolates one user from another. Isolation is an advantage when the
users are completely independent but is a disadvantage when the users want to
cooperate on some task and to access one another's files as shown in Fig. 5.27.
c) Tree-Structured Directories
A tree structure is the most common directory structure.
The tree has a root directory, and every file in the system has a unique path name.
A directory (or subdirectory) contains a set of files or subdirectories.
One bit in each directory entry defines the entry as a file (0) or as a subdirectory (1).
Special system calls are used to create and delete directories.
Path names can be of two types: absolute and relative. An absolute path begins at the
root and follows a down to the specified file, giving the directory names on the path.
A relative path defines a path from the current directory (Fig. 5.28).
For example, in the tree-structured file system of Fig. below if the current directory is
root/spell/mail, then the relative path name is prt/first and the files absolute path name
root/spell/mail/prt/jirst.
Directories are stored the same as any other file in the system, except there is a bit
that identifies them as directories, and they have some special structure that the OS
understands.
One question for consideration is whether or not to allow the removal of directories
that are not empty - Windows requires that directories be emptied first, and UNIX
provides an option for deleting entire sub-trees.
d) Acyclic-Graph Directories
When the same files need to be accessed in more than one place in the directory
structure (e.g. because they are being shared by more than one user), it can be useful
to provide an acyclic-graph structure. (Note the directed arcs from parent to child. )
o UNIX provides two types of links (pointer to another file) for implementing the
acyclic-graph structure.
o A hard link (usually just called a link) involves multiple directory entries that
both refer to the same file. Hard links are only valid for ordinary files in the same
filesystem.
o A symbolic link, that involves a special file, containing information about where
to find the linked file. Symbolic links may be used to link directories and/or files
in other filesystems, as well as ordinary files in the current filesystem as shown in
Fig. 5.29.
Hard links require a reference count, or link count for each file, keeping track of
how many directory entries are currently referring to this file. Whenever one of the
references is removed the link count is reduced, and when it reaches zero, the disk
space can be reclaimed.
For symbolic links there is some question as to what to do with the symbolic
links when the original file is moved or deleted:
o One option is to find all the symbolic links and adjust them also.
o Another is to leave the symbolic links dangling, and discover that they are no
longer valid the next time they are used.
o What if the original file is removed, and replaced with another file having the
same name before the symbolic link is next used?
Another approach to deletion is to preserve the file until all references to it are deleted. To
implement this approach, we must have some mechanism for determining that the last
reference to the file has been deleted.
When a link or a copy of the directory entry is established, a new entry is added to the file-
reference list. When a link or directory entry is deleted, we remove its entry on the list. The
file is deleted when its file-reference list is empty.
If cycles are allowed in the graphs, then several problems can arise:
o Search algorithms can go into infinite loops. One solution is to not follow links in
search algorithms. (Or not to follow symbolic links, and to only allow symbolic links
to refer to directories)
o Sub-trees can become disconnected from the rest of the tree and still not havetheir
reference counts reduced to zero. Periodic garbage collection is required to detect and
resolve this problem. (chkdsk in DOS and fsck in UNIX search for these problems,
among others, even though cycles are not supposed to be allowed in either system.
Disconnected disk blocks that are not marked as free are added back to the file
systems
Any files (or sub-directories) that had been stored in the mount point directory prior to
mounting the new filesystem are now hidden by the mounted filesystem, and are no longer
available. For this reason, some systems only allow mounting onto empty directories (Fig.
5.31 a).
Filesystems can only be mounted by root, unless root has previously conFig.d certain
filesystems to be mountable onto certain pre-determined mount points. (E.g. root may allow
users to mount floppy filesystems to /mnt or something like it) Anyone can run the mount
command to see what filesystems are currently mounted (Fig. 5.31 b).
The traditional Windows OS runs an extended two-tier directory structure, where the first tier
of the structure separates volumes by drive letters, and a tree structure is implemented below
that level.
Macintosh runs a similar system, where each new volume that is found is automatically mounted
and added to the desktop when it is found. More recent Windows systems allow filesystems to be
mounted to any directory in the filesystem, much like UNIX.
Fig. 5.32 shows the effects of mounting the volume residing on /device/dsk over /users. If the volume is
unmounted, the file system is restored to the situation depicted in Fig. 5.31.
Fig. 5.31 File system, (a) Existing system, (b) Unmounted volume. Fig. 5.32 Mount point
The advent of the Internet introduces issues for accessing files stored on remote
computers The original method was ftp, allowing individual files to be transported across
systems as needed. Ftp can be either account and password controlled, or anonymous,
not requiring any user name or password.
Various forms of distributed file systems allow remote file systems to be mounted onto
a local directory structure, and accessed using normal file access commands. (The actual
files are still transported across the network as needed, possibly using ftp as the
underlying transport mechanism.)
The WWW has made it easy once again to access files on remote systems without
mounting their filesystems, generally using (anonymous) ftp as the underlying file
transport
mechanism.
When one computer system remotely mounts a filesystem that is physically located
on another system, the system which physically owns the files acts as a server, and
the system which mounts them is the client.
User IDs and group IDs must be consistent across both systems for the system to
work properly. (I.e. this is most applicable across multiple computers managed by
the same organization, shared by a common group of users. )
The same computer can be both a client and a server. (E.g. cross-linked file systems.
)
There are a number of security concerns involved in this model:
o Servers commonly restrict mount permission to certain trusted systems only.
Spoofing (a computer pretending to be a different computer) is a potential
security risk.
o Servers may restrict remote access to read-only.
o Servers restrict which filesystems may be remotely mounted. Generally, the
information within those subsystems is limited, relatively public, and
protected by frequent backups.
The Domain Name System, DNS, provides for a unique naming system across
all of the Internet.
Domain names are maintained by the Network Information System, NIS, which
unfortunately has several security issues. NIS+ is a more secure version, but has
not yet gained the same widespread acceptance as NIS.
Microsoft's Common Internet File System, CIFS, establishes a network login
for each user on a networked system with shared file access. Older
Windowssystems used domains, and newer systems (XP, 2000), use active
directories. User names must match across the network for this system to be
valid.
A newer approach is the Lightweight Directory-Access Protocol, LDAP, which
provides a secure single sign-on for all users to access all resources on a
network. This is a secure system which is gaining in popularity, and which has
the maintenance advantage of combining authorization information in one central
location.
c) Failure Modes
When a local disk file is unavailable, the result is generally known immediately,
and is generally non-recoverable. The only reasonable response is for the response
to fail.
However, when a remote file is unavailable, there are many possible reasons, and
whether or not it is unrecoverable is not readily apparent. Hence most remote
access
IV SEMESTER OPERATING SYSTEM (18CS43) 26
systems allow for blocking or delayed response, in the hopes that the remote
system (or the network) will come back up eventually.
Consistency Semantics deals with the consistency between the views of shared files on
a networked system. When one user changes the file, when do other users see the
changes? The series of accesses between the open () and close () operations of a file is
called the file session.
a) UNIX Semantics
b) Session Semantics
c) Immutable-Shared-Files Semantics
Under this system, when a file is declared as shared by its creator, then the name
cannot be re-used by any other process and it cannot be modified.
5.6 Overview
Magnetic disks provide a bulk of secondary storage. Disks come in various sizes and speed. Here
the information is stored magnetically. Each disk platter has a flat circular shape like CD. The
two surfaces of a platter are covered with a magnetic material. The surface of a platter is logically
divided into circular tracks, which are subdivided into sectors. Sector is the basic unit ofstorage.
The set of tracks that are at one arm position makes up a cylinder (Fig. 5.1).
The number of cylinders in the disk drive equals the number of tracks in each platter. There
maybe thousands of concentric cylinders in a disk drive, and each track may contain
hundreds
of sectors. The storage capacity of disk drives is measured in gigabytes. The head moves from
the inner track of the disk to the outer track. When the disk drive is operating the disks is rotating
at a constant speed. To read or write the head must be positioned at the desired track and at the
beginning of the desired sector on that track.
5.6.1 Seek Time: -Seek time is the time required to move the disk arm
to the required track.
5.6.2 Rotational Latency (Rotational Delay):-Rotational latency is the
time taken for the disk to rotate so that the required sector comes
under the r/w head.
5.6.3 Positioning time or random-access time is the summation of seek
time and rotationaldelay.
5.6.4 Disk Bandwidth: -Disk bandwidth is the total number of
bytes transferred divided by total
5.6.5 time between the first request for service and the completion of
last transfer.
5.6.6 Transfer rate is the rate at which data flow between the drive and
the computer.
Magnetic tape is a secondary-storage medium. It is a permanent memory and can hold large
quantities of data. The time taken to access data (access time) is large compared with that of
magnetic disk, because here data is accessed sequentially. When the nth data has to be read, the
tape starts moving from first and reaches the nth position and then data is read from nth position.
It is not possible to directly move to the nth position. So tapes are used mainly for backup, for
storage of infrequently used information.
Modern disk drives are addressed as a large one-dimensional array. The one-dimensional array of
logical blocks is mapped onto the sectors of the disk sequentially. Sector 0 is the first sector of
the first track on the outermost cylinder. The mapping proceeds in order through that track, then
through the rest of the tracks in that cylinder, and then through the rest of the cylinders from
outermost to innermost.
i) CLV - The density of bits per track is uniform. The farther a track is from the center of
the disk, the greater its length, so the more sectors it can hold. As we move from outer
zones to inner zones, the number of sectors per track decreases. This architecture is
used in CD-ROM and DVD-ROM.
ii) CAV – There is same number of sectors in each track. The sectors are densely packed in
the inner tracks. The density of bits decreases from inner tracks to outer tracks to keep
the data rate constant.
i) Host-Attached Storage
Host-attached storage is storage accessed through local I/O ports. Example: the typical
desktop PC uses an I/O bus architecture called IDE or ATA. This architecture supports a
maximum of two drives per I/O bus. The other cabling systems are – SATA (Serially Attached
Technology Attachment), SCSI (Small Computer System Interface) and fiber channel (FC).
SCSI is a bus architecture. Its physical medium is usually a ribbon cable. FC is a high-
speed serial architecture that can operate over optical fiber or over a four-conductor copper cable.
An improved version of this architecture is the basis of storage-area networks (SANs).
i) Network-Attached Storage
Network-attached storage provides a convenient way for all the computers on a LAN to share a
pool of storage with the same ease of naming and access enjoyed with local host-attached
storage. However, it tends to be less efficient and have lower performance than some direct-
attached Storage options.
A storage-area network (SAN) is a private network connecting servers and storage units.
The power of a SAN lies in its flexibility. Multiple hosts and multiple storage arrays can attach to
the same SAN, and storage can be dynamically allocated to hosts. A SAN switch allows or
prohibits access between the hosts and the storage (Fig. 5.3). Fiber Chanel is the most common
SAN interconnect.
If the disk head is initially at 53, it will first move from 53 to 98 then to 183 and then to 37,
5.9.2 SSTF (Shortest Seek Time First) algorithm: This selects the request with
minimum seek time from the current head position. SSTF chooses the pending request
closest to the current head position. Eg:- consider a disk queue with request for i/o to
blocks on cylinders ( Fig. 5.5). 98, 183, 37, 122, 14, 124, 65, 67.
If the disk head is initially at 53, the closest is at cylinder 65, then 67, then 37 is closer
then 98 to 67. So, it services 37, continuing we service 14, 98, 122, 124 and finally 183.
The total head movement is only 236 cylinders. SSTF is a substantial improvement over
FCFS, it is not optimal.
5.9.3 SCAN algorithm: In this the disk arm starts moving towards
one end, servicing the request as it reaches each cylinder until it gets
to the other end of the disk. At the other end, the direction of the head
movement is reversed and servicing continues. The initial direction is
chosen depending upon the direction of the head. Eg: -:-consider a
disk queue with request for i/o to blocks on cylinders. 98, 183, 37,
122, 14, 124, 65, 67
If the disk head is initially at 53 and if the head is moving towards the outer track, it
services 65, 67, 98, 122, 124 and183. At cylinder 199 the arm will reverse and will move
towardsthe other end of the disk servicing 37 and then 15. The SCAN is also called as
elevator algorithm. (Fig 5.6).
If the disk head is initially at 53 and if the head is moving towards 0th track, it services 37
and then 15. At cylinder 0 the arm will reverse and will move towards the other end of the
disk servicing 65, 67, 98, 122, 124 and 183.
C-SCAN is a variant of SCAN designed to provide a more uniform wait time. Like SCAN, C-
SCAN moves the head from end of the disk to the other servicing the request along the way.
When the head reaches the other end, it immediately returns to the beginning of the disk,
without servicing any request on the return (Fig. 5.7). Eg: - consider a disk queue with request
for i/o to blocks on cylinders. 98, 183, 37, 122, 14, 124,65,67
If the disk head is initially at 53 and if the head is moving towards the outer track, it services 65,
67, 98, 122, 124 and 183. At cylinder 199 the arm will reverse and will move immediately
towards
Note: If the disk head is initially at 53 and if the head is moving towards track 0, it services 37
and 14 first. At cylinder 0 the arm will reverse and will move immediately towards the other end
of the disk servicing 65, 67, 98, 122, 124 and183.
5.9.5 Look Scheduling algorithm: Look and C-Look scheduling are different version of
SCAN and C-SCAN respectively. Here the arm goes only as far as the final request in each
direction. Thenit reverses, without going all the way to the end of the disk. The Look and C-Look
scheduling look for a request before continuing to move in a given direction (Fig. 5.8).
Eg: -:-consider a disk queue with request for i/o to blocks on cylinders. 98, 183, 37, 122,
14, 124, 65, 67
If the disk head is initially at 53 and if the head is moving towards the outer track, it services 65,
67, 98, 122, 124 and183. At the final request 183, the arm will reverse and will move towards the
first request 14 and then serves 37.
SSTF is commonly used algorithms has it has a less seek time when compared with other
algorithms. SCAN and C-SCAN perform better for systems with a heavy load on the disk, (ie.
more read and write operations from disk).
Selection of disk scheduling algorithm is influenced by the file allocation method, if contiguous
file allocation is choosen, then FCFS is best suitable, because the files are stored in contiguous
blicks and there will be limited head movements required. A linked or indexed file, in contrast,
may include blocks that are widely scattered on the disk, resulting in greater head movement.
The location of directories and index blocks is also important. Since every file must be opened to
be used, and opening a file requires searching the directory structure, the directories will be
accessed frequently. Suppose that a directory entry is on the first cylinder and a file's data are on
the final cylinder. The disk head has to move the entire width of the disk. If the directory entry
were on the middle cylinder, the head would have to move, at most, one-half the width. Caching
the directories and index blocks in main memory can also help to reduce the disk-arm movement,
particularly for read requests.
Because of these complexities, the disk-scheduling algorithm is very important and is written
as a separate module of the operating system.
The process of dividing the disk into sectors and filling the disk with a special data structure is
called low-level formatting. Sector is the smallest unit of area that is read/written by the disk
controller. The data structure for a sector typically consists of a header, a data area (usually 512
bytes in size) and a trailer. The header and trailer contain information used by the disk controller,
such as a sector number and an error-correcting code (ECC).
When the controller writes a sector of data during normal I/O, the ECC is updated with a value
calculated from all the bytes in the data area. When a sector is read, the ECC is recalculated and
is compared with the stored value. If the stored and calculated numbers are different, this
mismatch indicates that the data area of the sector has become corrupted and that the disk sector
may be bad.
Most hard disks are low-level-forniatted at the factory as a part of the manufacturing process.
This formatting enables the manufacturer to test the disk and to initialize the mapping from
logical block numbers to defect-free sectors on the disk.
When the disk controller is instructed for low-level-formatting of the disk, the size of data block
of all sector sit can also be told how many bytes of data space to leave between the header and
trailer of all sectors. It is of sizes, such as 256, 512, and 1,024 bytes. Formatting a disk with a
larger sector size means that fewer sectors can fit on each track; but it also means that fewer
headers and trailers are written on each track and more space is available for user data.
The operating system needs to record its own data structures on the disk. It does so in two steps.
partition and logical formatting.
Partition – is to partition the disk into one or more groups of cylinders. The operating system
can treat each partition as though it were a separate disk. For instance, one partition can hold a
copy of the operating system's executable code, while another holds user files.
logical formatting (or creation of a file system) - Now, the operating system stores the initial
file- system data structures onto the disk. These data structures may include maps of free and
allocated space (a FAT or modes) and an initial empty directory.
To increase efficiency, most file systems group blocks together into larger chunks, frequently
called clusters.
When a computer is switched on or rebooted—it must have an initial program to run. This
is called the bootstrap program. The bootstrap program –
initializes the CPU registers, device controllers, main memory, and then starts the
operating system.
Locates and loads the operating system from the disk
jumps to beginning the operating-system execution.
The bootstrap is stored in read-only memory (ROM). Since ROM is read only, it cannot be
infected by a computer virus. The problem is that changing this bootstrap code requires changing
the ROM, hardware chips. So most systems store a tiny bootstrap loader program in the boot
ROM whose only job is to bring in a full bootstrap program from disk. The full bootstrap
program can be changed easily: A new version is simply written onto the disk. The full bootstrap
program is stored in ''the boot blocks" at a fixed location on the disk. A disk that has a boot
partition is called a boot disk or system disk.
The Windows 2000 system places its boot code in the first sector on the hard disk (master boot
record, or MBR). The code directs the system to read the boot code from, the MBR. In addition
to containing boot code, the MBR contains a table listing the partitions for the hard disk and a
flag indicating which partition the system is to be booted from.
Table
Disk are prone to failure of sectors due to the fast movement of r/w head. Sometimes the
whole disk will be changed. Such group of sectors that are defective are called as bad blocks.
In MS-DOS format command, scans the disk to find bad blocks. If format finds a bad block, it
writes a special value into the corresponding FAT entry to tell the allocation routines not to use
that block.
In SCSI disks, bad blocks are found during the low-level formatting at the factory and is updated
over the life of the disk. Low-level formatting also sets aside spare sectors not visible to the
operating system. The controller can be told to replace each bad sector logically with one of the
spare sectors. This scheme is known as sector sparing or forwarding.
Some controllers replace bad blocks by sector slipping. Here is an example: Suppose that logical
block 17 becomes defective and the first available spare follows sector 202. Then, sector slipping
A swap space can reside in one of two places: It can be carved out of the normal
file system, or it can be in a separate disk partition. If the swap space is simply a large file
within the file system, normal file-system routines can be used to create it, name it, and allocate
its space. External fragmentation can greatly increase swapping times by forcing multiple seeks
during reading or writing of a process image. We can improve performance by caching the block
location information in physical memory.
Alternatively, swap space can be created in a separate raw partition. A separate swap-
space storage manager is used to allocate and deallocate the blocks from the raw partition.
Solaris allocates swap space only when a page is forced out of physical memory, rather
than when the virtual memory page is first created.
Linux is similar to Solaris in that swap space is only used for anonymous memory or for
regions of memory shared by several processes (Fig. 5.10). Linux allows one or more swap areas
to be established. A swap area may be in either a swap file on a regular file system or a raw swap
partition. Each swap area consists of a series of 4-KB page slots, which are used to hold swapped
pages. Associated with each swap area is a swap map—an array of integer counters, each
corresponding to a page slot in the swap area. If the value of a counter is 0, the corresponding
page slot is available. Values greater than 0 indicate that the page slot is occupied by a swapped
page. The value of the counter indicates the number of mappings to the swapped page; for
example, a value of 3 indicates that the swapped page is mapped to three different processes. The
data structures for swapping on Linux systems are shown in below Fig..
5.10
Protection
A key, time-tested guiding principle for protection is the ‗principle of least privilege‘.
It dictates that programs, users, and even systems be given just enough privileges to perform
their tasks. An operating system provides mechanisms to enable privileges when they are needed and
to disable them when they are not needed.
A computer system is a collection of processes and objects. Objects are both hardware
objects (such as the CPU, memory segments, printers, disks, and tape drives) and software
objects (such as files, programs, and semaphores). Each object (resource) has a unique name that
differentiates it from all other objects in the system.
The operations that are possible may depend on the object. For example, a CPU can only
be executed on. Memory segments can be read and written, whereas a CD-ROM or DVD-ROM
can only be read. Tape drives can be read, written, and rewound. Data files can be created,
opened, read, written, closed, and deleted; program files can be read, written, executed, and
deleted.
A process should be allowed to access only those resources
[Link] for which it has authorization
[Link] currently requires to complete process
A domain is a set of objects and types of access to these objects. Each domain is an
ordered pair of <object-name, rights-set>. Example, if domain D has the access right <file F,
{read, write}>, then all process executing in domain D can both read and write file F, and cannot
perform any other operation on that object.
Domains do not need to be disjoint (Fig. 5.11). They may share access rights. For
example, in below Fig., we have three domains: D1 D2, and D3. The access right < O4, (print}>
is shared by D2 and D3,it implies that a process executing in either of these two domains can
print object O5.
When a user creates a new object Oj, the column Oj is added to the access matrix with the
appropriate initialization entries, as dictated by the creator.
The process executing in one domain and be switched to another domain. When we switch a
process from one domain to another, we are executing an operation (switch) on an object (the
domain). Domain switching from domain Di to domain Dj is allowed if and only if the access right
switch € access (i,j). Thus, in the given Fig. 5.13, a process executing in domain D2 can switch to
domain D3 or to domain D5. A process in domain D4 can switch to D1, and one in domain D1 can
switch to domain D2.
Allowing controlled change in the contents of the access-matrix entries requires three additional
operations: copy, owner, and control.
The ability to copy an access right from one domain (or row) of the access matrix to another is
denoted by an asterisk (*) appended to the access right. The copy right allows the copying of the
access right only within the column for which the right is defined. In the below Fig., a process
executing in domain D2 can copy the read operation into any entry associated with file F2. Hence,
the access matrix of Fig. 5.15. can be modified to the access matrix shown in Fig. (b). This scheme
has two variants: A right is copied from access (i, j) to access (k, j); it is then removed from access
(i,j). This action is a transfer of a right, rather than a copy.
1) Propagation of the copy right- limited copy. Here, when the right R* is copied from access (i,j)
to access(k,j), only the right R (not R*) is created. A process executing in domain Dk cannot
further copy the right R.
We also need a mechanism to allow addition of new rights and removal of some rights. The
owner right controls these operations. If access(i,j) includes the owner right, then a process
executing in domain Di, can add and remove any right in any entry in column j.
For example, in below Fig. (a), domain D1 is the owner of F1, and thus can add and delete any
valid right in column F1. Similarly, domain D2 is the owner of F2 and F3 and thus can add and
remove any valid right within these two columns. Thus, the access matrix of Fig.(a) can be
modified to the access matrix shown in Fig.(b) as follows.
A mechanism is also needed to change the entries in a row. If access(i,j) includes the control
right, then a process executing in domain Di, can remove any access right from row j. For
example, in Fig., we include the control right in access(D3, D4). Then, a process executing in
domain D3 can modify domain D4
This is the simplest implementation of access matrix. A set of ordered triples <domain,
object, rights-set> is maintained in a file. Whenever an operation M is executed on an object Oj,
within domain Di, the table is searched for a triple <Di, Oj, Rk>. If this triple is found, the
operation is allowed to continue; otherwise, an exception (or error) condition is raised.
Drawbacks -
The table is usually large and thus cannot be kept in main memory.
Additional I/O is needed
5.11.2 Access Lists for Objects
Each column in the access matrix can be implemented as an access list for one object.
The empty entries are discarded. The resulting list for each object consists of ordered pairs
<domain, rights-set>. It defines all domains access right for that object. When an operation M is
executed on object Oj in Di, search the access list for object Oj, look for an entry <Di, Rj > with
M e Kj. If the entry is found, we allow the operation; if it is not, we check the
default set. If M is in the default set, we allow the access. Otherwise, access is
denied, and an exception condition occurs. For efficiency, we may check the
default set first and then search the access list.
A capability list for a domain is a list of objects together with the operations allowed on
those objects. An object is often represented by its name or address, called a capability. To
execute operation M on object Oj, the process executes the operation M, specifying the capability
for object Oj as a parameter. Simple possession of the capability means that access is allowed.
Capabilities are usually distinguished from other data in one of two ways:
Each object has a tag to denote its type either as a capability or as accessible
data.
Alternatively, the address space associated with a program can be split into two parts. One part is
accessible to the program and contains the program's normal data and instructions. The other
part, containing the capability list, is accessible only by the operatingsystem.
The lock-key scheme is a compromise between access lists and capability lists. Each
object has a list of unique bit patterns, called locks. Similarly, each domain has a list of unique
bit patterns, called keys. A process executing in a domain can access an object only if that
domain has a key that matches one of the locks of the object.
Solaris 10 advances the protection available in the Sun Microsystems operating system
by explicitly adding the principle of least privilege via role-based access control (RBAC). This
facility revolves around privileges. A privilege is the right to execute a system call or to use an
option within that system call (such as opening a file with write access). Privileges can be
assigned to processes, limiting them to exactly the access they need to perform their work.
Privileges and programs can also be assigned to roles. Users are assigned roles or can take
Since the capabilities are distributed throughout the system, we must find them before we
can revoke them. Schemes that implement revocation for capabilities include the following:
5.13.1 Reacquisition. Periodically, all capabilities are deleted from each domain. If
a process wants to use a capability, it may find that that capability has been deleted. The
process may then try to reacquire the capability. If access has been revoked, the process
will not be able to reacquire the capability.
5.13.3 Indirection. The capabilities point indirectly to the objects. Each capability
points to a unique entry in a global table, which in turn points to the object. We
implement revocation by searching the global table for the desired entry and deleting it.
Then, when an access is attempted, the capability is found to point to an illegal table
entry.
5.13.4 Keys. A key is a unique bit pattern that can be associated with a capability.
This key is defined when the capability is created, and it can be neither modified nor
inspected by the process owning the capability. A master key is associated with each
object; it can be defined or replaced with the set-key operation.
5.13.5 When a capability is created, the current value of the master key is associated
with the capability. When the capability is exercised, its key is compared with the
master key. If the keys match, the operation is allowed to continue; otherwise, an
exception condition is raised.
In key-based schemes, the operations of defining keys, inserting them into lists, and
deleting them from lists should not be available to all users.
1) An Example: Hydra
Operations on objects are defined procedurally. The procedures that implement such
operations are themselves a form of object, and they are accessed indirectly by capabilities. The
names of user-defined procedures must be identified to the protection system if it is to deal with
objects of the user defined type. When the definition of an object is made known to Hydra, the
names of operations on the type become auxiliary rights.
Hydra also provides rights amplification. This scheme allows a procedure to be certified
as trustworthy to act on a formal parameter of a specified type on behalf of any process that
holds a right to execute the procedure. The rights held by a trustworthy procedure are
independent of, and may exceed, the rights held by the calling process.
When a user passes an object as an argument to a procedure, we may need to ensure that
the procedure cannot modify the abject. We can implement this restriction readily by passing an
access right that does not have the modification (write) right.
The procedure-call mechanism of Hydra was designed as a direct solution to the problem
of mutually suspicious subsystems.
A Hydra subsystem is built on top of its protection kernel and may require protection of
its own components. A subsystem interacts with the kernel through calls on a set of kernel-
defined primitives that define access rights to resources defined by the subsystem.
A different approach to capability-based protection has been taken in the design of the
Cambridge CAP system. CAP's capability system is simpler and superficially less powerful than
that of Hydra. It can be used to provide secure protection of user-defined objects. CAP has two
kinds of capabilities.
The ordinary kind is called a data capability. It can be used to provide access to objects,
but the only rights provided are the standard read, write, and execute of the individual storage
segments associated with the object.
The second kind of capability is the software capability, which is protected, but not
interpreted, by the CAP microcode. It is interpreted by a protected (that is, a privileged)
procedure, which may be written by an application programmer as part of a subsystem. A
particular kind of rights amplification is associated with a protected procedure.
5.15.1 History
Version 0.01 (May 1991) had no networking, ran only on 80386-compatible Intel
processors and on PC hardware, had extremely limited device-drive support, and
supported only the Minix file system.
Linux 1.0 (March 1994) included these new features:
Version 1.2 (March 1995) was the final PC-only Linux kernel.
Released in June 1996, 2.0 added two major new capabilities:
[Link].1 Support for multiple architectures, including a fully 64-bit native Alpha port.
[Link].2 Support for multiprocessor architectures
Other new features included:
[Link].1 Improved memory-management code
[Link].2 Improved TCP/IP performance
[Link].3 Support for internal kernel threads, for handling dependencies between
loadable modules, and for automatic loading of modules on demand.
[Link].4 Standardized configuration interface
Available for Motorola 68000-series processors, Sun Sparc systems, and for PC and
PowerMac systems .
Linux uses many tools developed as part of Berkeley‘s BSD operating system,
MIT‘s X Window System, and the Free Software Foundation's GNU project.
The min system libraries were started by the GNU project, with improvements
provided by the Linux community.
Linux networking-administration tools were derived from 5.3BSD code;
recent BSD derivatives such as Free BSD have borrowed code from Linux in
return.
The Linux system is maintained by a loose network of developers collaborating
over the Internet, with a small number of public ftp sites acting as de facto
standard repositories.
Standard, precompiled sets of packages, or distributions, include the basic Linux
system, system installation and management utilities, and ready-to-install packages of
common UNIX tools.
The first distributions managed these packages by simply providing a means of
unpacking all the files into the appropriate places; modern distributions include
advanced package management.
Early distributions included SLS and Slackware. Red Hat and Debian are popular
distributions from commercial and noncommercial sources, respectively.
The RPM Package file format permits compatibility among the various Linux
distributions.
The Linux kernel is distributed under the GNU General Public License (GPL), the
terms of which are set out by the Free Software Foundation.
Anyone using Linux, or creating their own derivative of Linux, may not make the
derived product proprietary; software released under the GPL may not be redistributed
as a binary- only product.
5.16 Design Principles
Like most UNIX implementations, Linux is composed of three main bodies of code;
the most important distinction between the kernel and all other components.
The kernel is responsible for maintaining the important abstractions of the
operating system.
Kernel code executes in kernel mode with full access to all the physical resources of
the computer.
All kernel code and data structures are kept in the same single address space (Fig. 5.18) .
The system libraries define a standard set of functions through which applications
interact with the kernel, and which implement much of the operating-system functionality
that does not need the full privileges of kernel code.
The system utilities perform individual specialized management tasks.
Sections of kernel code that can be compiled, loaded, and unloaded independent of the
rest of the kernel.
A kernel module may typically implement a device driver, a file system, or a
networking protocol.
The module interface allows third parties to write and distribute, on their own terms,
device drivers or file systems that could not be distributed under the GPL.
Kernel modules allow a Linux system to be set up with a standard, minimal kernel,
without any extra device drivers built in.
Supports loading modules into memory and letting them talk to the rest of the kernel.
Module loading is split into two separate sections:
Managing sections of module code in kernel memory
Handling symbols that modules are allowed to reference
The module requestor manages loading requested, but currently unloaded, modules; it
also regularly queries the kernel to see whether a dynamically loaded module is still in
use, and will unload it when it is no longer actively needed.
5.19.1 Allows modules to tell the rest of the kernel that a new driver has become available.
5.19.2 The kernel maintains dynamic tables of all known drivers, and provides a set
of routines to allow drivers to be added to or removed from these tables at any
time.
5.19.3 Registration tables include the following items:
[Link] Device drivers
[Link] File systems
[Link] Network protocols
[Link] Binary format
5.21.1 UNIX process management separates the creation of processes and the running of
a new program into two distinct operations.
[Link] The fork system call creates a new process.
[Link] A new program is run after a call to execve.
5.21.2 Under UNIX, a process encompasses all the information that the operating
system must maintain t track the context of a single execution of a single program.
5.21.3 Under Linux, process properties fall into three groups: the
process‘s identity, environment, and context.
_ Process ID (PID). The unique identifier for the process; used to specify processes to the
operating system when an application makes a system call to signal, modify, or wait for another
process.
_ Credentials. Each process must have an associated user ID and one or more group IDs that
determine the process‘s rights to access system resources and files.
_ Personality. Not traditionally found on UNIX systems, but under Linux each process has an
associated personality identifier that can slightly modify the semantics of certain system calls.
Used primarily by emulation libraries to request that system calls be compatible with certain
specific flavors of UNIX.
5.23.1 The process‘s environment is inherited from its parent, and is composed of two null-
terminated vectors:
[Link] The argument vector lists the command-line arguments used to invoke the
running program; conventionally starts with the name of the program itself
[Link] The environment vector is a list of ―NAME=VALUE‖ pairs that associates
named environment variables with arbitrary textual values.
5.24.1 The (constantly changing) state of a running program at any point in time.
5.24.2 The scheduling context is the most important part of the process context; it is the
information that the scheduler needs to suspend and restart the process.
5.24.3 The kernel maintains accounting information about the resources currently being
consumed by each process, and the total resources consumed by the process in its lifetime
so far.
5.24.4 The file table is an array of pointers to kernel file structures. When making file I/O
system calls, processes refer to files by their index into this table.
5.24.5 Whereas the file table lists the existing open files, the file-system context applies to
requests to open new files. The current root and default directories to be used for new file
searches are stored here.
5.24.6 The signal-handler table defines the routine in the process‘s address space to be
called when specific signals arrive.
5.24.7 The virtual-memory context of a process describes the full contents of the its private
address space.
5.25.1 Linux uses the same internal representation for processes and threads; a thread is
simply a new process that happens to share the same address space as its parent.
5.25.2 A distinction is only made when a new thread is created by the clone system call.
[Link] fork creates a new process with its own entirely new process context
[Link] clone creates a new process with its own identity, but that is allowed to
share the data structures of its parent
5.25.3 Using clone gives an application fine-grained control over exactly what is
shared between two threads.
5.26 Scheduling
5.26.1 The job of allocating CPU time to different tasks within an operating system.
5.26.2 While scheduling is normally thought of as the running and interrupting of
processes, in Linux, scheduling also includes the running of the various kernel tasks.
5.26.3 Running kernel tasks encompasses both tasks that are requested by a running process and
Interrupt service routines are separated into a top half and a bottom half.
The top half is a normal interrupt service routine, and runs with
recursive interrupts disabled.
The bottom half is run, with all interrupts enabled, by a miniature scheduler that
ensures that bottom halves never interrupt themselves.
This architecture is completed by a mechanism for disabling selected
bottom halves while executing normal, foreground kernel code.
5.23.2. Interrupt Protection Levels
Each level may be interrupted by code running at a higher level, but will never
be interrupted by code running at the same or a lower level.
User processes can always be preempted by another process when a time-sharing
scheduling interrupt occurs ( Fig. 5.19).
Linux 2.0 was the first Linux kernel to support SMP hardware; separate processes or threads can execute in p
To preserve the kernel‘s non preemptible synchronization requirements, SMP imposes the restriction, via a s
kernel-mode code.
Linux‘s physical memory-management system deals with allocating and freeing pages, groups of pages, and s
It has additional mechanisms for handling virtual memory, memory mapped into the
address space of running processes.
The page allocator allocates and frees all physical pages; it can allocate ranges of
The VM system maintains the address space visible to each process: It creates pages of
virtual memory on demand, and manages the loading of those pages from disk or their
swapping back out to disk as required.
The VM manager maintains two separate views of a process‘s address space:
A logical view describing instructions concerning the layout of the address space.
The address space consists of a set of nonoverlapping regions, each representing a
continuous, page-aligned subset of the address space.
A physical view of each address space which is stored in the hardware page tables
for the process.
Virtual memory regions are characterized by:
The backing store, which describes from where the pages for a region come;
regions are usually backed by a file or by nothing (demand-zero memory)
The region‘s reaction to writes (page sharing or copy-onwrite).
The kernel creates a new virtual address space
1. When a process runs a new program with the exec system call
2. Upon creation of a new process by the fork system call
On executing a new program, the process is given a new, completely empty virtual-
address space; the program loading routines populate the address space with virtual
memory regions.
Creating a new process with fork involves creating a complete copy of the existing
process‘s virtual address space.
The kernel copies the parent process‘s VMA descriptors, then creates a new set
Linux maintains a table of functions for loading programs; it gives each function
the opportunity to try loading the given file when an exec system call is made.
The registration of multiple loader routines allows Linux to support both the ELF
and [Link] binary formats.
Initially, binary-file pages are mapped into virtual memory; only when a program tries
to access a given page will a page fault result in that page being loaded into physical
memory.
An ELF-format binary file consists of a header followed by several page-aligned
sections; the ELF loader works by reading the header and mapping the sections of the file
into separate regions of virtual memory.
Fig. 5.21 shows the typical layout of memory regions set up by the ELF loader. In a reserved
region at one end of the address space sits the kernel, in its own privileged region of virtual
memory inaccessible to normal user-mode programs. The rest of virtual memory is available to
applications, which can use the kernel's memory-mapping functions to create regions that map a
portion of a file or that are available for application data.
The loader's job is to set up the initial memory mapping to allow the execution of the program to
start. The regions that need to be initialized include the stack and the program's text and data
regions. The stack is created at the top of the user-mode virtual memory; it grows downward
toward lower-numbered addresses.
A set of definitions that define what a file object is allowed to look like
Ext2fs (Fig. 5.22) uses a mechanism similar to that of BSD Fast File System (ffs)
for locating data blocks belonging to a specific file.
The main differences between ext2fs and ffs concern their disk allocation policies.
In ffs, the disk is allocated to files in blocks of 8Kb, with blocks being subdivided
into fragments of 1Kb to store small files or partially filled blocks at the end of a
file.
Ext2fs does not use fragments; it performs its allocations in smaller units. The
default block size on ext2fs is 1Kb, although 2Kb and 4Kb blocks are also
supported.
Ext2fs uses allocation policies designed to place logically adjacent blocks of a file
into physically adjacent blocks on disk, so that it can submit an I/O request for
several disk blocks as a single operation.
The proc file system does not store data, rather, its contents are computed on
demand according to user file I/O requests.
proc must implement a directory structure, and the file contents within; it must then define
It uses this inode number to identify just what operation is required when a user
tries to read from a particular file inode or perform a lookup in a particular
directory inode.
When data is read from one of these files, proc collects the appropriate
information, formats it into text form and places it into the requesting process‘s
read buffer.
Character Devices
A device driver which does not offer random access to fixed blocks of data.
A character device driver must register a set of functions which implement the
driver‘s various file I/O operations.
The kernel performs almost no preprocessing of a file read or write request to a
character device, but simply passes on the request to the device.
The main exception to this rule is the special subset of character device drivers
which implement terminal devices, for which the kernel maintains a standard
interface.
Like UNIX, Linux informs processes that an event has occurred via signals.
There is a limited number of signals, and they cannot carry information: Only the fact
that a signal occurred is available to a process.
The Linux kernel does not use signals to communicate with processes with are running in
The pipe mechanism allows a child process to inherit a communication channel to its
parent, data written to one end of the pipe can be read a the other.
Shared memory offers an extremely fast way of communicating; any data written by one
process to a shared memory region can be read immediately by any other process that has
mapped that region into its address space.
To obtain synchronization, however, shared memory must be used in conjunction with
another Interprocess communication mechanism.