Multimedia Notes
Multimedia Notes
media which means ―middle or centre. Multimedia therefore means ―multiple intermediaries or ―multiple
means. Multimedia is the medium which provides information to the users in the form of text, audio, video,
animation and graphics. The information is delivered to the users by digital or electronic means. When a user
is allowed to control the elements of multimedia then it becomes interactive multimedia. Interactive
multimedia is called hypermedia when the user is given the structure of linked elements to control it.
Multimedia System : Any system that incorporates different flavours of media can be termed as a multimedia
system.
Multimedia information : Multimedia information can be defined as information that consists of one or more
different media types. Multimedia information consists of text, audio, video, 2D graphics, and 3D graphics.
Categories of Multimedia
a) Linear Multimedia - Linear active content progresses without any navigation control for the viewer
such as a cinema presentation.
b) Non Liner Multimedia - In a multimedia presentation the user can instantly navigate to different parts
of the presentation and display the frames in any way, without appreciable delay, due to which it is
called a nonlinear presentation. Non-linear content offers user interactivity to control progress as used
with a computer game or used in self-paced computer based training. Hypermedia is an example of
non-linear content.
Impact of Multimedia
Multimedia can be used in a variety of ways. Some of the most common applications are given below:
a) Business :- Multimedia is used for advertising and selling products on the Internet. Some businesses
use multimedia tools such as CD-ROMs, DVDs or online tutorials for training or educating staff
members about things the employer want them to learn or know. It also saves money of the employers
as now they don’t have to pay extra expenses on training or education of their human resources.
b) Research and Medicine :
Multimedia is increasingly used in research in the fields of science, medicine and mathematics. It is
mostly used for modelling and simulation. For instance, a scientist can look at a molecular model of a
particular substance and work on it to arrive at a new substance. In Medicine, doctors acquire training
by watching a virtual surgery or they can simulate how the human body is affected by diseases spread
by germs and then develop techniques to prevent it.
c) Public Access : Public Access is an area of application where several multimedia applications will be
available very soon. One of the application is the tourist information system, where a travel enthusiast
will be shown glimpse of the place he would like to visit. With the help of multimedia various source
providing applications could be created.
d) Entertainment : Multimedia is used to create special effects in films, TV serials, radio shows, games and
animations. Multimedia games are popular software programs that available online as well as on DVDs
and CD-ROMs. Use of special technologies such as virtual reality turn these games into real life
experiences. These games allow uses to fly aeroplanes, drive cars, do wrestling, etc
e) Industry : In the Industrial sector, multimedia is used to present information to all people related to
the industry such as shareholders, senior level managers and co-workers. Multimedia is also helpful in
advertising and selling products all over the world over internet.
f) Commercial : Creative presentations are used to grab the attention of the masses in advertising.
Industrial, business to business and interoffice communications are mostly developed by firms
providing creative services. They work on advanced multimedia presentations rather than simple slide
shows to sell ideas or make training programs more interesting.
g) Education : Multimedia is used as a source of information in the field of education. Pupils can research
on various topics such as solar system or information technology using different multimedia
presentations. To make teaching more interesting and fun for pupils, teachers can make multimedia
presentations of chapters. Visual images, animation, diagrams, etc., have more effect on pupils. Various
computer-based training (CBT) courses are also available online for study.
h) Multimedia in Public : Places In railway stations, hotels, museums, grocery stores and shopping malls
multimedia will become available at stand-alone terminals to provide information and help. Such
installation reduce demand on traditional information booths and personnel, add value, and they can
work around the clock, even in the middle of the night, when live help is off-dut
i) Engineering : Software engineers may use multimedia in computer simulations for anything from
entertainment to training such as military or industrial training. Multimedia for software interfaces are
often done as collaboration between creative professionals and software engineers.
Different Media Type / Components and Its Applications
a) Text - Text has been commonly used to express information not only for today but from the early days.
Today literature, news, and any information including internet contains a larger number of text. Today,
hypertext is commonly used in digital documents, allowing nonlinear access to information. The text in
the multimedia is used to communicate information to the user. Proper use of text and words in
multimedia presentation will help the content developer to communicate the idea and message to the
user. Even text is used in Film title and advertisement. Font is the part of any text document. Size of
the character is measured in point. One point is approximately 1/72 of an inch i.e., 0.0138. Typeface
of a font is defined as serif and sans serif. The serif is the little decoration at the end of a letter
stroke. Times, Times New Roman, Bookman are some fonts which come under serif category. Arial,
Optima, Verdana are some examples of sans serif font.
Hypermedia refers to the presentation of video, animation, and audio, which are often referred to as
“dynamic” or “time based” content or as “multimedia.” Non-web forms of hypertext and hypermedia
include CD-ROM and DVD encyclopaedias (such as Microsoft’s Encarta), e-books, and the online help
systems that we find in software products.
b) Image - Images consist of a set of units called pixels organized in the form of a twodimensional array.
The two dimensions specify the width and height of the images. Each pixel has bit depth, which
defines how many bits are used to represent an image
Bit depth / Color Depth—Bit depth represents the number of bits assigned to each pixel.
Accordingly, images are categorized by the bit depth as binary images where every pixel is
represented by one bit or gray-level images where every pixel is represented by a number of
bits (typically 8) or color images, where each pixel is represented by three color channels.
Formats – Format is varying application specific. Fax one image format, where as digital image
has another format.
Dimensionality : Stereo images are commonly used for depth-perception effects. Images can
also be stitched together to form mosaics and panoramas
Digital Representation of Images :-
All images are represented digitally as pixels. Each pixel is further represented by a number of bits,
which is commonly called the pixel depth.
c) Video : - Video is represented as a sequence of images. Each image in the sequence typically has the
same properties of width, height, and pixel depth. All of these parameters can be termed as spatial
parameters. Additionally, there is one more temporal parameter known as frames per second or fps.
This parameter describes how fast the images need to be shown per second for the user to perceive
continuous motion. Video can be classified in the following ways:
Aspect ratio—A common aspect ratio for video is 4:3, which defines the ratio of the width to
height.
Scanning format—Scanning helps convert the frames of video into a one-dimensional signal
for broadcast. Now a display devices can support progressive scanning.
****************************************************************************
Note :-
CCD (charge coupled device)- In the digital camera, there could be an image sensor CCD (charge
coupled device) array. Each sensor releases an electric charge that is proportional to the amount of
light energy falling on it; the more energy, the higher the charge (within a range). The released charge
is then converted into a digital representation in terms of bits, which are ultimately used to display the
image information on a rendering device.
Interlacing : A technique called interlacing is used to provide a flicker-free image without increasing
the bandwidth requirement. In the interlacing technique, each image frame is divided into two fields,
each consisting of alternate horizontal lines. An even field (half-frame) comprises all the even-
numbered horizontal scan lines. An odd field comprises all the odd-numbered horizontal scan lines.
***************************************************************************
d) Audio : Digital audio is characterized by a sampling rate in hertz, which gives the number of samples
per second. A sample can be defined as an individual unit of audio information. Each sample also has a
size, the sample size, which typically is anywhere from 8-bits to 16-bits depending on the application.
Audio is also describe as :
Dimensionality—The dimensions of an audio signal signify the number of channels that are
contained in the signal. These may be mono (one channel), stereo (two channels), which is by
far the most common. Recent standards also use surround sound which consists of many
channels
Frequency Range—Audio signals are also described by the frequency range or frequency band
that they contain. For example, audio voice signals are referred to as narrow band because
they contain lower frequency content. Music is normally referred to as wide band.
2D Graphics : 2D graphical elements have become commonplace in multimedia presentations to
enhance the message to be conveyed. A 2D graphic element is represented by 2D vector coordinates
and normally has properties such as a fill color, boundary thickness.
3D Graphics:- 3D graphics are primarily used today for high-end content in movies, computer games,
and advertising. 3D graphics largely make use of vector coordinate spaces. 3D graphics concepts and
practices have advanced considerably as a science but, until recently, were not a commonplace media
type.
When the object or image is two dimensional, When the object or image is three dimensional, then
then it is known as 2-D. it is known as 3-D.
2-D is more affordable as compared to 3-D. 3-D is more expensive as compared to 2-D.
It is all about the edges, boundaries or frames It is all about the movements of the images.
of the images.
Here we prefer the traditional drawing method Here we prefer computer software to create an
to create an image. image.
Here the image includes height and weight. Here the image includes height, weight and depth.
Circle, Rectangle, square, triangle, polygon, Cylinder, cube, primed, spear, etc are examples of
etc. are examples of 2-D. 3-D.
2-D is not fit for conceptual drawing. 3-D is fit for conceptual drawing.
----------------------------------------------------------------------------------------------------------------------------------
Text: Types of Text, Ways to Present Text, Aspects of Text Design, Character, Character Set,
Codes, Unicode,
Text - Text has been commonly used to express information not only for today but from the early days.
The text in the multimedia is used to communicate information to the user. Proper use of text and
words in multimedia presentation will help the content developer to communicate the idea and
message to the user. Even text is used in Film title and advertisement. Font is the part of any text
document. Size of the character is measured in point. One point is approximately 1/72 of an inch i.e.,
0.0138. Typeface of a font is defined as serif and sans serif. The serif is the little decoration at the end
of a letter stroke. Times, Times New Roman, Bookman are some fonts which come under serif
category. Arial, Optima, Verdana are some examples of sans serif font.
Unformatted Text: Also known as plaintext, this comprise of fixed sized characters from a limited
character set. The character set is called ASCII table which is short for American Standard Code for
Information Interchange and is one of the most widely used character sets. It basically consists of a
table where each character is represented by a unique 7-bit binary code. The characters include a to z,
A to Z, 0 to 9, and other punctuation characters like parenthesis, ampersand, single and double quotes,
mathematical operators, etc. All the characters are of the same height. In addition, the ASCII character
set also includes a number of control characters. These include BS (backspace), LF (linefeed), CR
(carriage return), SP (space), DEL (delete), ESC (escape), FF (form feed) and others.
****Plain text can be written in Text format(.txt), word processor format(.doc) or Adobe PDF
format(.pdf) or Rich Text Format.***
Hypertext is commonly used in digital documents, allowing nonlinear access to information. Text in
multimedia is generally combined with other types of media such as video, graphics, photography,
animation and sounds. Text is the most widely used and flexible means of presenting information on
screen and conveying ideas.
Formatted Text: Formatted text are those where apart from the actual alphanumeric characters, other
control characters are used to change the appearance of the characters, e.g. bold, underline, italics,
varying shapes, sizes and colors etc., Most text processing software use such formatting options to
change text appearance. It is also extensively used in the publishing sector for the preparation of
papers, books, magazines, journals, and so on. Rich Text on of the example of formatted text.
Rich Text Format (RTF) : Plain text files contain only the information codes, such as ASCII and
EBCDIC; formatted text files contain the information codes as well as formatting information.
Different word processors use different formatting information storage techniques. This makes
transfer of data from one system to another difficult at times, and at other times impossible. The
Rich Text Format (RTF) aims to bridge the gap between the different file formats used by different
word processors and desktop publishers. The main elements of the RTF are:
Character set
Font table
Color table
Document formatting
Section formatting
Paragraph formatting
General formatting
Character formatting: Bold, italic, underline, etc.
Special characters: Hyphens, backslash, etc.
The RTF aims to be a common format for exchanging information among var- ious word processors
and desktop publication systems. But, due to nonstandard implementations, transfer of data
sometimes does not occur with complete success (i.e., some loss of formatting may occur).
Animating Text: There are plenty of ways to retain a viewer’s attention when displaying text. For
example, you can animate bulleted text and have it “fly” onto the screen. You can “grow” a headline a
character at a time. For public speakers, simply highlighting the important text works well as a pointing
device.
Symbols and Icons: Symbols are concentrated text in the form of stand-alone graphic constructs.
Symbols convey meaningful messages. Icons are symbolic representations of objects and processes
common to the graphical user interfaces of many computer operating system. On the other hand,
pictures, icons, moving images, and sounds are more easily recalled and remembered by viewers. With
multimedia, you have the power to blend both text and icons to enhance the overall impact and value
of your message.
Today literature, news, and any information including internet contains a larger number of text. Text
can be used in many ways in multimedia:
• In a website
• In films such as titles and credits
• As subtitles in a film or documentary that provide a translation
• It may be used in advertisements
• it is used in text messaging
Character Sets and Alphabets
ASCII Character Set :
The American Standard Code for Information Interchange (ASCII) is the 7-bit character coding system
most commonly used by computer systems in the United States and abroad. ASCII assigns a number or
value to 128 characters, including both lower- and uppercase letters, punctuation marks, Arabic
numbers, and math symbols. Also included are 32 control characters used for device control messages,
such as carriage return, line feed, tab, and form feed. ASCII code numbers always represent a letter or
symbol of the English alphabet, so that a computer or printer can work with the number that
represents the letter, regardless of what the letter might actually look like on the screen or printout.
To a computer working with the ASCII character set, the number 65, for example, always represents an
uppercase letter A. Later, when displayed on a monitor or printed, the number is turned into the letter
Unicode Character :
A16-bit architecture for multilingual text and character encoding called Unicode. The original standard
accommodated up to about 65,000 characters to include the characters from all known languages and
alphabets in the world. It support all of the International Character set for writing purpose.
AUDIO
(Audio: Basic Sound Concepts, Types of Sound, Digitizing Sound, Computer Representation of Sound
(Sampling Rate, Sampling Size, Quantization), Audio Formats, Audio tools, MIDI)
Sound is a wave phenomenon like light, but it is macroscopic and involves molecules of air being
compressed and expanded under the action of some physical device. It is in one Dimension in nature.
For example, a speaker in an audio system vibrates back and forth and produces a longitudinal
pressure wave that we perceive as sound. Without air there is no sound—for example, in space. Since
sound is a pressure wave, it takes on continuous values, as opposed to digitized ones with a finite
range. Sound pressure levels (loudness or volume) are measured in decibels (dB). Given a sound with
fundamental frequency f , we define harmonics as any musical tones whose frequencies are integral
multiples of the fundamental frequency, i.e., 2 f , 3 f , 4 f , ..., etc. All the harmonics are periodic at the
fundamental frequency, their linear combinations are also periodic at the fundamental frequency.
Analog Signal :- A signal is analog if it can be represented by a continuous function. For instance, it
might encode the changing amplitude with respect to an input dimension(s). So an analog or analogue
signal is any continuous signal for which the time varying feature (variable) of the signal is a
representation of some other time varying quantity
Digital Signal :- Digital signals, on the other hand, are represented by a discrete set of values defined
at specific (and most often regular) instances of the input domain, which might be time, space, or
both.
Digitization : The conversion of signals from analog to digital signal is called digitization. Or Digitization
means conversion to a stream of numbers—preferably integers for efficiency.
Digitized sound is sampled sound. Every nth fraction of a second, a sample of sound is taken and stored
as digital information in bits and bytes. The quality of this digital recording depends upon how often
the samples are taken (sampling rate or frequency, measured in kilohertz, or thousands of samples
per second) and how many numbers are used to represent the value of each sample (bit depth,
sample size, resolution, or dynamic range).
So Sampling means measuring the quantity we are interested in, usually at evenly spaced intervals
and the rate at which it is perform is called the sampling frequency. Or we can say a sample refers to
a value or set of values at a point in time and/or space. A sampler is a subsystem or operation that
extracts samples from a continuous signal.
Sample sizes are either 8 bits or 16 bits. The larger the sample size, the more accurately the data will
describe the recorded sound. An 8-bit sample size provides 256 equal measurement units to describe
the level and frequency of the sound in that slice of time. A 16-bit sample size, on the other hand,
provides 65,536 equal units to describe the sound in that same slice of time.
Quantization : The amplitude values obtained after sampling may be long real numbers which are
usually rounded to the nearest predefined discrete values. This process of converting the real numbers
to the predefined discrete numbers is called quantization. Quantization corresponds to a discretization
of the intensity values. That is, of the co-domain of the function.
If the amplitude is greater than the intervals available, clipping of the top and bottom of the wave
occurs. Quantization can produce an unwanted background hissing noise, and clipping may severely
distort the sound.
****Audacity is a free open-source sound editing application for Windows, Macintosh, and Linux***
An electronic device called an Analog-to-Digital Converter (ADC) is used to convert analog signals to
digital values. A Digital-to-Analog Converter (DAC) is used to convert digitized signals back to their
original analog form.
Nyquist Sample:
As we know now, each sound is just made from sinusoids. Below Fig.a shows a single sinusoid: it is a
single, pure, frequency (only electronic instruments can create such boring sounds).
Now if the sampling rate just equals the actual frequency, we can see from Fig. b that a false signal is
detected: it is simply a constant, with zero frequency.
If, on the other hand, we sample at 1.5 times the frequency, Fig. c shows that we obtain an incorrect
(alias) frequency that is lower than the correct one—it is half the correct one (the wavelength, from
peak to peak, is double that of the actual signal). An alias is any artifact that does not belong to the
original signal. Thus, for correct sampling we must use a sampling rate equal to at least twice the
maximum frequency content in the signal. This is called the Nyquist rate.
The low frequency distortion produced due to lowest sampling rate, known as under-sampling, is
referred to as aliasing, because the input wave is not truly represented at the output.
The sampling frequency that is half of the input wave frequency, is referred as Nyquist frequency.
When sampling is done at much higher rate, it is called over-sampling. Both under and over sampling,
in general, should be avoided.
Mathematica Representation :
The conversion of signals from analog to digital occurs via two main processes: sampling and
quantization. The reverse process of converting digital signals to analog is known as interpolation.
i. Sampling : - Assume that we start with a one-dimensional analog signal in the time t
domain, with an amplitude given by x(t). The sampled signal is given by
xs(n) = x(nT), where T is the sampling period a
f = 1/T is the sampling frequency.
If you reduce T (increase f ), the number of samples increases; and correspondingly, so
does the storage requirement.
ii. Quantization :- Quantization deals with encoding the signal value at every sampled
location with a predefined precision, defined by a number of levels. ie. How many bits
do you use to represent the value of signal at each instance.
xq(n) = Q[xs(n)], where Q is the rounding function.
Q represents a rounding function that maps the continuous value x s (n) to the nearest
digital value using b bits. This actually depends on the type of signal and what its
intended use is. Audio signals, which represent music, must be quantized on 16 bits,
whereas speech only requires 8 bits.
Bit Rate :-It describes the number of bits being produced per second. Bit rate is of
critical importance when it comes to storing a digital signal, or transmitting it across
networks, which might have high, low, or even varying bandwidths
**** Audio CD - Audio CDs use a sampling rate of 44.1 KHz, 16-bit per sample, and two
channels to store the stereo sound.
**** DVD(Digital versatile disc) audio : Use a sampling rate of 192KHz (max), 24 (max)-
bit per sample, and up to 6 channels to store the sound.
MIDI Audio :
MIDI (Musical Instrument Digital Interface) is a communications standard developed in the early 1980s
for electronic musical instruments and computers. MIDI is a scripting language—it codes “events”
that stand for the production of certain sounds. Therefore, MIDI files are generally very small. For
example, a MIDI event might include values for the pitch of a single note, its volume, and what
instrument sound to play. MIDI represents a set of specifications used in instrument development so
that instruments from different manufacturers can easily exchange musical information.
MIDI Concepts
Music is organized into tracks in a sequencer. Each track can be turned on or off on
recording or playing back. Usually, a particular instrument is associated with a MIDI
channel. MIDI channels are used to separate messages. There are 16 channels, numbered
from 0 to 15. The channel forms the last four bits (the least significant bits) of that do
refer to the channel. The idea is that each channel is associated with a particular
instrument—for example, channel 1 is the piano, channel 10 is the drums. Nevertheless,
you can switch instruments midstream, if desired, and associate another instrument with
any channel.
Along with channel messages (which include a channel number), several other types of
messages are sent, such as a general message for all instruments indicating a change in
tuning or timing; these are called system messages. It is also possible to send a special
message to an instrument’s channel that allows sending many notes without a channel
specified. We will describe these messages in detail later.
The way a synthetic musical instrument responds to a MIDI message is usually by simply
ignoring any “play sound” message that is not for its channel. If several messages are for its
channel, say several simultaneous notes being played on a piano, then the instrument
responds, provided it is multi-voice—that is, can play more than a single note at once.
The MIDI protocol is an entire music description language in binary form. Each word describing an
action of a musical performance is assigned a specific binary code. The data in a MIDI status byte is
between 128 and 255; each of the data bytes is between 0 and 127. Actual MIDI bytes are 8 bit, plus
a 0 start and stop bit, making them 10-bit “bytes”.
The most important advantage of digital audio is its consistent playback quality.
A wider selection of application software and system support for digital audio is available for both the
Macintosh and Windows platforms.
The preparation and programming required for creating digital audio do not demand knowledge of
music theory, while working with MIDI data usually does require a modicum of familiarity with musical
scores, keyboards, and notation, as well as audio production.
You don’t have control over the playback hardware.
MP3/MPEG format:
MP3 (MPEG-1 Audio Layer 3) is a method to compress and store audio. An MP3 file can compress a music file
by up to 95% of its original CD-quality size while maintaining good enough audio quality.
Computers are digital devices, representing data as discrete numbers, while sound is analog, which can exist at
any value. For a computer or any digital device to store sound, it must encode the audio waveform as a series
of numbers. This encoding is typically done using 16 bits sampled at 44.1 kilohertz for CD audio. This
represents a lot of data -- about 10 megabytes (MB) per minute. This is known as CD-quality or uncompressed
audio and is often stored as a WAV file on a computer.
MP3 audio is a way to compress audio to use much less data and still sound good. MP3 is a lossy encoding
because some data is discarded and not recoverable later. First, a modified discrete cosine transform and a
fast Fourier transform technique are used to reduce the full audio data into smaller data representations of
the original values. Then, a psychoacoustic model is used to remove data that would not be heard by the
human ear, such as frequencies above the human hearing range or sounds too quiet to hear. This resulting
representation of the remaining audio data is then allocated and encoded to fit the data requirements. After
encoding, an MP3 file might only take about 1 MB per minute for good-quality music. Almost any modern
device can play an MP3 file. Computers and smartphones will have a built-in audio player that can play MP3s,
as well as other optional software, such as iTunes, that can play and organize MP3 files. Most modern CD
players can read and play MP3 files that have been burned to a CD. Likewise, most recent car entertainment
systems can play MP3 files stored on a CD or USB storage. MP3 files also contain ID3 tag information. This can
contain metadata such as the song title, artist name, album title and year.
MP3 was developed as an audio codec option for MPEG-1 video, or H.261. It could also be used for MPEG-2
video, or H.262. These combined video and audio files will often use the .mpg or .mpeg file extensions.
Because it was the third layer, or type, of audio codec for the MPEG standard, it used the .mp3 file extension
MP4 - MP4 is a format based on Apple’s QuickTime movie (.mov) “container” model and is similar to the MOV
format, which stores various types of media, particularly time-based streams such as audio and video. The
mp4 extension is used when the file streams audio and video together
ACC - The AAC (Advanced Audio Coding) format, which is part of the MP4 model, was adopted by Apple’s
iTunes store, and many music files are commercially available in this format.
RA — A Real Audio format designed for streaming audio over the Internet. The .ra format allows files to be
stored in a self-contained fashion on a computer, with all of the audio data contained inside the file itself.
AAC — The Advanced Audio Coding format is based on the MPEG4 audio standard owned by Dolby.
AU — The standard audio file format used by Sun, Unix and Java. The audio in au files can be PCM or
compressed with the ulaw, alaw or G729 codecs
Codec : A codec (compressor-decompressor) is software that compresses a stream of audio or video data for
storage or transmission, then decompresses it for playback. There are many codecs that do this with special
attention to the quality of music or voice after decompression. Some are “lossy” and trade quality for
significantly reduced file size and transmission speed; some are “lossless,” so original data is never altered.
Different codecs are optimized for different methods of delivery (for example, from a hard drive, from a DVD,
or over the Web). Codecs such as Theora and H.264 compress digital video information at rates that range
from 50:1 to 200: 1. Some codecs store only the image data that changes from frame to frame instead of the
data that makes up each and every individual frame. Other codecs use computation intensive methods to
predict what pixels will change from frame to frame and store the predictions to be deconstructed during
playback. These are all lossy codecs where image quality is (somewhat) sacrificed to significantly reduce file
size.
=======================================================================================
IMAGE
( Image: Formats, Image Color Scheme, Image Enhancement )
----------------------------------------------------------------------------------------------------------
---------- -
All images are represented digitally as pixels. An image is defined by image width, height, and pixel depth. The
image width gives the number of pixels that span the image horizontally and the image height gives the
number of lines in the image. Each pixel is further represented by a number of bits, which is commonly called
the pixel depth. The pixel depth is the same for all pixels of a given image. The number of bits used per pixel in
an image depends on the color space representation (gray or color) and is typically segregated into channels.
The total number of bits per pixel is, thus, the sum of the number of bits used in each channel.
Images are generated by the computer in two ways: as bitmaps and as vector-drawn graphics. Bitmaps may
also be called “raster” images.
Bitmap : A bitmap, is a simple matrix of the tiny dots that form an image and are displayed on a computer
screen. The spatial two-dimensional matrix representing an image is made up of pixels— the smallest image
resolution elements. Each pixel has a numerical value, that is, the number of bits available to code a pixel—
also called amplitude depth or pixel depth. A numerical value may represent either a black (numerical value 0)
or a white (numerical value 1) dot in bitonal (binary) images, or a level of gray in continuous-tone
monochromatic images, or the color attributes of the picture element in color pictures. BMP files are device-
independent bitmap files most frequently used in Windows systems. The BMP format is based on the RGB
color model. BMP does not compress the original image. The BMP format defines a header and a data region.
The header region (BITMAPINFO) contains information about size, color depth, color table, and compression
method. The data region contains the value of each pixel in a line. Valid color depth values are 1, 4, 8, and 24.
The BMP format uses the run-length encoding algorithm (lossless) to compress images with a color depth of
4 or 8bits/pixel.
8 bit greys : In this case each pixel takes 1 byte (8 bits) of storage resulting in 256 (0-255)different
states. If these states are mapped onto a ramp of greys from black to white, the bitmap is referred to
as a greyscale image. By convention, 0 is normally black and 255 white. So total 256 (2 8) grey shades
are available.
24 bit RGB : In color images, each R, G, B channel may be represented by 8 bits each, or 24 bits for a
pixel. Sometimes a additional fourth channel called the alpha channel is used. When the alpha channel
is present, it is represented by an additional 8 bits, bringing the total bit depth of each pixel to 32 bits.
The size of the image can, thus, vary depending on the representations used. For example, a color
image has a width of 640 and height of 480. If the R, G, B color channels are represented by 8 bits
each, the size of color image = 640 x 480 x 3 x 8 = 7.37 Mbits (921.6 Kbytes). If this were a gray image,
its size would be 640 x 480 x 8 =2.45 Mbits (307.2 Kbytes).
The color or continuous tone images are printed, printing technologies often prefer to print halftone images
where the number of colors used is minimized to lower printing costs. The halftone printing process creates
ranges of grays or colors by using variable-sized dots. The resolution of halftone images is measured in terms
of the frequency of the halftone dots, typically in dots per inch(DPI), instead of pixels. The dots control how
much ink is deposited at a specific location while printing on paper. For a process color image, four halftone
channels are used: cyan, magenta, yellow, and black (CMYK – printing color) —one for each ink used.
Alpha channel : Sometimes, an additional channel, called the alpha channel, is also used. In such cases, gray-
level images have two channels (one gray channel and one alpha channel), whereas color images have four
channels (one for R, one for G, one for B, and one for alpha). The alpha channel suggests a measure of the
transparency for that pixel value and is used in image compositing applications, such as chroma keying or blue
screen matting, and in cell animation. The alpha channel typically has the same bit depth as all the other
channels, for example, 8 bits for the alpha channel, resulting in each pixel having 32 bits (8 bits for R, G, B and
8 bits for alpha). For example, images created in digital film postproduction use 16 bits per channel, with a
total of 48 bits per pixel (or 64 if the alpha channel is present).
Aspect Ratios : Image aspect ratio refers to the width/height ratio of the images, and plays an important role
in standards. Different applications require different aspect ratios. Some of the commonly used aspect ratios
for images are 3:2 (when developing and printing photographs), 4:3 (television images), 16:9 (high-definition
images), and 47:20 (anamorphic formats used in cinemas). The ability to change image aspect ratios can
change the perceived appearance of the pixel sizes, also known as the pixel aspect ratio (PAR) or sample
aspect ratio (SAR).
Image Color
RGB Color Model
The RGB color space is a linear color space that formally uses single wavelength primaries. The different
spectra are generated by voltage gains applied to these primaries, and the resulting colors are usually
represented as a unit cube—usually called the RGB cube—whose edges represent the R, G, and B weights.
Here, red is usually shown as the x-axis, green being the y-axis, and blue being the z-axis, as in Figure 4-11. The
diagonal line, if you imagine it, connecting the black color (0,0,0) to white color (1,1,1) is made up of all gray
colors.
The RGB color is the system used in almost all color CRT monitors, and is device dependent (that is, the actual
color displayed depends on what monitor you have and what its settings are). It is called additive because the
three different primaries are added together to produce the desired color. Correspondingly, there is also a
subtractive color space known as the CMY space.
Roughly, the maximum and minimum value of a∗ correspond to red and green, while b∗ ranges from
yellow to blue. The chroma is a scale of colorfulness, with more colorful (more saturated) colors occupying the
outside of the CIELAB solid at each L∗ brightness level, and more washed-out (desaturated) colors nearer the
central achromatic axis. The CIELAB model is used by several high-end products, including Adobe Photoshop.
HSV
Along this same direction, in order to tie such perceptual concepts into camera dependent color, the HSV color
system tries to generate similar quantities. While there are many commonly used variants, HSV is by far the
most common. H stands for hue; S stands for ‘saturation’ of a color, defined by chroma divided by its
luminance— the more desaturated the color is the closer it is to gray; and V stands for “value”, meaning a
correlate of brightness as perceived by humans. The HSV color model is commonly used in image processing
and editing software.
Printing Color
CMY or CMYK Color Space
The CMY color space stands for cyan, magenta, and yellow, which are the complements of red, green, and
blue, respectively. This system is used for printing. Sometimes, this is also referred to as the CMYK color space,
when it refers to black as part of it. The CMY colors are called subtractive primaries; white is at (0,0,0) and
black is at (1,1,1). If you start with white and subtract no colors, you get white. If you start with white and
subtract all colors equally, you get black. It is important to understand why the CMY color space is used in
printing. In the CMY systems, the C, M, and Y combine subtractively to form black. While printing, to print
white in CMY is very trivial and can be obtained by setting C=M=Y=0 (that is, no pigment is printed).
Conversely, equal amounts of C, M, and Y should produce black. Producing black color is very common in
printing, and it is impractical to use all three C, M, and Y pigments each time to produce black. First, consuming
all three pigments to produce black can prove expensive. Second, while in theory black is produced this way, in
practice, the resulting dark color in neither fully black nor uniform. Hence, a fourth pigment K is commonly
used in the CMY system to produce true black. Use of K along with CMY generates a superior final printed
result with greater contrast.
VECTOR IMAGE : A vector is a line that is described by the location of its two endpoints. Vector drawing uses
Cartesian coordinates where a pair of numbers describes a point in two-dimensional space as the intersection
of horizontal and vertical lines (the x and y axes). The numbers are always listed in the order x,y. In three-
dimensional space, a third dimension—depth— is described by a z axis (x,y,z).
Vector graphics drawing code can be written into a text editor and save it as plain text with a .svg extension.
This is a Scalable Vector Graphics file. Open it in an HTML5-capable browser the drawing.
Graphics Interchange Format (GIF) :- The Graphics Interchange Format (GIF) was developed by CompuServe
Information Service in 1987. The GIF standard is limited to 8-bit (256) color images only. Three variations of
the GIF format are in use. The original specification, GIF87a, became a de facto standard because of its many
advantages over other formats. GIF images are compressed to 20 to 25 percent of their original size with no
loss in image quality using a compression algorithm called LZW(Lempel-Ziv-Welch) .The next update to the
format was the GIF89a specification. GIF89a added some useful features, including transparent GIFs and
supports simple animation via a Graphics Control Extension block in the data. Unlike the original GIF
specifications, which support only 256( 8-bit) colors, the GIF24 update supports true 24-bit colors, which
enables you to use more than 16 million colors. One drawback to using 24-bit color is that, before a 24-bit
image can be displayed on an 8-bit screen, it must be dithered, which requires processing time and may also
distort the image. GIF24 uses a compression technique called PNG.
The Tagged Image File Format (TIFF) :- It was designed by Aldus Corporation and Microsoft in 1987 to allow
portability and hardware independence for image encoding. It has become a de facto standard format. It can
save images in an almost infinite number of variations. As a result, no available image application can claim to
support all TIF/TIFF file variations, but most support a large number of variations. TIFF documents consist of
two components. The baseline part describes the properties that should support display programs. The second
part are extensions used to define properties, that is, the use of the CMYK color model to represent print
colors. An important basis to be able to exchange images is whether or not a format supports various color
models. TIFF offers binary levels, gray levels, palettes, RGB, and CMYK colors. Whether or not an application
supports the color system specified in TIFF extensions depends on the respective implementation. TIFF
supports a broad range of compression methods, including run-length encoding (which is called PackBits
compression in TIFF jargon), LZW compression, FAX Groups 3 and 4, and JPEG. In addition, various encoding
methods, including Huffman encoding, can be used to reduce the image size. TIFF differs from other image
formats in its generics. In general, the TIFF format can be used to encode graphical contents in different ways,
for example, to provide
JPEG :- The standard was created by a working group of the International Organization for Standardization
(ISO) that was informally called the Joint Photograph Experts Group. JPEG uses a loss compression which
means that image quality is lost in the process of compressing the image. The JPEG compression works by first
converting the image from RGB to YUV which stores information about each pixel using brightness, hue and
saturation. Then it reduces the amount of information it stores for hue and saturation since differences are
less noticeable to the human eye. In trying to decrease the file size of the JPEG (for example, when using the
quality slider in Photoshop), you will tend to notice artefacts occuring flat colour areas and especially near
edges. As a result, JPEG is best used for images that have more of a variation in colours. For example, images
with gradients or photographs can handle a lower quality setting with little noticeable loss in quality. Images
with text or large solid backgrounds are best left for GIF or PNG.
PNG : - The motivation for a new standard was in part the patent held by UNISYS and Compuserve on the LZW
compression method. It is similar to GIF in many ways but even better in others. It is lossless like GIF but
supports 24 bit colour, unlike GIF which only supports 8. The PNG supports alpha transparency, whereas GIF
only supports one-colour transparency. The PNG uses various compression filters to minimize overall image
size and can apply different filters on a per-line basis to achieve higher compression. The big attraction to PNGs
is its ability to do alpha transparency. Instead of a progressive display based on row-interlacing as in GIF
images, the display progressively displays pixels in a two-dimensional interlacing over seven passes through
each 8 × 8 block of an image. It supports both lossless and lossy compression with performance better than
GIF.
Windows BMP BitMap (BMP) : is one major system standard image file format for Microsoft Windows. It uses
raster graphics. BMP supports many pixel formats, including indexed color (up to 8 bits per pixel), and 16, 24,
and 32-bit color images. It makes use of Run-Length Encoding (RLE) compression (see Chap. 7) and can fairly
efficiently compress 24-bit color images due to its 24-bit RLE algorithm. BMP images can also be stored
uncompressed. In particular, the 16-bit and 32-bit color images (with α-channel information) are always
uncompressed.
PS and PDF: PostScript is an important language for typesetting, and many high-end printers have a PostScript
interpreter built into them. PostScript is a vector-based, rather than pixelbased, picture language: page
elements are essentially defined in terms of vectors. With fonts defined this way, PostScript includes
vector/structured graphics as well as text; bit-mapped images can also be included in output files.
Encapsulated PostScript files add some information for including PostScript files in another document. Several
popular graphics programs, such as Adobe Illustrator, use PostScript. Note, however, that the PostScript page
description language does not provide compression; in fact, PostScript files are just stored as ASCII. Therefore
files are often large, and in academic settings, it is common for such files to be made available only after
compression by some Unix utility, such as compress or gzip.
Therefore, another text + figures language has largely superseded PostScript in non-academic settings: Adobe
Systems Inc. includes LZW compression in its Portable Document Format (PDF) file format. As a consequence,
PDF files that do not include images have about the same compression ratio, 2:1 or 3:1, as do files compressed
with other LZW-based compression tools, such as the Unix compress or gzip, or the PC-based winzip (a variety
of pkzip) or WinRAR. For files containing images, PDF may achieve higher compression ratios by using separate
JPEG compression for the image content (depending on the tools used to create original and compressed
versions). A useful feature of the Adobe Acrobat PDF reader is that it can be configured to read documents
structured as linked elements, with clickable content and handy summary tree-structured link diagrams
provided.
=======================================================================================
Video
(Video: Analogue and Digital Video, Recording Formats and Standards (JPEG, MPEG, H.261) Transmission of
Video Signals, Video Capture)
+++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++++
Video, whether analog or digital, is represented by a sequence of discrete images shown in quick succession.
Each image in the video is called a frame, which is represented as a matrix of pixels defined by a width, height,
and pixel depth. The pixel depth is represented in a standardized color space such as RGB. These image
attributes remain constant for all the images in the length of the video. Thus, as with all images, video has the
same properties such as width, height, and aspect ratio.
Analog Video and Television:
An analog signal f (t) samples a time-varying image. So-called progressive scanning traces through a complete
picture (a frame) row-wise for each time interval. A high-resolution computer monitor typically uses a time
interval of 1/72s. In TV and in some monitors and multimedia standards, another system, interlaced scanning,
is used. Here, the odd-numbered lines are traced first, then the even-numbered lines. This results in “odd” and
“even” fields—two fields make up one frame.
In fact, the odd lines (starting from 1) end up at the middle of a line at the end of the odd field, and the even
scan starts at a half-way point. Above figure shows the scheme used. First the solid (odd) lines are traced—P to
Q, then R to S, and so on, ending at T—then the even field starts at U and ends at V. The scan lines are not
horizontal because a small voltage is applied, moving the electron beam down over time.
CRT (Cathode Ray Tube) displays are built like fluorescent lights and must flash 50–70 times per second to
appear smooth. In Europe, this fact is conveniently tied to their 50 Hz electrical system, and they use video
digitized at 25 frames per second (fps); in North America, the 60 Hz electric system dictates 30 fps. The jump
from Q to R and so on is called the horizontal retrace, during which the electronic beam in the CRT is blanked.
The jump from T to U or V to P is called the vertical retrace.
NTSC : NTSC stands for National Television Systems Committee and is the oldest and most widely used
television standard. The NTSC TV standard is mostly used in North America and Japan. It uses a familiar 4:3
aspect ratio (i.e., the ratio of picture width to height) and 525 scan lines per frame at 30 fps. More exactly, for
historical reasons NTSC uses 29.97 fps—or, in other words, 33.37 ms per frame. NTSC follows the interlaced
scanning system, and each frame is divided into two fields, with 262.5 lines/field. The “vertical retrace and
sync” and “horizontal retrace and sync” on the NTSC video raster. Blanking information is placed into 20 lines
reserved for control information at the beginning of each field. Hence, the number of active video lines per
frame is only 485. Similarly, almost 1/6 of the raster at the left side is blanked for horizontal retrace and sync.
The non blanking pixels are called active pixels. Pixels often fall between scanlines. Therefore, even with
noninterlaced scan, NTSC TV is capable of showing only about 340 (visually distinct) lines, —about 70% of the
485 specified active lines. With interlaced scan, it could be as low as 50 %. Image data is not encoded in the
blanking regions, but other information can be placed there, such as V-chip information, stereo audio channel
data, and subtitles in many languages.
PAL Video : The Phase Alternate Line (PAL) system was used in the United Kingdom, Western Europe,
Australia, South Africa, China, and South America. It uses 625 scan lines per frame, at 25 fps (or 40 ms/frame),
with a 4:3 aspect ratio and interlaced fields. Its broadcast TV signals are also used in composite video. Because
it has higher resolution than NTSC (625 vs. 525 scan lines), the visual quality of its pictures is generally better.
PAL uses the YUV color model with an 8 MHz channel, allocating a bandwidth of 5.5 MHz to Y and 1.8 MHz
each to U and V. To improve picture quality, chroma signals have alternate signs (e.g., +U and −U) in successive
scan lines; hence the name “Phase Alternating Line.”2 This facilitates the use of a (line-rate) comb filter at the
receiver—the signals in consecutive lines are averaged so as to cancel the chroma signals (which always carry
opposite signs) for separating Y and C and obtain high-quality Y signals.
SECAM Video : SECAM, which was invented by the French, is the third major broadcast TV standard. SECAM
stands for Systeme Electronique Couleur Avec Memoire. SECAM also uses 625 scan lines per frame, at 25 fps,
with a 4:3 aspect ratio and interlaced fields. The original design called for a higher number of scan lines (over
800), but the final version settled for 625. SECAM and PAL are similar, differing slightly in their color coding
scheme. In SECAM, U and V signals are modulated using separate color subcarriers at 4.25 MHz and 4.41 MHz,
respectively. They are sent in alternate lines—that is, only one of the U or V signals will be sent on each scan
line.
High-Definition Television (HDTV) : The High Definition Television (HDTV) initiative of the Federal
Communications Commission in the 1980s changed first to the Advanced Television (ATV) initiative and then
finished as the Digital Television (DTV) initiative by the time the FCC announced the change in 1996. This
standard, which was slightly modified from both the Digital Television Standard (ATSC Doc. A/53) and the
Digital Audio Compression Standard (ATSC Doc. A/52), moved U.S. television from an analog to a digital
standard. It also provided TV stations with sufficient bandwidth to present four or five Standard Television
(STV, providing the NTSC’s resolution of 525 lines with a 3:4 aspect ratio, but in a digital signal) signals or one
HDTV signal (providing 1,080 lines of resolution with a movie screen’s 16:9 aspect ratio).
The main thrust of High-Definition TV (HDTV) is not to increase the “definition” in each unit area, but rather
to increase the visual field, especially its width. HDTV provides high resolution in a 16:9 aspect ratio. This
aspect ratio allows the viewing of Cinemascope and Panavision movies. There was contention between the
broadcast and computer industries about whether to use interlacing or progressive-scan technologies . The
broadcast industry promulgated an ultra-high-resolution, 1920 × 1080 interlaced format (1080i) to become the
cornerstone of the new generation of high-end entertainment centers, but the computer industry wanted a
1280 × 720 progressive-scan system (720p) for HDTV. While the 1920 × 1080 format provides more pixels than
the 1280 × 720 standard, the refresh rates are quite different. The higherresolution interlaced format delivers
only half the picture every 1/60 of a second, and because of the interlacing, on highly detailed images there is
a great deal of screen flicker at 30 Hz. The computer people argue that the picture quality at 1280 × 720 is
superior and steady. Both formats have been included in the HDTV standard by the Advanced Television
Systems Committee (ATSC).
Analog Video : In an analog system, the output of the CCD is processed by the camera into three channels of
color information and synchronization pulses (sync) and the signals are recorded onto magnetic tape. There
are several video standards for managing analog CCD output, each dealing with the amount of separation
between the components—the more separation of the color information, the higher the quality of the image
(and the more expensive the equipment). If each channel of color information is transmitted as a separate
signal on its own conductor, the signal output is called component (separate red, green, and blue channels),
which is the preferred method for higher-quality and professional video work. Lower in quality is the signal
that makes up Separate Video (S-Video), using two channels that carry luminance and chrominance
information. The least separation (and thus the lowest quality for a video signal) is composite, when all the
signals are mixed together and carried on a single cable as a composite of the three color channels and the
sync signal. The composite signal yields less-precise color definition, which cannot be manipulated or color-
corrected as much as S-Video or component signals.
a) Component Video
High end video system uses three separate video signals for the red, green, and blue image planes. This is
referred to as component video. This kind of system has three wires (and connectors) connecting the camera
or other devices to a TV or monitor.
b) Composite Video
In composite video, color (“chrominance”) and intensity (“luminance”) signals are mixed into a single carrier
wave. Chrominance is a composite of two color components (I and Q, or U and V). This is the type of signal
used by broadcast color TV; it is downward compatible with black-and-white TV. In NTSC TV, for example, I and
Q are combined into a chroma signal, and a color subcarrier then puts the chroma signal at the higher
frequency end of the channel shared with the luminance signal. When connecting to TVs or VCRs, composite
video uses only one wire and audio signal is another addition to this one signal.
c)S-Video
As a compromise, S-video (separated video, or super-video, e.g., in S-VHS) uses two wires: one for luminance
and another for a composite chrominance signal. As a result, there is less crosstalk between the color
information and the crucial grayscale information.
Anti Alising : Antialiasing is a computer graphics method that removes the aliasing
effect. The aliasing effect occurs when rasterised images have jagged edges, sometimes called
"jaggies" (an image rendered using pixels). Technically, jagged edges are a problem that arises
when scan conversion is done with low-frequency sampling, also known as under-sampling, this
under-sampling causes distortion of the image. Moreover, when real-world objects made of
continuous, smooth curves are rasterised using pixels, aliasing occurs.
In contrast to that, some very general-purpose container types like AVI (.avi) and Quicktime (.mov) can contain
video and audio in almost any format, and have file extensions named after the container type, making it very
hard for the end user to use the file extension to derive which codec or program to use to play the files.
WMF: [Windows Media Format]. These are audio-video files comprising WMA and video codecs. They provide
high quality and media security for streaming and download and play applications on computers.
WMV: [Windows Media Video] Used in the Windows media Player, this is used to stream and download and
play audio and video content.
COMPRESSION
Compression reduces file size of media like text, image,
audio, video etc. to a process called CODEC. The
amount of compression which is to be achieved
depends on original media data and as well as the
compression technique applied.
Compression can be applied in following areas:
Data compression, the process of encoding digital
information using fewer bits
Audio compression (data), the compression of digital
audio streams and files
Bandwidth compression, a reduction in either the time
to transmit or in the amount of bandwidth required to transmit.
Compression artifact, noticeable defects in audio or video that has been compressed.
Image compression, the application of data compression on digital images.
Video compression, the compression of digital video streams and files.
Dynamic range compression, a compression process that reduces the dynamic range of an audio signal.
TYPES OF COMPRESSION
LOSSLESS V/S LOSSY COMPRESSION
Compression can be broadly divided into two types: lossy and lossless. Lossy compression implies the original
data changed permanently during compression unlike lossless compression. CODECS are used for lossless
compression to represent the existing information in a more compact form without actually discarding any
data.
The advantage of lossless compression is the original data stays intact without degradation of quality. E.g.:-
medical images like X-ray plates and ultra sonograph.
In lossy compression, parts of original data are completely discarded. Lossy compression is generally used
where media quality may be sacrificed to a certain extent for reducing space requirements like in multimedia
presentation and web page content.
Interframe compression exploits the redundancy between adjacent frames in a video sequence, referred to as
temporal redundancy.
Huffman Coding:
Huffman coding algorithm determines the optimal coding using the minimum number of bits. Hence, the
length (number of bits) of the coded characters will differ. The most frequently occurring characters are
assigned to the shortest code words. A Huffman code can be determined by successively constructing a binary
tree, whereby the leaves represent the characters that are to be encoded. Every node contains the relative
probability of occurrence of the characters belonging to the subtree beneath the node. The left edges are
labeled with 0 and right branches are labelled with 1.
Huffman Algorthim :
1. Initialization: put all symbols on the list sorted according to their frequency counts.
2. Repeat until the list has only one symbol left.
(a) From the list, pick two symbols with the lowest frequency counts. Form a Huffman subtree that has these
two symbols as child nodes and create a parent node for them.
(b) Assign the sum of the children’s frequency counts to the parent and insert it into the list, such that the
order is maintained.
(c) Delete the children from the list.
3. Assign a codeword for each leaf based on the path from the root
Example :
The letters A, B, C, D, and E are to be encoded and have relative probabilities of occurrence as follows:
p(A)=0.16, p(B)=0.51, p(C)=0.09, p(D)=0.13, p(E)=0.11.
a) The two characters with the lowest probabilities, C and E, are combined in the first binary tree, which has
the characters as leaves. The combined probability of their root node CE is 0.20. The edge from node CE to C is
assigned a 0 and the edge from CE to E is assigned a 0.
---------------------------------------------------------------------------------------------------------------------------------
JPEG Compression Steps
Step 1: Image Block Preparation
Step 2: Discrete Cosine Transform
Step 3: Quantization
Step 4: Run-Length Encoding
Step 5: Statistical Encoding
Step 1: Image Block Preparation The source image consists of a matrix of pixels. Each pixel is represented by
three color components (such as RGB or YUV). The JPEG compression technique works on each color
component independently. The three color components are first separated into three color planes. This
produces three matrices of pixels, one for each color component. Each color plane is divided into blocks of 8 x
8 pixels. For example, a 640 x 480 image will end up with:
(640/8) x (480/8) x 3 = 80 x 60 x 3 = 14,400 blocks
The remaining operation is performed block by block. The following explanation is given for the operations
performed on one block. The same sequence of operations is repeated for all the blocks in the image.
Step 2: Discrete Cosine Transform The second processing step is the Discrete Cosine Transform (DCT). The DCT
converts pixel amplitudes in the spatial domain into DCT coefficients in the frequency domain. The formula for
converting spatial domain signals into frequency domain signals is called the Forward Discrete Cosine
Transform (FDCT).
Step 3: Quantization The quantization step is included to give different levels of importance to the different
frequency components. The level of importance given to a frequency component is decided by the
quantization table. The DCT coefficient value in a specific position in the coefficient matrix is divided by the
corresponding entry in the quantization table. Thus, DCT coefficients that have a 1 entry in the quantization
table get the highest importance. It can be seen that the high-frequency coefficients have been given high
values in the quantization table. Since the values in the quantized table are all integers, any DCT coefficient
that is smaller than the quantization table entry will be reduced to zero. These zero values are further
compressed in the run-length encoding step. The largest DCT coefficient in each block is the DC coefficient; i.e.,
the average of the pixel values. To reduce the space required to store the DC coefficients, these are stored by
using DPCM. That is, the value of the DC coefficient of block n + 1 is stored relative to the value of the DC
coefficient of block n. The DC coefficient of block n is assumed to be the predicted value of the block n + 1 DC
coefficient. Therefore, only the prediction error must be stored, in place of the full DC coefficient of block n +
1.
Step 4: Run-Length Encoding The size of the memory required to store the quantized DCT coefficients can be
reduced by using run-length encoding. The DCT coefficients are stored in a zig-zag manner. This zig-zag
sequencing increases the chance of the same amplitude coefficients appearing next to each other. The high-
frequency components are the most likely to give large sequences of zeros or other low values.
Step 5: Statistical Encoding Run-length encoding is followed by a statistical encoding step. The JPEG standard
allows a choice of two statistical encoding methods: Huffman encoding and arithmetic encoding. Huffman
encoding produces variable-length codes that require a code-book. As explained earlier, Huffman encoding
does not allow a non integer number of bits to be used for representing input values. It therefore produces
less than ideal entropy encoding. Arithmetic encoding makes it possible to use a non integer number of bits to
represent input values, though it requires more processing. Thus, the choice between Huffman encoding and
arithmetic encoding must be made as a compromise between the “goodness” of encoding and the processing
power required to achieve it.
The last two steps use entropy encoding techniques; i.e., these steps produce lossless compression.
Information loss occurs mainly in the quantization step. Some loss may occur in the DCT step also because of
the finite word length of the digital storage system used. The JPEG standard can produce compression ratios
on the order of 20:1 for lossy compression and 2:1 for lossless compression.
------------------------------------------------------------
MPEG Encoding
Spatial and Temporal Compression : A moving picture varies in spatial as well as temporal domains. The
spatial domain variations occur in the image components (e.g., YUV values) when a single frame is scanned.
Temporal variance occurs when a sequence of image frames is captured and processed as time advances.
Below Figure shows a block diagrammatic representation of moving image compression. This figure highlights
the fact that moving image compression consists of two main aspects: intraframe and interframe
compression. Intraframe compression relates to the removal of spatial redundancies within individual
frames. Interframe compression relates to the removal of temporal redundancies that exist between
consecutive frames. Interframe compression is also called motion compensation.
Spatial Compression : The spatial compression aspects of the MPEG standard are similar to those of the JPEG
standard. Each frame is taken as an independent image and compressed by using the steps described
previously for the JPEG standard. A moving image compression method called Motion-JPEG (M-JPEG) is also
available. The M-JPEG method uses only spatial compression. The output of a video stream compressed by M-
JPEG is nothing but a sequence of frames, each compressed independently with the JPEG standard. The M-
JPEG method does not use the fact that there are common elements between consecutive frames. The
advantage is its simplicity; the disadvantage is the low compression ratio, compared to the compression ratio
that can be achieved by removing the temporal redundancies.
Temporal Compression : The MPEG compression procedure removes spatial as well as temporal redundancies
in a sequence of
consecutive frames.
Examples of temporal
redundancy are shown in
Figure 1 Figure 1(a) shows
three consecutive frames of
a shot in which the camera
is panning from left to right.
As this panning proceeds,
building-1 disappears to the
left of the frame and
building-4 appears on the
right side. A very small part
of the moon appears in the
first frame; it appears more
fully in the second and third
frames. We can identify the
picture elements that are
common between adjacent frames. The area marked with down sloping hashed lines is common between the
first 1(b). In this example, a car is moving into the frame from the right-hand side. The only variation from one
frame to the next is that more of the car becomes visible. The level of temporal redundancy is high in the two
examples described here. Removal of this temporal redundancy is achieved and second frames, and the area
marked with upsloping hashed lines is replicated from frame-2 to frame-3. It would be wasteful to retransmit
replicated image areas for every frame. These replicated image areas constitute the temporal redundancy in
the shot. Another example of temporal redundancy is shown in Figure through interframe encoding.
Interframe Encoding
The basic idea in interframe encoding is to save only the changes in the moving image from one frame to the
next, rather than saving the entire image data. This is done by using reference frames for coding a sequence of
frames.
Reference Frame A reference frame is an image frame that is used for coding and decoding one or more of the
other frames. An encoded reference frame often contains enough information to be decoded without
reference to any other frame.
I-frame : A frame that can be decoded without reference to any other frame is called an intra coded frame, or
I-frame. I-frames are usually used as reference frames, though every reference frame need not be an I-frame.
P-and B-frames: Interframe coding uses predictive coding. Data taken from a reference frame are used as the
predicted values to which error terms are added to regenerate the current frame data. Two prediction
techniques are used: forward prediction and bidirectional prediction. In forward predictive coding, a frame is
coded with reference to a past frame. In bidirectional coding, also called motion-compensated interpolation,
the current frame is coded with respect to past and future frames. A frame that uses only past frame(s) as
reference frame(s) is called a predicted frame (Pframe). A frame that is coded by referring to past as well as
future frames is called a bi-directional frame (B-frame).
Repetitive Frame Sequences Various repetitive frame sequences are used for motion picture encoding. Some
of the commonly used repetitive frame sequences are shown in Figure 2. The IBBBPBBB IBBBP...sequence is a
general frame sequence, the IBBPBBPBB IBBP...sequence is used for encoding PAL and SECAM video, and the
IBBPBBPBBPBB IBBP...sequence is used for encoding NTSC video.
The I-frames take up the most memory, because they are compressed just like still images, with the JPEG
standard. The P-frames can be better compressed than the I-frames, because their contents are described in
terms of the reference I-frames. The B-frames are the most well compressed frames, because their contents
are described in terms of reference I- and P-frames in both directions.
It may seem that by using more B-frames between reference I- and P-frames, the compression ratio can be
increased. Not always, because a B-frame will compress well only if it has good correlation with its reference
frames. Thus, if too many B-frames are interleaved between the reference frames, the expected benefit may
not occur due to reduced correlation between the B-frame being compressed and the I- and P-frames being
used as the reference frames.
Another important point is the order in which the frames should be transmitted over a network. As an I-frame
is used for decoding the next P-frame, the I-frame must be transmitted first. The P-frame is used for decoding
the preceding B-frames; thus, the P-frame must be transmitted before the B-frames.
Macroblock: The atomic object used for motion compensation operations is called a macroblock. A
macroblock is quite different from the 8 x 8 block used for spatial compression in the JPEG and MPEG
standards. Macroblocks are used for
motion compensation. A
macroblock covers a 16 x 16 pixel
area. It consists of four 8 x 8 blocks
of luminance (Y) values and two 8
x 8 blocks of subsampled color
difference (U and V) values.
Motion Vector : One of the main techniques used in motion compensation is to identify macroblocks whose
contents do not change from one frame to the other; only their position within the frame changes. A motion
vector (MV) describes the change in the position of a macroblock. This concept is shown in the above Figure.
The macroblock covering the lampshade in frame-1 is called MB x(f1). The same macroblock in frame-2 is called
MBx(f2). A motion vector MV2 describes the movement of this macroblock. The motion vector operation is
indicated by a motion operator represented by the symbol ->. The translation of the macroblock MB x from its
original position in frame-1 to its position in frame-2 can be described by the following formula:
MBx(f2) = MVx —> MBx(f2)
the macroblock MBx in frame f2 is nothing but the macroblock MBx in frame f2 translated by the motion vector
MVx. There are three main techniques for motion estimation: block matching, gradient matching, and phase
correlation.
Block Matching The motion vector is determined by taking a macroblock in the current frame and locating a
matching macroblock in the next frame. For the block matching process, various search algorithms are
possible, such as exhaustive search, three-step search, and logarithmic search . The simplest search algorithm
is exhaustive search, but it is also the most computeintensive search. In this algorithm, first a search area is
chosen. In Figure 2 the original macroblock is a block of m x n pixels in frame-1. The search area is chosen to be
of (m + 2p) x (n + 2p) pixels in frame-2. For a typical search process in MPEG, m= n= 16, and p = 6. In
exhaustive search, the reference
macroblock is panned through the search
area. At each step, the level of match
between the reference macroblock and
the test block is calculated by using a cost
function. The various cost functions that
have been proposed include Mean-
Absolute Difference, Mean-Squared
Difference, and Cross Correlation
Function.
Exhaustive Search : Exhaustive search is done by panning the reference block in a search area, as illustrated in
Figure 14.13. In the first step, the reference macroblock is kept to the top left corner of the search area. Then
the reference block is panned to the right one pixel at a time, giving a total of thirteen steps when the
reference block is at the top of the search area. In the next search panning, the reference macroblock is moved
down by one pixel, giving thirteen more search steps. There are a total of thirteen left-to-right panning
operations of the reference block. It can be seen that the total number of search steps will be (2p + 1) 2 = 169.
The chosen cost function is calculated at each step. The step giving the min- imum value for the cost function is
taken as the new position for the macroblock under consideration. The motion vector can be calculated from
the original and the new positions of the macroblock. For motion pictures of fast-moving objects, such as
sports events, the value of p may have to be increased to ensure that the reference macroblock has not moved
out of the search area. Because the number of search steps increases to the square of the value of Pp, the
exhaustive search algorithm can become too slow, especially for real-time operation.
Motion Compensation: The motion compensation operation uses macroblocks and motion vectors. If an exact
match is found between a reference macroblock and a matching block in the search area, only the motion
vector is transmitted. If the match between the matching block and the reference macroblock is not exact,
then error terms are transmitted along with the motion vector. The motion vectors are represented by the
DPCM technique and subjected to run-length and Huffman coding.
b) Frame in rate in Digital television in between 24, 30, or 60 entire frames or pictures every second,
depending upon settings; the speed with which each frame is replaced by the next one makes the images
appear to blend smoothly into movement. Movies on film are typically shot at a shutter rate of 24 frames per
second.
c) Quickly changing the viewed image is the principle of an animatic, a flip-book, or a zoetrope. To make an
object travel across the screen while it changes its shape, just change the shape and also move, or translate, it
a few pixels for each frame. Then, when you play the frames back at a faster speed, the changes blend
together and you have motion and animation.
Path animation : Path animation in 2-D space increases the complexity of an animation and provides motion,
changing the location of an image along a predetermined path (position) during a specified amount of time
(speed).
Animation Techniques
Cel Animation : The term cel derives from the clear celluloid sheets that were used for drawing each frame,
which have been replaced today by layers of digital imagery. Cels of famous animated cartoons have become
sought-after, suitable-for-framing collector’s items. Cel animation artwork begins with keyframes (the first and
last frame of an action). A minute of animation may thus require as many as 1,440 separate frames, and each
frame may be composed of many layers of cels.
Tweening : The series of frames in between the keyframes are drawn in a process called tweening. Tweening is
an action that requires calculating the number of frames between keyframes and the path the action takes,
and then actually sketching with pencil the series of progressively different outlines. As tweening progresses,
the action sequence is checked by flipping through the frames.
Kinematics : Kinematics is the study of the movement and motion of structures that have joints, such as a
walking man. Animating a walking step is tricky: you need to calculate the position, rotation, velocity, and
acceleration of all the joints and articulated parts involved—knees bend, hips flex, shoulders swing, and the
head bobs.
Morphing: Morphing is a popular (if not overused) effect in which one image transforms into another.
Morphing applications and other modeling tools that offer this effect can transition not only between still
images but often between moving images as well.
VRML : The Virtual Reality Modeling Language (VRML) is a format for describing three dimensional interactive
worlds and objects that can be used together with the World Wide Web. For example, VRML can be used to
generate three-dimensional representations of complex scenes such as illustrations, product definitions, or
Virtual Reality presentations. VRML is capable of representing static and animated objects as well as hyperlinks
to other media such as sound, motion pictures, and still pictures. Interpreters (browsers) for VRML are widely
available for many different platforms, as are authoring tools for generating VRML files. VRML is a model that
allows for the definition of new objects as well a registration process that makes it possible for application
developers to define common extensions to the base standard. Furthermore, there are mappings between
VRML elements and generally used features of 3-D Application Programmer Interfaces (API).
-------------------------------------------------------------------------------------------------------------------
Synchronization, Storage models and Access Techniques
Synchronisation:
The word synchronization refers to time. Synchronization in multimedia systems refers to the temporal
relations between media objects in the multimedia system. In a more general and widely used sense some
authors use synchronization in multimedia systems as comprising content, spatial and temporal relations
between media objects. We differentiate between time-dependent media object are equal, it is called
continuous media object. A video consists of a number of ordered frames; each of these frames has fixed
presentation duration. A time-independent media object is any kind of traditional media like text and images.
The semantic of the respective content does not depend upon a presentation according to the time domain.
Synchronization between media objects comprises relations between time dependent media objects and time-
independent media objects
For each layer, typical objects and operations on these objects are described in the following. The semantics of
the objects and operations are the main criteria for assigning them to one of the layers.
Media Layer: At the media layer, an application operated on a single continuous media stream, is treated as a
sequence of LDUs.
Stream Layer: The stream layer operates on continuous media streams, as well as on groups of media streams.
In a group, all streams are presented in parallel by using mechanisms for interstream synchronization.
The abstraction offered by the stream layer is the notion of streams with timing parameters concerning the
QoS for intrastream synchronization in a stream and interstream synchronization between streams of a group.
Continuous media is seen in the stream layer as a data flow with implicit time constraints; individual LDUs are
not visible. The streams are executed in a Real-Time Environment (RTE), where all processing is constrained by
well-defined time specifications.
Object Layer: The object layer operates on all types of media and hides the differences between discrete and
continuous media.
The abstraction offered to the application is that of a complete, synchronized presentation. This layer takes a
synchronization specification as input and is responsible for the correct schedule of the overall presentation.
The task of this layer is to close the gap between the needs for the execution of a synchronized presentation
and the stream-oriented services. The functions located at the object layer are to compute and execute
complete presentation schedules that include the presentation of the no-continuous media objects and the
calls to the stream layer.
Specification Layer: The specification layer is an open layer. It does not offer an explicit interface. This layer
contains applications and tools that are allowed to create synchronization specifications. Such tools are
synchronization editors, multimedia documents editors and authoring systems. Also located at the
specification layer are tools for converting specifications to an object layer format. The specification layer is
also responsible for mapping QoS requirements of the user level to the qualities offered at the object layer
interface.
Synchronization specification methods can be classified into the following main categories:
Interval-based specifications, which allow the specification of temporal relations between the time intervals
of the presentations of media objects.
Axes-based specifications, which relate presentation events to axes that are shared by the objects of the
presentation.
Control flow-based specifications, in which at given synchronization points, the flow of the presentations is
synchronized.
Event-based specifications, in which events in the presentation of media trigger presentation actions.
Difference between CLV and CAV
CLV CAV
CD data is recorded in Constant Linear velocity. Constant angular velocity (CAV) is a qualifier for
The disk must spin more quickly when inner track the rated speed of any disc containing
area is read than when the outer track area is information, and may also be applied to the
read. writing speed of recordable discs. A drive or disc
operating in CAV mode maintains a
constant angular velocity.
A typical CD-ROM drive operates in CLV mode Generally a floppy or hard disk drive, or
gramophone, which operates in CAV mode
CLV mode, the spindle motor speed varies so that In contrast, in CLV mode, the spindle motor
the medium passes by the head at the same speed varies so that the medium passes by the
speed regardless of where on the disk the head is head at the same speed regardless of where on
positioned. the disk the head is positioned.
The data rate is higher for the outer tracks than If the disk is recorded at the same areal density
for the inner tracks, whereas in CLV mode, the throughout, then when read or written in CAV
data rate is the same everywhere. mode,
Another advantage is that a device can switch A device can switch from reading one part of a
from reading one part of a disk to reading disk to reading another part quite slow.
another part more quickly,
Digital Camera-Based System Technology for end-to-end digital video processing is also available. A digital
camera is used to capture the images. Digital cameras use Charge Coupled Device (CCD) arrays to capture the
image, and a built-in ADC to convert the signal to digital format. The digital image can be stored on a magnetic
disk and/or transferred to a computer via an interface circuit, as shown in
The light now is parallel to the second filter and comes out wholly
through it to the eye of the observer. This constitutes a lighted pixel
on the screen. A battery connected across the liquid crystal container
generates a current through the liquid crystal, and re-orients its
molecules according to the direction of the current flow. This
distributes the orderly pattern of the liquid crystal molecules; so that
now the molecules at the grooved surfaces are no longer emerge. An
observer on the other side of the filter does not see any light coming
out. This arrangement creates a dark pixel
CD-ROM
CD-ROM is an excellent method for distributing multimedia projects. A CD may contain one or more tracks.
These are areas normally allocated for storing a single song in the Red Book format. Both Macintosh and
Windows support commands to access both Red Book Audio and the data tracks on a CD, but you cannot
access both at the same time. Though a CD contains tracks, the primary logical unit for data storage on a CD is
a sector, which is 1/75 second in length. Each sector of a CD contains 2,352 bytes of data
Due to the Constant Linear Velocity (CLV) playback of a CD, the rotational velocity at single speed is about
530 revolutions per second on the inside, but only about 200 revolutions per second on the outside. The
rotation delay describes the time it takes to find the desired sector within a maximum of one rotation and to
correctly set the rotation speed. Depending on the device, this time can be about 300 ms. For a CD-ROM
device with a real 40-time data transfer rate and about 9,000 revolutions per minute, the maximum rotation
delay is about 6.3ms. This translates to about 1.3 meters (51 inches) of travel along the data track each
second. CD players use very sensitive motors so that no matter where the read head is on the disc,
approximately the same amount of data is read in each second.
The seek time refers to the adjustment to the exact radius, whereby the laser must first find the spiral track
and adjust itself. The seek time frequently amounts to about 100ms.
The CD’s rotational speed and the density of the pits and lands on the CD allow data to be read at a sustained
rate of 150 Kbps in a single-speed reader. This is sufficient for good audio, but it is very slow for large image
files, motion video, and other multimedia resources, especially when compared to the high data-transfer rates
of hard disk drives. New drives that spin many times faster when reading computer data, and slower for Red
Book Audio, have been designed specifically for computers. In any case, CD access speed and transfer rate
from CD-ROM is much slower than from a hard disk.
DVD :
The Digital Versatile Disc (DVD) is, particularly in view of its larger storage space. A capacity of a double-layer
DVD is less than that of a double-sided DVD because the crosstalk that occurs when reading through the outer
layer must be reduced. DVDs achieve a higher capacity than CD-ROMs by using smaller pits (which yields a
higher track density), combined with a larger data area, more efficient coding of bits, more efficient error
correction and lower sector overhead. With Dolby AC-3 Digital Surround Sound as part of the DVD
specifications, six discrete audio channels can be programmed for digital surround sound, and with a separate
subwoofer channel, devel
From the standpoint of information technology, a DVD consists of a number of blocks of 37,856 bytes each.
Each block contains 16 sectors plus additional data for error detection and correction. Individual sectors
consist of 2,064 bytes divided into 12 rows as shown in Table 8-8. The first 12 bytes in the first row contain the
sector header (sector ID, sector ID error correction, and six reserved bytes). The rest of the block, except the
last four bytes, which contain the error detection code, holds user data.
In order to transfer parallel streams better, DVD interleaves 16 sectors together. This also provides better
robustness against errors. The result of the interleaving is a block of 192 rows (16 sectors × 12 rows per sector
= 192 rows). At the end of each row ten bytes are added for further error correction, resulting in an additional
16 rows at the end of each block. Thus, only 33,024 bytes of each 37,856-byte block are available for user data,
yielding a payload of only 87 percent.
SCANNER :
Principle: For images, digitization involves physical devices like the scanner or digital camera. The scanner is a
device used to convert analoge images into the digital form. The most common type of scanner for the office
environment is called the flatbed scanner.
The traditional way of attaching a scanner to the computer is through an interface cable connected to the
parallel port of the PC
To start a scanning operation, the paper document to be scanned is placed face down on the glass panel of the
scanner, and the scanner is activated using a software from the host computer. The light on getting reflected
by the paper image is made to fall on a grid of electronic sensors, by an arrangement of mirrors and lenses.
The electronic sensors are called Charge Coupled Devices (CCD) are converters of the light energy into voltage
pulses. After a complete scan, the image is converted from a continuous entity into a discrete form
represented by a series of voltage pulses. This process is called sampling. The voltage signals are temporarily
stored in a buffer inside the scanner. The next step called quantization involves representing the voltage pulses
as binary numbers and carried out by an ADC inside the scanner in conjunction with a software bundled with
the scanner called the scanning software. Since each number has been derived from the intensity of the
incident light, these essentially represent brightness values at different points of the image and are known as
pixels.
OCR : Optical Character Recognition (OCR). The OCR software traditionally works by a method called pattern
matching. Recent research on OCR is based on another technology called feature extraction. Using this
method, the software attempts to extract the core features of the characters and compare them to a table
stored within itself for recognition.
----------------------------------------------------------------------------------------------------------------
KD Tree
A K-D Tree(also called as K-Dimensional Tree) is a binary search tree where data in each node is a K-
Dimensional point in space. In short, it is a space partitioning(details below) data structure for organizing
points in a K-Dimensional space. A non-leaf node in K-D tree divides the space into two parts, called as half-
spaces. Points to the left of this space are represented by the left subtree of that node and points to the right
of the space are represented by the right subtree. We will soon be explaining the concept on how the space is
divided and tree is formed. For the sake of simplicity, let us understand a 2-D Tree with an example. The root
would have an x-aligned plane, the root’s children would both have y-aligned planes, the root’s grandchildren
would all have x-aligned planes, and the root’s great-grandchildren would all have y-aligned planes and so on.
How to determine if a point will lie in the left subtree or in right subtree? If the root node is aligned in plane
A, then the left subtree will contain all points whose coordinates in that plane are smaller than that of root
node. Similarly, the right subtree will contain all points whose coordinates in that plane are greater-equal to
that of root node.
Creation of a 2-D Tree: Consider following points in a 2-D plane: (3, 6), (17, 15), (13, 15), (6, 12), (9, 1), (2, 7),
(10, 19)
1. Insert (3, 6): Since tree is empty, make it the root node.
2. Insert (17, 15): Compare it with root node point. Since root node is X-aligned, the X-coordinate
value will be compared to determine if it lies in the right subtree or in the left subtree. This point
will be Y-aligned.
3. Insert (13, 15): X-value of this point is greater than X-value of point in root node. So, this will lie in
the right subtree of (3, 6). Again Compare Y-value of this point with the Y-value of point (17, 15)
(Why?). Since, they are equal, this point will lie in the right subtree of (17, 15). This point will be X-
aligned.
4. Insert (6, 12): X-value of this point is greater than X-value of point in root node. So, this will lie in
the right subtree of (3, 6). Again Compare Y-
value of this point with the Y-value of
point (17, 15) (Why?). Since, 12 < 15, this
point will lie in the left subtree of (17, 15).
This point will be X-aligned.
5. Insert (9, 1):Similarly, this point will lie in
the right of (6, 12).
6. Insert (2, 7):Similarly, this point will lie in
the left of (3, 6).
7. Insert (10, 19): Similarly, this point will lie in
the left of (13, 15).
All 7 points will be plotted in the X-Y plane as follows:
1. Point (3, 6) will divide the space into two parts: Draw line X = 3
2. Point (2, 7) will divide the space to the left of line X = 3 into two parts horizontally. Draw line Y = 7 to
the left of line X = 3
3. Point (17, 15) will divide the space to the right of line X = 3 into two parts horizontally. Draw line Y = 15
to the right of line X = 3
4. Point (6, 12) will divide the space below line Y = 15 and to the right of line X = 3 into two parts. Draw
line X = 6 to the right of line X = 3 and below line Y = 15.
5. Point (13, 15) will divide the space below line Y = 15 and to the right of line X = 6 into two parts. Draw
line X = 13 to the right of line X = 6 and below line Y = 15.
6. Point (9, 1) will divide the space between lines X = 3, X = 6 and Y = 15 into two parts. Draw line Y = 1
between lines X = 3 and X = 13.
7. Point (10, 19) will divide the space to the right of line X = 3 and above line Y = 15 into two parts. Draw
line Y = 19 to the right of line X = 3 and above line Y = 15.
--------------------------
Quadtrees
Quadtrees are a knowledge structure that encodes a two-dimensional space into adaptable cells. Similar
to binary trees, quadtrees are a tree structure where every non-leaf node has four children. In the Quadtree
data structure, every internal node has four children. They are usually a bi-dimensional analogue of octrees
that are used for the division of a two-dimensional space by repetitively subdividing it into four quadrants.
NorthWest
NorthEast
SouthWest
SouthEast
During the technique of quadtrees, nodes are recursively divided into pieces, with successive subdivisions
leading to smaller and smaller cells.
Quadtrees are used in image compression, where each node contains the average colour of each of its
children. The deeper you traverse in the tree, the more the detail of the image. Quadtrees are also used in
searching for nodes in a two-dimensional area. For instance, if you wanted to find the closest point to given
coordinates, you can do it using quadtrees.
Insert Function:
The insert functions is used to insert a node into an existing Quad Tree. This function first checks whether the
given node is within the boundaries of the current quad. If it is not, then we immediately cease the insertion. If
it is within the boundaries, we select the appropriate child to contain this node based on its location. This
function is O(Log N) where N is the size of distance.
Search Function:
The search function is used to locate a node in the given quad. It can also be modified to return the closest
node to the given point. This function is implemented by taking the given point, comparing with the
boundaries of the child quads and recursing. This function is O(Log N) where N is size of distance.
-----------------------
Quality of Service
QoS for multimedia data transmission depends on many parameters. We now list the most important ones as
below:
• Bandwidth A measure of transmission speed over digital links or networks, often in kilobits per second (kbps)
or megabits per second (Mbps) As shown before, the data rate of a multimedia stream can vary dramatically,
and both the average and the peak rates should be considered when planning for bandwidth for transmission.
• Latency (maximum frame/packet delay) The maximum time needed from transmission to reception, often
measured in milliseconds (msec, or ms). In voice communication, for example, when the round-trip delay
exceeds 50 msec, echo becomes a noticeable problem; when the one-way delay is longer than 250 msec,
talker overlap would occur, since each caller will talk without knowing the other is also talking.
• Packet loss or error A measure (in percentage) of the loss- or error rate of the packetized data transmission.
The packets can get lost due to network congestion or garbled during transmission over the physical links.
They may also be delivered late or in the wrong order. For real-time multimedia, retransmission is often
undesirable, and therefore alternative solutions like forward error correction (FEC), interleaving, or error-
resilient coding are to be used.
• Jitter (or delay jitter) A measure of smoothness (along time axis) of the audio/video playback. Technically,
jitter is related to the variance of frame/packet [Link] examples of high and low jitters in frame
playbacks. A large buffer (jitter buffer) can be used to hold enough frames to allow the frame with the longest
delay to arrive, so as to reduce playback jitter. However, this increases the latency and may not be desirable in
real-time and interactive applications.
• Sync skew A measure of multimedia data synchronization, often measured in milliseconds (msec). For a
good lip synchronization, the limit of sync skew is ±80 msec between audio and video. In general, ±200 msec is
still acceptable. For a video with voice the limit of sync skew is 120 msec if video precedes voice and 20 msec if
voice precedes video. The discrepancy is because we are used to have sound lagging image at a distance.
Virage : The Visual Information Retrieval (VIR) image search engine operates on objects within images. Image
indexing is performed after several preprocessing operations, such as smoothing and contrast enhancement.
The details of the feature vector are proprietary; however, it is known that the computation of each feature is
made by not one but several methods, with a composite feature vector composed of the concatenation of
these individual computations.
Multimedia Data Base Management Systems (MDBS)
MDBMS is embedded in the multimedia system domain, located between the application domain
(applications, documents) and the device domain (storage, compression and computer technology). The
MDBMS is integrated into the system domain through the operating system and communication components.
Characteristics of an MDBMS
An MDBMS can be characterized by its objectives when handling multimedia data:
1. Corresponding Storage : Media Multimedia data must be stored and managed according to the
specific characteristics of the available storage media. Here, the storage media can be both computer
integrated components and external devices. Additionally, readonly (such as a CD-ROM), write-once
and write-many storage media can be used.
2. Descriptive Search Methods : During a search in a database, an entry, given in the form of text or a
graphical image, is found using different queries and the corresponding search methods. A query of
multimedia data should be based on a descriptive, content oriented search in the form, for example,
of “The picture of the woman with a red scarf”. This kind search of relates to all media, including
video and audio.
4. Format-independent Interface : Database queries should be independent from the underlying media
format, meaning that the interfaces should be format-independent. The programming itself should
also be format-independent, although in some cases, it should be possible to access details of the
concrete formats.
5. View-specific and Simultaneous : Data Access The same multimedia data can be accessed (even
simultaneously) through different queries by several applications. Hence, consistent access to shared
data (e.g., shared editing of a multimedia document among several users) can be implemented.
6. Management of Large Amounts of Data : The DBMS must be capable of handling and managing large
amounts of data and satisfying queries for individual relations among data or attributes of relations.
7. Relational Consistency of Data Management : Relations among data of one or different media must
stay consistent corresponding to their specification. The MDBMS manages these relations and can use
them for queries and data output. Therefore, for example, navigation through a document is
supported by managing relations among individual parts of a document.
8. Real-time Data Transfer : The read and write operations of continuous data must be done in real-
time. The data transfer of continuous data has a higher priority than other database management
actions. Hence, the primitives of a multimedia operating system should be used to support the real-
time transfer of continuous data.
9. Long Transactions :The performance of a transaction in a MDBMS means that transfer of a large
amount of data will take a long time and must be done in a reliable fashion. An example of a long
transaction is the retrieval of a movie.
In the architecture model, the system components around MDBMS and MDBMS itself have the following
functions:
The operating system provides the management interface for MDBMS to all local devices.
The MDBMS provides an abstraction of the stored data and their equivalent devices, as is the case in
DBMS without multimedia.
The communication system provides for MDBMS abstractions for communication with entities at
remote computers. These communication abstractions are specified through interfaces according to,
for example, the Open System Interconnection (OSI) architecture.
A layer above the DBMS, operating system and communication system can unify all these different
abstractions and offer them, for example, in an object-oriented environment such as a toolkit. Thus, an
application should have access to each abstraction at different levels.
<html>
<head>
……
</head>
<body>
</body>
</html>
The HEAD describes document definitions, which are parsed before any document rendering is done. These
include page title, resource links, and meta-information the author decides to specify. The BODY part describes
the document structure and content. Common structure elements are paragraphs, tables, forms, links, item
lists, and buttons.
A very simple HTML page is as follows:
<html>
<head>
<title> Title Page </title>
</head>
<body>
<p> This is a simple html page </p>
</body>
</html>
Naturally, HTML has more complex structures and can be mixed with other standards. The standard has
evolved to allow integration with script languages, dynamic manipulation of almost all elements and
properties after display on the client side (dynamic HTML), and modular customization of all rendering
parameters using a markup language called Cascading Style Sheets (CSS). Nonetheless, HTML has rigid, non
descriptive structure elements, and modularity is hard to achieve
---------------------------------------------------------
SGML - The Standard Generalized Markup Language (SGML) was supported mostly by American publisher.
Authors prepare the text, i.e., the content. They specify in a uniform way the title, tables, etc., without a
description of the actual representation. The publisher specifies the resulting layout. The basic idea is that the
author uses tags for marking certain text parts. SGML determines the form of tags. But it does not specify their
location or meaning. User groups agree on the meaning of the tags. SGML makes a frame available with which
the user specifies the syntax description in an object-specific system. Here, classes and objects, hierarchies of
classes and objects, inheritance and the link to methods (processing instructions) can be used by the
specification. SGML specifies the syntax, but not the semantics.
For example,
This example shows an application of SGML in a text
document. The following figure shows the
processing of an SGML document. It is divided into
two processes:
Only the formatter knows the meaning of the tag
and it transforms the document into a formatted
document. The parser uses the tags, occurring in the
document, in combination with the corresponding
document type. Specification of the document
structure is done with tags. Here, parts of the layout
are linked together. This is based on the joint
context between the originator of the document and
the formatter process. It is one defined through
SGML.
Multimedia data are supported in the SGML standard only in the form of graphics. A graphical image as a CGM
(Computer Graphics Metafile) is embedded in an SGML document. The standard does not refer to other
media.
A link to concrete data can be specified through #NDATA. The data are stored mostly externally in a separate
file. The above example shows the definition of video which consists of audio and motion pictures. Multimedia
information units must be presented properly. The synchronization between the components is very
important here.
--------------------------------------------------
Details of ODA
The main property of ODA is the distinction among content, logical structure and layout structure. This is in
contrast to SGML where only a logical structure and the contents are defined. ODA also defines semantics.
Following figure shows these three aspects linked to a document. One can imagine these aspects as three
orthogonal views of the same document. Each of these views represent on aspect, together we get the actual
document. The content of the document consists of Content Portions. These can be manipulated according to
the corresponding medium.
A content architecture describes for each medium: (1) the
specification of the elements, (2) the possible access functions and,
(3) the data coding. Individual elements are the Logical Data Units
(LDUs), which are determined for each medium. The access functions
serve for the manipulation of individual elements. The coding of the
data determines the mapping with respect to bits and bytes. ODA
has content architectures for media text, geometrical graphics and
raster graphics. Contents of the medium text are defined through the
Character Content Architecture. The Geometric Graphics Content
Architecture allows a content description of still images. It also takes
into account individual graphical objects. Pixel-oriented still images
are described through Raster Graphics Content Architecture. It can
be a bitmap as well as a facsimile.
Logical structure: --> Like SGML, a sequential order of objects in the file.
Layout: --> The placement of content on the page
Content: --> Text, geometric graphics, and raster graphics (raster graphics are also called facsimile
images)
For example, a memo document (where layout may not be critical) could have logical and content components
in ODA format.
All ODA files can be classified as Formatted, Formatted Processable, or Processable. Formatted files are not to
be edited. Processable and Formatted Processable files can be edited.
DTD -A Document Type Definition (DTD) is a file associated with SGML and XML documents that defines how
markup tags should be interpreted by the application reading the document. The DTD uses SGML syntax to
explain precisely which elements and attributes may appear in a document and the context in which they may
be used. DTDs were briefly introduced earlier in this chapter. In this section, we’ll take a closer look.
A DTD is a text document that contains a set of rules, formally known as element
declarations , attlist (attribute) declarations, and entity declarations. DTDs are most often stored in a separate
file (with the .dtd suffix) and shared by multiple documents; however, DTD information can be
included inside the XML document as well. Both methods are demonstrated later in this section.
====================================================================
Multimedia Applications
Virtual Reality Virtual reality is another interesting offshoot of multimedia technology. The aim of a virtual
reality system is to create a very close interaction between the user’s senses and the computer system. The
user will usually have a head-mounted display system and various types of clothing with electronic interfaces
to monitor the user’s responses. User commands are input through the “electronic clothing,” and the
computer generates moving picture and sound through the head-mounted display system. Virtual reality
systems can be used for training as well as entertainment. The majority of virtual reality systems are
standalone systems, though networked virtual reality systems are also available.
An 8-bit word is most prevalent and is called a byte. A 4-bit word is called a nibble.
ASCII code (American Standard Code for Information Interchange) - 7-bit code and has 128 (2 7) distinct code
values.
When light reflected from an object passes through a video camera lens, that light is converted into an
electronic signal by a special sensor called a charge-coupled device (CCD). Top-quality broadcast cameras and
even camcorders may have as many as three CCDs (one for each color of red, green, and blue) to enhance the
resolution of the camera and the quality of the image.
***The values of R,G,B at a pixel depend on what camera sensors are used to image a scene
***RGB is additive color.
***CMYK is subtractive color. For example if you want to print yellow color in a paper it will add red and green
color and subtracts blue color.
**** Extension of MPEG audio format is .mp3
Which type of system is the PA system in audio devices? – Electroacoustic
(PA (Public Address) system is an electroacoustic system, in which sound is first converted into electrical signals by a
microphone. The electrical audio signals are amplified, processed, and applied to another transducer the loudspeaker, which
converts the audio signals into sound waves.
A public address system (or PA system) is an electronic system comprising microphones, amplifiers, loudspeakers, and
related equipment. It increases the apparent volume (loudness) of a human voice, musical instrument, or other acoustic
sound source or recorded sound or music.
PA systems are used in any public venue that requires that an announcer, performer, etc. be sufficiently audible at a
distance or over a large area.
Typical applications include sports stadiums, public transportation vehicles and facilities, and live or recorded music
venues and events
A PA system may include multiple microphones or other sound sources, a mixing console to combine and modify multiple
sources, and multiple amplifiers and loudspeakers for louder volume or wider distribution.)
Keyframe : A key frame (or keyframe) in animation and filmmaking is a drawing or shot that defines
the starting and ending points of any smooth transition. These are called frames because their position
in time is measured in frames on a strip of film or on a digital video editing. A symbol is a graphic,
button, or movie clip that you create once in the Animate (formerly Flash Professional CC) authoring
environment or by using the SimpleButton (AS 3.0) and MovieClip classes. You can then reuse
the symbol throughout your document or in other documents.A scene is a single event or conversation
between characters, occurring during one period of time and in one single place, that moves the story
forward toward a climax and resolution. A keyframe is a frame where a new symbol instance appears
in the timeline. You can also add a blank keyframe to the timeline as a placeholder for symbols you
plan to add later or to explicitly leave the frame blank