Multimedia Systems Course Module
Multimedia Systems Course Module
Developed By:
1. Aderajew Kassie (MSc in Information Technology, Lecturer)
2. Betelhem Zewdu (MSc in Information Technology, Lecturer)
3. Tsegay Mullu (MSc in Information Technology, Lecturer)
Reviewed By
1. Kidst Ergetie (MSc in Information Technology, Lecturer)
2. Wubetu Barud (MSc in Information Technology, Assistant Professor)
December 2021
Module Preface
This resource module is designed and developed in support of the Multimedia Systems for
Information Technology Program. It provides learning resources and teaching ideas.
Dear students,
Chapter One: In this chapter, you have been studied about the concept of Multimedia, History
of Multimedia, The difference between Multimedia and Hypermedia, and finally how to relate
Multimedia and Web.
Chapter Two: In this chapter, you will learn about Multimedia Authoring Tools. It contains
Multimedia Authoring, Some Useful Editing and Authoring Tools and Authoring Paradigm.
Chapter Three: In this chapter you will have been studied about Data Representations like
Graphics/Image Data Representation, Digital audio and MIDI and different Popular File Formats.
Chapter Four: Talks about Image and Video. It includes Color Science, Color Models in Images
and Color Models in Video.
Chapter Five: Fundamental Concepts in Video like Types of Video Signals, Analogue Video,
Digital Video, Different TV standards
Chapter Six: Basics of Digital Audio about Digitization of Sound, Quantization and Transmission
of Audio.
Chapter Seven: is about Lossless Compression Algorithms. It includes Introduction about
compression, Run-Length Coding, Huffman Coding, Shannon-Fano Encoding and Dictionary
Based Coding.
Chapter Eight: Talks Loss compression Technique, Distortion Measures, The Rate Distortion
Theory, Quantization, and Transform encoding
Chapter Nine: Basics of Image Compression Standards, and the JPEG Standard
Basics of Video Compression and Basics of Audio Compression.
Chapter Ten: About Basic Video Compression Techniques and Video Compression Based on
Motion Compensation.
Chapter Eleven: About MPEG Video and Audio Coding like Video Compression and MPEG
Audio Compressions
i
Contents
Module Preface ................................................................................................................................ i
ii
COLORS IN IMAGE AND VIDEO ............................................................................................ 44
iii
7.6. The Shannon-Fano Encoding Algorithm ........................................................................... 86
iv
Review Question ......................................................................................................................... 111
v
List of Figures
Figure 1. 1 Hypertext is non linear ................................................................................................. 5
Figure 1. 2 Example of Hypermedia ............................................................................................... 6
Figure 1. 3 Multimedia Software requirement ................................................................................ 9
Figure 1. 4 Multimedia information flow ..................................................................................... 17
Figure 2. 1 Iconic/Flow control .................................................................................................... 21
Figure 2. 2 Card and Page based Paradigm .................................................................................. 22
Figure 3. 1 Representation of Pixel............................................................................................... 30
Figure 3. 2 Monochrome image .................................................................................................... 32
Figure 3. 3 Gray-scale Images ...................................................................................................... 33
Figure 3. 4 8-bit Color Images ...................................................................................................... 33
Figure 3. 5 Color LUT for 8-bit color images. ............................................................................. 34
Figure 4. 1 Isaac Newton's experiments. ...................................................................................... 44
Figure 4. 2 RGB color Model ....................................................................................................... 46
Figure 4. 3 CMY Color Model ..................................................................................................... 47
Figure 4. 4 RGB and CMY Cube.................................................................................................. 48
Figure 4. 5 Additive and subtractive color. (a): RGB is used to specify additive color. (b): CMY
is used to specify subtractive color ............................................................................................... 50
Figure 4. 6 YUV decomposition of color image. Top image (a) is original color image; (b) is Y ;
(c,d) are (U; V ) ............................................................................................................................. 51
Figure 4. 7 I and Q components of color image ........................................................................... 52
Figure 5. 1 Scanning video .......................................................................................................... 55
Figure 5. 2 Progressive scan ........................................................................................................ 57
Figure 5. 3 Interlaced Scanning ................................................................................................... 58
Figure 6. 1 n analog signal: continuous measurement of pressure wave ..................................... 66
Figure 6. 2 Digitization ................................................................................................................ 68
Figure 6. 3 Digitization process (sampling, quantization, and coding) ....................................... 69
Figure 7.1 A general data compression scheme............................................................................ 76
Figure 7. 2 Lossless and Lossy compression technique .............................................................. 80
Figure 7. 3 Huffman Coding Tree................................................................................................ 84
Figure 7. 4 Shannon-Fano coding tree ......................................................................................... 88
vi
Figure 7. 5 Lempel-Ziv algorithm. .............................................................................................. 90
Figure 8. 1 Typical rate-distortion function. ................................................................................. 98
Figure 9. 1 Three phase of JPEG image .................................................................................... 102
Figure 9. 2 Block diagram for JPEG encoder ............................................................................ 102
Figure 9. 3 Displaying Quantization .......................................................................................... 103
Figure 11. 1 Relation between codec, data containers and compression algorithms. ................. 108
Figure 1 The overview of Adobe Flash Professional CS6.......................................................... 112
Figure 2 The document property of Macro Media Flash ............................................................ 113
Figure 3 The layout of Macromedia Flash .................................................................................. 114
Figure 4 The layout of Adobe Flash Professional ...................................................................... 114
Figure 5 Toolbox......................................................................................................................... 115
Figure 6 Covert to symbol dialogue box..................................................................................... 116
Figure 7 To create motion tween of an object ............................................................................ 116
Figure 8 To create an animation to represent the growing moon ............................................... 118
Figure 9 To create an animation to indicate a ball bouncing on steps ........................................ 119
Figure 10 To create the background with the text as you want .................................................. 120
Figure 11 To create an animation with the following features. .................................................. 121
Figure 12 To display the background given image through your name using mask. ................. 122
Figure 13 To design a visiting card............................................................................................. 124
Figure 14 To design a photographic image give a title and design border ................................. 125
Figure 15 To design a cover page for the book .......................................................................... 126
Figure 16 To design the image from different pictures and change the background color ........ 127
Figure 17 To adjust the brightness and contrast of the picture ................................................... 127
Figure 18 To design a picture background color, rotating and scaling. ...................................... 128
Figure 19 To design a word and apply different effects shadow emboss ................................... 129
Figure 20 To design different file and organize them in a single file and apply feather effects. 130
Figure 21 To change the image color as Black and White ......................................................... 131
Figure 22 To design a wallet photo............................................................................................. 132
Figure 23 The layout of ProShow gold ....................................................................................... 133
Figure 24 Add contents in the ProShow gold ............................................................................. 134
Figure 25 Overview of ProShow gold ........................................................................................ 135
vii
Figure 26 Add blank Page........................................................................................................... 136
Figure 27 To adjust the text or captions in text effect ................................................................ 137
Figure 28 Different slide styles ................................................................................................... 137
Figure 29 To add audio files in the ProShow gold ..................................................................... 138
Figure 30 To select transition and effects ................................................................................... 139
Figure 31 To display preview area.............................................................................................. 140
Figure 32 To publish a video ...................................................................................................... 141
viii
List of Tables
ix
CHAPTER ONE
1
Multi- multiple/many
Media- source
Source refers to different kind of information that we use in multimedia.
This includes text, graphics, audio, video and images.
Multimedia refers to multiple sources of information. It is a system, which integrates all the above
types.
Definitions:
Multimedia means computer information can be represented in audio, video and animated format
in addition to traditional format. The traditional formats are text and graphics.
General and working definition:
Multimedia is the field concerned with the computer controlled integration of text, graphics,
drawings, still and moving images (video), animation, and any other media where every type of
information can be represented, stored, transmitted, and processed digitally.
Activity 1.1.
What is multimedia? What do you think about multimedia application?
2
Digital video editing and production systems
Interactive movies, and TV
Video conferencing
Virtual reality(the creation of artificial environment that you can explore, e.g. 3-D images)
Augmented reality (placing real-appearing computer graphics and video objects into scenes
so as to take the physics of objects and lights (e.g., shadows) into account
Distributed lectures for higher education
Digital libraries
World Wide Web
On-line reference works e.g. encyclopedias, games, etc.
Electronic Newspapers/Magazines
Games
Groupware (enabling groups of people to collaborate on projects and share information)
Cooperative work environments that allow business people to edit a shared document or
schoolchildren to share a single game using two mice that pass control back and forth.
Making multimedia components editable - allowing the user side to decide what
components, video, graphics, and so on are actually viewed and allowing the client to move
components around or delete them - making components distributed.
Features of Multimedia
Multimedia has three aspects:
Content: movie, production, etc.
Creative Design: creativity is important in designing the presentation
Enabling Technologies: Network and software tools that allow creative designs to be presented.
Activity 1.2.
1.2.
WhereHistory of Multimedia
it is applied this multimediaSystems
applications?
Newspaper was perhaps the first mass communication medium, which used mostly text, graphics,
and images. In 1895, Gugliemo Marconi sent his first wireless radio transmission at Pontecchio,
Italy. A few years later (in 1901), he detected radio waves beamed across the Atlantic. Initially
invented for telegraph, radio is now a major medium for audio broadcasting. Television was the
new media for the 20th century. It brings the video and has since changed the world of mass
3
communications.
Motion pictures were originally conceived of in the 1830s to observe motion too rapid for
perception by the human eye. Thomas Alva Edison 'commissioned the invention of a motion
picture camera in 1887. Silent feature films appeared from 1910 to 1927; the silent era effectively
ended with the release of the jazz singer 1927.
4
Nuclear Research)
1990 - K. Hooper Woolsey, Apple Multimedia Lab gave education to 100 people
1992 - The first M-Bone audio multicast on the net (MBONE- Multicast Backbone)
1993 - U. Illinois National Center for Supercomputing Applications introduced NCSA
Mosaic (a web browser)
1994 - Jim Clark and Marc Andersen introduced Netscape Navigator (web browser)
1995 - Java for platform-independent application development.
1996 - DVD video was introduced; high-quality, full-length movies were distributed on a
single disk. The DVD format promised to transform the music, gaming and computer industries.
1998 - XML 1.0 was announced as a W3C Recommendation.
1998 - Handheld MP3 devices first made inroads into consumer tastes in the fall, with the
introduction of devices holding 32 MB of flash memory.
2000 - World Wide Web (WWW) size was estimated at over 1 billion pages.
Hypertext
5
Hypertext is therefore usually non-linear (as indicated above).
Hypermedia
Hypermedia is the application of hypertext principles to a wider variety of media, including audio,
animations, video, and images.
As we have seen, multimedia fundamentally means that computer information can be represented
through audio, graphics, images, video, and animation in addition to traditional media (text and
graphics). Hypermedia can be considered one particular multimedia application.
Examples of typical multimedia applications include: digital video editing and production systems;
electronic newspapers and magazines; the World Wide Web; online reference works, such as
encyclopedias; games; groupware; home shopping; interactive TV; multimedia courseware; video
conferencing; video-on-demand; and interactive movies.
Desirable Features for a Multimedia System
These are the following features of desirable for a Multimedia System:
1. Very high processing speed processing power. Why? Because there are large data to be
processed. Multimedia systems deals with large data and to process data in real time, the hardware
should have high processing capacity.
2. It should support different file formats. Why? Because we deal with different data types (media
types).
3. Efficient and High Input-output: input and output to the file subsystem needs to be efficient and
6
fast. It has to allow for real-time recording as well as playback of data.
4. Special Operating System: to allow access to file system and process data efficiently and
quickly. It has to support direct transfers to disk, real-time scheduling, fast interrupt processing,
I/O streaming, etc.
5. Storage and Memory: large storage units and large memory are required. Large
Caches are also required.
6. Network Support: Client-server systems common as distributed systems common.
7. Software Tools: User-friendly tools needed to handle media, design and develop applications,
deliver media.
Activity 1.3.
What are the challenges of multimedia? What is multimedia systems?
7
and the connection is terminated - no information is carried over for the next request.
The basic request format is
Method URI Version
Additional-Headers
Message-body
The Uniform Resource Identifier (URI) identifies the resource accessed, such as the host name,
always preceded by the token ''[Link] A URI could be a Uniform Resource Locator CURL), for
example. Here, the URI can also include query strings (some interactions require submitting data).
Method is a way of exchanging information or performing tasks on the URI. Two popular methods
are GET and POST. GET specifies that the information requested is in the request string itself, while
the POST method specifies that the resource pointed to in the URI should consider the message
body. POST is generally used for submitting HTML forms. Additional-Headers specifies additional
parameters about the client. For example, to request access to this textbook's web site, the
following HTTP message might be generated:
GET [Link] HTTP 1.1
The basic response format is
Version Status-Code Status-Phrase
Additional-Headers
Messagebody
Status-Code is a number that identifies the response type (or error that occurs), and
Status-Phrase is a textual description of it. Two commonly seen status codes and phrases are 200
OK when the request was processed successfully and 404 Not Found when the URI does not exist.
For example, in response to the example request above the web server may return something like:
HTTP/l.l 200 OK Server:
[No~plugs~here~please] Date: Wed, 25 July 2002
20 : 04 : 30 GMT
Content-Length: 1045 Content-Type: text/html
<HTML>
….
</HTML>
8
1.5. Multimedia System Requirement
1) Software tools
2) Hardware Requirement
Software Requirement
3 D m o d e lin g a n d
a n im a tio n to o ls T e x t e d itin g a n d
M u ltim e d ia
A u th o rin g to o ls w o r d p r o c e s s in g
Im a g e e d itin g
to o ls
O C R s o ftw a r e M u tlim e d ia p ro je c t
P a in tin g &
d r a w in g to o ls
s o u n d e d itin g
to o ls
a n im a tio n , a u d io ,
v id e o & d ig ita l to o ls
Word processors are used for writing letters, invoices, project content, etc. They include features
like:
9
spell check
table formatting
Thesaurus
templates ( e.g. letters, resumes, & other common documents)
Examples: Microsoft Word, Word perfect, Note pad
They are used to edit sound (music, speech, etc.) The user can see the representation of sound in
fine increment, score or waveform. User can cut, copy, and paste any portion of the sound to edit
it. You can also add other effects such as distort, echo, pitch, etc. Examples: -sound forge
E.g., Sound Forge Sound Forge is a sophisticated PC-based program for editing WAV files.
Sound can be captured from a CD-ROM drive or from tape or microphone through the sound
card, then mixed and edited. It also permits adding complex special effects.
Activity 1.4.
Discuss some multimedia authoring tools?
Multimedia authoring tools provide important framework that is needed for organizing and editing
objects included in the multimedia project (e.g. graphics, animation, sound, video, etc.). They
provide editing capability to limited extent.
Macromedia Flash: Flash allows users to create interactive movies by using the score metaphor
- a timeline arranged in parallel event sequences, much like a musical score consisting of musical
notes. Elements in the movie are called symbols in Flash. Symbols are added to a central
repository, called a library, and can be added to the movie's timeline. Once the symbols are present
at a specific time, they appear on the Stage, which represents what the movie looks like at a certain
time, and can be manipulated and moved by the tools built into Flash. Finished Flash movies are
commonly used to show movies or games on the web.
10
Macromedia Director: Director uses a movie metaphor to create interactive presentations. This
powerful program includes a built-in scripting language, Lingo, which allows creation of complex
interactive movies. The "cast" of characters in Director includes bitmapped sprites, scripts, music,
sounds, and palettes. Director can read many bitmapped file formats. The program itself allows a
good deal of interactivity, and Lingo, with its own debugger, allows more control, including
control over external devices, such as VCRs and videodisc players. Director also has web-
authoring features available, for creation of fully interactive Shockwave movies playable over the
web.
Authorware: is a mature, well-supported authoring product that has an easy learning curve for
computer science students because it is based on the idea of flowcharts (the so-called iconic/flow-
control metaphor). It allows hyperlinks to link text, digital movies, graphics, and sound. It also
provides compatibility between files produced in PC and Mac versions. Shockwave Authorware
applications can incorporate Shockwave files, including Director Movies, Flash animations, and
audio.
OCR software
These soft wares convert printed document into electronically recognizable ASCII character. It is
used with scanners. Scanners convert printed document into bitmap. Then these software’s break
the bitmap into pieces according to whether it contains text or graphics. This is done by examining
the texture and density of the bitmap and by detecting edges.
Use:
11
Painting and Drawing Tools
to create graphics for web and other purposes, painting and editing tools are crucial.
Painting Tools: are also called image-editing tools. They are used to edit images of different
format. They help us to retouch and enhance bitmap images. Some painting tools allow to edit
vector based graphics too. Some of the activities of editing include:
Adobe Illustrator: illustrator is a powerful publishing tool for creating and editing vector
graphics, which can easily be exported to use on the web.
Adobe Photoshop: Photoshop is the standard in a tool for graphics, image processing, and image
manipulation. Layers ofimages, graphics, and text can be separately manipulated for maximum
flexibility, and its "filter factory" permits creation of sophisticated lighting effects.
12
Macromedia Fireworks: Fireworks is software for making graphics specifically for the web. It
includes a bitmap editor, a vector graphics editor, and a JavaScript generator for buttons and
rellovers.
Video Editing
Animation and digital video movie are sequence of bitmapped graphic frames rapidly played
back. Some of the tools to edit video include:
Hardware Requirement
Three groups of hardware for multimedia:
1) Memory and storage devices
2) Input and output devices
3) Network devices
1) Memory and Storage Devices
Multimedia products require high storage capacity than text-based data. Huge drives are essential
for the enormous files used in multimedia and audiovisual creation.
I) RAM: is the primary requirement for multimedia system. Why?
Reasons:
- you have to store authoring software itself. E.g Flash takes 20MB of memory, Photoshop 16-
20MB, etc.
- digitized audio and video is stored in memory
- Animated files, etc.
To store this at the same time, you need large amount of memory
II) Storage Devices: large capacity storage devices are necessary to store multimedia data.
Floppy Disk: not sufficient to store multimedia data. Because of this, they are not used to store
multimedia data.
Hard Disk: the capacity of hard disk should be high to store large data.
CD: is important for multimedia because they are used to deliver multimedia data to users. A
wide variety of data like:
13
Music (sound, & video)
Multimedia Games
Educational materials
Tutorials that include multimedia
Utility graphics, etc
DVD: have high capacity than CDs. Similarly, they are also used to distribute multimedia data to
users. Some of the characteristics of DVD:
2) Input-Output Devices
I) Interacting with the system: to interact with multimedia system, we use either keyboard,
mouse, track ball, or touch screen, etc.
Mouse: multimedia project is typically designed to be used with mouse as an input pointing device.
Other devices like track ball and touch screen could be used in place of mouse. Track ball is similar
with mouse in many ways.
Wireless mouse: important when the presenter has to move around during presentation
Touch Screen: we use fingers instead of mouse to interact with touch screen computers.
There are three technologies used in touch screens:
i. Infrared light: such touch screens use invisible infrared light that are projected across the
surface of screen. A finger touching the screen interrupts the beams generating electronic
signal. Then it identifies the x-y coordinate of the screen where the touch occurred and sends
signals to the operating system for processing.
ii. Texture-coated: such monitors are coated with texture material that is sensitive towards
pressure. When user presses the monitor, the texture material on the monitor extracts the x-
y coordinate of the location and send signals to operating system
iii. Touch mate:
Use: touch screens are used to display/provide information in public areas such as airports,
Museums, transport service areas, hotels, etc.
14
Advantage:
user friendly
easy to use even for non-technical people
easy to learn how to use
II) Information Entry Devices: the purpose of these devices is to enter information to be
included in our multimedia project into our computer.
OCR: they enable us to use OCR softwares convert printed document into ASCII file.
Graphical Tablets/ Digitizer: both are used to convert points, lines, and curves from sketch into
digital format. They use a movable device called stylus.
Scanners: enable us to convert printed images into digital format.
Microphones: they are important because they enable us to record speech, music, etc. The
microphone is designed to pick up and amplify incoming acoustic waves or harmonics precisely
and correctly and convert them to electrical signals. You have to purchase a superior, high-quality
microphone because your recordings will depend on its quality.
Digital Camera and Video Camera (VCR): are important to record and include image and video
in MMS respectively. Digital video cameras store images as digital data, and they do not record
on film. You can edit the video taken using video camera and VCR using video editing tools.
Remark: video takes large memory space.
Output Devices
Depending on the content of the project, & how the information is presented, you need different
output devices. Some of the output hardwares are:
Speaker: if your project includes speeches that are meant to convey message to audience, or
background music, using speaker is obligatory.
Projector: when to use projector:
if you are presenting on meeting or group discussion,
if you are presenting to large number of audience
Plotter/printer: when the situation arises to present using papers, you use printer and/or plotters.
In such cases, print quality of the device should be taken into consideration.
Impact printers: not good quality graphics/poor quality
Non-impact printers: good quality graphics
15
3) Network Devices
Why do we require network devices?
The following network devices are required for multimedia presentation:
i) Modem: which stands for modulator demodulator, is used to convert digital signal into analog
signal for communication of the data over telephone line which can carry only analog signal. At
the receiving end, it does the reverse action i.e. converts analog to digital data. Currently, the
standard modem is called v.90, which has the speed of 56kbps (kilobits per second). Older
standards include v.34, which has the speed of 28kbps. Data is transferred through modem in
compressed format to save time and cost.
ii) ISDN: stands for Integrated Services Digital Network. It is circuit switched telephone network
system, designed to allow digital transmission of voice and data over ordinary telephone copper
wires. This has the advantage of better quality and higher speeds than available with analog
systems.
It has higher transmission speed i.e faster data transfer rate.
They use additional hardware hence they are more expensive.
iii) Cable modem: uses existing cables stretched for television broadcast reception. The data
transfer rate of such devices is very fast i.e. they provide high bandwidth. They are primarily used
to deliver broadband internet access, taking advantage of unused bandwidth on a cable television
network.
iv) DSL: provide digital data transmission over the telephone wires of local telephone network.
The speed of DSL is faster than using telephone line with modem. How? They carry a digital signal
over the unused frequency spectrum (analog voice transmission uses limited range of spectrum)
available on the twisted pair cables running between the telephone company's central office and
the customer premises.
Summary
Multimedia Information Flow
16
Figure 1. 4 Multimedia information flow
Review Questions
1. What is multimedia?
2. What are the desirable feature of multimedia?
3. Discuss some application are of multimedia.
4. What are the different hardware and software requirements of multimedia?
5. What is the difference between hypertext and hypermedia?
6. How web is related to multimedia?
17
CHAPTER TWO
In a slide show, interactivity generally consists of being able to control the pace (e.g., click to
advance to the next slide). The next level of interactivity is being able to control the sequence and
choose where to go next. Next is media control: start/stop video, search text, scroll the view, and
zoom. More control is available if we can control variables, such as changing a database search
query. The level of control is substantially higher if we can control objects - say, moving objects
around a screen, playing interactive games, and so on. Finally, we can control an entire simulation:
move our perspective in the scene, control scene objects
Activity 2.1.
What is authoring and Multimedia authoring?
18
Simple presentation packages such as PowerPoint
Powerful RAD tools such as Delphi, .Net, JBuilder.
True authoring environments, which lie somewhere in between in terms of technical
complexity.
Authoring systems vary widely in:
Orientation
Capabilities, and
Learning curve: how easy it is to learn how to use the application
Activity 2.2.
Why we use authoring system?
integrate text, graphics, video, and audio to create a single multimedia presentation
control interactivity by the use of menus, buttons, hotspots, hot objects etc.
publish as a presentation or a self-running executable; on CD/DVD, Intranet, WWW
Be extended through the use of pre-built or externally supplied components, plug-ins etc
let you create highly efficient, integrated workflow
Have a large user base.
19
2.2. Multimedia Authoring Paradigms
The authoring paradigm, or authoring metaphor, is the methodology by which the authoring system
accomplishes its task. Most authoring programs use one of several authoring metaphors, also
known as authoring paradigms. There are various paradigms:
Scripting Language
Icon-Based Control Authoring Tool
Card and Page Based Authoring Tool
Time Based Authoring Tool
Tagging Tools
Scripting Language
The idea here is to use a special language to enable interactivity (buttons, mouse, etc.) and allow
conditionals, jumps, loops, functions/macros, and so on. Closest in form to traditional
programming. The paradigm is that of a programming language, which specifies:
multimedia elements,
sequencing of media elements,
hotspots (e.g links to other pages),
synchronization, etc.
Usually use a powerful, object-oriented scripting language. Multimedia elements and events
become objects that live in a hierarchical order. In-program editing of elements (still graphics,
video, audio, etc.) tends to be minimal or non-existent. Most authoring tools provide visually
programmable interface in addition to scripting language. Media handling can vary widely
Examples
Activity 2.3.
What is Authoring Paradigm? Can you mention different authoring paradigm?
20
Iconic/Flow Control Tools
In these authoring systems, multimedia elements and interaction cues (or events) are organised as
objects in a structural framework.
21
Card and page Based Tools
In these authoring systems, elements are organized as pages of a book or a stack of cards. The
authoring system lets you link these pages or cards into organized sequences. You can jump, on
command, to any page you wish in a structured navigation pattern.
Well suited for Hypertext applications, and especially suited for navigation intensive
applications.
They are best suited for applications where the bulk of the content consist of elements
that can be viewed individually.
Extensible via XCMDs (External Command) and DLLs (Dynamic Link Libraries).
All objects (including individual graphic elements) to be scripted;
Many entertainment applications are prototyped in a card/scripting system prior to
compiled-language coding.
Each object may contain programming script that is activated when an event occurs.
Examples:
- Hypercard (Macintosh)
- SuperCard(Macintosh)
- ToolBook (Windows), etc.
22
Time Based Authoring Tools
In these authoring systems elements are organized along a time line with resolutions as high as
1/30th second. Sequentially organized graphic frames are played back at a speed set by developer.
Other elements, such as audio events, can be triggered at a given time or location in the sequence
of events.
Examples
o Macromedia Director
o Macromedia Flash
Macromedia Director
Director is a powerful and complex multimedia authoring tool, which has broad set of features to
create multimedia presentation, animation, and interactive application. You can assemble and
sequence the elements of project using cast and score. Three important things that director uses to
arrange and synchronize media elements:
Cast
Cast is multimedia database containing any media type that is to be included in the project. It
imports wide range of data type and multimedia element formats directly into the cast. You can
also create elements from scratch and add to cast. To include multimedia elements in cast into the
stages, you drag and drop the media on the stage.
Score
This is where the elements in the cast are arranged. It is sequence for displaying, animating, and
playing cast members. Score is made of frames and frames contain cast member. You can set frame
rate per second.
Lingo
23
It enables interactivity and programmed control of elements
It enables to control external sound and video devices
It also enables you to control operations of internet such as sending mail, reading
documents, images, and building web pages.
Macromedia Flash
Library: a place where objects that are to be re-used are stored. The Library window shows all
the current symbols in the scene and can be toggled by the Window > Library command. A symbol
can be edited by double-clicking its name in the library, which causes it to appear on the stage.
Symbols can also be added to a scene by simply dragging the symbol from the Library onto the
stage.
Timeline: used to organize and control a movie content over time. Manages the layers and
timelines of the scene. The left portion of the Timeline window consists of one or more layers of
the Stage, which enables you to easily organize the Stage's contents. Symbols from the Library can
be dragged onto the Stage, into a particular layer. For example, a simple movie could have
two layers, the background and foreground. The background graphic from the library can
be dragged onto the stage when the background layer is selected.
Layer: helps to organize contents. Timeline is divided into layers.
24
ActionScript: enables interactivity and control of movies. Action scripts allow you to trigger
events such as moving to a different keyframe or requiring the movie to stop. Action scripts can
be attached to a keyframe or symbols in a keyframe. Right clicking on the symbol and pressing
Actions in the list can modify the actions of a symbol. Similarly, by right clicking on the keyframe
and pressing Actions in the pop-up, you can apply actions' to a keyframe. A Frame Actions
window will come up, with a list of available actions on the left and the current actions
being applied symbol on the right.
Activity 2.4.
What are the main differences between macromedia director and macro media flash?
Tagging
Tags in text files (e.g. HTML) to:
link to pages,
provide interactivity, and
Integrate multimedia elements.
Examples:
o SGML/HTML
o SMIL (Synchronized Media Integration Language)
o VRML
o 3DML
Most of them are displayed in web browsers using plug-ins or the browser itself can
understand them.
This metaphor is the basis of WWW.
It is limited but can be extended by the use of suitable multimedia tags
Multimedia Production
A multimedia project can involve a host of people with specialized skills. Multimedia production
can easily involve an art director, graphic designer, production artist, producer, project manager,
writer, user interface designer, sound designer, videographer, and 3D and 2D animators, as well as
programmers.
25
2.3. Some Useful Editing and Authoring Tools
Since the first step in creating a multimedia application is probably creation of interesting video
clips, we start looking at a video editing tool. This is not really an authoring tool, but video creation
is so important that we include a small introduction to one such program. The tools we look at are
the following:
Adobe Premiere 6
Macromedia Director
Macromedia Flash
Dreamweaver MX
Adobe Premiere
Adobe Premiere is a very simple video editing program that allows you to quickly create a simple
digital video by assembling and merging multimedia components. It effectively uses the score
authoring metaphor, in that components are placed in tracks horizontally, in a Timeline window.
The File > New Project command opens a window that displays a series of presets - assemblies of
values for frame resolution, compression method, and frame rate. There are many preset options,
most of which conform to some NTSC or PAL video standard. Start by importing resources, such
as AVI (Audio Video Interleave) video files and WAV sound files and dragging them from the
Project window onto tracks 1 or 2.
Activity 2.5.
Can you mention the criteria to select a multimedia tool?
The multimedia project you are developing has its own underlying structure and purpose.
When selecting tools for your project you need to consider that purpose. Some of the features that
you have to take into consideration when selecting authoring tools are:
1) Editing Feature: editing feature for multimedia data especially image and text are often
included in authoring tools. The more editors in your authoring system, the less
26
specialized editing tools you need. The editors that come with authoring tools offer only
subset of features found in dedicated in editing tool. If you need more capability, still you
have to go to dedicated editing tools (e.g. sound editing tools for sound editing).
2) Organizing feature: the organization of media in your project involves navigation
diagrams, or flow charts, etc. Some authoring tools provides a visual flowcharting facility.
Such features help you for organizing the project. e.g IconAuthor, and AuthorWare use
flowcharting and navigation diagram method to organize media.
3) Programming feature: there are different types of programming approach:
a. Visual programming: this is programming using cues, icons, and objects. It is
done using drag and drop. To include sound in your project, drag and drop it in
stage.
Advantage: the simplest and easiest authoring process. It is particularly useful for
slide show and presentation.
b. Programming with scripting language: Some authoring tool provide very high
level scripting language and interpreted scripting environment. This helps for
navigation control and enabling user input.
c. Programming with traditional language such as Basic or C. Some authoring
tools provide traditional programming tools like program written in C. We can call
these programs to authoring tools. Some authoring tools allow to call DLL
(Dynamic Link Library).
d. Document development tools
4) Interactivity feature: interactivity offers to the end user of the project to control the
content and flow of information. Some of interactivity levels:
a. Simple branching: enables the user to go to any location in the presentation using
key press, mouse click, etc.
b. Conditional branching: branching based on if-then decisions
c. Structured branching: support complex programming logic such as nested if-then
subroutines
5) Performance-tuning features: accomplishing synchronization of multimedia is
sometimes difficult because performance varies with different computers. In such cases
27
you need to use authoring tools own scripting language to specify time and sequence on
system.
6) Playback feature: easy testing of the project. Testing enables you to debug the system
and find out how the user interacts with it. Not waste time in assembling and testing the
project
7) Delivery feature: delivering your project needs building run-time version of the project
using authoring tools. Why run time version (executable format):
It does not require the full authoring software to play
It does not allow users to access or change the content, structure, and programming
of the project. Distributerun-time version
8) Cross platform feature: multimedia projects should be compatible with different platform
like Macintosh, Windows, etc. This enables the designer to use any platform to design the
project or deliver it to any platform.
9) Internet playability: web is significant delivery medium for multimedia. Authoring tools
typically provide facility so that output can be delivered in HTML or DHTML format.
10) Ease of learning: is it easy to learn? The designer should not waste much time learning
how to use it. Is it easy to use?
Review Questions
1. Explain briefly about Multimedia Authoring. Is that important for Multimedia? How?
2. Can you mention different Multimedia Authoring tool with their importance?
3. Explain briefly the characteristics of authoring tool.
4. What does it mean time based authoring tool?
28
CHAPTER THREE
Activity 3.1.
What is an image?
An image could be described as two-dimensional array of points where every point is allocated its
own color. Every such single point is called pixel, short form of picture element. Image is a
collection of these points that are colored in such a way that they produce meaningful information
/data. Pixel (picture element) contains the color or hue and relative brightness of that point in the
image. The number of pixels in the image determines the resolution of the image.
29
Bitmap resolution most graphics applications let you create bitmaps up to 300 dots per inch
(dpi). Such high resolution is useful for print media, but on the screen most of the
information is lost, since monitors usually display around 72 to 96 dpi.
A bit-map representation stores the graphic/image data in the same manner that the
computer monitor contents are stored in video memory.
Most graphic/image formats incorporate compression because of the large size of
the data.
P ix e l
Types of Images
There are two basic forms of computer graphics: bit-maps and vector graphics. The kind you use
determines the tools you choose. Bitmap formats are the ones used for digital photographs. Vector
formats are used only for line drawings.
They are formed from pixels - a matrix of dots with different colors. Bitmap images are defined
by their dimension in pixels as well as by the number of colors they represent. For example, a
640X480 image contains 640 pixels and 480 pixels in horizontal and vertical direction
respectively. If you enlarge a small area of a bit-mapped image, you can clearly see the pixels that
are used to create it.
30
Each of the small pixels can be a shade of gray or a color. Using 24-bit color, each pixel can be set
to any one of 16 million colors. All digital photographs and paintings are bitmapped, and any other
kind of image can be saved or exported into a bitmap format. In fact, when you print any kind of
image on a laser or ink-jet printer, it is first converted by either the computer or printer into a
bitmap form so it can be printed with the dots the printer uses.
To edit or modify bitmapped images you use a paint program. Bitmap images are widely used but
they suffer from a few unavoidable problems. They must be printed or displayed at a size
determined by the number of pixels in the image. Bitmap images also have large file sizes that are
determined by the images dimensions in pixels and its color depth. To reduce this problem, some
graphic formats such as GIF and JPEG are used to store images in compressed format.
Activity 3.2.
Explain briefly the type of images with example. And what is Dithering?
Vector graphics
They are really just a list of graphical objects such as lines, rectangles, ellipses, arcs, or curves -
called primitives. Draw programs, also called vector graphics programs, are used to create and edit
these vector graphics. These programs store the primitives as a set of numerical coordinates and
mathematical formulas that specify their shape and position in the image. This format is widely
used by computer-aided design programs to create detailed engineering and design drawings. It is
also used in multimedia when 3D animation is desired. Draw programs have a number of
advantages over paint-type program.
These include:
31
Text can be wrapped around objects.
1. Monochrome/Bit-Map Images
Images consist of pixels, or pels - picture elements in digital images. A 1-bit image consists
of on and off bits only and thus is the simplest type of image. Each pixel is stored as a
single bit (0 or 1). Hence, such an image is also referred to as a binary image. It is also
called a I-bit monochrome image, since it contains no color.
The value of the bit indicates whether it is light or dark
A 640 x 480 monochrome image requires 37.5 KB of storage.
Dithering is used to calculate patterns of dots such that values from 0 to 255 correspond to
patterns that are more and more filled at darker pixel values, for printing on a 1-bit printer.
Dithering is often used for displaying monochrome images
2. Gray-scale Images
Each pixel is usually stored as a byte (value between 0 to 255). The entire image can be
thought of as a two-dimensional array of pixel values. We refer to such an array as a bitmap,
a representation of the graphics/image data that parallels the manner in which it is stored
in video memory.
This value indicates the degree of brightness of that point. This brightness goes from black
to white
A 640 x 480 grayscale image requires over 300 KB of storage.
32
Figure 3. 3 Gray-scale Images
Such image files use the concept of a lookup table to store color information. Basically, the image
stores not color but instead just a set of bytes, each of which is an index into a table with 3-byte
values that specify the color for a pixel with that lookup table index. In a way, it is a bit, as a paint-
by-number children's art set, with number 1 perhaps standing for orange, number 2 for green, and
so on - there is no inherent pattern to the set of actual colors.
33
Color lookup Tables (LUTs)
It used in 8-bit color images is to store only the index, or code value, for each pixel. Then, if a
pixel stores, say, the value 25, the meaning is to go to row 25 in a color lookup table (LUT). While
images are displayed as two-dimensional arrays of values, they are usually stored in row-column
order as simply a long series of values. For an 8-bit image, the image file can store in the file header
information just what 8-bit values for R, G, and B correspond to each index. Figure 3.8 displays
this idea. The LUT is often called a palette.
Activity 3.3.
What is a bit map image? Can you list out the different type of bitmap images?
Image Resolution
34
Image resolution refers to the spacing of pixels in an image and is measured in pixels per inch, ppi,
sometimes called dots per inch, dpi. The higher the resolution, the more pixels in the image. A printed
image that has a low resolution may look pixelated or made up of small squares, with jagged edges and
without smoothness. Image size refers to the physical dimensions of an image.
Activity 3.4.
What is image resolution? How number of pixel and image resolution is directly proportional?
The most common formats used on internet are the GIF, JPG, and PNG.
Graphics Interchange Format (GIF) devised CompuServe, initially for transmitting graphical
images over phone lines via modems.
Uses the Lempel-Ziv Welch algorithm (a form of Huffman Coding), modified slightly for
image scan line packets (line grouping of pixels).
Limited to only 8-bit (256) color images, suitable for images with few distinctive colors
(e.g., graphics drawing)
Supports one-dimensional interlacing (downloading gradually in web browsers. Interlaced
images appear gradually while they are downloading. They display at a low blurry
resolution first and then transition to full resolution by the time the download is complete.)
35
Supports animation multiple pictures per file (animated GIF)
GIF format has long been the most popular on the Internet, mainly because of its small size
GIFs allow single-bit transparency, which means when you are creating your image, you
can specify one color to be transparent. This allows background colors to show through the
image.
PNG
JPEG/JPG
36
Uses complex lossy compression, which allows user to set the desired level of quality
(compression). A compression setting of about 60% will result in the optimum balance of
quality and file size.
Though JPGs can be interlaced, they do not support animation and transparency unlike
GIF.
TIFF
Tagged Image File Format (TIFF), stores many different types of images (e.g.,
monochrome, grayscale, 8-bit & 24-bit RGB, etc.)
Uses tags, keywords defining the characteristics of the image that is included in the file.
For example, a picture 320 by 240 pixels would include a 'width' tag followed by the
number '320' and a 'depth' tag followed by the number '240'.
Developed by the Aldus Corp. in the 1980s and later supported by the Microsoft
TIFF is a lossless format (when not utilizing the new JPEG tag which allows for JPEG
compression)
It does not provide any major advantages over JPEG and is not as user-controllable.
Do not use TIFF for web images. They produce big files, and more importantly, most
web browsers will not display TIFFs.
Bit Map (BMP) is the major system standard graphics file format for Microsoft Windows, used in
Microsoft Paint and other programs. It makes use of run-length encoding compression and can
efficiently store 24-bit bitmap images. Note, however, that BMP has many different modes,
including uncompressed 24-bit images.
37
PAINT was originally used in Mac Paint program, initially only for 1-bit monochrome
images.
PICT is a file format that was developed by Apple Computer in 1984 as the native format
for Macintosh graphics.
The PICT format is a meta-format that can be used for both bitmap images and vector
images though it was originally used in MacDraw (a vector based drawing program) for
storing structured graphics.
Still an underlying Mac format (although PDF on OS X).
Activity 3.5.
What does it mean system dependent and system independent file format?
Sound is produced by a rapid variation in the average density or pressure of air molecules above
and below the current atmospheric pressure. We perceive sound as these pressure fluctuations
cause our eardrums to vibrate. These usually minute changes in atmospheric pressure are referred
to as sound pressure and the fluctuations in pressure as sound waves. Sound waves are produced
by a vibrating body, be it a guitar string, loudspeaker cone or jet engine. The vibrating sound source
causes a disturbance to the surrounding air molecules, causing them bounce off each other with a
force proportional to the disturbance. The back and forth oscillation of pressure produces a sound
waves.
38
Activity 3.6.
What is digital audio? How to record and play a digital audio? What does it mean
digitization?
In order to play digital audio (i.e WAVE file), you need a card with a Digital to Analog Converter
(DAC) circuitry on it. Most sound cards have both an ADC (Analog to Digital Converter) and a
DAC so that the card can both record and play digital audio. This DAC is attached to the Line Out
jack of your audio card, and converts the digital audio values back into the original analog audio.
This analog audio can then be routed to a mixer, or speakers, or headphones so that you can hear
the recreation of what was originally recorded. Playback process is almost an exact reverse of the
recording process. First, to record digital audio, you need a card, which has an Analog to Digital
Converter (ADC) circuitry. The ADC is attached to the Line In (and Mic In) jack of your audio
card, and converts the incoming analog audio to a digital signal.
First, to record digital audio, you need a card, which has an Analog to Digital Converter (ADC)
circuitry. The ADC is attached to the Line In (and Mic In) jack of your audio card, and converts
the incoming analog audio to a digital signal. Your computer software can store the digitized audio
on your hard drive, visually display on the computer's monitor, mathematically manipulate in order
to add effects, or process the sound, etc. While the incoming analog audio is being recorded, the
ADC is creates many digital values in its conversion to a digital audio representation of what is
being recorded. These values must be stored for later playback.
Digitizing Sound
This creates a need to convert Analog audio to Digital audio — specialized hardware. This is also
known as Sampling.
39
1. The Traditional Discrete Audio File:
In traditional audio file, you can save to a hard drive or other digital storage medium.
WAV: The WAV format is the standard audio file format for Microsoft Windows applications,
and is the default file type produced when conducting digital recording within Windows. It
supports a variety of bit resolutions, sample rates, and channels of audio. This format is very
popular upon IBM PC (clone) platforms, and is widely used as a basic format for saving and
modifying digital audio data.
AIF: The Audio Interchange File Format (AIFF) is the standard audio format employed by
computers using the Apple Macintosh operating system. Like the WAV format, it supports a
variety of bit resolutions, sample rates, and channels of audio and is widely used in software
programs used to create and modify digital audio.
AU: The AU file format is a compressed audio file format developed by Sun Microsystems and
popular in the unix world. It is also the standard audio file format for the Java programming
language. Only supports 8-bit depth thus cannot provide CD-quality sound.
MP3: MP3 stands for Motion Picture Experts Group, Audio Layer 3 Compression. MP3 files
provide near-CD-quality sound but are only about 1/10th as large as a standard audio CD
file. Because MP3 files are small, they can easily be transferred across the Internet and
played on any multimedia computer with MP3 player software.
MIDI/MID: MIDI (Musical Instrument Digital Interface), is not a file format for storing or
transmitting recorded sounds, but rather a set of instructions used to play electronic music
on devices such as synthesizers. MIDI files are very small compared to recorded audio
file formats. However, the quality and range of MIDI tones is limited.
Streaming is a network technique for transferring data from a server to client in a format that can
be continuously read and processed by the client computer. Using this method, the client computer
can start playing the initial elements of large time-based audio or video files before the entire file
is downloaded. As the Internet grows, streaming technologies are becoming an increasingly
important way to deliver time-based audio and video data.
40
For streaming to work, the client side has to receive the data and continuously feed it to the player
application. If the client receives the data more quickly than required, it has to temporarily store
or buffer the excess for later play. On the other hand, if the data does not arrive quickly enough,
the audio or video presentation will be interrupted. There are three primary streaming formats that
support audio files: RealNetwork's RealAudio (RA, RM), Microsoft’s Advanced Streaming
Format (ASF) and its audio subset called Windows Media Audio 7 (WMA) and Apple s
QuickTime 4.0+ (MOV).
RA/RM
For audio data on the Internet, the de facto standard is RealNetwork's RealAudio (.RA) compressed
streaming audio format. These files require a RealPlayer program or browser plug-in. The latest
versions of RealNetworksserver and player software can handle multiple encodings of a single
file, allowing the quality of transmission to vary with the available bandwidth. Webcast radio
broadcast of both talk and music frequently uses RealAudio. Streaming audio can also be provided
in conjunction with video as a combined RealMedia (RM) file.
ASF
The pure audio file format used in Windows Media Technologies is Windows Media Audio 7
(WMA files). Like MP3 files, WMA audio files use sophisticated audio compression to reduce file
size. Unlike MP3 files, however, WMA files can function as either discrete or streaming data and
can provide a security mechanism to prevent unauthorized use.
MOV
Apple QuickTime movies (MOV files) can be created without a video channel and used as a sound-
only format. Since version 4.0, QuickTime provides true streaming capability. QuickTime also
accepts different audio sample rates, bit depths, and offers full functionality in both Windows as
well as the Mac OS. Popular audio file formats are:
41
au (Unix)
aiff (MAC)
wav (PC)
mp3
MIDI
Definition of MIDI: MIDI is a protocol that enables computer, synthesizers, keyboards, and other
musical device to communicate with each other. This protocol is a language that allows
interworking between instruments from different manufacturers by providing a link that is capable
of transmitting and receiving digital data. MIDI transmits only commands; it does not transmit an
audio signal. It was created in 1982.
Activity 3.7.
What is MIDI? What are the different components of MIDI?
2. Sequencer: It can be a stand-alone unit or a software program for a personal computer. (It
used to be a storage server for MIDI data. Nowadays it is more a software music editor on
the computer. It has one or more MIDI INs and MIDI OUTs.
Track: Track in sequencer is used to organize the recordings. Tracks can be turned on or off on
recording or playing back.
Channel: MIDI channels are used to separate information in a MIDI system. There are 16 MIDI
channels in one cable. Channel numbers are coded into each MIDI message.
Timbre: The quality of the sound, e.g., flute sound, cello sound, etc.
42
Multi-timbral: capable of playing many different sounds at the same time (e.g., piano, brass,
drums,..)
Voice: Voice is the portion of the synthesizer that produces sound. Synthesizers can have many
(1Two, Two0, Two4, 36, etc.) voices. Each voice works independently and simultaneously to
produce sounds of Different timbre and pitch.
Review Questions
1. What is the difference bit map and vectors graphics? Which one is better? Why.
2. What is grayscale Image?
3. What are the different popular file formats of images?
4. Explain briefly the common audio formats. Where we used?
43
CHAPTER FOUR
In 1672, Isaac Newton discovered that white light could be split into many colors by a prism. The
colors produced by light passing through prism are arranged in precise array or spectrum. The
colors spectral signature is identified by its wavelength.
Activity 4.1.
What does it mean color model?
44
The Color of Objects
Here we consider the color of an object illuminated by white light. Color is produced by the
absorption of selected wavelengths of light by an object. Objects can be thought of as absorbing
all colors except the colors of their appearance, which are reflected back. A blue object illuminated
by white light absorbs most of the wavelengths except those corresponding to blue light. These
blue wavelengths are reflected by the object.
Activity 4.2.
Discuss different color spaces.
Red
45
Green
Blue
Absence of all colors create black and presence of the three colors form white. These colors are
called additive colors. Pure black (0,0,0). Pure white (255,255,255) all other colors are produced
by varying the intensity of these three primaries and mixing the colors.
These colors are called additive colors since they add together the way light adds to make colors,
and is a natural color space to use with video displays.
The subtractive color system reproduces colors by subtracting some wavelengths of light from
white. The three subtractive color primaries are cyan, magenta, and yellow (CMY). If none of
these colors is present, the color produced is white because nothing has been subtracted from the
white light. If all the colors are present at their maximum amounts, the color produced is black
because all of the light has been subtracted from the white light.
A color model used with printers and other peripherals. Three primary colors, cyan (C), magenta
(M), and yellow (Y), are used to reproduce all colors.
46
Activity 4.3.
What are the main difference between RGB and CMY color models?
The three colors together absorb all the light that strikes it, appearing black (as contrasted to RGB
where the three colors together made white). "Nothing" on the paper is white (as contrasted to
RGB where nothing was black). These are called the subtractive or "paint" colors. Cyan, Magenta,
and Yellow (CMY) are complementary colors of RGB. CMY model is mostly used in printing
devices where the color pigments on the paper absorb certain colors (e.g., no red light reflected
from cyan ink) and in painting.
In practice, it is difficult to have the exact mix of the three colors to perfectly absorb all light and
thus produce a black color. Expensive inks are required to produce the exact color, and the paper
must absorb each color in exactly the same way. To avoid these problems, a forth color is often
added - black - creating the CYMK color space, even though the black is mathematically not
required. Sometimes, an alternative CMYK model (K stands for Black) is used in color printing
(e.g., to produce darker black than simply mixing CMY).
Activity 4.4.
Why RGB colors are additive and CMY, colors are subtractive. Where it is applicable?
47
4.2. Color Models in Images
Colors models and spaces used for stored, displayed, and printed images.
Color images are encoded as integer triplet (R,G,B) values. These triplets encode how much the
corresponding phosphor should be excited in devices such as a monitor. For images produced from
computer graphics, we store integers proportional to intensity in the frame buffer.
Additive color: Namely, when two light beams affect on a target, their colors add; when two
phosphors on a CRT screen are turned on, their colors add. However, for ink deposited on paper,
the opposite situation holds: yellow ink subtracts blue from white illumination, but reflects red
and green; it appears yellow. Instead of red, green, and blue primaries, we need primaries that
amount to -red, -green, and -blue. i.e we need to subtract R, or G, or B. These subtractive color
primaries are Cyan (C), Magenta (M) and Yellow (Y) inks.
So far, we have effectively been dealing only with additive color. Namely, when two light
beams impinge on a target, their colors add; when two phosphors on a CRT screen are turned
48
on, their colors add. Therefore, for example, red phosphor + green phosphor makes yellow light.
However, for ink deposited on paper, in essence the opposite situation holds: yellow ink subtracts
blue from white illumination but reflects red and green; which is why it appears
yellow! Therefore, instead of red, green, and blue primaries, we need primaries that amount to ---
-red, -green, and -blue; we need to subtract R, G, or B. These subtractive color primaries are
cyan (C), magenta (M), and yellow (Y) inks. RGB and CMY are connected. In the additive (RGB)
system, black is "no light", RGB = (0,0, 0). In the subtractive CMY system, black arises from,
subtracting all the light by laying down inks with C = M = Y = 1.
Simplest model we can invent to specify what ink density to lay down on paper, to make a certain
desired RGB color: Then the inverse transform is:
Undercolor removal: Sharper and cheaper printer colors: calculate that part of the CMY mix that
would be black, remove it from the color proportions, and add it back as real black. The new
specification of inks is thus: K stands for black.
Color combinations that result from combining primary colors available in the two situations,
additive color and subtractive
49
(a) (b)
Figure 4. 5 Additive and subtractive color. (a): RGB is used to specify additive color. (b): CMY
is used to specify subtractive color
Methods of dealing with color in digital video derive largely from older analog methods
of coding color for TV. Typically, some version of the luminance is combined with color
information in a single signal. YIQ is used to transmit TV signals in North America and Japan.
This coding also makes its way into VHS videotape coding in these countries, since video tape
technologies also use YIQ. In Europe, videotape uses the PAL or SECAM coding’s, which are
based on TV that uses a matrix transform called YUV. Finally, digital video mostly uses a matrix
transform called YCbCr that is closely related to YUV.
Established in 1982 to build digital video standard. Video is represented by a sequence of fields
(odd and even lines). Two fields make a frame. Works in PAL (50 fields/sec) or NTSC (60
fields/sec). The luminance (brightness), Y, is retained separately from the chrominance (color).
The Y component determines the brightness of the color (referred to as luminance or luma), while
the U and V components determine the color itself (it is called chroma). U is the axis from blue to
50
yellow and V is the axis from magenta to cyan. Y ranges from 0 to 1 (or 0 to 255 in digital formats),
while U and V range from -0.5 to 0.5 (or -128 to 127 in signed digital form, or 0 to 255 in unsigned
form). The component of a television signal which carries information on the brightness of the
image. Y(Luminance) is the intensity of light emitted from a surface per unit area in a given
direction. Chrominance refers to the difference between a color and a reference white at the same
luminance.
Figure 4. 6 YUV decomposition of color image. Top image (a) is original color image; (b) is Y ;
(c,d) are (U; V )
YIQ is used in color TV broadcasting, it is downward compatible with Black and White TV. The
YIQ color space is commonly used in North American television systems. Note that if the
chrominance is ignored, the result is a "black and white" picture. I and Q are a rotated version of
U and V. Y in YIQ is the same as in YUV; U and V are rotated by 33 degree. Y (luminance). I is
red-orange axis, Q is roughly orthogonal to I. Eye is most sensitive to Y (luminance), next to I,
next to Q. YIQ is intended to take advantage of human color response characteristics. Eye is more
51
sensitive to (I) than in (Q). Therefore less bandwidth is required for Q than for I. NTSC limits I to
1.5 MHZ and Q to 0.6 MHZ. Y is assigned higher bandwidth, 4MHZ.
YCbCr
This is similar to YUV. This color space is closely related to the YUV space, but with the
coordinates shifted to allow all positive valued coefficients. The luminance (brightness), Y, is
retained separately from the chrominance (color). During development and testing of JPEG, it
became apparent that chrominance sub sampling in this space allowed a much better compression
than simply compressing RGB or CYM. Sub sampling means that only one half or one quarter as
much detail is retained for the color as for the brightness. It is used in MPEG and JPEG
compressions.
Y-Luma component
Cb
Cr Chrominace
Activity 4.5.
What are the different color models in image and video?
52
Summary of Color
Color images are encoded as (R,G,B) integer triplet values. These triplets encode how much
the corresponding phosphor should be excited in devices such as a monitor.
Three common systems of encoding in video are RGB, YIQ, and YcrCb(YUV).
Besides the hardware-oriented color models (i.e., RGB, CMY, YIQ, YUV), HSB (Hue,
Saturation, and Brightness, e.g., used in Photoshop) and HLS (Hue, Lightness, and Saturation)
are also commonly used.
YIQ uses properties of the human eye to prioritize information. Y is the black and white
(luminance) image; I and Q are the color (chrominance) images. YUV uses similar idea.
YUV is a standard for digital video that specifies image size, and decimates the chrominance
images (for 4:2:2 video)
A black and white image is a 2-D array of integers
Review Questions
1. Explain RGB and CMY color models with their color cube.
2. Is there any difference between CMY and CMYK colors?
3. Can you describe HSB color model that used in Photoshop?
53
CHAPTER FIVE
Activity 5.1.
What is Video? Is there any difference between image and video?
Each screen-full of video is made up of thousands of pixels. A pixel is the smallest unit of an
image. A pixel can display only one color at a time. Your television has 720 vertical lines of pixels
(from left to right) and 486 rows of pixels (top to bottom). A total of 349,920 pixels (720 x 486)
for a single frame. There are two types of video. Those are Analog Video and Digital Video.
54
industry. Distortion of images and noise are common problems for analog video. In an analogue
video signal, each frame is represented by a fluctuating voltage signal. This is known as an
analogue waveform. One of the earliest formats for this was composite video. Analog formats are
susceptible to loss due to transmission noise effects. Quality loss is also possible from one
generation to another. This type of loss is like photocopying, in which a copy of a copy is never as
good as the original. Most TV is still sent and received as an analog signal. Once the electrical
signal is received, we may assume that brightness is at least a monotonic function of voltage.
An analog signal f(t) samples a time-varying image. So-called progressive scanning traces through
a complete picture (a frame) row-wise for each time interval. A high-resolution computer monitor
typically uses a time interval of 1n2 second. In TV and in some monitors and multimedia standards,
another system, interlaced scanning, is used. Here, the odd-numbered lines are traced first, then
the even-numbered lines. This result in "odd" and "even" fields - two fields make up one frame.
In fact, the odd lines (starting from 1) end up at the middle of a line at the end of the odd field, and
the even scan starts at a halfway point. The following figure shows the scheme used. First the solid
(odd) lines are traced - P to Q, then R to S, and so on, ending at T – then the even field starts at U
and ends at V. The scan lines are not horizontal because a small voltage is applied, moving the
electron beam down over time.
55
5.3. Digital Video
Digital technology is based on images represented in the form of bits. A digital video signal is
actually a pattern of 1's and 0's that represent the video image. With a digital video signal, there is
no variation in the original signal once it is captured on to computer disc. Therefore, the image
does not lose any of its original sharpness and clarity. The image is an exact copy of the original.
A computer is the most common form of digital technology. The limitations of analog video led
to the birth of digital video. Digital video is just a digital representation of the analogue video
signal. Unlike analogue video that degrades in quality from one generation to the next, digital
video does not degrade. Each generation of digital video is identical to the parent. Even though the
data is digital, virtually all digital, formats are still stored on sequential tapes. There are two
significant advantages for using computers for digital video:
Computer-based digital video is defined as a series of individual images and associated audio.
These elements are stored in a format in which both elements (pixel and sound sample) are
represented as a series of binary digits (bits). Almost all digital video uses component video.
An analog video can be very similar to the original video copied, but it is not identical. Digital
copies will always be identical and will not lose their sharpness and clarity over time. However,
digital video has the limitation of the amount of RAM available, whereas this is not a factor with
analog video. Digital technology allows for easy editing and enhancing of videos.
Activity 5.2.
Discuss analog vs digital video? Where it is applied?
56
Displaying Video
Progressive scan
Interlaced scan
Progressive scan
Progressive scan updates all the lines on the screen at the same time. This is known as progressive
scanning. Today all PC screens write a picture like this
Interlaced Scanning
Interlaced scanning writes every second line of the picture during a scan, and writes the other half
during the next sweep. Doing that we only need 25/30 pictures per second. This idea of splitting
up the image into two parts became known as interlacing and the splitted up pictures as fields.
Graphically seen a field is basically a picture with every 2nd line black/white. Here is an image that
shows interlacing so that you can better imagine what happens.
57
Figure 5. 3 Interlaced Scanning
Activity 5.3.
How to display a video with interlaced vs progressive scan?
58
the three channels. This is not the case for S-Video or Composite Video. Component video,
however, requires more bandwidth and good synchronization of the three components.
2. Composite video/1 Signal: color (chrominance) and luminance signals are mixed into a
single carrier wave. Some interference between the two signals is inevitable. Composite
analog video has all its components (brightness, color, synchronization information, etc.)
combined into one signal. Due to the compositing (or combining) of the video components,
the quality of composite video is marginal at best. The results are color bleeding, low clarity
and high generational loss.
In NTSC TV, for example, I and Q are combined into a chroma signal, and a color
subcarrier then puts the chroma signal at the higher frequency end of the channel shared
with the luminance signal. The chrominance and luminance components can be separated
at the receiver end, and the two color components can be further recovered.
When connecting to TVs or VCRs, composite video uses only one wire (and hence one
connector, such as a BNC connector at each end of a coaxial cable or an RCA plug at each
end of an ordinary wire), and video color signals are mixed, not sent separately. The audio
signal is another addition to this one signal. Since color information is mixed and both
color and intensity are wrapped into the same signal, some interference between the
luminance and chrominance signals is inevitable.
3. S-Video/2 Signal (Separated video): a compromise between component analog video and
the composite video. It uses two lines, one for luminance and another for composite
chrominance signal.
As a compromise, S-video (separated video, or super-video, e.g" in S-VHS) uses two wires:
one for luminance and another for a composite chrominance signal. As a result, there is
less crosstalk between the color information and the crucial gray-scale information.
The reason for placing luminance into its own part of the signal is that black-and-white
information is crucial for visual perception. Humans are able to differentiate spatial
resolution in grayscale images much better than for the color part of color images (as
opposed to the "black-and-white" part). Therefore, color information sent can be much less
59
accurate than intensity information. We can see only large blobs of color, so it makes sense
to send less color detail.
Activity 5.3.
What are the different types of color video signal? Discuss briefly with their number of
wire, where it is applied, and compare their bandwidth and quality.
PAL is a TV standard originally invented by German scientists and uses 625 horizontal lines at a
field rate of 50 fields per second (or 25 frames per second). It is used in Australia, New Zealand,
United Kingdom, and Europe.
SECAM uses the same bandwidth as PAL but transmits the color information sequentially. It is
used in France, East Europe, etc. SECAM (System Electronic Pour Couleur Avec Memoire) is
60
very similar to PAL. It specifies the same number of scan lines and frames per second. SECAM
also uses 625 scan lines per frame, at 25 frames per second; it is the broadcast standard for France,
Russia, and parts of Africa and Eastern Europe.
SECAM and PAL are similar, differing slightly in their color-coding scheme. In SECAM
U and V, signals are modulated using separate color subcarriers at 4.25 MHz and 4.41 MHz,
respectively. They are sent in alternate lines - that is, only one of the U or V signals will be sent
on each scan line.
The NTSC TV standard is mostly used in North America and Japan. NTSC is a black-and-white
and color compatible 525-line system that scans a nominal 30 interlaced television picture frames
per second. Used in USA, Canada, and Japan.
525 scan lines per frame, 30 frames per second (or be exact, 29.97 fps, 33.37 sec/frame)
Interlaced, each frame is divided into 2 fields, 262.5 lines/field
20 lines reserved for control information at the beginning of each field
First-generation HDTV was based on an analog technology developed by Sony and NHK in Japan
in the late 1970s. HDTV successfully broadcast the 1984 Los Angeles Olympic Games in Japan.
Multiple sub-Nyquist Sampling Encoding (MUSE) was an improved NHK HDTV with hybrid
61
analog/digital technologies that was put in use in the 1990s. It has 1,125 scan lines, interlaced (60
fields per second), and a 16:9 aspect ratio. It uses satellite to broadcast ~ quite appropriate for
Japan, which can be covered with one or two satellites. The Direct Broadcast Satellite (DBS)
channels used have a bandwidth of 24 :MHz.
High-Definition television (HDTV) means broadcast of television signals with a higher resolution
than traditional formats (NTSC, SECAM, PAL) allow. Except for early analog formats in Europe
and Japan, HDTV is broadcasted digitally, and therefore its introduction sometimes coincides with
the introduction of digital television (DTV).
The HDTV signal is digital resulting in crystal clear, noise-free pictures and CD quality sound. It
has many viewer benefits like choosing between interlaced or progressive scanning.
62
per second). The latter provides slightly better picture quality but requires much higher
bandwidth.
Activity 5.5.
Which TV standard is good quality and why?
File formats in the PC platform are indicated by the 3 letter filename extension.
With digital video, four factors have to be kept in mind. These are:
Frame Rate
The standard for displaying any type of non-film video is 30 frames per second (film is
24 frames per second). This means that the video is made up of 30 (or 24) pictures or framesfor
every second of video. Additionally these frames are split in half (odd lines and even lines), to
form what are called fields.
Color Resolution
Color resolution refers to the number of colors displayed on the screen at one time. Computers
deal with color in an RGB (red-green-blue) format, while video uses a variety of formats. One of
the most common video formats is called YUV. Although there is no direct correlation between
63
RGB and YUV, they are similar in that they both have varying levels of color depth (maximum
number of colours).
Spatial Resolution
The third factor is spatial resolution - or in other words, “How big is the picture?” Since PC and
Macintosh computers generally have resolutions in excess of 640 by 480, most people assume that
this resolution is the video standard. A standard analogue video signal displays a full, over scanned
image without the borders common to computer screens. The National Television Standards
Committee ( NTSC) standard used in North America and Japanese Television uses a 768 by 484
display. The Phase Alternative system (PAL) standard for European television is slightly larger at
768 by 576. Most countries endorse one or the other, but never both.
Since the resolution between analogue video and computers is different, conversion of analogue
video to digital video at times must take this into account. This can often the result in the down-
sizing of the video and the loss of some resolution.
Image Quality
The last and most important factor is video quality. The final objective is video that looks
acceptable for your application.
Review Questions
1. Describe different TV Standards.
2. What are the different factors of digital video?
3. What is Progressive scan and interlacing scan? Is there any difference?
4. Explain the different video file formats.
64
CHAPTER SIX
E.g: When I record a speech given out by my mother on her anniversary, the input to the recorder
is "voice" while the output, when I play the speech on a player is audio. Animation is an art of
drawing sketches of object and then showing them in a series of frames so that it looks like a
moving and living thing to us while a video is a recording of either still or moving objects.
Activity 6.1.
What is audio? Is there any difference between voice and audio?
Sound is a wave phenomenon like light, but it is macroscopic and involves molecules of air
being compressed and expanded under the action of some physical device. For example, a speaker
in an audio system vibrates back and forth and produces a longitudinal pressure wave that we
perceive as sound. (As an example, we get a longitudinal wave by vibrating a Slinky along its
length; in contrast, we get a transverse wave by waving the Slinky back and forth perpendicular to
its length.) Without air, there is no sound - for example, in space. Since sound is a pressure wave,
it takes on continuous values, as opposed to digitized ones with a finite range. Nevertheless,
65
if we wish to use a digital version of sound waves, we must form digitized representations of audio
information.
Even though such pressure waves are longitudinal, they still have ordinary wave properties and
behaviors, such as reflection (bouncing), refraction (change of angle when entering a medium with
a different density), and diffraction (bending around an obstacle). This makes the design of
"surround sound" possible.
Sampling Audio
Analog Audio
Most natural phenomena around us are continuous; they are continuous transitions between two
different states. Sound is not exception to this rule i.e. sound also constantly varies. Since sound
consists of measurable pressures at any 3D point, we can detect it by measuring the pressure level
at a location, using a transducer to convert pressure to voltage levels.
66
Figure 6.1 shows the one-dimensional nature of sound. Values change over time in amplitude: the
pressure increases or decreases with time. The amplitude value is a continuous quantity. Since we
are interested in working with such data in computer storage, we must digitize the analog signals
(i.e., continuous-valued voltages) produced by microphones. For image data, we must likewise
digitize the time-dependent analog signals produced by typical video cameras. Digitization means
conversion to a stream of numbers – preferably integers for efficiency.
Activity 6.2.
What is analog and digital signal?
Continuously varying signals are represented by analog signal. Signal is a continuous function f in
the time domain. For value y=f(t), the argument t of the function represents time. If we graph f, it
is called wave. A wave has three characteristics:
Amplitude
Frequency, and
Phase
Amplitude:is the intensity of signal. This is can be determined by looking at the height of signal.
If amplitude increases, the sound becomes louder. Amplitude measures the how high or low the
voltage of the signal is at a given point of time.
Frequency: is the number of times the wave cycle is repeated. This can be determined by
counting the number of cycles in given time interval. Frequency is related with pitchness of the
sound. Increased frequencyhigh pitch.
67
Figure 6. 2 Digitization
Converting an analog audio to digital audio requires that the analog signal is sampled. Sampling
is the process of taking periodic measurements of the continuous signal. Samples are taken at
regular time interval, i.e. every T seconds. This is called sampling frequency/sampling rate.
Digitized audio is sampled audio. Many times each second, the analog signal is sampled. How
often these samples are taken is referred to as sampling rate. The amount of information stored
about each sample is referred to as sample size.
Analog signal is represented by amplitude and frequency. Converting these waves to digital
information is referred to as digitizing. The challenge is to convert the analog waves to numbers
(digital information. The more numbers on the scale the better the quality of the sample, but more
bits will be needed to represent that sample.
In digital form, the measure of frequency is referred to as how often the sample is taken. In the
graph below the sample has been taken 7 times (reading across). Frequency is talked about in
terms of Kilohertz (KHz).
KHz = 1000Hz
68
Music CDs use a frequency of 44.1 KHz. A frequency of 22 KHz for example, would mean that
the sample was taken less often.
Sampling means measuring the value of the signal at a given time period. The samples are then
quantized. Quantization is rounding the value of each sample to the nearest amplitude number
in the graph. For example, if amplitude of a specific sample is 5.6, this should be rounded either
up to 6 or down to 5. This is called quantization. Quantization is assigning a value (from a set) to
a sample. The quantized values are changed to binary pattern. The binary patterns are stored in
computer.
Activity 6.3.
How to digitizing sounds?
The following diagram shows digitization process (sampling, quantization, and coding)
Activity 6.4.
What is sampling rate?
69
Sample Rate
A sample is a single measurement of amplitude. The sample rate is the number of these
measurements taken every second. In order to accurately represent all of the frequencies in a
recording that fall within the range of human perception, generally accepted as 20Hz-20KHz,
we must choose a sample rate high enough to represent all of these frequencies. At first
consideration, one might choose a sample rate of 20 KHz since this is identical to the highest
frequency. This will not work, however, because every cycle of a waveform has both a positive
and negative amplitude and it is the rate of alternation between positive and negative amplitudes
that determines frequency. Therefore, we need at least two samples for every cycle resulting in
a sample rate of at least 40 KHz.
Activity 6.5.
How to quantize a signal?
70
The sampling rate of a real signal needs to be greater than twice the signal bandwidth. Audio
practically starts at 0 Hz, so the highest frequency present in audio recorded at 44.1 kHz is 22.05
kHz (22.05 kHz bandwidth).
Activity 6.6.
What are the implication of sampling rate?
Each sample can only be measured to a certain degree of accuracy. The accuracy is dependent on
the number of bits used to represent the amplitude, which is also known as the sample resolution.
How do we store each sample value (quantized value)?
71
Examples:
Abebe sampled audio for 10 seconds. How much storage space is required if
a) 22.05 KHz sampling rate is used, and 8 bit resolution with mono recording?
b) 44.1 KHz sampling rate is used, and 8 bit resolution with mono recording?
c) 44.1 KHz sampling rate is used, 16 bit resolution with stereo recording?
d) 11.025 KHz sampling rate, 16 bit resolution with stereo recording?
Solution:
a) m=22050*8*10*1
m= 1764000bits=220500bytes=220.5KB
b) m=44100*8*10*1
m= 3528000 bits=441000butes=441KB
c) m=44100*16*10*2
m= 14112000 bits= 1764000 bytes= 1764KB
d) m=1025*16*10*2
m= 3528000 bits= 441000 bytes= 441KB
Quantization and transformation of data are collectively known as coding of the data. Differences
in signals between the present and a previous time can effectively reduce the size of signal
values and, most important; concentrate the histogram of pixel values (differences, now) into a
much smaller range. The result of reducing the variance of values is that lossless compression
methods that produce a bit stream with shorter bit lengths for more likely values. In general,
producing quantized sampled output for audio is called pulse code modulation or PCM. The
differences version is called DPCM (and a crude but efficient variant is called DM).
72
Audio is analog - the waves we hear travel through the air to reach our eardrums. We know that
the basic techniques for creating digital signals from analog ones consist of sampling and
quantization. Sampling is invariably done uniformly – we select a sampling rate and produce one
value for each sampling time.
Boundaries for quantizer input intervals that will all be mapped into the same output level form a
coder mapping and the representative values that are the output values from a quantizer are a
decoder mapping.
Review Questions
4. Describe the following terms
a. Amplitude
b. Frequency, and
c. Phase
5. How to convert analog audio to digital audio? Is there any effect if it is not converted?
6. Discuss the different sampling rate and for which purpose we use it?
7. What is sampling resolution?
8. If Mr. X sampled audio for 20 seconds. How much storage space is required if
a. 22.05 KHz sampling rate is used, and 16 bit resolution with mono recording?
73
b. 44.1 KHz sampling rate is used, and 8 bit resolution with mono recording?
c. 44.1 KHz sampling rate is used, 16 bit resolution with stereo recording.
d. 44.1 KHz sampling rate is used, 16 bit resolution with mono recording.
e. 11.025 KHz sampling rate, 8 bit resolution with stereo recording?
f. 11.025 KHz sampling rate, 8 bit resolution with mono recording?
74
75
CHAPTER SEVEN
Take, for example, a video signal with resolution 320x240 pixels and 256 (8 bits) colors, 30 frames
per second. Raw bit rate = 320x240x8x30
= 18,432,000 bits
A 90 minute movie would take 2.3x60x90 MB = 12.44 GB. Without compression, data storage
and transmission would pose serious problems!
Figure 7.1 depicts a general data compression scheme, in which compression is performed by an
encoder and decompression is performed by a decoder. We call the output of the encoder codes or
code words. The intermediate medium could be either data storage or a communication/computer
network. If the compression and decompression processes induce no information loss, the
compression scheme is lossless; otherwise, it is Lossy. The next several chapters deal with Lossy
compression algorithms as they are commonly used for image, video, and audio compression.
Here, we concentrate on 10ssless compression.
Activity 7.1.
What is multimedia data compression?
76
7.2. Basics of Information Theory
Information theory is the scientific study of the quantification, storage, and communication of
digital information. ... A key measure in information theory is entropy. Entropy quantifies the
amount of uncertainty involved in the value of a random variable or the outcome of a random
process.
Information theory is the mathematical treatment of the concepts, parameters and rules governing
the transmission of messages through communication systems. It was founded by Claude Shannon
toward the middle of the twentieth century and has since then evolved into a vigorous branch of
mathematics fostering the development of other scientific fields, such as statistics, biology,
behavioral science, neuroscience, and statistical mechanics. The techniques used in information
theory are probabilistic in nature and some view information theory as a branch of probability
theory. In a given set of possible events, the information of a message describing one of these
events quantifies the symbols needed to encode the event in an optimal way. ‘Optimal’ means that
the obtained code word will determine the event unambiguously, isolating it from all others in the
set, and will have minimal length, that is, it will consist of a minimal number of symbols.
Information theory also provides methodologies to separate real information from noise and to
determine the channel capacity required for optimal transmission conditioned on the transmission
rate.
The foundation of information theory was laid in a 1948 paper by Shannon titled, “A Mathematical
Theory of Communication.” Shannon was interested in how much information a given
communication channel could transmit. In neuroscience, you are interested in how much
information the neuron’s response can communicate about the experimental stimulus.
Information theory is based on a measure of uncertainty known as entropy (designated “H”). For
example, the entropy of the stimulus S is written H(S) and is defined as follows:
H(S)=−∑SP(s)log2P(s)
The subscript S underneath the summation simply means to sum over all possible stimuli S=[1, 2
… 8]. This expression is called “entropy” because it is similar to the definition of entropy in
77
thermodynamics. Thus, the preceding expression is sometimes referred to as “Shannon entropy.”
The entropy of the stimulus can be intuitively understood as “how long of a message (in bits) do I
need to convey the value of the stimulus?” For example, suppose the center-out task had only two
peripheral targets (“left” and “right”), which appeared with an equal probability. It would take only
one bit (a 0 or a 1) to convey which target appeared; hence, you would expect the entropy of this
stimulus to be 1 bit. That is what the preceding expression gives you, as P(S)=0.5 and log2(0.5)=−1.
The center-out stimulus in the dataset can take on eight possible values with equal probability, so
you expect its entropy to be 3 bits.
Data compression is about finding ways to reduce the number of bits or bytes used to store or
transmit the content of multimedia data. It is the process of encoding information using fewer bits
Eg. ZIP file format. As with any communication, compressed data communication only works
when both the sender and receiver of the information understand the encoding scheme.
Is compression useful?
Compression is useful because it helps reduce the consumption of resources, such as hard disk
space or transmission bandwidth.
On the downside, compressed data must be decompressed to be used, and this extra processing
may be harmful to some applications. For instance, a compression scheme for video may require
expensive hardware for the video to be decompressed fast enough to be viewed as it's being
decompressed. The option of decompressing the video in full before watching it may be
inconvenient, and requires storage space for the decompressed video.
The design of data compression schemes therefore involves trade-offs among various factors,
including
78
The amount of distortion introduced: to what extent quality loss is tolerated.
Activity 7.2.
What is the difference between Lossless and Lossy compression technique?
Types of Compression
Lossless Compression
The original content of the data is not lost/changed when it is compressed (encoded). It is used
mainly for compressing symbolic data such as database records, spreadsheets, texts, executable
programs, etc., Lossless compression can recover the exact original data after compression
where exact replication of the original is essential and changing even a single bit cannot be
tolerated. Examples: Run Length Encoding (RLE), Lempel Ziv (LZ), Huffman Coding.
Lossy Compression
The original content of the data is lost to certain degree when compressed. For visual and audio
data, some loss of quality can be tolerated without losing the essential nature of the data. Lossy
compression is used for image compression in digital cameras like JPEG, audio compression like
mp3. Video compression in DVDs with MPEG format.
79
Figure 7. 2 Lossless and Lossy compression technique
GIF image files and WinZip use lossless compression. For this reason, zip software is popular for
compressing program and data files. Lossless compression does not lose any data in the
compression process.
The advantage is that the compressed file will decompress to an exact duplicate of the
original file, mirroring its quality.
The disadvantage is that the compression ratio is not all that high, precisely because no
data is lost.
To get a higher compression ratio -- to reduce a file significantly beyond 50% -- you must use
lossy compression.
Lossless & lossy compression have become part of our everyday vocabulary due to the popularity
of MP3 music file, JPEG image file, MPEG video file. A sound file in WAV format, converted to
a MP3 file will lose much data as MP3 employs a lossy compression. JPEG uses lossy
compression, while GIF follows lossless compression techniques.
80
An example of lossless vs. lossy compression is the following string: 25.888888888. This string
can be compressed as 25.9! 8 and interpreted as, "twenty five point 9 eights", the original string is
perfectly recreated, just written in a smaller form.
In which case, the original data is lost, at the benefit of a smaller file size. The above example is a
very simple example of run-length encoding.
Activity 7.3.
Why we compress the data?
This encoding scheme tries to tally occurrence of data value (Xi) along with its run length, i.e.(Xi
, Length_of_Xi).
It compress data by storing runs of data (that is, sequences in which the same data value occurs in
many consecutive data elements) as a single data value & count. For example, consider the
following image with long runs of white pixels (W) and short runs of black pixels (B).
WWWWWWWWWWBWWWWWWWWWBBBWWWWWWWWWWWW
The RLE data compression algorithm, the compressed code is: 10W1B9W3B12W (Interpreted as
ten W's, one B, nine W's, three B's, …)
81
Lossless compression schemes are reversible so that the original data can be reconstructed,
Lossy schemes accept some loss of data in order to achieve higher compression.
These Lossy data compression methods typically offer a three-way tradeoff between
E.g. Huffman coding- Estimate probabilities of symbols, code one symbol at a time, shorter codes
for symbols with high probabilities.
Dictionary-based coding: The previous algorithms (both entropy and Huffman) require the
statistical knowledge. Dictionary based coding, such as Lempel-Ziv (LZ), compression
techniques do not require prior information to compress strings. Rather, replace symbols with
a pointer to dictionary entries.
Activity 7.4.
What does it mean statistical based and dictionary based compression technique? Is there
any difference between the two?
82
which maps source symbols to a variable number of bits. Variable-length codes can allow sources
to be compressed and decompressed with zero error and still be read back symbol by symbol. In
this section we will discuss about the Shannon-Fano algorithm and Huffman coding.
The problem: Given a set of n symbols and their weights (or frequencies), construct a tree
structure (a binary tree for binary code) with the objective of reducing memory space and decoding
time per symbol. For instance, Huffman coding is constructed based on frequency of occurrence
of letters in text documents.
The output of the Huffman encoder is determined by the Model (probabilities). The higher the
probability of occurrence of the symbol, the shorter the code assigned to that symbol and vice
versa. This will enable to easily control the most frequently occurring symbols in a data and also
reduce the time taken during decoding each symbols.
83
Associate binary code: 1 with the right branch and 0 with the left branch
Step 4: Create a unique code word for each symbol by traversing the tree from the root to the leaf.
Example 1: Consider a 7-symbol alphabet given in the following table to construct the Huffman
coding.
Character a b c d e f g
The Huffman encoding algorithm picks each time two symbols (with the smallest frequency) to
combine.
Example 2:
Character A B c d e f
Frequency 5 9 12 13 16 45
84
Example 3: construct the tree & binary code by using Huffman coding
Character A B C d E F G
Frequency 37 18 29 13 30 17 6
85
7.6. The Shannon-Fano Encoding Algorithm
1. Calculate the frequency of each of the symbols in the list.
3. Divide the list into two half’s, with the total frequency counts of each half being as close as
possible to each other.
4. The right half is assigned a code of 1 and the left half with a code of 0.
5. Recursively apply steps 3 and 4 to each of the halves, until each symbol has become a
corresponding code leaf on the tree. That is, treat each split as a list and apply splitting and
code assigning till you are left with lists of single elements.
6. Generate code word for each symbol
Let us assume the source alphabet S={X1,X2,X3,Ö,Xn} and Associated probability P={P1,P2,P3
,Ö,Pn}. The steps to encode data using Shannon-Fano coding algorithm is as follows: Order the
source letter into a sequence according to the probability of occurrence in non-increasing order i.e.
decreasing order.
ShannonFano (sequence s)
Attach 0 to the codeword of one letter and 1 to the codeword of another; Else if s has more than
two letter. Divide s into two subsequences S1, and S2 with the minimal difference between
probabilities of each subsequence; extend the codeword for each letter in S1 by attaching 0, and
by attaching 1 to each codeword for letters in S2;
ShannonFano(S1);
ShannonFano(S2);
Example 1: Given five symbols A to E with their frequencies being 15, 7, 6, 6 & 5; encode them
using Shannon-Fano entropy encoding
Symbol A B C D E
Count 15 7 6 6 5
86
Solution:
Step1: Say, we are given that there are five symbols (A to E) that can occur in a source with their
frequencies being 15 7 6 6 and 5. First, sort the symbols in decreasing order of frequency.
Step2: Divide the list into two equal halves. That is, the counts of both halves are as close as
possible to each other. Therefore, in this case we split the list between B and C & assign 0 and 1.
Step3: We recursively repeat the steps of splitting and assigning code until each symbol become a
code leaf on the tree. That is, treat each split as a list, apply splitting, and code assigning until you
are left with lists of single elements.
Step 4: Note that we split the list containing C, D and E between C and D because the difference
between the split lists is 11 minus 6, which is 5, if we were to have divided between D and E we
would get a difference of 12-5 which is 7.
Step5: We complete the algorithm and as a result have codes assigned to the symbols.
87
S1={A,B} P={0.35,0.17}=0.52
S2={C,D,E} P={0.17,0.17,0.16}=0.46
The difference is only 0.52-0.46=0.06. This is the smallest possible difference when we divide
the message. Attach 0 to S1 and 1 to S2. Subdivide S1 into sub groups.
S11={A} attach 0 to this
S12={B} attach 1 to this
Again subdivide S2 into subgroups considering the probability again.
S21={C} P={0.17}=0.17
S22={D,E} P={0.16,0.15}=0.31
Attach 0 to S21 and 1 to S22. Since S22 has more than one letter in it, we have to subdivide it.
S221={D} attach 0
S222={E} attach 1
A=00 B=01
C=10 D=110
E=111
Instead of transmitting ABCDE, we transmit 000110110111.
88
7.7. Lempel-Ziv Encoding
Data compression up until the late 1970's mainly directed towards creating better methodologies
for Huffman coding. An innovative, radically different method was introduced in1977 by Abraham
Lempel and Jacob Ziv. The zip and unzip use the LZH technique while UNIX's compress methods
belong to the LZW and LZC classes.
Lempel-Ziv compression
The problem with Huffman coding is that it requires knowledge about the data before encoding
takes place. Huffman coding requires frequencies of symbol occurrence before codeword is
assigned to symbols
In Lempel-Ziv compression not rely on previous knowledge about the data rather builds this
knowledge in the course of data transmission/data storage. Lempel-Ziv algorithm (called LZ) uses
a table of code-words created during data transmission; and it transmits the index of the
symbol/word instead of the word itself. Each time it replaces strings of characters with a reference
to a previous occurrence of the string.
89
Lempel-Ziv Compression Algorithm
The multi-symbol patterns are of the form: C0C1 . . . Cn-1 Cn. The prefix of a pattern consists
of all the pattern symbols except the last: C0C1 . . . Cn-1
Lempel-Ziv Output: there are three options in assigning a code to each symbol in the list
If the last input symbol or the last pattern is in the dictionary, asign
(dictionaryPrefixIndex, )
Eg: Encode (i.e., compress) the string ABBCBCABABCAABCAAB using the LZ algorithm.
Note: The above is just a representation, the commas and parentheses are not transmitted
90
Codeword: (0, A) (0, B) (2, C) (3, A) (2, A) (4, A) (6, B)
Codeword index: 1 2 3 4 5 6 7
Activity 7.6.
Encode (i.e., compress) the following strings using the Lempel-Ziv algorithm.
SATATASACITASA
91
a single decimal number. The input symbols are processed one at each iteration. The interval
derived at the end of this division process is used to decide the code word for the entire sequence
of symbols.
UL= LL+d (ul)*d (f) Where LL: lower limit, d (u, l) difference of upper and lower & d (f) is
frequency of letter
For the first letter in B, the lower limit is zero, and the upper limit is 0.4.
UL= LL+d (ul)*d (f)
B =0 + (0.4 - 0) *0.4= 0+0.4*0.4=0.16
E= 0 + (0.4 - 0) *0.6= 0+0.4*0.6=0.24
L= 0 + (0.4 - 0) *0.8= 0+0.4*0.8=0.32
92
A= 0 + (0.4 - 0) *1= 0+0.4*1=0.4
Similar to others
A message is represented by a half-open interval [a, b) where a and b are real numbers between a
and 1. Initially, the interval is [0, 1). When the message becomes longer, the length of the interval
shortens, and the number of bits needed to represent the interval increases. Suppose the alphabet
is [A, B, C, D, E, F, $], in which $ is a special symbol used to terminate the message, and the
known probability distribution is listed below.
93
7.9. Lossless Image Compression
One of the most commonly used compression techniques in multimedia data compression is
differential coding. The basis of data reduction in differential coding is the redundancy in
consecutive symbols in a data stream. Audio is a signal indexed by one dimension, time. Here we
consider how to apply the lessons learned from audio to the context of digital image signals ~hat
are indexed by two, spatial, dimensions (x, y).
Let's consider differential coding in the context of digital images. In a sense, we move from signals
with domain in one dimension to signals indexed by numbers in two dimensions (x, y) - the rows
and columns of an image. Later, we'll look at video signals. These are even more complex, in that
they are indexed by space and time (x, y, t). Because of the continuity of the physical world, the
gray-level intensities (or color) of background and foreground objects in images tend to change
relatively slowly across the image frame. Since we were dealing with signals in the time domain
for audio, practitioners generally refer to images as signals in the spatial doma'in. The generally
slowly changing.
94
Lossless JPEG
Lossless IPEG is a special case of the JPEG image compression. It differs drastically from
other IPEG modes in that the algorithm has no lossy steps. Thus we treat it here and consider
the more used JPEG methods in Chapter 9. Lossless JPEG is invoked when the user selects
a 100% quality factor in an image tool. Essentially, lossless IPEG is included in the JPEG
compression standard simply for completeness. The following predictive method is applied on the
unprocessed original image (or each color band of the original color image). It essentially involves
two steps: forming a differential prediction and encoding.
Review Questions
1. Given the following symbols and their corresponding frequency of occurrence, find an
optimal binary code for compression
character A B C D E T
frequency 6 5 12 17 10 25
95
CHAPTER EIGHT
For example, if the reconstructed image were the same as original image except that it is shifted to
the right by one vertical scan line, an average human observer would have a hard time
distinguishing it from the original and would therefore conclude that the distortion is small.
However, when the calculation is carried out numerically, we find a large distortion, because of
the large changes in individual pixels of the reconstructed image. The problem is that we need a
measure of perceptual distortion, not a more naive numerical approach. Of the many numerical
distortion measures that have been defined, we present the three most commonly used in image
compression. If we are interested in the average pixel difference, the mean square error (MSE) 0'2
is often used. It is defined as:
96
Where Xn, Yn, and N are the input data sequence, reconstructed data sequence, and length
of the data sequence, respectively.
Activity 8.1.
What is distortion measure?
If we are interested in the size of the error relative to the signal, we can measure the signal to
noise ratio (SNR) by taking the ratio of the average square of the original data sequence and the
mean square error (MSE). In decibel units (dB), it is defined as:
Where σx2 is the average square value of the original data sequence and σd2 is the MSE.
Another commonly used measure for distortion is the peak-signal-to-noise ratio (PSNR), which
measures the size of the error relative to the peak value of the signal x peak. It is given by:
97
algorithms.
Figure 8.1 shows a typical rate-distortion function. Notice that the minimum possible rate at D =
0, no loss, is the entropy of the source data. The distortion corresponding to a rate R(D) = 0 is the
maximum amount of distortion incurred when "nothing" is coded. Finding a closed-form analytic
description of the rate-distortion function for a given source is difficult, if not impossible.
8.4. Quantization
Quantization in some form is the heart of any lossy scheme. Without quantization, we would
indeed be losing little information. The source we are interested in compressing may contain a
large number of distinct output values (or even infinite, if analog). To efficiently represent the
source output, we have to reduce the number of distinct values to a much small set, via
quantization.
Each algorithm (each quantizer) can be uniquely determined by its partition of the input range on
the encoder side and the set of output values, on the decoder side. The input and
output of each quantizer can be either scalar values or vector values, thus leading to scalar
quantizer and vector quantizer.
98
8.5. Transform Coding
From basic principles of information theory, we know that coding vectors is more efficient than
coding scalars. To carry out such an intention, we need to group blocks of consecutive samples
from the source input into vectors.
Let X = {X1, X2, ... , Xk}T be a vector of samples. Whether our input data is an image, a piece of
music, an audio or video clip, or even a piece of text, there is a good chance that a substantial
amount of correlation is inherent among neighboring samples Xi. The rationale behind transform
coding is that if Y is the result of a linear transform T of the input vector X in such a way that the
components of Yare much less correlated, then Y can be coded more efficiently than X.
For example, if most information in an RGB image is contained in a main axis, rotating so that
this direction is the first component means that luminance can be compressed differently from
color information. This will approximate the luminance channel in the eye. In higher dimensions
than three, if most information is accurately described by the first few components of a transformed
vector, the remaining components can be coarsely quantized, or even set to zero, with little signal
distortion. The more decorrelated – that is, the less effect one dimension has on another (the more
orthogonal the axes), the more chance we have of dealing differently with the axes that store
relatively minor amounts of information ,without affecting reasonably accurate reconstruction of
the signal from its quantized or truncated transform coefficients.
Generally, the transform T itself does not compress any data. The compression comes from the
processing and quantization of the components of Y. Discrete Cosine Transform (DCT) as a tool
to decorrelate the input signal.
The Discrete Cosine Transform (DCT), a widely used transform coding technique, is able
to perform decorrelation of the input signal in a data-independent manner. Because of this,
it has gained tremendous popularity. We will examine the definition of the DCT and discuss
some of its properties, in particular the relationship between it and the more familiar Discrete
Fourier Transform (DFT).
99
Definition of DCT: Let's start with the two-dimensional DCT. Given a function f (i, j)
over two integer variables i and j (a piece of an image), the 2D DCT transforms it into a
new function F(u, v), with integer u and v running over the same range as i and j.
Review Questions
1. What is distortion?
2. What does it mean D=0?
3. How to Quantize the images? Explain it
4. By which mechanism to transform encoding is done? Explain it.
100
CHAPTER NINE
COMPRESION
Recent years have seen an explosion in the availability of digital images, because of the increase
in numbers of digital imaging devices, such as scanners and digital cameras. The need to efficiently
process and store images in digital form has motivated the development of many image
compression standards for various applications and needs.
In 1986, the Joint Photographic Experts Group (JPEG) was formed to standardize algorithms for
compression of still images, both monochrome and color. JPEG was formally accepted as an
international standard in 1992. JPEG is a Lossy image compression method.
Activity 9.1.
How to Compress JPEG images?
As we know, unlike one-dimensional audio signals, a digital image f (i, j) is not defined over the
time domain. Instead, it is defined over a spatial domain - that is, an image is a function of the two
dimensions i and j (or, conventionally, x and y). The 2D DCT is used as one step in JPEG, to yield
a frequency response that is a function F (u, v) in the spatial frequency domain, indexed by two
integers II and v. JPEG is a Lossy image compression method. The effectiveness of the DCT
transform coding method in JPEG relies on three major observations:
101
DPCM on DC component and RLE on AC Components
Entropy Coding — Huffman or Arithmetic
Activity 9.2.
What does
Discrete it mean
cosine Quantization?
transform (DCT)
102
In this step, each block of 64 pixels goes through a transformation called the discrete cosine
transform (DCT). The transformation changes the 64 values so that the relative relationships
between pixels are kept but the redundancies are revealed. The formula is given in Appendix G.
P(x, y) defines one value in the block, while T(m, n) defines the value in the transformed block.
Quantization
After the table is created, the values are quantized to reduce the number of bits needed for
encoding. Quantization divides the number of bits by a constant and then drops the fraction. This
reduces the required number of bits even more. In most implementations, a quantizing table (8 by
8) defines how to quantize each value. The divisor depends on the position of the value in the table.
This is done to optimize the number of bits and the number of 0s for each particular application.
Compression
After quantization, the values are read from the table, and redundant 0s are removed. However, to
cluster the 0s together, the process reads the table diagonally in a zigzag fashion rather than row
by row or column by column. JPEG usually uses run-length encoding at the compression phase to
compress the bit pattern resulting from the zigzag linearization.
Activity 9.3.
Figure 9. 3 Displaying Quantization
Discuss briefly JPEG Modes?
103
9.2. JPEG Modes
The JPEG standard supports numerous modes (variations). Some of the commonly used ones
are:
• Sequential Mode
• Progressive Mode
• Hierarchical Mode
• Lossless Mode
Sequential Mode: This is the default JPEG mode. Each gray-level image or color image
component is encoded in a single left-to-right, top-to-bottom scan. We implicitly assumed this
mode in the discussions so far. The "Motion JPEG" video codec uses Baseline Sequential
JPEG, applied to each image frame in the video.
Progressive Mode: Progressive JPEG delivers low-quality versions of the image quickly,
followed by higher-quality passes, and has become widely supported in web browsers. Such
multiple scans of images are of course most useful when the speed of the communication line
is low. In Progressive Mode, the first few scans carry only a few bits and deliver a rough picture
of what is to follow. After each additional scan, more data is received, and image quality is
gradually enhanced. The advantage is that the user-end has a choice whether to continue
receiving image data after the first scan(s).
Progressive JPEG can be realized in one of the following two ways. The main steps (DCT,
quantization, etc.) are identical to those in Sequential Mode.
Spectral selection: This scheme takes advantage of the spectral (spatial frequency spectrum)
characteristics of the DCT coefficients: the higher AC components provide only detail
information.
Lossless Mode: Lossless JPEG is a very special case of JPEG, which indeed has no loss in its
image quality. As discussed in the previous chapter, however, it employs only a simple
differential coding method, involving no transform coding. It is rarely used, since its
compression ratio is very low compared to other, Lossy modes. On the other hand, it meets a
special need, and the newly developed JPEG-LS standard is specifically aimed at lossless
image compression.
104
CHAPTER TEN
Video compression techniques take advantage of the repetition of portions of the picture from one
image to another by concentrating on the changes between neighboring images. The two types of
redundancy in video frames.
The Moving Picture Experts Group (MPEG) method is used to compress video. In principle, a
motion picture is a rapid sequence of a set of frames in which each frame is a picture. In other
words, a frame is a spatial combination of pixels, and a video is a temporal combination of frames
that are sent one after another. Compressing video, then, means spatially compressing each frame
and temporally compressing a set of frames.
Spatial compression: The spatial compression of each frame is done with JPEG, or a modification
of it. Each frame is a picture that can be independently compressed.
105
Temporal compression: In temporal compression, redundant frames are removed. When we
watch television, for example, we receive 30 frames per second. However, most of the consecutive
frames are almost the same.
A video can be viewed as a sequence of images stacked in the temporal dimension. Since
the frame rate of the video since is often relatively high (e.g.: >15 frames per second) and the
camera parameters (focal length, position, viewing angle, etc.) usually do not change rapidly
between-frames, the contents of consecutive frames are usually similar, unless certain objects
in the scene move extremely fast. In other words, the video has temporal redundancy.
Temporal redundancy is often significant and it is exploited, so that not every frame of
the video needs to be coded independently as a new image. Instead, the difference between
the current frame and other frame(s) in the sequence is coded. If redundancy between them
is great enoug1t, the difference images could consist mainly of small values and low entropy,
which is good for compression.
Activity 10.1.
What is Compression in Video?
106
adopted. Motion compensation is not performed at the pixel level, nor at the level of video
object, as in later video standards (such as MPEG-4). Instead, it is at the macroblock level.
The current image frame is referred to as the Target frame. A match is sought between
the macroblock under consideration in the Target frame and the most similar macroblock in
previous and/or future frame(s) [referred to as Reference frame(s). In that sense, the Target
macroblock is predicted from the Reference macroblock.
107
CHAPTER ELEVEN
MPEG VIDEO AND AUDIO CODING
The Moving Pictures Experts Group (MPEG) was established in 1988 to create a standard for
delivery of digital video and audio. Membership grew from about 25 experts in 1988 to a
community of more than 350, from about 200 companies and organizations. It is appropriately
recognized that proprietary interests need to be maintained within the family of MPEG standards.
This is accomplished by defining only a compressed bit stream that implicitly defines the decoder.
The compression algorithms, and thus the encoders, are completely up to the manufacturers.
Compression is a reversible conversion (encoding) of data that contains fewer bits. This allows a
more efficient storage and transmission of the data. The inverse process is called decompression
(decoding). Software and hardware that can encode and decode are called decoders. Both
combined form a codec and should not be confused with the terms data container or compression
algorithms.
Figure 11. 1 Relation between codec, data containers and compression algorithms.
As we discuss in chapter 7, Lossless compression allows a 100% recovery of the original data. It
is usually used for text or executable files, where a loss of information is a major damage. These
compression algorithms often use statistical information to reduce redundancies. Huffman-Coding
and Run Length Encoding are two popular examples allowing high compression ratios depending
on the data.
Using lossy compression does not allow an exact recovery of the original data. Nevertheless, it can
be used for data, which is not very sensitive to losses and which contains a lot of redundancies,
108
such as images, video or sound. Lossy compression allows higher compression ratios than lossless
compression.
The MPEG-1 Standard was published 1992 and its aim was it to provide VHS quality with a
bandwidth of 1,5 Mb/s, which allowed to play a video in real time from a 1x CD-ROM. The frame
rate in MPEG-1 is locked at 25 (PAL) fps and 30 (NTSC) fps respectively. Further MPEG-1 was
designed to allow a fast forward and backward search and a synchronization of audio and video.
A stable behavior, in cases of data loss, as well as low computation times for encoding and
decoding was reached, which is important for symmetric applications, like video telephony.
In 1994 MPEG-2 was released, which allowed a higher quality with a slightly higher bandwidth.
MPEG-2 is compatible to MPEG-1. Later it was also used for High Definition Television (HDTV)
and DVD, which made the MPEG-3 standard disappear completely. The frame rate is locked at 25
(PAL) fps and 30 (NTSC) fps respectively, just as in MPEG-1. MPEG-2 is more scalable than
MPEG-1 and is able to play the same video in different resolutions and frame rates.
MPEG-4 was released 1998 and it provided lower bit rates (10Kb/s to 1Mb/s) with a good quality.
It was a major development from MPEG-2 and was designed for the use in interactive
environments, such as multimedia applications and video communication. It enhances the MPEG
family with tools to lower the bit-rate individually for certain applications. It is therefore more
adaptive to the specific area of the video usage. For multimedia producers, MPEG-4 offers a better
reusability of the contents as well as a copyright protection. The content of a frame can be grouped
into object, which can be accessed individually via the MPEG-4 Syntactic Description Language
(MSDL). Most of the tools require immense computational power (for encoding and decoding),
which makes them impractical for most “normal, non- professional user” applications or real time
applications. The real-time tools in MPEG-4 are already included in MPEG-1 and MPEG-2.
109
predictive encoding, the differences between samples are encoded instead of encoding all the
sampled values. This type of compression is normally used for speech.
The most common compression technique used to create CD-quality audio is based on the
perceptual encoding technique. This type of audio needs at least 1.411 Mbps, which cannot be sent
over the Internet without compression. MP3 (MPEG audio layer 3) uses this technique.
A high quality digital audio representation can be achieved by directly measuring the analogue
waveform at equidistant time steps along a linear scale. The distance between the grid markers on
the time axis is defined by the sample frequency, and on the amplitude axis by the maximal
amplitude divided by the number of bits. Therefore, higher sample frequencies give narrower grids
along the time axis and more bits (narrower grids) along the amplitude axis. The individual values
to be stored on computers and representing the sound wave can only be located on the cross points
of the underlying grid. Since the actual sound pressure will never be exactly on these cross-points,
errors are made, i.e. digitization noise is introduced. It is obvious that the narrower the grids are
the smaller the error will be. Therefore, a higher sample frequency and a larger number of bits will
increase the quality.
However, for speech signals we know that the highest frequency a child’s voice can create is in
the order of 8 kHz. According to the Nyquist rule, a double as high sample frequency is sufficient
to represent a frequency component nicely, so 20 kHz is sufficient. Human ear is sensitive up to
20 kHz, therefore the HiFi norm requires sampling frequencies above 40 kHz. Also for normal
speech signals and singing, a dynamic range of 96 dB is said to be sufficient which is achieved by
16 bits. For certain pieces of classic music, however, 96 dB is not sufficient to hear the low sound
level signals as well as the high level ones. Therefore, for classic music 24 bits are recommended
that yield a dynamic range of up to 144 dB.
110
perceptually irrelevant parts of the audio signal. Removal of such parts results in inaudible
distortions, thus MPEG/audio can compress any signal meant to be heard by the human ear. In
keeping with its generic nature, MPEG/audio offers a diverse assortment of compression modes:
Review Question
1. What compression in
b. Image,
c. Audio and
d. Video?
2 What is the difference between spatial redundancy and temporal redundancy?
111
Lab Manual
112
Figure 2 The document property of Macro Media Flash
Setting the Frame Rate to 24 will allow animations to play smoothly, if the computer is fast
enough to render the movements. Twenty-four frames per second is a good guideline to
follow for using your animations on the internet.
Layout - Stage and Timeline:
Below is the default layout for Macromedia Flash MX. The timeline indicates what frame
you are at and indicates the number of frames in your movie. Within the timeline you will
find layers - you can have any number of layers within a movie and it is within these layers
that you put your graphics, text, and sounds. The Work Area is not viewable when you
play your movie, so it is a place to work on objects or if you want your objects to “fly in”
to your movie then start them from the Work Area. The Stage is where all viewable objects
lie. Anything on the stage is seen by the user.
113
Figure 3 The layout of Macromedia Flash Menu
Here is the layout for Adobe Flash Professional
Toolbox
Stage
Work Area
Timeline
Toolbox:
114
The toolbox contains all tools necessary for drawing, viewing, coloring and modifying your
objects. Each tool in the toolbox comes with a specific set of options to modify that tool.
The diagram below outlines the grouping of tools.
Figure 5 Toolbox
Motion tweening is the ability to move an object in either linear or non-linear fashion. This
must be created carefully in order to achieve the desired results. This tutorial will outline
the steps in order to achieve a motion tween.
115
Figure 6 Covert to symbol dialogue box
In order to complete a motion tween you must first convert any object
to a graphic symbol.
5. Now your ball is a Symbol (Button) and is ready for movement.
6. Right click on Frame 15 of layer 1 and choose Insert Frame. This will create a length for
your animation (15 frames).
7. Highlight the first frame by a single left mouse click on the frame. The frame should
turn black.
8. Go to Insert - Create Motion Tween. You will notice that a dotted line will fill in between
frame 1 and frame 15.
9. You are now ready to preview your animation. Goto Control - Test Movie
(Keyboard shortcut = Ctrl + Enter)
Output:
116
Example 2: To create an animation to represent the growing moon.
1. Open Adobe Flash Professional-> click file -> new > Action script 3.0 ->ok. Go to
windows->properties->select the properties tool-> choose the Background to black.
2. Go to fill color under tool bar-> select the white color.
3. Select the oval tool in order to draw the moon. u will get a white circle.
4. Select the oval tool in order to draw the moon. u will get a white circle.
5. Select the white circle on the worksheet using the selection tool->right click- >convert
to symbol->select movie clip->give suitable name eg: moon->click ok.
6. Go to filter->click on the + symbol->select glow to apply glowing effect-> select the
color to white under glow and adjust the blur x/blur y values.
7. Click on the + symbol again and chose blur-> again adjust the blur x/blur y values.
8. Place the moon where ever you want on the work area. double click on layer 1 and
rename as MOON.
9. Insert another layer->rename it as Animation.
10. Select the fill color to black-> select oval tool and draw a circle on the moon to cover
the moon->select the newly added circle-> right click-> convert to symbol-> movie clip-
> name it as Animation.
11. Go to filter-> select + symbol->give the glow and blur effect as did for moon.
12. Select the 150th frame in moon layer->right click->insert key frame. repeat the same
for Animation layer.
13. Click on the 149th keyframe of animation layer ->right click->press create motion->
select the animation movie clip and move slowly across the moon.
14. Finally go to control-> test movie-> u will get a growing moon as the output.
Output:
117
Figure 8 To create an animation to represent the growing moon
1. Open Adobe Flash Professional-> click file -> new > Action script 3.0 ->ok
2. Select the line tool and draw the steps. Color it using the paint bucket tool
3. Select the circle from the tool bar and create a circle on the work area.
4. Now fill the color to the circle using the paint bucket tool from the tool bar.
5. Go to frames right click on the first frame and choose insert key frame. Slightly move
the ball. Repeat the same procedure by adding new key frames to show the ball change the
shape of the ball slightly when it touches the surface.
6. In order to change the shape use the free transform tool.
7. Go to control and click on test movies .you will observe the ball bouncing on steps.
Output:
118
Figure 9 To create an animation to indicate a ball bouncing on steps
1. Open Adobe Flash Professional-> click file -> new > Action script 3.0 ->ok
2. Go to file-> import->open external library-> select a background image Click open.
3. The selected image will be stored in your library. Open library and drag the image on
the work area by selecting the image.
4. Resize the image to fit on the work area.
5. Select the text tool from the tool bar.
6. Type your name. Select the text and go to property to apply appropriate font effects like
font size, style and Color etc.
7. Go to control-> movie clip to see the final output.
Output:
119
Figure 10 To create the background with the text as you want
1. Go to start->Adobe Flash Professional -click File -> new -> Action Script 3.0.
2. Choose the textbox from the tool bar. Type the word as ‘WELCOME’ on layer1.
3. Select the complete word, increase its Font size and change the colour.
4. In the timeline window, select the 1st frame-Right click on it-choose insert key frame.
Now delete a last letter {E} and change the colour of the remaining word.
5. Repeat the above procedure till you delete every letter in ‘WELCOME’.
6. Now select all the key frames->Right click ->choose ‘Reverse key frames’.
7. After reversing the frames copy the last frame and paste on its next. now in the new
frame Select all the complete word ‘welcome’ and change the colour.
8. Finally, go to ‘control’-click on ‘test movie’ you will get the required animation.
Welcome letters should appear one by one. The fill color of the text should change to a
different color, then display of the full word.
120
Output:
Example: 6 To display the background given image through your name using mask.
1. Go to start->Adobe Flash Professional -click File -> new -> Action Script 3.0.
2. Go to file-> import->open external library-> select a background image Click open.
3. The selected image will be stored in your library. Open library and drag the image on
the work area by selecting the image.
4. Go to view->zoom out->resize the picture such that it should fit the work area.
[Link] layer2. choose the text tool from the toolbar and type your name.
[Link] the text to change its font size and color of your choice. place the text on the left
of the work area.
[Link] click on the 70th keyframe of layer 2 and insert a keyframe. move the text to the
right side of the work area->right click on the 69th frame of layer 2-> choose create
motion tween.
[Link] click on layer 2 choose the option mask.
121
[Link] to control->test movie to see the animation.
Output:
Figure 12 To display the background given image through your name using mask.
ADOBE PHOTOSHOP:
Example 1: To design a visiting card containing at least one Graphic and text
information.
1. Open Adobe Photoshop CS 6-> file-> new-> enter height 300 and width 400 for the
Visiting card.
122
You can set the width and height of the main area in pixel or centimeter and set the
resolution value. Most of the time for banner or photo it is better to use cm. You can also
set the color mode (RGB for computer and CMY for printable uses)
2. Select the rectangle tool in the tool bar and draw on the half of the work area-> color
it. Repeat the same for remaining half-> use different colors to color.
3. Copy any picture of your choice and place it on the work area-> Resize it using free
transform tool.
4. Select the text tool and type text of your choice.
5. Apply the text font size, color and style of your choice.
Output:
123
Figure 13 To design a visiting card
Example: 2 To design a photographic image, give a title for the image and design
border.
1. Open Adobe Photoshop CS 6-> file-> new-> enter height 400 and width 400
2. Open an image file and copy the image->paste the copied image on the new file.
3. Right click on the rectangle tool->custom shape-> select the shape->select the color-
>drag on your image.
4. Select the text tool->type the text as you want.
5. Save the file as the file extension psd (photoshop).
Output:
124
Figure 14 To design a photographic image give a title and design border
1. Open Adobe Photoshop CS 6-> file-> new-> enter height 500 and width 400 for the
cover page.
2. Select the rectangle tool in the tool bar and draw on the half of the work area-> color
it. Repeat the same for remaining area-> use different colors to color it.
3. Copy any picture of your choice and place it on the work area->resize it using free
transform tool (edit-> free transform and you can resize it).
4. Select the text tool and type text of your choice.
5. Apply the text font size, color and style of your choice.
6. Go to layer->layer style->blended option-> select glow options of your choice.
7. Apply the effects using blended options.
Output:
125
Figure 15 To design a cover page for the book
Example 4: To design the image that extracted from different pictures and change
the background color.
1. Open Adobe Photoshop CS 6-> file->open-> choose a file and open it.
2. Select the picture you want to using the lasso tool.
3. Go to edit-> copy->Again go to file->new->give height 500 and width 500.
4. Choose appropriate background and foreground color from the tool bar.
5. Go to edit->fill->under use select background color->ok.
6. Go to edit->paste.
Output:
126
Figure 16 To design the image from different pictures and change the background color
1. Open Adobe Photoshop CS 6-> file->open-> choose a file and open it.
2. Go to image->adjustments->Brightness/Contrast.
3. After getting the Brightness/Contrast window adjust the brightness and contrast by
Dragging the appropriate bar setting.
4. Finally save the image file.
Output:
127
1. Open Adobe Photoshop CS 6-> file->open-> choose a file and open it.
2. Select the flower from the image using the lasso tool.
3. Go to edit-> copy->Again go to file->new->give height 500 and width 500.
4. Choose appropriate background and foreground color from the tool bar.
5. Go to edit->fill->under use select background color->ok.
6. Go to edit->paste->again go to edit->free transform tool-> you will get a box around
the image for scaling and rotating.
7. Rotate and scale as per your requirement->and press apply.
8. Save the image.
Output:
Program No: 7 To design a word and apply different effects shadow emboss
1. Open Adobe Photoshop CS 6-> file->open-> choose a file and open it.
2. Select the text tool and place on the work area-> type your institute name
3. Select the typed text go to layer->layer style->blended option-> tick drop shadow,
inner shadow, bevel and emboss->contour->satin->gradient overlay.
4. Finally save the image.
128
Output:
Example: 8 To design cut images from different file and organize them in a single file
and apply feather effects.
Output:
129
Figure 20 To design different file and organize them in a single file and apply feather effects.
Output:
130
Figure 21 To change the image color as Black and White
In first page put size 4.5cmx7cm then to move wallet photo press ctrl+alt+move the cropped
image in to the new page (10cmx15cm).
131
Figure 22 To design a wallet photo
Step 2: go to file click new show click create and click continue to wizard
Step 3: select the theme you want to insert or select All effects to add the video
132
Figure 23 The layout of ProShow gold
Step 4: click continue and add some content like pictures by clicking add content on the top of
the screen and the select the picture that you want to add. Next to this click continue + preview.
After this, the selected picture is ready for in the editing page. Click Apply + exit wizard if you
finish select the picture, or you can go back the wizard if you edit something.
133
Figure 24 Add contents in the ProShow gold
Step 5: you can add more pictures or video by dragging from the file list to the editing Page.
134
Menu
Preview
File list Area
Editing
page
135
Add black
136
Figure 27 To adjust the text or captions in text effect
Step 7: or select the slide, go to the effect menu and click it.
Step 8: you can add sound track or music in the picture or video by go to in the music menu.
Click + add sign in the soundtrack. Then click add sound file and select the music from your
computer. Select the audio and click OK.
137
+ add (1)
138
Select
the one
(2)
Click
Apply
139
Preview area
Click here
140
Publish (1)
Click
video file
(2)
141
References
142