Module2 MMC
Module2 MMC
• 2.1 Introduction
• 2.2 Digitization principles
• 2.3 Text
• 2.4 Images
• 2.5 Audio
• 2.6 Video
2
Introduction
3
Multimedia Information Representation
4
2.2 Digitization principles
2.2.1 Analog signals
• As mentioned earlier the amplitude of the signal varies
continuously with time
• The Fourier analysis can be used to show that any time varying
signal is made up of infinite number of single-frequency
sinusoidal components
• The range of frequencies of the sinusoidal components that make
up the signal is called the signal bandwidth
• Speech bandwidth: 50Hz – 10kHz
• Music Bandwidth: 15Hz – 20kHz 5
Analog Signals –Signal Properties
6
Analogue Signals –Signal Properties fig 2.1 contd..
8
2.2.2 Encoder design
• A bandlimiting filter and an analog-to-digital converter(ADC), the latter
comprising a sample-and-hold and a quantizer
• Remove selected higher-frequency components from the source signal (A)
• (B) is then fed to the sample-and-hold circuit
• Sample the amplitude of the filtered signal at regular time intervals (C)
and hold the sample amplitude constant between samples (D)
• Quantizer circuit which converts each sample amplitude into a binary
value known as a codeword (E)
• The signal to be sampled at a rate which is higher than the maximum rate
of change of the signal amplitude
• The number of different quantization levels used to be as large as possible
9
Encoder Design
• The most significant bit of the codeword represents the sign of the
sample
• A binary 0 indicates a positive value and a binary 1 indicates a negative
value
• The signal must be sampled at a much higher rate than the maximum
rate of change of the signal amplitude
• The number of quantization levels should be as large as possible to
represent the signal accurately
11
Sampling Rate
12
Alias signal generation due to undersampling
14
Quantization Intervals
• Representation of the analogue samples would require an infinite
number of digits
16
Quantization noise polarity
17
Dynamic Range
18
Digitization Principles (Analog Digital)
As we can see from these values, with 6 bits the quantization noise is greater than Vmin and hence is
unacceptable. With 10 bits, however, the quantization noise is an order of magnitude less than Vmin and
hence will have a much-reduced effect.
19
Decoder Design
Encoder+decode = Codec
21
Text
22
Unformatted Text – The basic ASCII character set
• Printable characters
(alphabetic, numeric, and
punctuation)
•Control characters
(Format control: Back space,
escape, delete, form feed,
line feed, CR etc.)
(Transmission control
characters)
(Information Separators: FS
& RS)
• The American Standard Code for Information Interchange is one of the most widely
used character sets and the table includes the binary codewords used to represent each
character (7 bit binary code) 23
Unformatted Text – Supplementary set of Mosaic
characters
• The characters in columns 010/011 and 110/111 are replaced with the set of
mosaic characters; and then used, together with the various uppercase
characters illustrated, to create relatively simple graphical images
24
Unformatted Text – Examples of
Videotext/Teletext
• Although in practice the total page is made up of a matrix of symbols and characters
which all have the same size, some simple graphical symbols and text of larger sizes can
be constructed by the use of groups of the basic symbols
•Application of this is in Videotext or Teletext 25
Formatted Text
Formatted
Text string
• Hypertext can be used to create an electronic version of documents with the index,
descriptions of departments, courses on offer, library, and other facilities all written in
hypertext as pages with various defined hyperlinks 27
Hypertext – Electronic Document in
hypertext
29
Images: Graphics
• Images include computer-generated images (referred to as computer graphics
or simply graphics) and include digitized images of both documents and
pictures
• All types of images are displayed in the form of a two-dimensional matrix of
individual picture elements (pixels or pels), but represented differently within
the computer memory (file)
• Each type of these images are created differently.
•Various s/w packages or programs are used to generate these graphical
images.
•E.g. paintbrush ( select objects or draw by pencil), clipart images
•Textual information or overlapping objects shall also be added in these 30
Graphics: Display screen
• VGA is a common type of display that consists of a matrix of 640 horizontal pixels by 480
vertical pixels with for example, 8 bits per pixel which allows each pixel to have one of 256
different colours
31
Graphics
• All objects are made up of a series of lines that are connected to each other
and, what appear as a curved line, in practice is a series of short lines each
made up of a string of pixels
• Each object has a number of attributes associated with it. These include
its shape, size in terms of pixel position, colour of the border etc..
32
Graphics - Conclusions
33
Digitized Documents: Fax Principles
34
Digitized Documents- Digitization format
• Fax machines uses a single binary digit to represent each pel, a 0 for a white pel and a 1
for a black pel. Hence the digital representation of a scanned page produces a stream about
2 million bits.
• Single binary digit per pel means fax machines are best suited for bitonal images. 35
Digitised Pictures
• For digitizing continuous-tone mono-chromatic images - more than a single bit is used to
digitize each picture element.
• For example, good quality black-and-white pictures can be obtained by using 8 bits per
picture element.
• This yields 256 different levels of gray per element-varying between white and black-
which gives a substantially improved picture quality over a facsimile image when
reproduced.
• In the case of color images, in order to understand the digitization format used, it is
necessary first to obtain an understanding of the principles of how color is produced and
how the picture tubes used in computer monitors (on which the images are eventually
displayed) operate.
36
Colour Derivative Principles – additive
colour mixing ( R + G + B)
•Black is produced when all three primary colours (R,G,B) are zero.
• Useful for producing a colour image on a black surface as is the case in display
applications
37
Colour Derivative Principles - Subtractive
colour mixing
• The picture tubes used in most television sets operate using what is
known as a raster-scan; this involves a finely-focussed electron
beam being scanned over the complete screen 39
Digitized Pictures- Raster Scan
41
Digitized Pictures – Raster scan display architecture
•Frame: Each complete set of horizontal scan lines (either 525 for North & South America
and most of Asia, or 625 for Europe and other countries)
•Flicker: Caused by the previous image fading from the eye retina before the following
image is displayed, after a low refresh rate ( to avoid this a refresh rate of 50 times per
second is required) 42
Raster scan display architecture
•Pixel depth: Number of
bits per pixel that
determines the range of
different colours that can
be produced
• Colour Look-up Table
(CLUT): Table that
stores the selected
colours in the subsets as
an address to a location
reducing the amount of
memory required to
store an image 43
Digitized Pictures – Concepts
• Aspect Ratio: This is the ratio of the screen width to the screen
height ( television tubes and PC monitors have an aspect ratio of 4/3
and wide screen television is 16/9)
44
Digitized Pictures – Screen Resolutions
• Aspect ratio
• Both the number of pixels per scanned line and the number of
lines per frame
• The ratio of the screen width to the screen height
• National Television Standards Committee (NTSC), PAL(UK),
CCIR(Germany), SECAM (France)
• Table 2.1
46
2.4.3 Digitized pictures
47
Digitized Pictures(5)
Example 2.3
Derive the time to transmit the following digitized images at both 64Kbps and 1.5Mbps networks
a 6404808 VGA-compatible image
a 102476824 SVGA-compatible image
Solution
The size of each image in bit is as follows
a VGA image = 6404808 = 2.46Mbits
an SVGA image = 102476824 =18.88Mbits
The time to transmit each image is given as follows
49
Digitized Pictures – Colour Image Capture: Schematic
•Charge-coupled devices
(CCD): Image sensor that
converts the level of light
intensity on each photosites
into an equivalent electrical
charge
[Link] color associated with each photosite is determined by O/p of R,G or B together with
nearby 8 neighbours.
[Link] uses three separate exposure filters for obtaining charges with each RGB filter
[Link] exposure comprising of split three light beams for a separate image sensor(RGB)51
Audio
• Two types of audio
– Speech
– Music type audio
• An audio produced by
– Naturally with Microphone Analog
– Electronically using some synthesizers Digital
• Nyquist rate
– Speech is 20 KHz (20Kbps)
– Audio is 40 KHz(40Kbps)
• To limit the quantization noise use of 12 bits for speech and 16 bits for music are
tested to prove this consideration
• In case of stereophonic sounds a double bit rate is used that of mono signal.
52
Audio
Example 2.4
Assuming the bandwidth of a speech signal is from 50 Hz through to 10kHz and that of a music signal is
from 15 Hz through to 20 kHz, derive the bit rate that is generated by the digitization procedure in each
case assuming the Nyquist sampling rate is used with 12 bits per sample for the speech signal and 16 bits
per sample for the music signal. Derive the memory required to store a 10-minute passage of stereophonic
music.
• Solution
(i) Bit rates: Nyquist sampling rate = 2 fmax
Speech: Nyquist rate = 2 10 kHz = 20 kHz or 20 ksps
Hence with 12 bits per sample, bit rate generated = 20 k 12 = 240 kbps
Music: Nyquist rate = 2 20 kHz = 40 kHz or 40 ksps
Hence bit rate generated = 40 k 16 = 640kbps (mono)
or 2 640k = 1280 kbps (stereo)
(ii) Memory required: Memory required = bit rate (bps) time (s)/8 bytes
Hence at 1280 kbps and 600 s,
1280 × 103 × 600
𝑀𝑒𝑚𝑜𝑟𝑦 𝑟𝑒𝑞𝑢𝑖𝑟𝑒𝑑 = = 96 𝑀𝑏𝑦𝑡𝑒𝑠
8 53
AUDIO:PCM speech
• It is a digitization process.
• Defined in ITU-T recommendations G.711
• PCM consists of encoder and decoder
• It consists of expander and compressor
• As compared to earlier where linear quantization is used – noise level same for both
loud and low signals.
• As ear is more sensitive to noise on quite signals than loud signals, PCM system
consists of non-linear quantization with narrow intervals through compressor
• At the destination expander is used.
• The overall operation is companding.
• Before sampling and using ADC, signal passed through compressor first and passed to
ADC and quantized.
• At the receiver, codeword is first passed to DAC and expander
• Two compressor characteristics – A-law and µ-law
54
PCM Principles
• Figure 2.17 PCM principles: (a) signal encoding and decoding schematic;
56
Figure 2.17 Continued (c) expander characteristic;
• Note that in the G.711 standard a 3-bit segment code and 4-bit quantization code are used.
57
2.5.2 CD-quality audio
Standard for CD players and CD ROMS –CD-DA(CD digital Audio) standard
• Music –audible BW of 15Hz to 20KHz and min sampling rate of 40ksps.
• Actual rate is higher than this to allow imperfections in band limiting filter used,
and the resulting bit rate is then compatible with one of the higher transmission
channel bit rates available in public networks.
• One of the sampling rates used is 44.1ksps which means that the signal is sampled at
23 microsecond intervals.
• BW of recording channel on a CD is large, a high number of bits per sample can be
used.
• The standard defines 16 bits per sample, which is the minimum requirement with
music to avoid the effect of quantization noise.
• Linear quantization can be used with these number of bits that yields 65536 equal
quantization intervals.
• For stereophonic music, two separate channels are required and hence the total bit
58
61
MIDI: Musical Instrument Digital Interface
• Use the sound card's defaults for sounds: use a simple scripting language and hardware
setup called MIDI.
• MIDI Overview
(a) MIDI is a scripting language it codes "events" that stand for the production of sounds.
E.g., a MIDI event might include values for the pitch of a single note, its duration, and its
volume.
(b) MIDI is a standard adopted by the electronic music industry for controlling devices, such
as synthesizers and sound cards, that produce music.
(c) The MIDI standard is supported by most synthesizers, so sounds created on one
synthesizer can be played and manipulated on another synthesizer and sound reasonably
close.
(d) Computers must have a special MIDI interface, but this is incorporated into most sound
cards. The sound card must also have both D/A and A/D converters.
62
MIDI: Concepts
• MIDI channels are used to separate messages.
(a) There are 16 channels numbered from 0 to 15. The channel forms the last 4 bits (the least significant
bits) of the message.
(b) Usually a channel is associated with a particular instrument: e.g., channel 1 is the piano, channel 10 is
the drums, etc.
(c) Nevertheless, one can switch instruments midstream, if desired, and associate another instrument with
any channel.
• A. Channel messages: can have up to 3 bytes:
a) The first byte is the status byte (the opcode, as it were); has its most significant bit set to 1.
b) The 4 low-order bits identify which channel this message belongs to (for 16 possible channels).
c) The 3 remaining its hold the message. For a data byte, the most significant bit is set to 0.
A.1. Voice messages:
a) This type of channel message controls a voice, i.e., sends information specifying which note to play or
to turn off, and encodes key pressure.
b) Voice messages are also used to specify controller effects such as sustain, vibrato, tremolo, and the
pitch wheel.
63
2.6 Video
2.6.1 Broadcast television
• As per previous discussion the screen is coated with three different color
phosphorous triods each of it is activated by an electronic beam.
• Scanning sequence of the screen is left to right and top to bottom with
resolution of 525 lines (NTSC) and 625 for PAL/CCIR/SECAM.
• It is necessary to use a minimum refresh rate of 60/50 times per second to
avoid flicker
• A refresh rate of 25 times per second is sufficient for smooth production of
motion sequences.
• Field: the first transmission comprising only the odd scan lines and the
second transmission the even scan lines .
• The two field are then integrated together in the television receiver using a
technique known as interlaced scanning 64
2.6.1 Broadcast television
66
2.6.1 Broadcast television
67
• Example 2.6:Derive the scaling factors used for both the U and V (as used in PAL)
and I and Q (as used in NTSC) color difference signals in terms of the three R, G,
B color signals.
• Solution: PAL: Y = 0.299R + 0.587G + 0.114B
• U = 0.493 (B – Y) and V = 0.877 (R – Y)
• Hence U = 0.493B – 0.493 (0.299R + 0.587 G + 0.114B)
• = –0.147R – 0.289G + 0.437B
• and V = 0.877R – 0.877 (0.299R + 0.587G + 0.114B)
• = 0.615R – 0.515G – 0.100B
• NTSC: I = 0.74 (R – Y) – 0.27 (B – Y)
• = 0.74R – 0.27B – 0.47Y
• = 0.599R – 0.276G – 0.324B
• Q = 0.48 (R – Y) + 0.41 (B – Y)
• = 0.48R + 0.41B – 0.89Y
• = 0.212R – 0.528G + 0.311B
68
2.6.1 Broadcast television: Signal Bandwidth
• Luminance is occupied by
lower frequency band
• And Chrominance is
transmitted in upper frequency
band with two separate
subscribers, This is to avoid
interference.
• Audio is transmitted separately
with two or more subscribers.
69
Digital Video
• The advantages of digital representation for video are many. For example:
(a) Video can be stored on digital devices or in memory, ready to be
processed (noise removal, cut and paste, etc.), and integrated to various
multimedia applications;
(b) Direct access is possible, which makes nonlinear video editing
achievable as a simple, rather than a complex, task;
(c) Repeated recording does not degrade image quality;
(d) Ease of encryption and better tolerance to channel noise.
70
Chroma Subsampling
• Since humans see color with much less spatial resolution than
they see black and white, it makes sense to "decimate" the
chrominance signal.
• Interesting (but not necessarily informative!) names have
arisen to label the different schemes used.
• To begin with, numbers are given stating how many pixel
values, per four original pixels, are actually sent:
(a) The chroma subsampling scheme "4:4:4" indicates that no
chroma subsampling is used: each pixel's Y, Cb and Cr values
are transmitted, 4 for each of Y, Cb, Cr.
71
Chroma Subsampling
• Eye have shown that the resolution of the eye is less sensitive for
color than it is for luminance
• 4:2:2 format
• The original digitization format used in Recommendation CCIR-601
• A line sampling rate of 13.5MHz for luminance and 6.75MHz for the
two chrominance signals
• Sampling rate of 13.5 MHz yields 52 micro seconds of active sweep
time
• 52 × 10-6 × 13.5 × 106= 702 samples per line
• The number of samples per line is increased to 720
74
Figure 2.21 Sample positions with 4:2:2
digitization format.
75
2.6.2 Digital video
76
Example 2.7: Solution
• 525-line system: The number of samples per line is 720 and the number of visible lines is
480. Hence the resolution of the luminance (Y) and two chrominance (Cb and Cr) signals
are: Y = 720 × 480
• Cb = Cr = 360 × 480
• Bit rate: Line sampling rate is fixed at 13.5 MHz for Y and 6.75 MHz for both Cb and Cr, all
with 8 bits per sample.
• Hence: Bit rate = 13.5 × 106 × 8 + 2 (6.75 × 106 × 8) = 216Mbps
• Memory required: Memory required per line = 720 × 8 + 2 (360 × 8)
• = 11 520 bits or 1440 bytes
• Hence memory per frame, each of 480 lines = 480 × 11520
• = 5.5296Mbits or 691.2kbytes
• and memory to store 1.5 hours assuming 60 frames per second:
• = 691.2 × 60 × 1.5 × 3600kbytes
• = 223.9488 Gbytes
77
Example 2.7: Solution contd..
• 625-line system: Resolution: Y = 720 × 576
• Cb = Cr = 360 × 576
• Bit rate = 13.5 × 106 × 8 + 2 (6.75 × 106 × 8) = 216Mbps
• Memory per frame = 576 × 11 520 = 6.63555 Mbits or 829.44 kbytes
• and memory to store 1.5 hours assuming 50 frames per second:
• = 829.44 × 50 × 1.5 × 3600 kbytes
• = 223.9488 Gbytes
• It should be noted that, in practice, the bit rate figures are less than the computed values
since they include samples during the retrace times when the beam is switched off.
Nevertheless, as we can deduce from the computed values, both the bit rate and the
memory requirements are very large for both systems and it is for this reason that the
various lower resolution formats have been defined.
78
Quiz on Circuit
switching and
packet
switching
79
2.6.2 Digital video
• 4:2:0 format is used in digital video broadcast applications
• Interlaced scanning is used and the absence of chrominance samples in alternative
lines
• The same luminance resolution but half the chrominance resolution
• Fig2.22
80
Figure 2.22 Sample positions in 4:2:0 digitization
format.
81
2.6.2 Digital video
525-line system Y 720 480
Cb Cr 360 240
625-line system
Y 720 480
Cb Cr 360 288
82
2.6.2 Digital video:HDTV
• HDTV formats: the resolution to the newer 16/9 wide-screen tubes can
be up to 1920 × 1152 pixels. (4/3 screen resolution 1440 × 1152 pixels).
• No. of visible lines per frame are 1080.
• Uses 4:2:2 digitization format for studio application & 4:2:0 format for
broad cast applications.
• Frame refresh time 50/60Hz with 4:2:2 and half of this for 4:2:0.
• Worst case bit rate is 4 times the values specified for other formats.
83
2.6.2 Digital video
84
Figure 2.23 Sample positions for SIF and CIF.
85
Figure 2.24 Sample positions for QCIF.
S-QCIF: 𝑌 = 128 × 96
𝐶𝑏 = 𝐶𝑟 = 64 × 48
86
87
88
89
2.6.3 PC video
90
2.6.4 Video content