0% found this document useful (0 votes)
14 views37 pages

Multimedia Data Processing Overview

The document provides an overview of multimedia data processing, defining multimedia as a combination of various content forms like text, audio, and video. It categorizes multimedia into linear and non-linear types, discusses its applications across different fields such as education, entertainment, and journalism, and explains the basic elements of multimedia. Additionally, it covers digital images, including types like raster and vector graphics, and their characteristics and uses.

Uploaded by

tekangglory
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views37 pages

Multimedia Data Processing Overview

The document provides an overview of multimedia data processing, defining multimedia as a combination of various content forms like text, audio, and video. It categorizes multimedia into linear and non-linear types, discusses its applications across different fields such as education, entertainment, and journalism, and explains the basic elements of multimedia. Additionally, it covers digital images, including types like raster and vector graphics, and their characteristics and uses.

Uploaded by

tekangglory
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MULTIMEDIA DATA PROCESSING

Code: SWE 116

Introduction

Multimedia is a form of communication that combines different content forms such as text, audio, images,
animations, or video into a single presentation, in contrast to traditional mass media, such as printed material
or audio recordings. Popular examples of multimedia include video podcasts, audio slideshows, animated
shows, and movies.

Multimedia presentations may be viewed by person on stage, projected, transmitted, or played locally with
a media player. A broadcast may be a live or recorded multimedia presentation. Broadcasts and recordings
can be either analog or digital electronic media technology. Digital online multimedia may be downloaded
or streamed. Streaming multimedia may be live or on-demand.

Multimedia games and simulations may be used in a physical environment with special effects, with
multiple users in an online network, or locally with an offline computer, game system, or simulator.

Categorization.

Multimedia may be broadly divided into linear and non-linear categories:

 Linear active content progresses often without any navigational control for the viewer such as a
cinema presentation;
 Non-linear uses interactivity to control progress as with a video game or self-paced computer-based
training. Hypermedia is an example of non-linear content.

Multimedia presentations can be live or recorded:

 A recorded presentation may allow interactivity via a navigation system;


 A live multimedia presentation may allow interactivity via an interaction with the presenter or
performer.

Basic Elements of Multimedia

 Texts

They are characters that are used to create words, sentences and paragraphs.

Page | 1
 Graphics

A digital representation of non-text information, such as a drawing, chart, or photograph.

 Animation

Flipping through a series of still images. It is a series of graphics that create an illusion of motion.

 Audio

It is a music, speech, or any other sound.

 Video

It is photographic images that are played back at speeds of 15 to 30 frames a second and it provide the
appearance of full motion.

Page | 2
Usage/application
Multimedia finds its application in various areas including, but not limited to, advertisements, art, education,
entertainment, engineering, medicine, mathematics, business, scientific research and spatial temporal
applications. Several examples are as follows:

 Creative industries

Creative industries use multimedia for a variety of purposes ranging from fine arts, to entertainment, to
commercial art, to journalism, to media and software services provided for any of the industries listed below.
An individual multimedia designer may cover the spectrum throughout their career. Request for their skills
range from technical, to analytical, to creative.

 Commercial uses

Much of the electronic old and new media used by commercial artists and graphic designers is multimedia.
Exciting presentations are used to grab and keep attention in advertising. Business to business, and
interoffice communications are often developed by creative services firms for advanced multimedia
presentations beyond simple slide shows to sell ideas or liven up training. Commercial multimedia
developers may be hired to design for governmental services and nonprofit services applications as well.

 Entertainment and fine arts

Multimedia is heavily used in the entertainment industry, especially to develop special effects in movies
and animations (VFX, 3D animation, etc.). Multimedia games are a popular pastime and are software
programs available either as CD-ROMs or online. Some video games also use multimedia features.
Multimedia applications that allow users to actively participate instead of just sitting by as passive recipients
of information are called interactive multimedia. In the arts there are multimedia artists, whose minds are
able to blend techniques using different media that in some way incorporates interaction with the viewer.
One of the most relevant could be Peter Greenaway who is melding cinema with opera and all sorts of digital
media. Another approach entails the creation of multimedia that can be displayed in a traditional fine arts
arena, such as an art gallery. Although multimedia display material may be volatile, the survivability of the
content is as strong as any traditional media. Digital recording material may be just as durable and infinitely
reproducible with perfect copies every time.

 Education

In education, multimedia is used to produce computer-based training courses (popularly called CBTs) and
reference books like encyclopedia and almanacs. A CBT lets the user go through a series of presentations,
text about a particular topic, and associated illustrations in various information formats. Edutainment is the
combination of education with entertainment, especially multimedia entertainment.

Learning theory in the past decade has expanded dramatically because of the introduction of multimedia.
Several lines of research have evolved, e.g. cognitive load and multimedia learning.

From multimedia learning (MML) theory, David Roberts has developed a large group lecture practice using
PowerPoint and based on the use of full-slide images in conjunction with a reduction of visible text (all text
can be placed in the notes view’ section of PowerPoint). The method has been applied and evaluated in 9
disciplines. In each experiment, students’ engagement and active learning have been approximately 66%
greater, than with the same material being delivered using bullet points, text, and speech, corroborating a

Page | 3
range of theories presented by multimedia learning scholars like Sweller and Mayer. The idea of media
convergence is also becoming a major factor in education, particularly higher education. Defined as separate
technologies such as voice (and telephony features), data (and productivity applications), and video that
now share resources and interact with each other, media convergence is rapidly changing the curriculum in
universities all over the world. Higher education has been implementing the use of social media applications
such as Twitter, YouTube, Facebook, etc. to increase student collaboration and develop new processes in
how information can be conveyed to students.

 Educational technology
 Interactive multimedia educational game.

Multimedia provides students with an alternate means of acquiring knowledge designed to enhance teaching
and learning through various mediums and platforms. In the 1960s, technology began to expand into the
classrooms through devices such as screens and telewriters. This technology allows students to learn at their
own pace and gives teachers the ability to observe the individual needs of each student. The capacity for
multimedia to be used in multi-disciplinary settings is structured around the idea of creating a hands-on
learning environment through the use of technology. Lessons can be tailored to the subject matter as well
as be personalized to the students' varying levels of knowledge on the topic. Learning content can be
managed through activities that utilize and take advantage of multimedia platforms. This kind of learning
encourages interactive communication between students and teachers and opens feedback channels,
introducing an active learning process especially with the prevalence of new media and social media.
Technology has impacted multimedia as it is largely associated with the use of computers or other electronic
devices and digital media due to its capabilities concerning research, communication, problem-solving
through simulations and feedback opportunities. The innovation of technology in education through the use
of multimedia allows for diversification among classrooms to enhance the overall learning experience for
students.

 Social work

Multimedia is a robust education and research methodology within the social work context. The five
different multimedia which supports the education process are narrative media, interactive media,
communicative media, adaptive media, and productive media. Contrary to long-standing belief, multimedia
technology in social work education existed before the prevalence of the internet. It takes the form of images,
audio, and video into the curriculum.

Mayer's Cognitive Theory of Multimedia Learning suggests, "people learn more from words and pictures
than from words alone." According to Mayer and other scholars, multimedia technology stimulates people's
brains by implementing visual and auditory effects, and thereby assists online users to learn efficiently.
Researchers suggest that when users establish dual channels while learning, they tend to understand and
memorize better. Mixed literature of this theory are still present in the field of multimedia and social work.

 Language communication

With the spread and development of the English language around the world, it has become an important
way of communicating between different people and cultures. Multimedia Technology creates a platform
where language can be taught. The traditional form of teaching English as a Second Language (ESL) in
classrooms have drastically changed with the prevalence of technology, making easier for students to obtain
language learning skills. Multimedia motivates students to learn more languages through audio, visual and
animation support. It also helps create English contexts since an important aspect of learning a language is
developing their grammar, vocabulary and knowledge of pragmatics and genres. In addition, cultural

Page | 4
connections in terms of forms, contexts, meanings and ideologies have to be constructed. By improving
thought patterns, multimedia develops students’ communicative competence by improving their capacity to
understand the language.

 Journalism

Newspaper companies all over are trying to embrace the new phenomenon by implementing its practices in
their work. To keep up with the changing world of multimedia, journalistic practices are adopting and
utilizing different functions of multimedia through the inclusions of visuals such as varying audio, video,
text, etc. in their writings.

News reporting is not limited to traditional media outlets. Freelance journalists can make use of different
new media to produce multimedia pieces for their news stories. It engages global audiences and tells stories
with technology, which develops new communication techniques for both media producers and consumers.
The Common Language Project, later renamed to The Seattle Globalist, is an example of this type of
multimedia journalism production.

Multimedia reporters who are mobile (usually driving around a community with cameras, audio and video
recorders, and laptop computers) are often referred to as mojos, from mobile journalist.

 Engineering

Software engineers may use multimedia in computer simulations for anything from entertainment to training
such as military or industrial training. Multimedia for software interfaces are often done as a collaboration
between creative professionals and software engineers. Multimedia helps expand the teaching practices that
can be found in engineering to allow for more innovated methods to not only educated future engineers, but
to help evolve the scope of understanding of where multimedia can be used in specialized engineer careers
like software engineers.

 Mathematical and scientific research

In mathematical and scientific research, multimedia is mainly used for modeling and simulation. For
example, a scientist can look at a molecular model of a particular substance and manipulate it to arrive at a
new substance. Representative research can be found in journals such as the Journal of Multimedia.

 Medicine

In medicine, doctors can get trained by looking at a virtual surgery or they can simulate how the human
body is affected by diseases spread by viruses and bacteria and then develop techniques to prevent it.
Multimedia applications such as virtual surgeries also help doctors to get practical training.

 Virtual reality

Virtual reality is a new platform for multimedia in which it merges all categories of multimedia into one
virtual environment. Virtual reality was first introduced in 1957 by cinematographer Morton Heilig in the
form of an arcade style booth called Sensorama. The first virtual reality headset was created by American
computer scientist Ivan Sutherland and Bob Sproull, his student, in 1968. Virtual reality is used for
educational and also recreational purposes like watching movies, interactive video games, simulations etc.
Ford Motor Company uses this technology to show customers the interior and exterior of their cars via their
Immersion Lab. In Pima County, Arizona their police force is trained by using Virtual Reality to create

Page | 5
scenarios for police to practice in. Many video game platforms now support VR technology, including
Sony's PlayStation, Nintendo's Switch, as well as the Oculus VR headsets that can be used for PC gaming.

DIGITAL IMAGES
A digital image is a numeric representation of a two-dimensional image. It is a type of image which is formed
from pixels. Each pixel has some finite size and represented by some finite intensity to show the image. The
pixels are set of a properly arranged rectangular array. An image resolution is the quantity of pixels per inch
(ppi) or dots per inch (dpi) of that image. Larger the ppi or dpi the higher the resolution and the better the
quality.
Digital imaging acquisition is the creation of a digitally encoded representation of the visual characteristics of
an object, such as a physical scene or the interior structure of an object. The term is often assumed to imply or
include the processing, compression, storage, printing, and display of such images. A key advantage of a digital
image, versus an analog image such as a film photograph, is the ability to make copies and copies of copies
digitally indefinitely without any loss of image quality.

Types of Digital Images


Digital images are classified depending on whether the image resolution is fixed and it principally has two
types; raster (bitmapped images) or vector type.

1. Bitmapped (Raster) Images

A bitmap describes a type of image that web-users encounter all the time without realizing it. Basically, it’s
a grid where each individual square is a pixel that contains color information. The key characteristics are
the number of pixels (or squares in the grid), and the amount of information in each grid square (pixel).

Bitmaps and raster images are interchangeable terms that refer to the same concept defines as a grid full of
pixels that when arranged densely enough form a clear image.

What You Need to Know About Bitmaps

Like raster graphics, bitmaps are made up of individual, tiny points that blend together to form a unified
image. Unlike vector graphics, which are infinitely scalable, you can’t stretch or enlarge them without
compromising the quality.

Page | 6
a. How It’s Created and Stored

When you break down an image into a grid made of thousands of squares, you get a bitmap. Each square in
that grid holds a little bit of color data and displays (or doesn’t display) a color based on that data. Like a
color-by-numbers sheet, a key correlates each point’s data assignment with a color. In the end, it provides
the literal map that tells you what that image should look like once it’s put together.

b. Bitmap Graphics vs. Raster Graphics

Raster graphics and bitmaps are closely related. Though they’re not exactly the same thing, the two phrases
are often used interchangeably. If you’re curious about the subtleties: * A bitmap refers to a specific type of
data storage—a map of bits. * A pixmap is, similarly, map of pixels. * A raster image can be either of the
above, depending on how complicated the encoded data is.

c. Bitmap File Formats

There are several file formats to choose from, and each has advantages and disadvantages. You’ve likely
heard of (or used) some of these file types—BMP (BitMaP), GIF (Graphics Interchange Format), JPEG
(Joint Photographic Experts Group), EXIF (Exchangeable Image File Format), PNG (Portable Network
Graphics), and TIFF (Tag Image File Format). Note: All of them except BMP files can be compressed and
transferred via the web.

You can determine a file’s format by where it came from or how the image will be used. Windows
exclusively uses BMPs; GIFs and JPEGs are designed for web transfer; and EXIF files come from digital
cameras and carry photo-specific information and camera settings.

BMPs are large, full files that can’t be compressed. Other formats, like GIF, JPEG, and PNG use
compression algorithms that make the files smaller and easier to upload and download via the internet,
making them extra convenient when working with online projects or designs.

d. Problems with Bitmaps

First, remember that the size of the file correlates directly with its quality. The higher the quality, the bigger
the file - which can be tough if you want to use a high-quality image on the web (and you probably do!).

Second, keep in mind that a bitmap graphic can become pixelated. Pixelation happens when you stretch an
image until the pixels become visible, making it blurry or blocky. But if you’re aware of these problems and
take them into consideration during your project, you can work with bitmaps without too much hassle.

Page | 7
e. When to Use Bitmaps in Your Design

Bitmap images can streamline your design process by saving you time, effort, and energy. Knowing what
you need beforehand will save yourself stress and complications down the line. They’re ideal when you
need realistic, easily-editable images, or when working with photos.

 When Creating Realistic Graphics and Images

Bitmaps are perfect for creating detailed images (like photographs) because of the amount of data each pixel
can store. The greater the amount of data, the broader the range of colors it can display. And, it’s much
easier to create a colorful image with realistic, transitioning color tones when you have access to a full range
of colors.

 When You Need Detailed Images That Can Be Easily Altered

When you’re looking to alter, edit, or add effects to your images, bitmaps are the way to go. You can change
the color profiles of individual pixels to affect the look of the entire image. You can also smooth or feather
out lines, create a drop shadow, and even up the contrast between the background and foreground.

 When You’re Working with Photos

Because photos naturally use pixels to form images, it makes sense to translate them to an analogous bitmap.
You can make subtle, transitional edits and alterations by targeting the color makeup, which then affects the
image composition. Just watch for pixelation.

2. Vector graphics
Vector graphics are computer images created using a sequence of commands or mathematical statements
that place lines and shapes in a two-dimensional or three-dimensional space.

In vector graphics, a graphic artist's work, or file, is created and saved as a sequence of vector statements.
A vector graphic file describes a series of points to be connected.

Page | 8
These files are sometimes called geometric files. Images created with tools such as Adobe Illustrator and
Corel's CorelDRAW are usually vector image files.

Simplified, vector graphics are like connect-the-dots drawings.

What are vector graphics used for?


Graphic artists, illustrators and designers use vector graphics for a variety of reasons, including the
following:

 Scalability. Vector formats are good for projects that require scalable graphics, including scalable type
and text. For example, company and brand logos are displayed at different sizes; they show up in the
corner of a mobile application or on a roadside billboard. A logo created with vector graphics can be
scaled up or down without loss of quality or creating a large file.

It was the scalability feature of vector graphics that resulted in its return, after falling out of favor
to raster graphics in the 1980s. Vector graphics were originally used in computer displays in the 1960s
and 1970s. World Wide Web Consortium worked on Vector Markup Language, which evolved into the
Scalable Vector Graphics open source language that contains vector and raster elements.

 App and web development. Vector graphics are useful in application and web development because
web apps and the graphics they contain must work with various screen sizes and device types. For
instance, Amazon WorkLink is a mobile app that enables a fully interactive representation of corporate
data on an employee's mobile device.

 Animation. Animated images are also usually created as vector files, which provide for cleaner and
smoother images.

 Computer-aided design (CAD). CAD programs frequently use vector files for manufacturing,
engineering and design because of their scalability and ease when it comes to editing the mathematical
formulas.

Types of vector files


There are several commonly used vector file types. They include the following:

 .ai -- Adobe Illustrator File

 .cdr -- CorelDRAW Image File

Page | 9
 .dxf -- Drawing Exchange Format File

 .eps -- Encapsulated PostScript File

 .svg -- Scalable Vector Graphics File

 .wmf -- Windows Metafile

Vector vs. raster


A raster graphics image maps bits directly to a display space, also called a bitmap. Raster graphics are made
up of a fixed number of pixels, which makes them less scalable than vector graphics. At a certain point,
when the raster image is enlarged enough, the edges become ragged, and it appears pixelated -- i.e., when
the pixels become visible. Raster graphics cannot be scaled up without sacrificing image quality.

Vector and raster images can look different. This is because vector graphics must have a separate shape
for each color shade, while raster images can have every pixel be a different color, showing subtle color
gradations and depth more clearly. At larger sizes, the edges of raster images become ragged and the images
pixilate. Vector images are more scalable.

There is also a one-to-one relationship between each pixel and the memory raster graphics occupies on a
computer. Computers must store information for every pixel of a raster image, whereas vector images only
store the series of points that need to be connected by lines, curves, etc. Consequently, vector files are
usually smaller than raster files. Vector image files are easier to modify than raster image files for this
reason.

Vector and raster images can be converted into one another with the right software. Adobe Illustrator
and Adobe Photoshop are examples of software that enable users to convert one image format to the other.

Raster files are particularly good for portraying color depth, as each pixel can be a different color. And there
are more pixels that can be unique colors than vectors that can be unique colors. This makes raster file
formats good for editing digital photographs.

Certain file types can include vector and raster elements -- PDF and SVG files are two examples.

Advantages and disadvantages of vector graphics


It is important to consider both the benefits and drawbacks of using vector files.

Page | 10
Advantages

 Scalability. As previously mentioned, this is the main advantage of vector graphics. Because vector
graphics are derived from mathematical vector relationships, or relationships between points that create
lines and curves, they appear clean and exact at any size.

 Small file size. Vector graphics generally have a small file size because they only store a small number
of points and the mathematical relationships between them. Those relationships are expressed in code,
which is less memory-intensive compared to storing pixels.

 Easy to edit. Vector files are easy to edit because users can change vector relationships fast to swap out
colors or change line shapes, for example. This is useful in an iterative process, like graphic design, that
requires a lot of editing.

 Easy to load. Because file sizes are smaller, it is easy to port and load vector files to different devices
and programs.

 Easy to duplicate. It is also easy to create clones of a vector image and copy certain features of one
graphic to another.

 Precision. The ability to scale vector graphics up or down means they have a precise look and feel.

Disadvantages

 Less detail. Vector files are limited in dealing with complex images. For instance, photographs require
color shading and blending that vector files cannot provide as well as raster files.

 Skill and time requirements. Vector files can require more skill and time to create.

 Limited browser support. There is less support for vector graphics on web browsers than for raster
graphics.

 Inconsistency. Vector images can vary from one application to another, depending on how compatible
the rendering and creating applications are, among other factors.

Page | 11
IMAGE COMPRESSION
Image compression is an application of data compression that encodes the original image with few bits. The
objective of image compression is to reduce the redundancy of the image and to store or transmit data in an
efficient form.

At its core, image compression is when you remove or group together certain parts of an image file in
order to reduce its size. Why do that? Here are a few reasons.

 For website optimization. Sites with uncompressed images can take longer to load, and can cause
your visitors to bounce because of this.
 For sending and uploading images. Uploading an uncompressed image can take a while, and some
email servers have a file size limit.
 For reducing the storage impact on your hard drive.

How does image compression work?

There are two kinds of image compression methods - lossless vs lossy. Let's take a quick look at them both.

Lossless compression

Lossless compression is a method used to reduce the size of a file while maintaining the same quality as
before it was compressed. For example, in a DSLR camera, you probably have the option to save photos
as either RAW or JPEG. RAW files have no compression and are great if you're a professional photo editor.
But they take up more space. JPEG, on the other hand, won't fill up your hard drive as fast, but some of the
data is lost in the conversion.

Types of lossless images include:

 RAW - Found in many DSLRs, and keeps all the light data received from the camera's sensor. For
a professional, this great news. However, these files types tend to be quite large in size. Additionally,
there are different versions of RAW, and you may need certain software to edit the files.
 PNG - Compresses images to keep their small size by looking for patterns on a photo, and
compressing them together. The compression is reversible, so once you open a PNG file, the image
recovers exactly.
 BMP - A format found exclusively to Microsoft. It's lossless, but not frequently used.

It should also be noted that converting a lossy photo back to lossless won't restore the photo's data.

Lossy compression

In order to give the photo an even smaller size, lossy compression discards some parts of a photo. However,
this doesn't mean the photo will look bad. Here are the two main types of lossy compression.

JPG

Also known as JPEG, this format gets rid of bits and pieces of a photo that you may notice depending upon
the level of compression you apply. A normal amount of compression will not be noticeable, while extreme
compression may be obvious.

Page | 12
There are also other ways a JPG image's quality may be reduced. If you rotate the JPG too much, you'll
notice a difference in quality. This is because the photo has to recompress itself with every rotation, losing
some data in the process. There are however programs out there that rotate a JPG losslessly. The same
degradation applies if you save a JPG multiple times.

GIF

GIF compresses files by reducing the number of colors it has. If the photo has more than 256 colors (the
maximum amount of colors older computers could have) this format will make the image look less
appealing. The best use for GIFs are for images that are animated.

The example below shows a comparison between GIF images which range from 8 colors to 256 colors.

Methods of compression

Now that we've discussed various image formats, the following explains a few image compression
methods used to achieve either lossless or lossy compression. These algorithms, or variations of these
algorithms, are also what is used in image compression tools and services.

Deflate

Deflate is a lossless data compression algorithm used for PNG images. It uses a combination of LZ77 and
Huffman coding to achieve compression results that do not affect the quality of the image.

Run-length

Run-length encoding is a form of lossless compression that takes redundant strings or runs of data and stores
them as one unit. Say you have a picture of red and white stripes, and there are 12 white pixels and 12 red
pixels. Normally, the data for it would be written as WWWWWWWWWWWWRRRRRRRRRRRR, with
W representing the white pixel and R the red pixel. Run length would put the data as 12W and 12R. Much
smaller and simpler while still keeping the data unaltered.

Transform

Transform encoding is a lossy compression commonly used for JPEGs. There are millions of shades of
colors, and transform encoding takes colors that have similar shades and makes them one single value.
Depending upon the compression value you define (i.e. the number of shades of colors you group together)
you may or may not notice a difference in the image's quality.

Page | 13
3. SOUND

The sensation felt by our ears is called sound. It is a form of energy which makes us hear. We hear several
sounds around us in our everyday life. It travels in the form of wave. A sound is a form of energy, just like
electricity, heat or light. When you strike a bell, it makes a loud ringing noise. Now instead of just listening
to the bell, put your finger on the bell after you have struck it. Can you feel it shaking? This movement or
shaking, i.e. the to and fro motion of the body is termed as Vibration. The sound is a vibration that moves
as an audible form of energy through a medium. The sound moves through a medium by alternately
contracting and expanding parts of the medium it is travelling through. The movement of molecules of a
medium is essential for the propagation of sound waves. Hence sound waves cannot travel through the
emptiness of vacuum. Sound wave can be described by five characteristics.

We know that sound travels in the form of wave. A wave is a vibratory disturbance in a medium which
carries energy from one point to another without there being a direct contact between the two points.
We can say that a wave is produced by the vibrations of the particles of the medium through which it passes.
There are two types of waves: Longitudinal waves and Transverse waves.

Longitudinal Waves: A wave in which the particles of the medium vibrate back and forth in the ‘same
direction’ in which the wave is moving. Medium can be solid, liquid or gases. Therefore, sound waves are
longitudinal waves.

Transverse Waves: A wave in which the particles of the medium vibrate up and down ‘at right angles’ to
the direction in which the wave is moving. These waves are produced only in a solids and liquids but not in
gases.
Sound is a longitudinal wave which consists of compressions and rarefactions travelling through a medium.

Sound wave can be described by five characteristics: Wavelength, Amplitude, Time-Period,


Frequency and Velocity or Speed.

Page | 14
1. Wavelength

The minimum distance in which a sound wave repeats itself is called its wavelength. That is it is the length
of one complete wave. It is denoted by a Greek letter λ (lambda). We know that in a sound wave, the
combined length of a compression and an adjacent rarefaction is called its wavelength. Also, the distance
between the centres of two consecutive compressions or two consecutive rarefactions is equal to its
wavelength.

Note: The distance between the centres of a compression and an adjacent rarefaction is equal to half of its
wavelength i.e. λ/2. The S.I unit for measuring wavelength is metre (m).

2. Amplitude
When a wave passes through a medium, the particles of the medium get displaced temporarily from their
original undisturbed positions. The maximum displacement of the particles of the medium from their
original undisturbed positions, when a wave passes through the medium is called amplitude of the wave. In
fact the amplitude is used to describe the size of the wave. The S.I unit of measurement of amplitude is
metre (m) though sometimes it is also measured in centimetres. Do you know that the amplitude of a wave
is the same as the amplitude of the vibrating body producing the wave?

3. Time-Period
The time required to produce one complete wave or cycle or cycle is called time-period of the wave. Now,
one complete wave is produced by one full vibration of the vibrating body. So, we can say that the time
taken to complete one vibration is known as time-period. It is denoted by letter T. The unit of measurement
of time-period is second (s).

4. Frequency

Page | 15
The number of complete waves or cycles produced in one second is called frequency of the wave. Since one
complete wave is produced by one full vibration of the vibrating body, so we can say that the number of
vibrations per second is called frequency. For example: if 10 complete waves or vibrations are produced in
one second then the frequency of the waves will be 10 hertz or 10 cycles per second. Do you know that the
frequency of a wave is fixed and does not change even when it passes through different substances?

The S.I unit of frequency is hertz or Hz. A vibrating body emitting 1 wave per second is said to have a
frequency of 1 hertz. That is 1 Hz is equal to 1 vibration per second.
Sometimes a bigger unit of frequency is known as kilohertz (kHz) that is 1 kHz = 1000 Hz. The frequency
of a wave is denoted by the letter f.
The frequency of a wave is the same as the frequency of the vibrating body which produces the wave.

What is the relation between time-period and frequency of a wave?


The time required to produce one complete wave is called time-period of the wave. Suppose the time-
period of a wave is T seconds.
In T seconds number of waves produced = 1
So, in 1 second, number of waves produced will be = 1/T
But the number of waves produced in 1 second is called its frequency.
Therefore, F = 1/Time-period
f = 1/T
where f = frequency of the wave
T = time-period of the wave

5. Velocity of Wave (Speed of Wave)


The distance travelled by a wave in one second is called velocity of the wave or speed of the wave. It is
represented by the letter v. The S.I unit for measuring the velocity is metres per second (m/s or ms-1).

What is the relationship between Velocity, Frequency and Wavelength of a Wave?


Velocity = Distance travelled/ Time taken
Let v = λ / T
Where T = time taken by one wave.
v=fXλ
This formula is known as wave equation.
Where v = velocity of the wave
f = frequency
λ = wavelength
Velocity of a wave = Frequency X Wavelength
This applies to all the waves like transverse waves like water waves, longitudinal waves like sound waves
and the electromagnetic waves like light waves and radio waves
Therefore we have learnt various characteristics of sound waves.

Digitization of Sound

Digitization is a process of converting the analog signals to a digital signal. There are three steps of
digitization of sound.

Page | 16
 Sampling - Sampling is a process of measuring air pressure amplitude at equally spaced moments
in time, where each measurement constitutes a sample. A sampling rate is the number of times the
analog sound is taken per second. A higher sampling rate implies that more samples are taken
during the given time interval and ultimately, the quality of reconstruction is better. The sampling
rate is measured in terms of Hertz, Hz in short, which is the term for Cycle per second. A sampling
rate of 5000 Hz(or 5kHz,which is more common usage) implies that mt uj vu8i 9ikuhree sampling
rates most often used in multimedia are 44.1kHz(CD-quality), 22.05kHz and 11.025kHz.
 Quantization - Quantization is a process of representing the amplitude of each sample as integers
or numbers. How many numbers are used to represent the value of each sample known as sample
size or bit depth or resolution. Commonly used sample sizes are either 8 bits or 16 bits. The larger
the sample size, the more accurately the data will describe the recorded sound. An 8-bit sample size
provides 256 equal measurement units to describe the level and frequency of the sound in that slice
of time. A 16-bit sample size provides 65,536 equal units to describe the sound in that sample slice
of time. The value of each sample is rounded off to the nearest integer (quantization) and if the
amplitude is greater than the intervals available, clipping of the top and bottom of the wave occurs.
 Encoding - Encoding converts the integer base-10 number to a base-2 that is a binary number. The
output is a binary expression in which each bit is either a 1(pulse) or a 0(no pulse).

Quantization of Audio

Quantization is a process to assign a discrete value from a range of possible values to each sample. Number
of samples or ranges of values are dependent on the number of bits used to represent each sample.
Quantization results in stepped waveform resembling the source signal.
 Quantization Error/Noise - The difference between sample and the value assigned to it is known
as quantization error or noise.
 Signal to Noise Ratio (SNR) - Signal to Ratio refers to signal quality versus quantization error.
Higher the Signal to Noise ratio, the better the voice quality. Working with very small levels often
introduces more error. So instead of uniform quantization, non-uniform quantization is used as

Page | 17
companding. Companding is a process of distorting the analog signal in controlled way by
compressing large values at the source and then expanding at receiving end before quantization
takes place.

Transmission of Audio

In order to send the sampled digital sound/ audio over the wire that it to transmit the digital audio, it is first
to be recovered as analog signal. This process is called de-modulation.
 PCM Demodulation - PCM Demodulator reads each sampled value then apply the analog filters
to suppress energy outside the expected frequency range and outputs the analog signal as output
which can be used to transmit the digital signal over the network.

Page | 18
Mono vs Stereo Sound

Mono or monoaural sound only uses one channel when converting a signal into a sound. Even if there are
multiple speakers, the same signal will go to both speakers. This then gives the effect that the sounds, even
if they are coming from separate speakers, are coming from one single position or source.

In today’s age of technology, most signals are more compatible with stereo, instead of mono sound, which
was widely used for radio broadcasts in the past.

Contrary to mono sound, stereo sound uses more than one channel when converting a signal into a sound.
This essentially means that each signal sent out is unique. For example, when a song sends different sounds
to the left earbud versus the right, like in the classic song “Bohemian Rhapsody”, it requires and is best used
with stereo sound.

Stereo sound, then, gives the effect of sound coming from different sources and positions, which is typical
and very common in today’s technology, especially in speakers that are produced for the ‘surround sound’
effect.

Which is Better?

There is no real answer. It all has to do with preference and situation. For example, stereo sound is great for
when you are watching a movie with lots of music and environmental sounds so you can really feel like you
are in the movie.

If you tend to only wear one earbud at a time, which--believe it or not--is a thing, you may prefer mono
sound. This is so you can still experience the entire song and all of its parts, even if you only have the one
earbud in your ear.

Mono sound is when only one channel is used to convert a signal to a sound. Stereo sound is when multiple
channels are used to convert multiple signals to sounds. Your preference for either one is entirely based on
you, because just like sound, everyone is different.

Audio File Format and Compression


An audio file format is a file format for storing digital audio data on a computer system. The bit layout of
the audio data (excluding metadata) is called the audio coding format and can be uncompressed,
or compressed to reduce the file size, often using lossy compression. The data can be a raw bitstream in an
audio coding format, but it is usually embedded in a container format or an audio data format with defined
storage layer.

Format types
It is important to distinguish between the audio coding format, the container containing the raw audio data,
and an audio codec. A codec performs the encoding and decoding of the raw audio data while this encoded

Page | 19
data is (usually) stored in a container file. Although most audio file formats support only one type of audio
coding data (created with an audio coder), a multimedia container format (as Matroska or AVI) may support
multiple types of audio and video data.
There are three major groups of audio file formats:

 Uncompressed audio formats, such as WAV, AIFF, AU or raw header-less PCM;


 Formats with lossless compression, such as FLAC, Monkey's Audio (filename
extension .ape ), WavPack (filename extension .wv ), TTA, ATRAC Advanced
Lossless, ALAC (filename extension .m4a ), MPEG-4 SLS, MPEG-4 ALS, MPEG-4 DST, Windows
Media Audio Lossless (WMA Lossless), and Shorten (SHN).
 Formats with lossy compression, such
as Opus, MP3, Vorbis, Musepack, AAC, ATRAC and Windows Media Audio Lossy (WMA lossy).
Uncompressed audio format
One major uncompressed audio format, LPCM, is the same variety of PCM as used in Compact Disc Digital
Audio and is the format most commonly accepted by low level audio APIs and D/A converter hardware.
Although LPCM can be stored on a computer as a raw audio format, it is usually stored in a .wav file
on Windows or in a .aiff file on macOS. The Audio Interchange File Format (AIFF) format is based on
the Interchange File Format (IFF), and the WAV format is based on the similar Resource Interchange File
Format (RIFF). WAV and AIFF are designed to store a wide variety of audio formats, lossless and lossy;
they just add a small, metadata-containing header before the audio data to declare the format of the audio
data, such as LPCM with a particular sample rate, bit depth, endianness and number of channels. Since
WAV and AIFF are widely supported and can store LPCM, they are suitable file formats for storing and
archiving an original recording.
BWF (Broadcast Wave Format) is a standard audio format created by the European Broadcasting Union as
a successor to WAV. Among other enhancements, BWF allows more robust metadata to be stored in the
file. See European Broadcasting Union: Specification of the Broadcast Wave Format (EBU Technical
document 3285, July 1997). This is the primary recording format used in many professional audio
workstations in the television and film industry. BWF files include a standardized timestamp reference
which allows for easy synchronization with a separate picture element. Stand-alone, file based, multi-track
recorders from AETA, Sound Devices, Zaxcom HHB Communications Ltd, Fostex, Nagra,
Aaton, and TASCAM all use BWF as their preferred format.
Lossless compressed audio format
A lossless compressed audio format stores data in less space without losing any information. The original,
uncompressed data can be recreated from the compressed version.
Uncompressed audio formats encode both sound and silence with the same number of bits per unit of time.
Encoding an uncompressed minute of absolute silence produces a file of the same size as encoding an
uncompressed minute of music. In a lossless compressed format, however, the music would occupy a
smaller file than an uncompressed format and the silence would take up almost no space at all.
Lossless compression formats include the common FLAC, WavPack, Monkey's Audio, ALAC (Apple
Lossless). They provide a compression ratio of about 2:1 (i.e. their files take up half the space of PCM).
Development in lossless compression formats aims to reduce processing time while maintaining a good
compression ratio.
Lossy compressed audio format
Lossy audio format enables even greater reductions in file size by removing some of the audio information
and simplifying the data. This, of course, results in a reduction in audio quality, but a variety of techniques

Page | 20
are used, mainly by exploiting psychoacoustics, to remove the parts of the sound that have the least effect
on perceived quality, and to minimize the amount of audible noise added during the process. The
popular MP3 format is probably the best-known example, but the AAC format found on the iTunes Music
Store is also common. Most formats offer a range of degrees of compression, generally measured in bit rate.
The lower the rate, the smaller the file and the more significant the quality loss.

List of formats

File Creation Description


Extension Company

.3gp Multimedia container format can contain proprietary formats as AMR, AMR-
WB or AMR-WB+, but also some open formats

.aa Audible ( A low-bitrate audiobook container format with DRM, containing audio encoded as
Amazon.c either MP3 or the ACELP speech codec.
om)

.aac The Advanced Audio Coding format is based on the MPEG-2 and MPEG-4 standards.
AAC files are usually ADTS or ADIF containers.

.aax Audible ( An Audiobook format, which is a variable-bitrate (allowing high quality) M4B file
Amazon.c encrypted with DRM. MPB contains AAC or ALAC encoded audio in an MPEG-
om) 4 container. (More details below.)

.act ACT is a lossy ADPCM 8 kbit/s compressed audio format recorded by most Chinese
MP3 and MP4 players with a recording function, and voice recorders

.aiff Apple A standard uncompressed CD-quality, audio file format used by Apple. Established 3
years prior to Microsoft's uncompressed version wav.

.alac Apple An audio coding format developed by Apple Inc. for lossless data compression of digital
music.

.amr AMR-NB audio, used primarily for speech.

.ape Matthew Monkey's Audio lossless audio compression format.


T.
Ashland

.au Sun The standard audio file format used by Sun, Unix and Java. The audio in au files can
Microsyst be PCM or compressed with the μ-law, a-law or G729 codecs.
ems

Page | 21
.awb AMR-WB audio, used primarily for speech, same as the ITU-T's G.722.2 specification.

.dss Olympus DSS files are an Olympus proprietary format. It is a fairly old and poor codec. GSM or
MP3 are generally preferred where the recorder allows. It allows additional data to be
held in the file header.

.dvf Sony A Sony proprietary format for compressed voice files; commonly used by Sony dictation
recorders.

.flac A file format for the Free Lossless Audio Codec, an open-source lossless compression
codec.

.gsm Designed for telephony use in Europe, gsm is a very practical format for telephone
quality voice. It makes a good compromise between file size and quality. Note that wav
files can also be encoded with the gsm codec.

.iklax iKlax An iKlax Media proprietary format, the iKlax format is a multi-track digital audio
format allowing various actions on musical data, for instance on mixing and volumes
arrangements.

.ivs 3D Solar A proprietary version with Digital Rights Management developed by 3D Solar UK Ltd
UK Ltd for use in music downloaded from their Tronme Music Store and interactive music and
video player.

.m4a An audio-only MPEG-4 file, used by Apple for unprotected music downloaded from
their iTunes Music Store. Audio within the m4a file is typically encoded with AAC,
although lossless ALAC may also be used.

.m4b An Audiobook / podcast extension with AAC or ALAC encoded audio in an MPEG-
4 container. Both M4A and M4B formats can contain metadata including chapter
markers, images, and hyperlinks, but M4B allows "bookmarks" (remembering the last
listening spot), whereas M4A does not.[7]

.m4p Apple A version of AAC with proprietary Digital Rights Management developed by Apple for
use in music downloaded from their iTunes Music Store and their music streaming
service known as Apple Music.

.mmf Yamaha, A Samsung audio format that is used in ringtones. Developed by Yamaha (SMAF stands
Samsung for "Synthetic music Mobile Application Format", and is a multimedia data format
invented by the Yamaha Corporation, .mmf file format).

.mp3 MPEG Layer III Audio. It is the most common sound file format used today.

Page | 22
.mpc Musepack or MPC (formerly known as MPEGplus, MPEG+ or MP+) is an open source
lossy audio codec, specifically optimized for transparent compression of stereo audio at
bitrates of 160–180 kbit/s.

.msv Sony A Sony proprietary format for Memory Stick compressed voice files.

.nmf NICE NICE Media Player audio file

.ogg, [Link] A free, open source container format supporting a variety of formats, the most popular
Foundatio of which is the audio format Vorbis. Vorbis offers compression similar to MP3 but is
.oga,
n less popular. Mogg, the "Multi-Track-Single-Logical-Stream Ogg-Vorbis", is the multi-
.mogg channel or multi-track Ogg file format.

.opus Internet A lossy audio compression format developed by the Internet Engineering Task Force
Engineeri (IETF) and made especially suitable for interactive real-time applications over the
ng Task Internet. As an open format standardised through RFC 6716, a reference implementation
Force is provided under the 3-clause BSD license.

.ra, .rm RealNetw A RealAudio format designed for streaming audio over the Internet. The .ra format
orks allows files to be stored in a self-contained fashion on a computer, with all of the audio
data contained inside the file itself.

.raw A raw file can contain audio in any format but is usually used with PCM audio data. It
is rarely used except for technical tests.

.rf64 One successor to the Wav format, overcoming the 4GiB size limitation.

.sln Signed Linear PCM format used by Asterisk. Prior to v.10 the standard formats were
16-bit Signed Linear PCM sampled at 8 kHz and at 16 kHz. With v.10 many more
sampling rates were added.[8]

.tta The True Audio, real-time lossless audio codec.

.voc Creative The file format consists of a 26-byte header and a series of subsequent data blocks
Technolog containing the audio information
y

.vox The vox format most commonly uses the Dialogic ADPCM (Adaptive Differential
Pulse Code Modulation) codec. Similar to other ADPCM formats, it compresses to 4-
bits. Vox format files are similar to wave files except that the vox files contain no
information about the file itself so the codec sample rate and number of channels must
first be specified in order to play a vox file.

Page | 23
.wav Standard audio file container format used mainly in Windows PCs. Commonly used for
storing uncompressed (PCM), CD-quality sound files, which means that they can be
large in size—around 10 MB per minute. Wave files can also contain data encoded with
a variety of (lossy) codecs to reduce the file size (for example the GSM or MP3 formats).
Wav files use a RIFF structure.

.wma Microsoft Windows Media Audio format, created by Microsoft. Designed with Digital Rights
Management (DRM) abilities for copy protection.

.wv Format for wavpack files.

.webm Royalty-free format created for HTML5 video.

.8svx Electronic The IFF-8SVX format for 8-bit sound samples, created by Electronic Arts in 1984 at the
Arts birth of the Amiga.

.cda Format for cda files for Radio.

4. VIDEO
Video is an electronic medium for the recording, copying, playback, broadcasting, and display of moving
visual media. Video was first developed for mechanical television systems, which were quickly replaced by
cathode ray tube systems which were later replaced by flat panel displays of several types.

Types of Video Signal

Video signals can be organized in three different ways: Component video, Composite video, and S - video.
Component Video
Component video is a video signal that has been split into two or more component channels. In popular use,
it refers to a type of component analog video (CAV) information that is transmitted or stored as three
separate signals. Component video can be contrasted with composite video (NTSC, PAL or SECAM) in
which all the video information is combined into a single line - level signal that is used in analog television.
Like composite, component - video cables do not carry audio and are often paired with audio cables.
When used without any other qualifications the term component video generally refers to analog YPBPR
component video with sync on luma.
Composite Video

Page | 24
Composite video (1 channel) is an analog video transmission (no audio) that carries standard definition
video typically at 480i or 576i resolution. Video information is encoded on one channel in contrast with
slightly higher quality S - video (2 channel), and even higher quality component video (3 channels).

Composite video is usually in standard formats such as NTSC, PAL, and SECAM and is often designated
by the CVBS initialism, meaning "Color, Video, Blanking, and Sync."
S - Video
Separate Video (2 channel), more commonly known as S - Video and Y/C, is an analog video transmission
(no audio) that carries standard definition video typically at 480i or 576i resolution. Video information is
encoded on two channels: luma (luminance, intensity, "Y") and chroma (colour, "C"). This separation is in
contrast with slightly lower quality composite video (1 channel) and higher quality component video (3
channels). It's often referred to by JVC (who introduced the DIN - connector pictured) as both an S - VHS
connector and as Super Video.
The four - pin mini - DIN connector (shown at right) is the most common of several S - Video connector
types. Other connector variants include seven - pin locking "dub" connectors used on many professional S
- VHS machines, and dual "Y" and "C" BNC connectors, often used for S - Video patch panels. Early Y/C
video monitors often used phono (RCA connector) that were switchable between Y/C and composite video
input. Though the connectors are different, the Y/C signals for all types are compatible.

Analog Video

Analog video is a video signal transferred by an analog signal. When combined in to one channel, it is called
composite video as is the case, among others with NTSC, PAL and SECAM.
Analog video may be carried in separate channels, as in two channel S - Video (YC) and multi - channel
component video formats.

Analog video is used in both consumer and professional television production applications. However, digital
video signal formats with higher quality have been adopted, including serial digital interface (SDI), Firewire
(IEEE 1394), Digital Visual Interface (DVI) and High - Definition Multimedia Interface (HDMI).

Most TV is still sent and received as an analog signal. Once the electrical signal is received, we may assume
that brightness is at least a monotonic function of voltage, if not necessarily linear, because of gamma
correction.

Page | 25
An analog signal f(t) samples a time - varying image. So - called progressive scanning traces through a
complete picture (a frame) row - wise for each time interval. A high - resolution computer monitor typically
uses a time interval of 1/72 second.
In TV and in some monitors and multimedia standards, another system, interlaced scanning, is used. Here,
the odd - numbered lines are traced first, then the even - numbered lines. This results in "odd" and "even"
fields — two fields make up one frame.
In fact, the odd lines (starting from 1) end up at the middle of a line at the end of the odd field, and the even
scan starts at a half - way point. The following figure shows the scheme used. First the solid (odd) lines are
traced— P to Q, then R to S, and so on, ending at T — then the even field starts at U and ends at V. The
scan fines are not horizontal because a small voltage is applied, moving the electron beam down over time.

Interlaced raster scan

Interlacing was invented because, when standards were being defined, it was difficult to transmit the amount
of information in a full frame quickly enough to avoid flicker. The double number of fields presented to the
eye reduces perceived flicker.

Because of interlacing, the odd and even lines are displaced in time from each other. This is generally not
noticeable except when fast action is taking place onscreen, when blurring may occur. For example, in the
video in the following figure, the moving helicopter is blurred more than the still background.

Page | 26
Since it is sometimes necessary to change the frame rate, resize, or even produce stills from an interlaced
source video, various schemes are used to de - interlace it. The simplest de - interlacing method consists of
discarding one field and duplicating the scan lines of the other field, which results in the information in one
field being lost completely. Other, more complicated methods retain information from both fields.
CRT displays are built like fluorescent lights and must flash 50 to 70 times per second to appear smooth. In
Europe, this fact is conveniently tied to their 50 Hz electrical system, and they use video digitized at 25
frames per second (fps); in North America, the 60 Hz electric system dictates 30 fps.

The jump from Q to R and so on is called the horizontal retrace, during which the electronic beam in the
CRT is blanked. The jump from T to U or V to P is called the vertical retrace.
Interlaced scan produces two fields for each frame:(a) the video frame; (b) Field 1; (c) Field 2; (d)
difference of Fields

Since voltage is one - dimensional — it is simply a signal that varies with time — how do we know when a
new video line begins? That is, what part of an electrical signal tells us that we have to restart at the left side
of the screen?

Page | 27
The solution used in analog video is a small voltage offset from zero to indicate black and another value,
such as zero, to indicate the start of a line. Namely, we could use a "blacker - than - black" zero signal to
indicate the beginning of a line.

The following figure shows a typical electronic signal for one scan line of NTSC composite video. 'White'
has a peak value of 0.714 V; 'Black' is slightly above zero at 0.055 V; whereas

Blank is at zero volts. As shown, the time duration for blanking pulses in the signal is used for
synchronization as well, with the tip of the Sync signal at approximately — 0.286 V. In fact, the problem
of reliable synchronization is so important that special signals to control sync take up about 30% of the
signal!

Electronic signal for one NTSC scan line

The vertical retrace and sync ideas are similar to the horizontal one, except that they happen only once per
field.

NTSC Video
NTSC, named for the National Television System Committee, is the analog television system that is used
in most of North America, parts of South America (except Brazil, Argentina, Uruguay, and French Guiana),
Myanmar, South Korea, Taiwan, Japan, the Philippines, and some Pacific island nations and territories.
Most countries using the NTSC standard, as well as those using other analog television standards, are
switching to newer digital television standards, of which at least four different ones are in use around the

Page | 28
world. North America, parts of Central America, and South Korea are adopting the ATSC standards, while
other countries are adopting or have adopted other standards.

The first NTSC standard was developed in 1941 and had no provision for color television. In 1953 a second
modified version of the NTSC standard was adopted, which allowed color television broadcasting
compatible with the existing stock of black - and - white receivers. NTSC was the first widely adopted
broadcast color system and remained dominant where it had been adopted until the first decade of the 21st
century, when it was replaced with digital ATSC. After nearly 70 years of use, the vast majority of over -
the - air NTSC transmissions in the United States were turned off on June 12, 2009 and August 31, 2011 in
Canada and most other NTSC markets.

Digital broadcasting permits higher - resolution television, but digital standard definition television in these
countries continues to use the frame rate and number of lines of resolution established by the analog NTSC
standard; systems using the NTSC frame rate and resolution (such as DVDs) are still referred to informally
as "NTSC". NTSC baseband video signals are also still often used in video playback (typically of recordings
from existing libraries using existing equipment) and in CCTV and surveillance video systems.

Video raster, including retrace and sync data

Samples per line for various analog video formats

Page | 29
Different video formats provide different numbers of samples per line, as listed in the above table. Laser
disks have about the same resolution as Hi - 8. (In comparison, mini DV 1/4 - inch tapes for digital video
are 480 lines by 720 samples per line.)

Interleaving Y and C signals in the NTSC spectrum

PAL Video
PAL (Phase Alternating Line) is a TV standard originally invented by German scientists. It uses 625 scan
lines per frame, at 25 frames per second (or 40 msec / frame), with a 4 : 3 aspect ratio and interlaced fields.
Its broadcast TV signals are also used in composite video. This important standard is widely used in Western
Europe, China, India and many other parts of the world.

Page | 30
PAL uses the YUV color model with an 8 MHz channel, allocating a bandwidth of 5.5 MHz to Y and 1.8
MHz each to U and V. The color subcarrier frequency is fsc ≈ 4.43 MHz. To improve picture quality, chroma
signals have alternate signs (e.g., +U and — U) in successive scan lines; hence the name "Phase Alternating
Line. This facilitates the use of a (line - rate) comb filter at the receiver — the signals in consecutive lines
are averaged so as to cancel the chroma signals (which always carry opposite signs) for separating Y and C
and obtain high - quality Y signals.
SECAM Video
SECAM, which was invented by the French, is the third major broadcast TV standard. SECAM stands for
Systeme Electronique Couleur Avec Memorie. SECAM also uses 625 scan lines per frame, at 25 frames per
second, with a 4:3 aspect ratio and interlaced fields. The original design called for a higher number of scan
lines (over 800), but the final version settled for 625.

SECAM and PAL are similar, differeing slightly in their color coding scheme. In SECAM, U and V signals
are modulated using separate color subcarriers at 4.25 MHz and 4.41 MHz, respectively. They are sent in
alternate lines - that is, only one of the U or V signals will be sent on each scan line.

Table Comparison of the analog broadcast TV systems.

Digital Video

Digital video comprises a series of orthogonal bitmap digital images displayed in rapid succession at a
constant rate. In the context of video these images are called frames. We measure the rate at which frames
are displayed in frames per second (FPS).

Since every frame is an orthogonal bitmap digital image it comprises a raster of pixels. If it has a width of
W pixels and a height of Hpixels we say that the frame size is WxH.
Pixels have only one property, their color. The color of a pixel is represented by a fixed number of bits. The
more bits the more subtle variations of colors can be reproduced. This is called the color depth (CD) of the
video.

An example video can have a duration (T) of 1 hour (3600sec), a frame size of 640 x 480 (W x H) at a color
depth of 24bits and a frame rate of 25fps. This example video has the following properties:
 pixels per frame = 640 * 480 = 307,200
 bits per frame = 307,200 * 24 = 7,372,800 = 7.37Mbits
 bit rate (BR) = 7.37 * 25 = 184.25Mbits / sec

Page | 31
 video size (VS) = 184Mbits / sec * 3600sec = 662,400Mbits = 82,800Mbytes = 82.8Gbytes
The advantages of digital representation for video are many. It permits
 Storing video on digital devices or in memory, ready to be processed (noise removal, cut and paste, and
so on) and integrated into various multimedia applications
 Direct access, which makes nonlinear video editing simple Repeated recording without degradation of
image quality
 Ease of encryption and better tolerance to channel noise
Table Comparison of analog broadcast TV systems

In earlier Sony or Panasonic recorders, digital video was in the form of composite video. Modem digital
video generally uses component video, although RGB signals are first converted into a certain type of color
opponent space, such as YUV. The usual color space is YCbCr.

Chroma Subsampling
Since humans see color with much less spatial resolution than black and white, it makes sense to decimate
the chrominance signal. Interesting but not necessarily informative names have arisen to label the different
schemes used. To begin with, numbers are given stating how many pixel values, per four original pixels,
are actually sent. Thus the chroma subsampling scheme "4:4:4" indicates that no chroma subsampling is
used. Each pixel's Y, Cb, and Cr values are transmitted, four for each of Y, Cb, and Cr.
The scheme "4:2:2" indicates horizontal subsampling of the Cb and Cr signals by a factor of 2. That is, of
four pixels horizontally labeled 0 to 3, all four 7s are sent, and every two Cbs and two Crs are sent,
as {CbO, Y0)(Cr0, Yl)(Cb2, Y2)(Cr2, Y3)(Cb4, Y4), and so on.
The scheme "4:1:1" subsamples horizontally by a factor of 4. The scheme "4:2:0" subsamples in both the
horizontal and vertical dimensions by a factor of 2. Theoretically, an average chroma pixel is positioned
between the rows and columns, as shown in the below figure. We can see that the scheme 4:2:0 is in fact
another kind of 4:1:1 sampling, in the sense that we send 4, 1, and 1 values per 4 pixels. Therefore, the
labeling scheme is not a very reliable mnemonic!

Page | 32
Scheme 4:2:0, along with others, is commonly used in JPEG and MPEG.

CCIR Standards for Digital Video


The CCIR is the Consultative Committee for International Radio. One of the most important standards it
has produced is CCIR - 601, for component digital video. This standard has since become standard ITU - R
- 601, an international standard for professional video applications. It is adopted by certain digital video
formats, including the popular DV video.
The NTSC version has 525 scan fines, each having 858 pixels (with 720 of them visible, not in the blanking
period).Because the NTSC version uses 4:2:2, each pixel can be

Chroma subsampling

Page | 33
represented with two bytes (8 bits for Y and 8 bits alternating between Cb and Cr). The CCIR 601. (NTSC)
data rate (including blanking and sync but excluding audio) is thus approximately 216 Mbps (megabits per
second):
525 x 858 x 30 x 2 bytes x 8 bits / byte≈ 216 Mbps

During blanking, digital video systems may make use of the extra data capacity to carry audio signals,
translations into foreign languages, or error - correction information.

The following table shows some of the digital video specifications, all with an aspect ratio of 4:3. The CCIR
601 standard uses an interlaced scan, so each field has only half as much vertical resolution (e.g., 240 lines
in NTSC).

Table Digital video specifications

CIF stands for Common Intermediate Format, specified by the International Telegraph and Telephone
Consultative Committee (CCITT), now superseded by the International Telecommunication Union, which
oversees both telecommunications (ITU - T) and radio frequency matters (ITU - R) under one United
Nations body. The idea of CIF, which is about the same as VHS quality, is to specify a format for lower
bitrate. CIF uses a progressive (noninterlaced) scan. QCIF stands for Quarter - CIF, and is for even lower

Page | 34
bitrate. All the CIF / QCIF resolutions are evenly divisible by 8, and all except 88 are divisible by 16; this
is convenient for block - based video coding in H.261 and H.263.
CIF is a compromise between NTSC and PAL, in that it adopts the NTSC frame rate and half the number
of active lines in PAL. When played on existing TV sets, NTSC TV will first need to convert the number
of lines, whereas PAL TV will require frame-rate conversion.

High Definition TV (HDTV)


The introduction of wide - screen movies brought the discovery that viewers seated near the screen enjoyed
a level of participation (sensation of immersion) not experienced with conventional movies. Apparently the
exposure to a greater field of view, especially the involvement of peripheral vision, contributes to the sense
of "being there". The main thrust of High Definition TV (HDTV) is not to increase the "definition" in each
unit area, but rather to increase the visual field, especially its width.

First - generation HDTV was based on an analog technology developed by Sony and NHK in Japan in the
late 1970s. HDTV successfully broadcast the 1984 Los Angeles Olympic Games in Japan. Multiple sub -
Nyquist Sampling Encoding (MUSE) was an improved NHK HDTV with hybrid analog / digital
technologies that was put in use in the 1990s. It has 1,125 scan lines, interlaced (60 fields per second), and
a 16:9 aspect ratio. It uses satellite to broadcast — quite appropriate for Japan, which can be covered with
one or two satellites.

The Direct Broadcast Satellite (DBS) channels used have a bandwidth of 24 MHz. In general, terrestrial
broadcast, satellite broadcast, cable, and broadband networks are all feasible means for transmitting HDTV
as well as conventional TV. Since uncompressed

Table Advanced Digital TV Formats Supported by ATSC

Page | 35
HDTV will easily demand more than 20 MHz bandwidth, which will not fit in the current 6 MHz or 8 MHz
channels, various compression techniques are being investigated. It is also anticipated that high - quality
HDTV signals will be transmitted using more than one channel, even after compression.

In 1987, the FCC decided that HDTV standards must be compatible with the existing NTSC standard and
must be confined to the existing Very High Frequency (VHF) and Ultra High Frequency (UHF) bands. This
prompted a number of proposals in North America by the end of 1988, all of them analog or mixed analog
/ digital.

In 1990, the FCC announced a different initiative — its preference for full - resolution HDTV. They decided
that HDTV would be simultaneously broadcast with existing NTSC TV and eventually replace it. The
development of digital HDTV immediately took off in North America.

Witnessing a boom of proposals for digital HDTV, the FCC made a key decision to go all digital in 1993.
A "grand alliance" was formed that included four main proposals, by General Instruments, MIT, Zenith, and
AT&T, and by Thomson, Philips, Sarnoff and others. This eventually led to the formation of the Advanced
Television Systems Committee (ATSC), which was responsible for the standard for TV broadcasting of
HDTV. In 1995, the U.S. FCC Advisory Committee on Advanced Television Service recommended that
the ATSC digital television standard be adopted.

The standard supports video scanning formats shown in Table. In the table, "I" means interlaced scan and
"P" means progressive (noninterlaced) scan. The frame rates supported are both integer rates and the NTSC
rates — that is, 60.00 or 59.94, 30.00 or 29.97, 24.00 or 23.98 fps.

For video, MPEG - 2 is chosen as the compression standard. As will be seen in Chapter, it uses Main Level
to High Level of the Main Profile of MPEG - 2. For audio, AC - 3 is the standard. It supports the so - called
5.1 channel Dolby surround sound — five surround channels plus a subwoofer channel.

The salient difference between conventional TV and HDTV [4, 6] is that the latter has a much wider aspect
ratio of 16:9 instead of 4:3. (Actually, it works out to be exactly one - third wider than current TV) Another
feature of HDTV is its move toward progressive (noninterlaced) scan. The rationale is that interlacing
introduces serrated edges to moving objects and flickers along horizontal edges.

Page | 36
The FCC has planned to replace all analog broadcast services with digital TV broadcasting by the year 2006.
Consumers with analog TV sets will still be able to receive signals via an 8 - VSB (8 - level vestigial
sideband) demodulation box. The services provided will include
 Standard Definition TV (SDTV)— the current NTSC TV or higher
 Enhanced Definition TV (EDTV) — 480 active lines or higher — the third and fourth rows
 High Definition TV (HDTV)— 720 active lines or higher. So far, the popular choices are 720P (720
lines, progressive, 30 fps) and 10801 (1,080 lines, interlaced, 30 fps or 60 fields per second). The latter
provides slightly better picture quality but requires much higher bandwidth.

Video Compression

Video compression is a process that reduces and removes redundant video information so that a digital video
file/ stream can be sent across a network and stored more efficiently. An encoding algorithm is applied to
the source video to create a compressed stream that is ready for transmission, recording, or storage. To
decode (play) the compressed stream, an inverse algorithm is applied. The time it takes to compress, send,
decompress and ultimately display a stream is known as latency.
A video codec (encoder/decoder) employs a pair of algorithms that work together. The process for encoding
and decoding must be matched; video content that is compressed using one standard cannot be
decompressed with a different standard. Different video compression standards utilize different methods of
reducing data, and hence, results may differ in bit rate (i.e. bandwidth), latency, and image quality.
Types of compression are often categorized by the amount of data that’s maintained through the stages of
processing. "Lossless" refers to a compression method in which there is no loss of data during the
transmission of a video signal from source to display. The displayed image is identical to the original source
image. "Visually lossless" means that the displayed image will appear identical to the original image, even
if some data may actually have been lost during compression. "Lossy" compression usually involves some
loss of data during the data-reduction process, but noticeable quality degradation may or may not be
apparent.

Page | 37

You might also like