lOMoARcPSD|62440031
Chapter 2 - Tracking and Computer Vision for AR and MR
Sem 7 INFT MU
Virtual and augmented reality (University of Mumbai)
messages.pdf_cover_qr_code_label
messages.studocu_not_sponsored_or_endorsed_by_college
messages.downloaded_by
lOMoARcPSD|62440031
Chapter 2 - Tracking and Computer
Vision for AR and MR
Tracking and Computer Vision for AR and MR
Multimodal Displays; Visual Perception; Spatial Display Model; Visual Displays; Tracking,
Calibration and Registration; Coordinate Systems; Characteristics of Tracking Technology;
Stationary Tracking Systems; Mobile Sensors; Optical Tracking; Sensor Fusion; Marker
Tracking; Multiple Camera Infrared Tracking; Natural Feature Tracking by Detection;
Incremental Tracking; Simultaneous Localization and Tracking; Outdoor Tracking
2.1 Introduction
2.1.1 Visual Perception
Visual perception is the ability to perceive our surroundings through the light that enters our
eyes. The visual perception of colors, patterns, and structures has been of particular interest
in relation to graphical user interfaces (GUIs) because these are perceived exclusively
through vision.
It is clear that individual perceptual fields significantly determine the nature and technical
principles of augmented reality. When creating an augmented reality classification system
based on the above, it is necessary to apply a perspective that reflects the given perceptual
field; one must consider the AR system’s method and capability of targeting the user’s
perceptual field with appropriate stimuli.
2.1.2 Displays
a) Multimodal Displays
Multimodal displays have gained considerable interest over the past 15 years. By distributing
information across multiple sensory channels, these interfaces can support a greater sense
of immersion in virtual-reality environments. They also have the potential to overcome
challenges in complex real-life workplaces, such as data overload and breakdowns in
attention management. To realize this potential, designers require a firm understanding of
the conceptual basis for multimodal information presentation.
b) Spatial Display Model
The Spatial Reality Display (SR Display) reproduces spatial images in three dimensions as if
they were real and can be viewed by the naked eye without special glasses or headsets. It
messages.downloaded_by
lOMoARcPSD|62440031
allows you to see the depth, texture, and appearance of the object with a real sense of
presence, so you can fully express the creator's vision when sharing product designs or
showing variations in color and shape in a showroom. This new spatial experience, makes it
feel as if there is another world behind the display, and can fully deliver the creator's intent to
the viewer.
1) High-speed, high-precision real-time sensing technology
High-speed vision sensors and eye-recognition technology always detects the position of the
viewer's eyes. It detects the position of each of the left and right eyes in real-time, not only
horizontally and vertically, but also in depth
.
2) Real-time video generation algorithm
Based on the positional information of the user's eyes, the light source video that comes out
of the display panel is generated in real time from 3DCG data.
3) Micro-optical lenses
Our unique micro-optical lens, which delivers stereoscopic images to the left and right eyes,
is attached to the entire panel surface with ultra-high precision, enabling natural naked-eye
stereopsis.
c) Visual Displays
An electronic visual display, informally a screen, is a display device for presentation of
images, text, or video transmitted electronically, without producing a permanent record.
Electronic visual displays include television sets, computer monitors, and digital signage. By
the above definition, an overhead projector (along with screen onto which the text, images,
or video is projected) could reasonably be considered an electronic visual display since it is
a display device for the presentation of an images, plain text, or video transmitted
electronically without producing a permanent record. They are also ubiquitous in mobile
computing applications like tablet computers, smartphones, and information appliances.
messages.downloaded_by
lOMoARcPSD|62440031
These are the technologies used to create the various displays in use today.
1) Electroluminescent (EL) display
2) Liquid crystal (LC) display with Light-emitting diode (LED)-backlit Liquid crystal (LC)
display
3) Light-emitting diode (LED) display
4) OLED display
5) AMOLED display
6) Plasma (P) display
7) Quantum dot (QD) display
2.2 Tracking, Calibration and Registration
Registration: alignment of spatial properties
Calibration: offline adjustment of measurements
- Spatial calibration yields static registration
- Offline: once in lifetime or once at startup
- Alternative: autocalibration
Tracking: dynamic sensing and measuring of spatial properties
- Tracking yields dynamic registration
- Tracking in AR/VR always means “in 3D”
2.2.1 Coordinate system
messages.downloaded_by
lOMoARcPSD|62440031
Model Transformation
The model transformation describes the relationship of 3D local object coordinates and 3D
global world coordinates. The model transformation determines where objects are placed in
the real world. Virtual objects are controlled by the application and do not require tracking,
except in very rare situations. For example, object tracking is needed when only an
augmented video stream is available from which to derive tracking. Real objects can be part
of a static real scene or can be allowed to move. Static real scenes do not require a model
transformation. For every moving real object in the scene with which we want to register
virtual information, we must track its model transformation. However, many AR scenarios
deal only with moving objects independent of any global coordinate system, in particular
when using markers. In this case, we do not need a separate world coordinate system and
can use one view transformation per tracked real object instead.
View Transformation
The view transformation describes the relationship of 3D global world coordinates and 3D
camera coordinates. Most AR scenarios allow an observer to move in the real world.
Therefore, tracking the view transformation is the most important objective. AR typically
requires a separate viewing transformation for the camera and the display of the user. If only
a single camera needs to be tracked in a video see-through device, no display calibration
may be necessary. However, other systems in particular, systems using stereoscopic
displays may require calibration of camera and display.
Projective Transformation
messages.downloaded_by
lOMoARcPSD|62440031
The projective transformation describes the relationship of 3D camera coordinates and 2D
device coordinates. Typically, the content of the view frustum (truncated pyramid) is mapped
to a unit cube and then projected onto the screen by dropping the Z component and applying
a viewport transformation (for obtaining screen units in the correct aspect ratio). The
projective transformation is usually calibrated offline. This needs to be done for each camera
and each display separately.
2.3 Sensors
2.3.1. Mobile sensors
Sensors is the device which is used in smartphones to detect various aspects of
environment. They sense data for which they are made and works according to that. There
are various sensors which are available nowadays in smartphones which is in-built and
helps in functioning of the smartphone. Basically, they work for better user experience.
1) Motion Sensors – Motion sensors are useful for monitoring device movement, such as
tilt, shake, rotation, or swing. Smartphones identify their orientation through use of an
accelerometer. The motion sensors present in accelerometer can be used to detect
earthquakes or in medical devices.
2) Environmental Sensors – The Environmental Sensors are used to detect temperature,
humidity, heat losses. Basically, it is used to monitor environmental parameters. They come
up with sensors like Gas Sensors, Humidity Sensors etc.
3) Position Sensors – The Android smartphone provides two sensors that let you determine
position of device- geomagnetic field sensor with combination of accelerometer sensor.
4) Ambient Light Sensor – This sensor works in controlling brightness level of screen. It is
available in almost every smartphone ranging from mid to high. If you have put your
smartphone to Auto-brightness mode, then when you move out in light, there your phone will
automatically boost brightness of the screen. When you come in dark, then with help of this
sensor, phone’s brightness will become dim. Depending on intensity of light, this sensor
manages brightness of the screen.
5) Proximity Sensor – They are available in almost every smartphone at top of the screen.
Infrared light flows through this sensor. When any physical object comes in contact with this
light, it detects it and reacts towards it. For example, when you talk on your phone and place
your phone on your ear, infrared light detects physical object i.e, your ear. Sensing that,
messages.downloaded_by
lOMoARcPSD|62440031
screen’s light automatically goes off. This saves both battery life and prevents accidental
screen touches.
6) Accelerometer Sensor – It is most important sensor which should be available in every
smartphone. It helps phone to check its orientation. For Example, if you rotate your phone in
landscape mode, then all icons present on screen also moves to landscape mode, and when
you want you can change it into portrait mode, this is because of these sensors.
7) Gyroscope Sensor – You must have heard about it’s name. Virtual Reality is possible
only because of these types of sensors. If you buy VR head set, put your phone inside, then
that is only possible because of gyroscope sensors. Even 360 degree pictures or videos and
AR (Augmented Reality) is possible only because of these sensors. These sensors helps
phone to know that which axis (Angles and Directions) it is using at that point of time in very
precise manner. Basically, it adjusts contents of phone according to user.
8) Barometer Sensor - These sensors are not available in every phone, it is available in
high range phone. This is used for detecting altitude (height) data. For Example, The health
app in smartphones also uses these sensors. Going up from stairs or moving from ground
level to floor level, every detail is given by barometer sensor precisely and data is sent to
GPS which then is calculated. It also helps in GPS.
9) Compass Sensor – Compass sensor is very normal and available in every phone, helps
in detecting direction like normal compass do.
There are various other sensors which are not too important for Smartphones but still exists.
10) Pedometer Sensor – It count your steps that how much you have walked. It is available
in high-end devices and some specific devices only.
11) Hall Sensor – It is basically used in tablets in comparison to phones. If you buy flip case
cover for your tablet, and when you open that cover, without pressing any button, light of the
screen will automatically starts and it will start tablet and when you close that flip-cover, light
will goes off.
12) IR Blaster – It is available in every phone of XIAOMI ranging from low to high. For other
companies, it is not available for every phone. These sensors are used to control electronic
devices. For Example, you can control TV, AC or any other electronic devices from your
smartphone if you have IR Blaster in your phone.
messages.downloaded_by
lOMoARcPSD|62440031
2.3.2. Sensor Fusion
A typical mobile device has multiple sensors: at least one camera plus GPS, inertial sensors,
and a compass. Given that individual tracking technologies -optical and non-optical—have
distinct advantages and disadvantages, the best results are obtained when tracking makes
use of the input provided by all available sensors. An obvious way to improve upon
single-sensor tracking is to use multiple types of sensors simultaneously. On the one hand,
such a combination of sensors in a hybrid tracking system increases the weight, cost, and
power consumption of the resulting system, and requires additional calibration effort for the
registration of the sensors with respect to one another. On the other hand, it provides for a
superior overall system performance, which overcomes individual limitations.
In signal processing and robotics, the combination of multiple sensors is often called sensor
fusion. This approach requires both sensor fusion algorithms and a software architecture to
support multiple sensors.
There are 3 Types of Sensor Fusions:
1) Complementary Sensor Fusion
Complementary sensor fusion occurs when multiple sensors supply different degrees of
freedom. No interaction between the sensors is necessary, other than combining the
resulting data. Of course, this combination can still be nontrivial, if the sensors are not
synchronized and use different individual update rates. Such a situation requires at least
some form of temporal interpolation or extrapolation. The most common use of
complementary sensor fusion is to combine a position-only sensor with an orientation-only
sensor to yield full 6DOF. For example, in a modern mobile phone, GPS delivers position
information, while the compass and accelerometer deliver orientation data.
2) Competitive Sensor Fusion
Competitive sensor fusion combines the data from different sensor types measuring the
same degree of freedom independently. The individual measurements are combined into a
measurement of superior quality using some form of mathematical fusion. Redundant sensor
fusion is a simple variant of competitive sensor fusion. When the primary sensor is delivering
measurements, secondary sensors are ignored. Only when the operation of a primary
sensor is not possible does a secondary sensor take over. For example, poor or intermittent
GPS reception.
3) Cooperative Sensor Fusion
In cooperative sensor fusion, a primary sensor relies on information from a secondary
sensor to obtain its measurements. For example, most modern phones contain assisted
messages.downloaded_by
lOMoARcPSD|62440031
GPS (A-GPS), which speeds up the initialization of GPS measurements by deriving a
position constraint from the ID of the cell- tower to which the phone has established a radio
link. Likewise, GPS and compass technologies or accelerometers may be used as an index
into a database of natural features, so that feature matching has a higher success rate. In a
more general sense, cooperative sensor fusion can be described as any measurement of a
property that cannot be derived from either sensor alone.
2.4 Tracking System
2.4.1. Introduction
Tracking term to describe dynamic sensing and measuring of AR systems. To display virtual
objects registered to real objects in three-dimensional space, we must know the position and
orientation of the AR display relative to the real objects. Performs measurements
continuously.
2.4.2. Characteristics of Tracking Technology
1) Physical Phenomenon - Measurements can exploit electromagnetic radiation (including
visible light, infrared light, laser light, radio signals, and magnetic flux), sound, physical
linkage, gravity, and inertia. Specialized sensors are available for each of these physical
phenomena.
2) Measurement Principle- We can measure signal strength, signal direction, and time of
flight (both absolute time and phase of a periodic signal). Note that time-of-flight
measurements require some form of secondary communication channel to confirm clock
synchronization between sender and receiver. Moreover, we can measure electromechanical
properties.
3) Sensor Arrangement - A common approach is to use multiple sensors together in a
known rigid geometric configuration, such as a stereo camera rig. Such a configuration can
either be sparse, if only a few sensors are used, or in the form of a dense 2D array, such as
a digital camera sensor with millions of pixels.
4) Signal Sources - Sources provide the signal that is picked up by the sensors. Like
sensors, sources must be positioned in a known geometric configuration. Sources can be
either passive or active.
- Passive sources rely on natural signals present in the environment, such as natural light or
the Earth's magnetic field. When no external source is apparent, such as in inertial sensing,
the signaling method is described as source less sensing.
- Active sources rely on some form of electronics to produce a physical signal.
5) Degrees of Freedom - In measuring systems, a degree of freedom (DOF) is an
independent dimension of measurement. Registering real and virtual objects in
messages.downloaded_by
lOMoARcPSD|62440031
three-dimensional space usually requires determining the pose of objects with six degrees of
freedom (6DOF).
6) Measurement Error - Real-world sensors are subject to both systematic and random
measurement errors. Systematic measurement error, such as a static offset, a scale factor
error, or a systematic deviation from ideal measurements because of predictable or
measurable influences of the environment can be addressed by improved calibration efforts.
7) Temporal Characteristics - There are two important temporal characteristics of tracking
systems: update rate and latency. The update rate (or temporal resolution) is the number of
measurements performed per given time interval. Latency is the time it takes from the
occurrence of a physical event, such as a motion, to a corresponding data record becoming
available to the AR application.
2.4.3. Stationary Tracking Systems
The stationary tracking systems were the first to become popular for virtual reality
applications, emerging in the 1990s [Meyer et al. 1992] [Rolland et al. 2001]. Mechanical
tracking, electromagnetic tracking, and ultrasonic tracking systems, because of their
stationary nature, are not very popular for AR today. Nevertheless, these systems are useful
to understand some basic principles of tracking
1) Mechanical tracking
Mechanical tracking, which is probably the oldest technique, builds on mechanical
engineering methods that are very well understood. Usually, the end-effector of an
articulated arm with two to four limbs is tracked. This requires knowledge of the extent of
every limb and measurement of the angles at every joint. Joints can have one, two, or three
degrees of freedom in orientation, which are measured using rotary encoders or
potentiometers. From the known length of the limbs and the angular measurements of the
joints, a mathematical formulation of a kinematic chain can be set up to determine the
position and orientation of the end-effector.
This approach delivers high precision and a fast update rate, but the freedom of operation is
severely limited by the mechanical structure. However, movement constraints of the limbs
may prevent the arm from reaching the full range of orientations. Thus, mechanical tracking
can be seen as an outside-in setup with a severely restricted workspace. For AR, it is
undesirable to have the articulated arm in the field of view, where virtual or real objects
should be placed.
2) Electromagnetic Tracking
messages.downloaded_by
lOMoARcPSD|62440031
Electromagnetic tracking uses a stationary source producing three orthogonal magnetic
fields (Figure 3.7). Position and orientation are measured simultaneously from magnetic field
strength and direction using small tethered sensors equipped with three orthogonal coils.
Decreasing field strength with distance and tether length of the sensors typically limit the
operating range to a hemisphere of 1-3 m diameter.
3) Ultrasonic Tracking
Ultrasonic tracking measures the time of flight of a sound pulse traveling from source to
sensor. If a separate (wired or infrared) synchronization channel is available, three
measurements are sufficient for trilateration. Otherwise, additional measurements are
necessary. Multiple ultrasonic sensors can pick up a signal simultaneously, but multiple
sources must send their pulses sequentially to avoid interference. This factor, together with
the modest speed of sound, limits the update rate to 10–50 measurements per second,
which must be shared for all tracked objects. Further limitations include the requirement of
an open line of sight for clear reception, susceptibility to disturbances from loud
environmental noises, and dependence of the speed of sound on air temperature
2.4.4 Optical Tracking
● Optical tracking is a means of determining in real-time the position of an object by
tracking the positions of either active or passive infrared markers attached to the
object.
● Optical tracking systems make use of visual information to track the user. There are
a number of ways this can be done. The most common is to make use of a video
camera that acts as an electronic eye to “watch” the tracked object or person.
Computer vision techniques are then used to determine the object's position based
on what the camera “sees.
● Illumination The first aspect of optical tracking to be discussed is the nature of the
light. We must distinguish approaches that rely on naturally occurring, passive
illumination and those that rely on active illumination.
● Passive Illumination Passive illumination comprises light sources that are not an
integral part of the tracking system. Passive illumination comes both from natural
light sources, in particular the sun, and artificial light sources, such as ceiling lights.
Like humans, conventional cameras see light in the visible spectrum (380-780 nm)
reflected off objects in the environment. Using a conventional digital camera with
passive illumination is the simplest approach to optical tracking in terms of physical
setup.
● Active Illumination Active illumination overcomes the dependence on external light
sources in the environment by combining the optical sensor with an active source of
messages.downloaded_by
lOMoARcPSD|62440031
illumination. Because active illumination in the visible spectrum changes how the
user perceives the environment and is therefore disturbing, a popular approach is to
rely on infrared illumination
● Structured Light Structured light goes one step further than active illumination with
unstructured light sources by projecting a known pattern onto a scene.
● The source of the structured light can be a conventional projector or a laser light
source. The observed reflections are picked up by a camera and used to detect the
geometry of the scene and the contained object.
2.4.5. Multiple Camera Infrared Tracking
In general, the known points in the world will not be constrained to a plane, as assumed in
the previous section on tracking of flat markers. For tracking arbitrary objects, we require
general pose estimation, which addresses the problem of determining the camera pose from
2D-3D correspondences between known points q, in world coordinates and their projections
p, in image coordinates.
In this section, we describe a simple infrared tracking system designed to track rigid body
markers composed of four or more retro-reflective spheres (an approach introduced in
Chapter 3). It uses an outside-in setup with multiple infrared cameras [Dorfmüller 1999]. A
minimum of two cameras in a known configuration a calibrated stereo camera rig is required.
With this strategy, the additional input and wider coverage of the scene from multiple viewing
angles will improve the tracking quality and the working volume. In practice, four cameras
set up in the corners of a laboratory space are a popular configuration. Use of more than two
cameras will improve the performance of the system, but is not fundamentally different from
the stereo case.
The stereo camera tracking pipeline consists of the following steps:
1. Blob detection in all images to locate the spheres of the rigid body markers
2. Establishment of point correspondences between blobs using epipolar geometry between
the cameras
3. Triangulation to obtain 3D candidate points from the multiple 2D points
4. Matching of 3D candidate points to 3D target points
5. Determination of the target's pose using absolute orientation (as described, for example,
by Horn [1987] and Umeyama [1991])
2.4.6. Natural Feature Tracking by Detection
● Unlike AR solutions that use markers as their basis for recognition, natural-feature
tracking solutions can be applied to almost any image as long as the image is
complex enough.
messages.downloaded_by
lOMoARcPSD|62440031
● An example of a natural-feature tracking application is a mobile application that can
recognize a movie poster.
● With natural-feature tracking, the application can analyze the poster and identify it by
comparing the poster image to similar images. In contrast, a marker-based solution
requires a special identifier to be included on the poster; it would be the marker that
provides the identification rather than the poster image.
2.4.7. Incremental Tracking
A tracking system that uses information from a previous step is said to use incremental
tracking or recursive tracking. If the last tracking iteration was successful, there is good
reason to believe that we can be successful again by searching for the inliers from the last
frame and searching close to their last known positions. Such an approach can significantly
facilitate both components of detection:
· Local search. Interest point extraction benefits from limiting the search to a small window
around the prior position.
· Direct matching. The matching can be done by simply comparing the image patch around
an interest point to the image patch in the image being searched. This avoids the costly
creation and comparison of descriptors, though it works only for simple tracking models and
small camera motions. In practice, incremental tracking typically relies on good prior
information of the camera pose.
Incremental tracking requires two components: an incremental search component and an
interest point matching component. Incremental (active) search is performed near the
location of the interest point in the last frame.
2.4.8. Simultaneous Localization and Tracking
The simplest form of model-free tracking, which can be seen as a precursor to simultaneous
localization and mapping (SLAM), is sometimes called visual odometry. In a nutshell, visual
odometry means continuous 6DOF tracking of a camera pose relative to an arbitrary starting
point. This approach originally comes from the field of mobile robotics. Visual odometry
computes a 3D reconstruction of the environment, but uses it just to support the incremental
tracking. A basic visual odometry pipeline encompasses the following steps:
1. Detect interest points in the first frame.
2. Track the interest point in 2D from the previous frame
3. Determine the essential matrix between the current and previous frames from the feature
correspondences with a five-point algorithm
4. Recover the incremental camera pose from the essential matrix.
messages.downloaded_by
lOMoARcPSD|62440031
5. Since the essential matrix determines the translation part of the pose only up to scale, this
scale must be estimated separately, so that it is consistent throughout the tracked image
sequence. To achieve this aim, 3D point locations are triangulated from multiple 3D
observations of the same image feature over time (see "Triangulation from More Than Two
Cameras"). This approach is called structure from motion (SFM).
6. Proceed to the next frame.
2.4.9. Outdoor Tracking
● Indoor environments are usually more predictable whereas the outdoor environments
are limitless in terms of location and orientation.
● As stated earlier, GPS is a good tracking option when working outdoors.
● A differential GPS and a compass was used for position and orientation judgement.
Latitudes and longitudes of several viewpoints were collected in a database along
with the set of images captured at different times of the year with varying light
conditions.
● Reference images were utilized for video tracking and matching was performed to
discover these reference images for the outdoor AR system.
● A video image was examined with the reference images and a matching score was
achieved. For the finest matching score, the 2D transformation was measured and
the current camera position and orientation were deducted.
● This transformation was utilized to register the model on the video frame. The
matching technique was based on Fourier Transformation to be robust against
variation in lighting conditions hence it was limited to only 2D transformations like
rotation and translation.
● This technique had a fixed number of computations therefore it was appropriate for
real-time operation without using markers, it yet worked on 10Hz which is a low rate
for real-time display.
2.4.10 Marker Tracking
● Marker-based Augmented Reality uses a designated marker to activate the
experience. Popular markers include Augmented Reality QR codes, logos, or product
messages.downloaded_by
lOMoARcPSD|62440031
packaging. The shapes or images must be distinctive and recognizable for the
camera to properly identify it in various surroundings.
● There is another important factor of marker-based Augmented Reality. The
marker-based AR experience is tied to the marker. This means that the placement of
digital elements depends on the location of the marker. In most cases, the
experience will display on top of the marker and move along with the marker as it is
turned or rotated.
● You’ll see exactly what we mean in the following two examples.
Marker-Based AR Example: Augmented Reality QR Code
Video : [Link]
Any standard Augmented Reality QR code will do the trick, but this top salesman cleverly
integrated it into his business card design. Real estate is a cut-throat industry, and this agent
certainly stands out among his competition with this memorable AR experience. Clients
simply point their phone’s cameras at his business card to see the Star Wars-esque
greeting.
Marker-Based AR Example: Logo
Video: [Link]
Here is an example that relies on a distinguishable logo design rather than an AR QR code.
We chose another business card for ease of comparison, but remember, the concept applies
to any medium. You can attach QR codes, logos, or other markers to anything from business
cards to billboards.
References:
1. [Link]
_of_the_Augmented_Reality_in_the_Context_of_Education
2. [Link]
edFrom=fulltext
3. [Link]
[Link]
4. [Link]
5. [Link]
Important questions:
1. Discuss briefly visual perception and spatial Display Model. List the characteristics of
Tracking Technology.
2. What do you mean by Augmentation? Describe the methods of Augmentation.
messages.downloaded_by