Advances in AI: RL, Audio Processing, GANs
Advances in AI: RL, Audio Processing, GANs
Supervised Learning
• The model is trained using labeled data (input-output pairs).
• The objective is to learn a mapping from inputs to the correct outputs.
• Examples: classification, regression.
Unsupervised Learning
• The model is trained using unlabeled data.
• The objective is to discover hidden patterns or structures in the data.
• Examples: clustering, dimensionality reduction.
RL
• No labeled data is provided.
• Learning occurs by interacting with the environment.
• The model receives rewards or penalties and learns the optimal
action policy over time.
• Examples: game-playing agents, robotic control, autonomous
driving.
Summary Table
Unsupervised Reinforcement
Feature Supervised Learning
Learning Learning
Data No labels; reward
Labeled Unlabeled
Type signal
Find Maximize
Goal Predict output
structure/pattern cumulative reward
Feedback Direct (correct
None Reward/penalty
Type answer provided)
Learning Trial-and-error
Mapping function Pattern discovery
Style decision-making
2 (a) Define audio signal processing and explain its main objectives. 2M
Sol: Audio signal processing is the discipline concerned with the acquisition,
digitization, representation, analysis, and modification of acoustic waveforms
using mathematical and algorithmic techniques. It involves converting real-
world sound pressure waves into discrete-time signals through sampling,
quantization, and encoding, and subsequently applying digital signal
processing (DSP) operations for analysis, enhancement, transmission, and
synthesis. In its technical form, audio signal processing sits at the intersection
of acoustics, Fourier analysis, statistical modeling, and digital systems, where
the goal is to manipulate audio in its digitized form for reliable storage,
transmission, and interpretation.
The first objective is to convert an analog audio waveform,
𝑥(𝑡)
into a discrete-time signal
𝑥[𝑛] = 𝑥(𝑛𝑇𝑠 )
using a sampling period 𝑇𝑠 , ensuring:
➢ Compliance with the Nyquist sampling theorem
➢ Minimum aliasing through anti-aliasing filters
➢ High-resolution quantization for low distortion
This stage defines the fidelity of all downstream processing.
Once digitized, audio is analysed using mathematical transforms to reveal
spectral and temporal characteristics:
✓ DFT/FFT for spectral analysis
✓ STFT, Wavelet Transform, Cepstrum for time–frequency features
✓ Extraction of technical descriptors: MFCCs, pitch contours,
formants, zero-crossing rate, spectral centroid
This enables speech recognition, audio classification, and source
identification.
Digital Enhancement and Restoration
DSP algorithms aim to improve signal quality in the digital domain:
✓ Adaptive filtering for noise cancellation
✓ Linear prediction (LPC) for speech enhancement
✓ De-reverberation using inverse filtering
✓ Echo suppression via least-mean-square (LMS) algorithms
The objective is to maximize SNR and perceptual quality in real-time
systems.
4M
2. Transformation:
o It uses a neural network to transform this noise into a
structured data sample (image, audio, text, etc.).
o The transformation attempts to map the noise distribution
to the real data distribution 𝑝𝑑𝑎𝑡𝑎 (𝑥).
3. Objective:
o The generator tries to fool the discriminator into
classifying its output as real.
o During training, it updates its parameters to maximize:
log(𝐷(𝐺(𝑧)))
meaning it wants the discriminator to output a probability close to 1 (real)
for generated samples.
4. Learning Behavior:
o The generator improves only when the discriminator
correctly or incorrectly identifies its fakes.
o Over time, it learns the complex patterns, structure, and
statistical properties of the real dataset.
Generator:
• Takes noise → generates fake data.
• Learns to approximate the real data distribution.
• Tries to fool the discriminator.
• Its success is measured by how realistic its outputs become.
Role of the Discriminator (D)
The discriminator acts as a binary classifier that distinguishes between:
• Real samples from the training dataset
• Fake samples produced by the generator
It ensures the generator does not drift away from generating realistic
content.
1. Input:
o It receives both real data samples 𝑥and generated samples 𝐺(𝑧).
2. Output:
o Produces a probability
𝐷(𝑥) ∈ [0,1]
where:
▪ 1 → real
▪ 0 → fake
3. Training Objective:
The discriminator tries to maximize:
log(𝐷(𝑥)) + log(1 − 𝐷(𝐺(𝑧)))
or equivalently maximize:
log(𝐷(𝐺(𝑧)))
Generator:
• Creates fake samples from random noise.
• Tries to mimic the real data distribution.
• Learns by trying to fool the discriminator.
Discriminator:
• Classifies inputs as real or fake.
• Provides feedback to the generator.
• Learns deeper representations to detect synthetic samples.
4 What are the key characteristics of the human vocal tract that influence 6M
speech signals?
Sol: The human vocal tract behaves as a dynamic, shape-varying acoustic
filter that significantly modifies the excitation generated by the vocal folds.
Its physical and physiological properties directly determine the spectral
structure, formant locations, resonance behaviour, and temporal variations
observed in speech signals.
The major characteristics are as follows:
1. Vocal Tract Shape and Length
The vocal tract is an irregular acoustic tube extending from the glottis to the
lips, and its effective length (≈16–18 cm) and continuously varying cross-
sectional shape determine the resonant behaviour of speech.
• Different articulatory configurations (tongue movement, jaw
opening, lip rounding) modify the acoustic transfer function.
• Changes in shape cause formant frequencies (F1, F2, F3...) to shift,
which uniquely identify vowels.
• Consonants arise from constrictions or closures at specific locations
(velar, alveolar, labial).
2. Resonances and Anti-Resonances
The vocal tract exhibits resonant frequencies where specific harmonic
components of the source are amplified.
• Formants: Peaks in the spectral envelope created by resonances of
the tract.
• Anti-resonances: Occur when side cavities (nasal or oral) trap
sound energy, creating spectral dips.
• These resonances determine the timbre, vowel identity, and
phoneme classification.
3. Dynamic Articulator Motion
The position of articulators (tongue, lips, teeth, velum) changes rapidly
during speech.
Key effects:
• Continuous time-variation causes coarticulation, where adjacent
phonemes influence each other.
• Dynamic transitions in formant trajectories carry important
phonetic and prosodic information.
• Temporal variations determine consonant–vowel transitions and
speech intelligibility.
4. Acoustic Coupling of Oral and Nasal Cavities
The velum controls whether the nasal cavity participates in speech
production.
• When the velum lowers, the nasal cavity forms an additional
resonant tract, leading to nasalized vowels and nasal consonants
(m, n, ŋ).
• The coupled oral–nasal tract creates:
o Extra resonances (nasal formants)
o Anti-resonances that attenuate specific frequency bands
This produces the characteristic spectral shape of nasal
speech.
5. Radiation Characteristics at the Lips
The lips behave as a high-pass filter, due to acoustic radiation into the air.
• Low-frequency components are attenuated; high-frequency
components are enhanced.
• This modifies the spectral tilt of the speech signal, making it more
intelligible.
6. Physiological Properties
The walls of the vocal tract are not rigid; they absorb and dissipate acoustic
energy.
• Soft tissues introduce frequency-dependent damping, reducing
higher formants.
• Absorption and viscous losses cause bandwidth widening and
lower spectral peaks.
These properties explain why real speech deviates from ideal tube-model
predictions.
7. Speaker-Specific Characteristics
Inter-speaker variations include:
• Vocal tract length (different for males, females, children)
• Oral cavity volume, tongue size, lip thickness
• Glottal excitation differences (pitch and spectral tilt)
Section B
6 (a) What is meant by an artificial neuron, and how does it mimic the 4M
biological neuron?
Sol: Biological neurons and artificial neurons have some similarities and
differences in their structure and function. Here are some of the main points
of comparison:
Structure: Biological neurons have a complex and organic structure,
consisting of dendrites, soma, axon, and synapses. Artificial neurons have a
simple and mathematical structure, consisting of inputs, weights, bias, and
activation function.
Function: Biological neurons process and transmit electrical and chemical
signals, using action potentials and neurotransmitters. Artificial neurons
process and transmit numerical values, using weighted sums and activation
functions.
Learning: Biological neurons learn and adapt through synaptic plasticity,
changing the strength and number of synapses based on experience and
stimuli. Artificial neurons learn and adapt through weight adjustment,
changing the value and number of weights based on error and feedback.
Efficiency: Biological neurons are highly efficient and parallel, processing
and transmitting signals at high speed and low energy consumption. Artificial
neurons are less efficient and sequential, requiring more time and power to
perform computations and communications.
ANN
Artificial neurons are usually organized into layers, forming a neural network.
The first layer receives the input data, the last layer produces the output, and
the intermediate layers are called hidden layers. Each layer performs a specific
transformation on the data, passing it to the next layer. The more layers and
neurons a neural network has, the more complex functions it can learn.
The input layer has three neurons, corresponding to three features of the data.
The hidden layer has four neurons, performing some computation on the input
data. The output layer has one neuron, producing the final prediction or
decision.
How do they work together?
The way neural networks work is by adjusting the weights of the connections
between the neurons based on the error of the network predictions compared
to the actual data. This is called the training process, where the network learns
from the data and improves its performance. The training process can be done
using various algorithms, such as gradient descent, backpropagation,
stochastic gradient descent, etc.
The goal of the training process is to minimize the error or loss function,
which measures how well the network fits the data. The lower the error or
loss, the better the network performs. The training process can be repeated
until the network reaches a satisfactory level of accuracy or meets some
predefined criteria.
Activation Functions Explained
One of the key components of artificial neurons is the activation function,
which determines if a neuron should be activated or not based on the input
signals. Activation functions introduce non-linearity into neural networks,
enabling them to solve complex problems beyond linear separability. Linear
separability means that the data can be separated by a straight line.
However, not all data sets are linearly separable. Some data sets are more
complex and require curved or nonlinear boundaries to separate them. This is
where activation functions come in handy. They allow the network to learn
nonlinear functions and create nonlinear boundaries. Different types of
activation functions serve various purposes depending on the network
architecture and the problem type. Some of the common activation functions
are:
Sigmoid: This function maps the input to a value between 0 and 1, creating a
smooth curve. It is useful for binary classification problems, such as
predicting whether an email is spam or not. However, it has some drawbacks,
such as being prone to saturation and vanishing gradients, meaning that the
network stops learning when the input is too large or too small.
ReLU: This function maps the input to either 0 or the input itself, creating a
linear and non-linear region. It is useful for speeding up the convergence and
avoiding the vanishing gradient problem. However, it has some drawbacks,
such as being prone to dying neurons, meaning that some neurons stop
responding to any input and become inactive.
Tanh: This function maps the input to a value between -1 and 1, creating a
symmetrical curve. It is useful for centring the data and creating a zero mean.
However, it has some drawbacks, such as being prone to saturation and
vanishing gradients, like the sigmoid function.
Sol:
First-Generation KBS
Second-Generation KBS
Aspect (Classical Expert
(Modern Intelligent Systems)
Systems)
• Layered, modular, and
• Monolithic distributed architecture.
architecture. • Integration of multiple
• System components intelligent modules (e.g.,
1. Architecture
tightly coupled. learning, uncertainty handling,
• Often rule-based with planning).
limited modularity. • Supports agent-based and
blackboard architectures.
• Uses hybrid representation:
• Knowledge stored as
rules + frames + ontologies +
rigid IF–THEN rules.
probabilistic models.
• Symbolic and
• Knowledge encoded in
2. Knowledge deterministic
semantic networks, conceptual
Representation representations.
graphs, fuzzy sets, Bayesian
• Limited ability to
models, and ML models.
handle incomplete or
• Can model uncertainty and
uncertain data.
ambiguity.
• Purely symbolic • Mixed reasoning: symbolic +
reasoning. probabilistic + learning-based
• Uses simple forward inference.
3. Reasoning or backward chaining. • Includes non-monotonic
• Weak at handling reasoning, case-based
uncertain, noisy, or reasoning, neural inference,
dynamic data. fuzzy reasoning.
• Better suited for real-time and
complex domains.
Sol: A recurrent neural network (RNN) is a deep learning model that is trained
to process and convert sequential data input into a specific sequential data
output. Sequential data is data such as words, sentences, or time-series data
where sequential components interrelate based on complex semantics and
syntax rules. An RNN is a software system that consists of many
interconnected components mimicking how humans perform sequential data
conversions, such as translating text from one language to another. RNNs are
largely being replaced by transformer-based artificial intelligence (AI) and
large language models (LLM), which are much more efficient in sequential
data processing.
Hidden layer
RNNs work by passing the sequential data that they receive to the hidden
layers one step at a time. However, they also have a self-looping
or recurrent workflow: the hidden layer can remember and use previous
inputs for future predictions in a short-term memory component. It uses the
current input and the stored memory to predict the next sequence.
For example, consider the sequence: Apple is red. You want the RNN to
predict red when it receives the input sequence Apple is. When the hidden
layer processes the word Apple, it stores a copy in its memory. Next, when it
sees the word is, it recalls Apple from its memory and understands the full
sequence: Apple is for context. It can then predict red for improved accuracy.
This makes RNNs useful in speech recognition, machine translation, and other
language modeling tasks.
Training
Machine learning (ML) engineers train deep neural networks like RNNs by
feeding the model with training data and refining its performance. In ML, the
neuron's weights are signals to determine how influential the information
learned during training is when predicting the output. Each layer in an RNN
shares the same weight.
One-to-many
This RNN type channels one input to several outputs. It enables linguistic
applications like image captioning by generating a sentence from a single
keyword.
Many-to-many
The model uses multiple inputs to predict multiple outputs. For example, you
can create a language translator with an RNN, which analyzes a sentence and
correctly structures the words in a different language.
Many-to-one
Several inputs are mapped to an output. This is helpful in applications like
sentiment analysis, where the model predicts customers’ sentiments
like positive, negative, and neutral from input testimonials.
Exploding gradient
An RNN can wrongly predict the output in the initial training. You need
several iterations to adjust the model’s parameters to reduce the error rate.
You can describe the sensitivity of the error rate corresponding to the model’s
parameter as a gradient. You can imagine a gradient as a slope that you take
to descend from a hill. A steeper gradient enables the model to learn faster,
and a shallow gradient decreases the learning rate.
Exploding gradient happens when the gradient increases exponentially until
the RNN becomes unstable. When gradients become infinitely large, the RNN
behaves erratically, resulting in performance issues such as overfitting.
Overfitting is a phenomenon where the model can predict accurately with
training data but can’t do the same with real-world data.
Vanishing gradient
The vanishing gradient problem is a condition where the model’s gradient
approaches zero in training. When the gradient vanishes, the RNN fails to
learn effectively from the training data, resulting in underfitting. An underfit
model can’t perform well in real-life applications because its weights weren’t
adjusted appropriately. RNNs are at risk of vanishing and exploding gradient
issues when they process long data sequences.
Self-attention
Transformers don’t use hidden states to capture the interdependencies of data
sequences. Instead, they use a self-attention head to process data sequences in
parallel. This enables transformers to train and process longer sequences in
less time than an RNN does. With the self-attention mechanism, transformers
overcome the memory limitations and sequence interdependencies that RNNs
face. Transformers can process data sequences in parallel and use positional
encoding to remember how each input relates to others.
Parallelism
Transformers solve the gradient issues that RNNs face by enabling parallelism
during training. By processing all input sequences simultaneously, a
transformer isn’t subjected to backpropagation restrictions because gradients
can flow freely to all weights. They are also optimized for parallel computing,
which graphic processing units (GPUs) offer for generative AI developments.
Parallelism enables transformers to scale massively and handle complex NLP
tasks by building larger models.
1. Feed-forward weights
1. Input → Hidden-1
5 (inputs) × 5 (hidden-1) = 25
2. Hidden-1 → Hidden-2
5 × 3 = 15
3. Hidden-2 → Output
3×1=3
3. Bias terms
Bias for every neuron except input layer:
• Hidden-1: 5
• Hidden-2: 3
• Output: 1
Total biases:
5+3+1= 9
Section C
9 16M
1. Weights
1. Input → Hidden 1
5 × 4 = 20 weights
2. Hidden 1 → Hidden 2
4 × 3 = 12 weights
3. Hidden 2 → Output
3 × 1 = 3 weights
Total weights
20 + 12 + 3 = 35
2. Biases
Bias in every non-input layer:
• Hidden 1: 4 biases
• Hidden 2: 3 biases
• Output: 1 bias
Total biases
4+3+1= 8
For LeafYield, the goal is to predict the crop yield in kilograms, which is
a continuous numeric value (a regression problem).
Therefore, the most appropriate loss function is:
✓ Mean Squared Error (MSE)
Ravi Shankar
AIDT, Amity University