0% found this document useful (0 votes)
11 views31 pages

Statistical Pattern Recognition Overview

Statistical Pattern Recognition involves classifying objects into groups based on features and probability distributions. Key components include feature extraction, classifiers, and evaluation metrics, with applications in classification and regression tasks. The document also discusses the curse of dimensionality and polynomial curve fitting as techniques for modeling relationships in data.

Uploaded by

gkgk93005
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views31 pages

Statistical Pattern Recognition Overview

Statistical Pattern Recognition involves classifying objects into groups based on features and probability distributions. Key components include feature extraction, classifiers, and evaluation metrics, with applications in classification and regression tasks. The document also discusses the curse of dimensionality and polynomial curve fitting as techniques for modeling relationships in data.

Uploaded by

gkgk93005
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Statistical Pattern Recognition

Statistical Pattern Recognition is the process of identifying which group an object belongs to based on
probability and past data.

Core Concepts of Statistical Pattern Recognition


Concept Simple Explanation

A pattern is just an item described using numbers or features


Pattern / Feature Vector
(like height, color, weight).

The groups into which we classify patterns (e.g., Spam vs Not


Classes / Categories
Spam, Apple vs Orange).

Assumes each class has some statistical behavior (e.g., apples


Probability Distribution
usually weigh between 150-200g).

Selecting important characteristics from raw data (like edges


Feature Extraction
from images or keywords from text).

The method used to decide which class a pattern belongs to


Decision Rule (Classifier)
(e.g., Bayes Rule, KNN, SVM).

The system learns from labeled examples to estimate


Training Phase
probabilities or boundaries.

The system predicts the class of new, unseen data using the
Testing / Classification Phase
learned model.

Used in some methods (like KNN) to measure how close two


Distance / Similarity Measure
patterns are.

Accuracy, Precision, Recall, Confusion Matrix — used to check


Evaluation Metrics model performance.

✅ Key Steps in Statistical


Pattern Recognition

Description (Simple Explanation)


Concept Simple Explanation

Step

Gather training data (e.g., images, text, signals) with known class
1. Data Collection
labels.

Convert raw data into useful numeric features (like height, color,
2. Feature Extraction / Selection
frequency, keywords).

Learn probability distributions or decision boundaries for each


3. Model Estimation / Training
class using training data.

For a new input, compute which class it most likely belongs to


4. Classification / Decision Making
(using Bayes Rule, KNN, SVM, etc.).

Measure accuracy using test data — via Confusion Matrix,


5. Performance Evaluation
Precision, Recall, etc.

6. Optimization / Feedback
Improve the model by tuning parameters or adding more data.
(Optional)

Classification and Regression


Classification and Regression are two fundamental types of supervised learning techniques
used in Statistical Pattern Recognition and Machine Learning. Both aim to predict outcomes
based on input data, but they differ in the type of output they generate.

Classification is the process of assigning input data to a predefined class.

1. Classification

Definition:
Classification is the process of assigning input data to one of several predefined categories or
classes. The output is discrete, meaning it consists of labels like Yes/No, Spam/Not Spam,
Disease/No Disease, etc.

Examples:

 Email filtering → Spam or Not Spam


 Medical diagnosis → Cancer or Healthy
 Handwriting recognition → Digits 0–9

Working Principle:

 During training, the model learns from labeled examples.


 During testing, it predicts the most probable class for new inputs.

Common Algorithms:

 Naive Bayes
 Decision Trees
 K-Nearest Neighbors (KNN)
 Support Vector Machines (SVM)

Evaluation Metrics:

 Accuracy, Precision, Recall, F1-Score, Confusion Matrix

2. Regression

Definition:
Regression is used when the output is a continuous numerical value. Instead of predicting
categories, it predicts quantities such as price, temperature, sales, or age.

Examples:

 Predicting house prices


 Estimating temperature for the next day
 Forecasting stock market values

Working Principle:

 The model fits a mathematical function (like a line or curve) to the training data.
 For new inputs, it outputs a numerical estimate.

Common Algorithms:

 Linear Regression
 Polynomial Regression
 Support Vector Regression (SVR)
 Neural Network Regression

Evaluation Metrics:
 Mean Squared Error (MSE)
 Mean Absolute Error (MAE)
 R-Squared (R²)

3. Key Differences Between Classification and Regression


Aspect Classification Regression

Output Type Categorical / Discrete Continuous / Numerical

Example Spam / Not Spam Predict Email Length

Decision Boundary Classes separated Best-fit line or curve

Loss Function Cross-Entropy, Log Loss MSE, MAE

Evaluation Accuracy, F1-Score RMSE, R²

4. Applications in Real World


Domain Classification Regression

Finance Fraud Detection Loan Amount Prediction

Healthcare Disease Detection Blood Sugar Level Prediction

Marketing Customer Segmentation Sales Forecasting

Autonomous Vehicles Object Identification Distance Estimation

Conclusion

Classification and regression are core components of predictive modeling.

 Classification is suited for decision-making problems where outcomes are categories,


 while Regression is used for forecasting and estimation where outputs are numerical
values.

Both play a vital role in machine learning, data science, and statistical pattern recognition,
enabling systems to learn from past data and make intelligent predictions.
1. Features

Definition:
A feature is an individual measurable property or characteristic of an object . Features are
the inputs that help a model understand and distinguish between different classes.

Examples of Features:

Domain Example Object Features

Face Recognition Human Face Eye distance, Skin tone, Nose width

Email Filtering Email Number of links, Keyword count ("free"), Sender domain

Medical Diagnosis Patient Age, Blood Pressure, Glucose Level

Types of Features:

 Numerical (Continuous): Height, temperature, income


 Categorical (Discrete): Gender (Male/Female), Color (Red/Blue)
 Binary: Yes/No, 0/1, Present/Absent
 Derived/Engineered: Body Mass Index (calculated from height & weight)

Importance:
Well-chosen features improve model accuracy, while irrelevant features may confuse the
classifier.

2. Feature Vectors

Definition:

A feature vector is a collection of these features that serve as an input to machine learning
models, classifiers, or statistical algorithms.

A feature vector is an ordered list (or array) of multiple features representing a single data
instance. It converts real-world objects into numerical format, which machine learning models
can process.

Example:

Suppose we want to classify fruits:


Fruit Weight (g) Color Code (Red=1, Green=2, Yellow=3) Sweetness (0-10)

Apple 150 1 7

Banana 120 3 8

So, Feature Vector = [150, 1, 7] for Apple.

Mathematical Representation:

X=[x1,x2,x3,...,xn]X = [x_1, x_2, x_3, ..., x_n]X=[x1,x2,x3,...,xn]

Where each xix_ixi is a feature.

Importance:
Feature vectors allow computers to compare, classify, or cluster objects mathematically
using distances, similarity measures, or probability.

3. Classifiers

Definition:
A classifier is a decision-making model that uses feature vectors to assign data to a specific
class or category.

They make decisions about which class a particular feature vector belongs to.

Working Principle:

1. Input: Feature Vector


2. Model: Learns patterns from labeled data
3. Output: Class Label (e.g., Spam / Not Spam, Disease / No Disease)

Types of Classifiers:

Type Example Algorithms Nature

Probabilistic Naive Bayes Based on probability

Distance-Based K-Nearest Neighbors Uses distance/similarity

Linear Logistic Regression, SVM Creates linear boundaries

Tree-Based Decision Trees, Random Forest Based on rule splitting


Type Example Algorithms Nature

Neural Networks Deep Learning Complex non-linear decision making

Conclusion

 Features describe objects using measurable attributes.


 Feature Vectors convert those attributes into numerical form for processing.
 Classifiers use these vectors to make intelligent decisions and assign classes.

These three components form the core foundation of pattern recognition and machine
learning systems. Without meaningful features and proper classification models, no machine
can learn or recognize patterns effectively.

Pre-Processing and Feature Extraction


Pre-processing is the stage of preparing raw data for machine learning models. It involves transforming
the data into a cleaner and more structured format, which improves the performance of the classifiers.

Pre-Processing Steps:

1. Data Cleaning: Removing noise, handling missing values, and correcting errors in the dataset.

2. Normalization/Standardization:

Normalization: Rescales the data to a range of [0, 1] or [-1, 1].

Standardization: Transforms the data to have zero mean and unit variance.

These techniques ensure that no feature dominates others because of its scale.

3. Dimensionality Reduction: Techniques like Principal Component Analysis (PCA) and Linear
Discriminant Analysis

(LDA) are used to reduce the number of features while preserving as much information as possible.

4. Noise Reduction: Smoothing or filtering techniques (e.g., Gaussian filter, Median filter) are used to
remove irrelevant noise from data.

Feature Extraction:
Definition: The process of transforming raw data into a set of features that better represent the
underlying patterns in the data. Feature extraction focuses on identifying key information from the raw
data that is most relevant for the learning task.

Common Feature Extraction Techniques (With Explanation)


1. PCA (Principal Component Analysis)

 Reduces the number of features by keeping only the most important information.
 Used for dimensionality reduction in images and numeric data.

2. LDA (Linear Discriminant Analysis)

 Similar to PCA but focuses on maximizing class separation.


 Common in face recognition and classification problems.

3. Histogram of Oriented Gradients (HOG)

 Extracts edge and shape information from images.


 Used in object detection like pedestrian or vehicle recognition.

4. SIFT (Scale-Invariant Feature Transform)

 Detects key points in images that remain stable under scaling or rotation.
 Useful in logo matching, image stitching, and object recognition.

5. TF-IDF (Term Frequency–Inverse Document Frequency)

 Converts text into numbers based on word importance.


 Used in text classification and search engines.

6. Word Embeddings (Word2Vec, GloVe, BERT)

 Converts words into dense vector representations with meaning.


 Helps in sentiment analysis, chatbots, NLP tasks.

7. MFCC (Mel-Frequency Cepstral Coefficients)

 Extracts frequency-based features from speech or audio.


 Used in speech recognition and music classification.

8. Fourier Transform / FFT

 Converts time-domain signals into frequency-domain features.


 Used in signal processing and vibration analysis.

9. Wavelet Transform

 Captures both time and frequency information from signals.


 Better than FFT when signal changes over time.

10. Statistical Features (Mean, Variance, Skewness)

 Simple numeric summaries of data.


 Commonly used in sensor data and time-series analysis.

The Curse of Dimensionality


The curse of dimensionality refers to various challenges and phenomena that arise when working with
high-dimensional data. As the number of dimensions (features) increases, the amount of data needed to
learn a model effectively grows exponentially, making it difficult for machine learning algorithms to
perform well. This is especially problematic in pattern recognition and classification tasks.

Key Issues Caused by High Dimensions

1. Data Sparsity
o In high dimensions, data points are spread out.
o More data is needed to cover the space sufficiently.
2. Distance Concentration
o Distances between points become almost the same, making similarity measures
like Euclidean distance less effective.
3. Increased Computational Complexity
o More dimensions → more calculations → slower algorithms.
4. Overfitting
o Models may fit noise instead of patterns because there are too many features.
5. Visualization Difficulty
o Hard to visualize or interpret data beyond 3 dimensions.

Examples

1. K-Nearest Neighbors (KNN)


o With many features, all points are “far apart,” so finding nearest neighbors
becomes unreliable.
2. Image Recognition
o An image with 256x256 pixels has 65,536 dimensions. Without dimensionality
reduction, algorithms are slow and prone to overfitting.
3. Text Classification
o Bag-of-Words representation can have thousands of features (words), making the
space extremely sparse.

Solutions / Mitigation Techniques

1. Dimensionality Reduction
o PCA (Principal Component Analysis)
oLDA (Linear Discriminant Analysis)
oAutoencoders (Neural Networks)
2. Feature Selection
o Remove irrelevant or redundant features.
3. Regularization
o Techniques like L1 or L2 regularization reduce overfitting.
4. Clustering / Manifold Learning
o Map high-dimensional data to lower-dimensional manifolds

Summary
Features and Feature Vectors: Features are individual measurable characteristics used to represent
data, while feature vectors are the collection of these features in a numerical form that serves as the
input for classifiers.

Classifiers: These are algorithms that categorize new data points based on patterns learned from
training data. They use the feature vectors to assign data points to different categories. Common
classifiers include Naive Bayes, Decision Trees, SVM, and k-NN.

Pre-Processing and Feature Extraction: Pre-processing involves cleaning and normalizing the data,
while feature extraction involves transforming raw data into a set of features that better represent the
patterns in the data. This process helps reduce the dimensionality and improves model performance.

The Curse of Dimensionality: As the number of features increases, data points become sparse, and
distance metrics become less meaningful, making it hard to learn useful patterns. Dimensionality
reduction and feature selection techniques are used to mitigate these issues.

Polynomial Curve Fitting


Definition

Polynomial curve fitting is a regression technique used to model the relationship between a
dependent variable (output) and an independent variable (input) using a polynomial
function.

 The goal is to fit a smooth curve through a set of data points.


 It is a type of non-linear regression, but can be treated as linear regression in terms of
coefficients.

Polynomial Function

A polynomial of degree nnn is expressed as:

y=a0+a1x+a2x2+a3x3+⋯+anxny
Where:

 y = dependent variable
 x = independent variable
 a0,a1,...,ana_0, a_1, ..., a_na0,a1,...,an = coefficients to be determined from data

Steps in Polynomial Curve Fitting

1. Collect Data Points


o Gather (xi,yi)(x_i, y_i)(xi,yi) pairs for i=1,2,...,mi = 1, 2, ..., mi=1,2,...,m.
2. Choose Polynomial Degree nnn
o Degree determines curve flexibility.
o Higher degree → fits data closely, may overfit.
3. Set Up Equations (Normal Equations)
o Use least squares method to minimize error:

E=∑i=1m(yi−y^i)2E = \sum_{i=1}^{m} (y_i - \hat{y}_i)^2E=i=1∑m(yi−y^i)2

4. Solve for Coefficients a0,a1,...,ana_0, a_1, ..., a_na0,a1,...,an


o Can use matrix methods or computational tools.
5. Fit Curve and Evaluate
o Plot the curve and check goodness of fit (R², residuals).

Advantages

 Can model complex non-linear relationships.


 Flexible: increasing degree allows better fit.
 Simple mathematical formulation using least squares.

Disadvantages / Challenges

 High-degree polynomials can overfit training data.


 Extrapolation outside data range may be unreliable.
 Sensitive to outliers, which can distort the curve.

Applications

 Engineering: Modeling stress-strain relationships.


 Economics: Predicting demand or stock prices.
 Physics / Chemistry: Fitting experimental data.
 Machine Learning: As a preprocessing step for regression tasks.

Example

Suppose you have data points:

x y

12

25

3 10

 Fit a quadratic curve: y=a0+a1x+a2x2y = a_0 + a_1 x + a_2 x^2y=a0+a1x+a2x2


 Solve using least squares → get coefficients → plot curve fitting the points.

Model Complexity
Definition

Model complexity refers to the capacity of a model to capture patterns in data, which
depends on the number of parameters, features, or the form of the model.

 A simple model has fewer parameters and captures only basic patterns.
 A complex model has more parameters and can capture intricate relationships.

Key Concepts

1. Underfitting
o Occurs when a model is too simple to capture the underlying pattern.
o Example: Fitting a linear model to a clearly nonlinear dataset.
o Effect: High bias → poor training and test accuracy.
2. Overfitting
o Occurs when a model is too complex and captures noise along with the pattern.
o Example: Fitting a high-degree polynomial to a few data points.
o Effect: Low training error but poor generalization → high test error.
3. Optimal Complexity
o Balance between underfitting and overfitting.
o The model is just complex enough to capture real patterns without memorizing
noise.
Factors Affecting Model Complexity

1. Number of Features / Input Variables


o More features → higher dimensionality → increased complexity.
2. Degree of the Model
o For polynomial regression: higher degree → more flexible curve → more
complex.
3. Type of Model
o Simple linear regression vs deep neural networks.
4. Amount of Data
o Complex models require more data to avoid overfitting.

Techniques to Control Model Complexity

1. Regularization
o L1 (Lasso) or L2 (Ridge) penalties to limit coefficient size.
2. Feature Selection / Dimensionality Reduction
o Reduce unnecessary or correlated features.
3. Cross-Validation
o Helps to choose a model with optimal complexity.
4. Pruning (for Decision Trees)
o Remove branches that add little predictive power.
5. Early Stopping (for Neural Networks)
o Stop training when test error starts increasing.

Advantages of Proper Complexity

 Achieves good generalization on unseen data.


 Balances bias and variance.
 Efficient use of computational resources.

Consequences of Poor Complexity


Issue Description

Underfitting Model too simple → cannot learn patterns → poor accuracy

Overfitting Model too complex → learns noise → poor generalization

Increased Computation Complex models require more memory and time


Issue Description

Difficult Interpretation High-complexity models are hard to explain

Example

 Predicting house prices:


o Linear model (simple) → may underfit if price depends on many nonlinear
factors.
o High-degree polynomial or deep neural network (complex) → may overfit small
dataset.
o Optimal: Medium complexity model with selected features → good prediction.

Multivariate non-linear functions.

When dealing with multiple input variables (features), a multivariate non-linear function describes the
relationship between the input variables and the output using non-linear transformations.

Definition

A multivariate non-linear function is a function that:

1. Depends on two or more independent variables (multivariate), and


2. Has a non-linear relationship between input variables and the output.

Mathematically:

y=f(x1,x2,...,xn)y = f(x_1, x_2, ..., x_n)y=f(x1,x2,...,xn)

where fff is non-linear, meaning the function cannot be expressed as a simple linear
combination of inputs:

y≠a0+a1x1+a2x2+...+anxny \neq a_0 + a_1 x_1 + a_2 x_2 + ... + a_n x_ny a0+a1x1+a2x2+...+anxn

Characteristics

1. Multiple Inputs:
o Depends on several independent variables (x1,x2,...,xnx_1, x_2, ..., x_nx1,x2
,...,xn).
2. Non-Linearity:
oCan include powers, products, exponential, logarithmic, trigonometric, or other
non-linear terms.
3. Complex Relationships:
o Output changes non-proportionally with inputs.
4. Harder to Model:
o Requires advanced techniques (like polynomial regression, neural networks, or
kernel methods).

Examples

1. Polynomial Non-Linear Function:

y=2x12+3x2x1−5x23+7y = 2x_1^2 + 3x_2 x_1 - 5x_2^3 + 7y=2x12+3x2x1−5x23+7

2. Exponential Function:

y=ex1+2x2+sin⁡(x3)y = e^{x_1 + 2x_2} + \sin(x_3)y=ex1+2x2+sin(x3)

3. Non-Linear Interaction Function:

y=x1x2+x22−ln⁡(x3)y = x_1 x_2 + x_2^2 - \ln(x_3)y=x1x2+x22−ln(x3)

Applications

 Engineering: Stress-strain modeling, control systems


 Economics: Predicting supply-demand relationships
 Machine Learning: Non-linear regression, neural networks, kernel-based SVM
 Physics: Modeling motion, fluid dynamics, or chemical reactions

Techniques to Handle Multivariate Non-Linear Functions

1. Polynomial Regression – Fit polynomial terms of multiple variables.


2. Neural Networks – Capture highly non-linear multivariate patterns.
3. Kernel Methods (SVM) – Map input to higher-dimensional space for linear separation.
4. Non-Linear Optimization – Estimate coefficients using methods like gradient descent.
5. Dimensionality Reduction – Reduce feature space before modeling complex non-linear
relationships.
Advantages

 Can model complex, real-world relationships that linear models cannot.


 Provides better accuracy if the underlying relationship is truly non-linear.

Challenges

 Requires more data to estimate parameters accurately.


 Can overfit easily if the model is too flexible.
 Interpretation is difficult compared to linear models.
 Computation may be time-consuming for high-dimensional data.

Bayes’ Theorem
Definition

Bayes’ Theorem is a fundamental rule in probability theory that describes how to update the
probability of a hypothesis based on new evidence.

 It allows us to calculate the posterior probability of an event given prior knowledge and
observed data.

Mathematical Formula
P(H∣E)=P(E∣H)⋅P(H)P(E)P(H|E) = \frac{P(E|H) \cdot P(H)}{P(E)}P(H∣E)=P(E)P(E∣H)⋅P(H)

Where:

 P(H∣E)P(H|E)P(H∣E) = Posterior probability: Probability of hypothesis HHH given


evidence EEE
 P(E∣H)P(E|H)P(E∣H) = Likelihood: Probability of observing evidence EEE if HHH is
true
 P(H)P(H)P(H) = Prior probability: Initial probability of hypothesis HHH before seeing
evidence
 P(E)P(E)P(E) = Marginal probability: Total probability of observing evidence EEE

Intuitive Explanation

 Start with a prior belief about an event.


 Observe new evidence.
 Update the probability of the event considering the new evidence.
In simple words:

“Bayes’ theorem tells us how to revise our beliefs after seeing new data.”

Example

Suppose a medical test for a disease is 99% accurate, and 1% of the population has the
disease.

 Let HHH = person has the disease


 Let EEE = person tests positive

Using Bayes’ Theorem:

P(H∣E)=P(E∣H)⋅P(H)P(E)P(H|E) = \frac{P(E|H) \cdot P(H)}{P(E)}P(H∣E)=P(E)P(E∣H)⋅P(H)

 P(H)=0.01P(H) = 0.01P(H)=0.01 (prior)


 P(E∣H)=0.99P(E|H) = 0.99P(E∣H)=0.99 (likelihood)
 P(E)=P(E∣H)P(H)+P(E∣No H)P(No H)=0.99⋅0.01+0.05⋅0.99=0.0594P(E) = P(E|H)P(H) +
P(E|\text{No H})P(\text{No H}) = 0.99 \cdot 0.01 + 0.05 \cdot 0.99 =
0.0594P(E)=P(E∣H)P(H)+P(E∣No H)P(No H)=0.99⋅0.01+0.05⋅0.99=0.0594

P(H∣E)=0.99⋅0.010.0594≈0.1667P(H|E) = \frac{0.99 \cdot 0.01}{0.0594} \approx


0.1667P(H∣E)=0.05940.99⋅0.01≈0.1667

Even with a positive test, the probability of having the disease is ~16.7%!

Applications

1. Medical Diagnosis – Updating disease probabilities based on test results


2. Spam Detection – Classifying emails using words as evidence (Naive Bayes)
3. Machine Learning – Feature-based probabilistic classification
4. Decision Making – Risk assessment and Bayesian inference
5. Fault Detection – Predicting system failures from sensor data

Advantages

 Provides a mathematically sound way to update probabilities


 Works well with uncertain or incomplete information
 Foundation for Naive Bayes classifier and other probabilistic models
Limitations

 Requires prior probabilities, which may be subjective


 Calculating marginal probabilities can be difficult for complex systems
 Assumes correct likelihood estimation for accuracy

Definition

A decision boundary is a hypothetical line, curve, or surface in the feature space that
separates different classes in a classification problem.

 It defines the regions where a classifier assigns data points to a particular class.
 Any new data point is classified based on which side of the boundary it falls.

Key Concepts

1. Feature Space
o An n-dimensional space where each axis represents a feature.
o Decision boundaries exist in this space to separate classes.
2. Types of Decision Boundaries
o Linear Boundary: Straight line (2D), plane (3D), or hyperplane (nD).
 Example: Logistic Regression, Linear SVM
o Non-Linear Boundary: Curved or complex surface.
 Example: Decision Trees, Kernel SVM, Neural Networks
3. Classification Rule
o A classifier assigns a point xxx to class CiC_iCi if xxx lies in the region defined
for that class by the decision boundary.

Illustration (2D Example)

 Suppose we have two features x1x_1x1 and x2x_2x2, and two classes, A and B.
 The classifier learns a line or curve separating the two classes.
 New points on one side → Class A, on the other side → Class B.

Linear Example:

w1x1+w2x2+b=0w_1 x_1 + w_2 x_2 + b = 0w1x1+w2x2+b=0

 This equation represents a linear decision boundary in 2D space.

Non-Linear Example:
x12+x22=r2x_1^2 + x_2^2 = r^2x12+x22=r2

 Represents a circular boundary separating classes inside vs outside the circle.

Importance

1. Helps visualize how a classifier separates classes.


2. Indicates complexity of the classifier (linear vs non-linear).
3. Guides feature engineering—well-chosen features make boundaries simpler.
4. Determines generalization ability—overfitted boundaries may capture noise.

Applications

 Pattern Recognition: Face detection, handwriting recognition


 Machine Learning: Visualizing SVM, decision tree, neural network classification
 Medical Diagnosis: Separating healthy vs diseased patients based on features
 Finance: Credit risk classification

Factors Affecting Decision Boundaries

1. Number of Features – More features → higher-dimensional boundary.


2. Classifier Type – Linear vs non-linear algorithms.
3. Training Data Quality – Noisy or overlapping classes → fuzzy boundaries.
4. Regularization – Reduces overfitting → smoother boundaries.

Parametric Methods
Definition

Parametric methods are statistical and machine learning techniques that assume a specific form
for the underlying probability distribution of the data.

 These methods estimate a finite set of parameters (like mean, variance, or coefficients)
from training data.
 Once parameters are estimated, the model can predict or classify new data.

Key Concepts

1. Assumption of Distribution
oParametric methods assume the data comes from a known distribution (e.g.,
Gaussian/Normal, Exponential, Binomial).
2. Parameter Estimation
o Parameters such as mean (μ\muμ), variance (σ2\sigma^2σ2), or regression
coefficients are estimated using training data.
o Methods: Maximum Likelihood Estimation (MLE), Method of Moments, Least
Squares.
3. Model Representation
o The model is fully defined by the estimated parameters.
o Examples:
 Linear Regression → coefficients a0,a1,…a_0, a_1, …a0,a1,…
 Gaussian Naive Bayes → mean and variance of each feature per class

Advantages of Parametric Methods

1. Simplicity and Efficiency


o Only a few parameters need to be estimated → computationally faster.
2. Less Data Required
o Works well even with small datasets if assumptions are valid.
3. Interpretable
o Model parameters have clear meaning, e.g., effect of a feature on output.
4. Predictive Power
o Can make predictions for unseen data efficiently once parameters are learned.

Disadvantages / Limitations

1. Assumption Dependent
o Accuracy depends on whether the assumed distribution fits the data.
2. Limited Flexibility
o Cannot capture complex or non-linear relationships if the model is too simple.
3. Poor Performance with Misfit
o If assumptions are wrong, predictions or classification results may be highly
inaccurate.

Examples of Parametric Methods


Method Key Assumptions Parameters

Linear relationship between input & Coefficients a0,a1,…a_0, a_1, …a0


Linear Regression
output ,a1,…
Method Key Assumptions Parameters

Non-linear relation modeled as Coefficients a0,a1,…,ana_0, a_1, …,


Polynomial Regression
polynomial a_na0,a1,…,an

Mean and variance per feature per


Gaussian Naive Bayes Features are normally distributed
class

Linear combination of inputs maps to Coefficients b0,b1,…b_0, b_1, …b0


Logistic Regression
probability ,b1,…

LDA (Linear Discriminant Classes follow Gaussian distribution


Class means, covariance matrix
Analysis) with same covariance

Applications

 Medical Diagnosis: Predict disease probability using Gaussian Naive Bayes


 Finance: Credit scoring with logistic regression
 Engineering: Predicting stress, temperature, or other measurements
 Pattern Recognition: Face recognition using LDA or parametric classifiers

Conclusion

Parametric methods are efficient, interpretable, and mathematically simple methods for
modeling and classification.

 Best suited when data fits the assumed distribution.


 If the assumptions are violated, non-parametric methods are often preferred.

Sequential Parameter Estimation


Definition

Sequential Parameter Estimation is a method of estimating the parameters of a statistical


model incrementally, as new data becomes available.

 Unlike traditional batch methods (which require all data at once), sequential estimation
updates the parameters continuously without storing all past data.
 It is particularly useful in real-time or online systems where data arrives over time.
Key Concepts

1. Parameter Estimation
o Parameters such as mean (μ\muμ) and variance (σ2\sigma^2σ2) are estimated
from observed data.
o In sequential estimation, these estimates are updated with each new
observation.
2. Incremental Update
o Suppose we want to estimate the mean μ\muμ of a population:
 After observing nnn samples, the estimate is:

μ^n=x1+x2+⋯+xnn\hat{\mu}_n = \frac{x_1 + x_2 + \dots + x_n}{n}μ^n=nx1+x2


+⋯+xn

 When a new sample xn+1x_{n+1}xn+1 arrives:

μ^n+1=μ^n+xn+1−μ^nn+1\hat{\mu}_{n+1} = \hat{\mu}_n + \frac{x_{n+1} -


\hat{\mu}_n}{n+1}μ^n+1=μ^n+n+1xn+1−μ^n

3. Variance Update
o Variance can also be updated sequentially using similar incremental formulas.
4. Advantages Over Batch Estimation
o No need to store entire dataset.
o Efficient in memory and computation.
o Suitable for streaming data or online learning.

Example

Suppose we want to estimate the average temperature in a city:

 Day 1: Temperature = 30°C → μ^1=30\hat{\mu}_1 = 30μ^1=30


 Day 2: Temperature = 32°C → μ^2=30+32−302=31\hat{\mu}_2 = 30 + \frac{32 -
30}{2} = 31μ^2=30+232−30=31
 Day 3: Temperature = 31°C → μ^3=31+31−313=31\hat{\mu}_3 = 31 + \frac{31 -
31}{3} = 31μ^3=31+331−31=31

The estimate updates sequentially as new data arrives.

Applications

1. Real-Time Systems
o Robotics: Updating sensor measurements
o
Weather forecasting: Streaming temperature or humidity data
2. Machine Learning
o Online learning algorithms
o Adaptive filtering
3. Finance
o Real-time stock price prediction
o Risk estimation as new transactions occur
4. Signal Processing
o Adaptive filters for noise reduction

Advantages

 Memory-efficient: No need to store all past data.


 Computationally efficient: Updates require simple calculations.
 Real-time applicability: Suitable for online and dynamic environments.
 Flexibility: Works with non-stationary data streams.

Disadvantages / Limitations

 Sensitive to outliers if not handled carefully.


 May require tuning of learning rate or weighting for stability.
 Estimation quality depends on the order of incoming data in some cases.

Conclusion

Sequential Parameter Estimation is a powerful technique for estimating model parameters in


dynamic or streaming environments.

 It allows continuous improvement of estimates as new information arrives, making it


essential for real-time applications in engineering, finance, and machine learning.

Linear Discriminant Functions


Definition

A Linear Discriminant Function (LDF) is a mathematical function used in pattern


recognition and machine learning to separate two or more classes of data using a linear
boundary in feature space.

 It is primarily used in classification problems where the classes are linearly separable.
 The function assigns a score to each class based on the input features, and the class with
the highest score is selected.

Mathematical Form

For a d-dimensional input vector x=[x1,x2,...,xd]Tx = [x_1, x_2, ..., x_d]^Tx=[x1,x2,...,xd]T,


the linear discriminant function is:

g(x)=wTx+w0g(x) = w^T x + w_0g(x)=wTx+w0

Where:

 w=[w1,w2,...,wd]Tw = [w_1, w_2, ..., w_d]^Tw=[w1,w2,...,wd]T → weight vector


(parameters)
 w0w_0w0 → bias or threshold
 g(x)g(x)g(x) → discriminant score for a class

Classification Rule:

 If g(x)>0g(x) > 0g(x)>0 → assign to class 1


 If g(x)<0g(x) < 0g(x)<0 → assign to class 2

Geometric Interpretation

 The equation wTx+w0=0w^T x + w_0 = 0wTx+w0=0 represents a hyperplane (line in


2D, plane in 3D) that separates classes in the feature space.
 Decision boundary is the hyperplane where the discriminant function equals zero.
 All points on one side of the hyperplane belong to one class, and points on the other side
belong to the other class.

Linear Discriminant Analysis (LDA)

 LDA is a popular method to find the optimal linear discriminant function that
maximizes class separability.
 Objective: Maximize the ratio of between-class variance to within-class variance.
 LDA can be used for dimensionality reduction while preserving class separability.
Advantages

1. Simple and Efficient – Easy to compute and implement.


2. Interpretable – Linear weights indicate feature importance.
3. Good for Linearly Separable Data – Works well when classes are clearly separated.
4. Less Data Required – Needs fewer samples compared to non-linear methods.

Disadvantages

1. Cannot Handle Non-Linearity – Fails if classes are non-linearly separable.


2. Sensitive to Outliers – Outliers can affect the linear boundary.
3. Assumes Homoscedasticity – Assumes same covariance for all classes (in LDA).

Applications

 Face Recognition – Linear boundaries between different people’s faces.


 Medical Diagnosis – Classifying patients as healthy or diseased.
 Finance – Predicting credit approval (good vs bad credit).
 Pattern Recognition – Handwriting and character recognition.

Example

Suppose we have two classes, A and B, in a 2D feature space (x1,x2)(x_1, x_2)(x1,x2):

g(x)=2x1+3x2−5g(x) = 2x_1 + 3x_2 - 5g(x)=2x1+3x2−5

 Decision Boundary: 2x1+3x2−5=02x_1 + 3x_2 - 5 = 02x1+3x2−5=0


 Points above the line → Class A, points below → Class B

Conclusion

Linear Discriminant Functions provide a simple and effective way to classify data when classes
are linearly separable.

 They are widely used in pattern recognition, machine learning, and statistical
classification, and form the basis for Linear Discriminant Analysis (LDA).
Fisher’s Linear Discriminant (FLD)
Definition

Fisher’s Linear Discriminant (FLD) is a dimensionality reduction and classification


technique used to project high-dimensional data onto a line such that the separation between
classes is maximized.

 Developed by R.A. Fisher in 1936.


 Particularly useful when there are two or more classes, and we want a linear
combination of features that best separates them.

Objective

Fisher’s method aims to:

1. Maximize the between-class variance (SBS_BSB) → makes class means far apart.
2. Minimize the within-class variance (SWS_WSW) → makes data in each class compact.

The optimal projection vector www maximizes the Fisher criterion:

J(w)=wTSBwwTSWwJ(w) = \frac{w^T S_B w}{w^T S_W w}J(w)=wTSWwwTSBw

Where:

 SBS_BSB = Between-class scatter matrix


 SWS_WSW = Within-class scatter matrix
 www = Projection vector (linear combination of features)

Formulation for Two Classes

Suppose we have two classes C1C_1C1 and C2C_2C2 with means μ1\mu_1μ1 and μ2\mu_2μ2,
and covariance matrices Σ1\Sigma_1Σ1 and Σ2\Sigma_2Σ2:

1. Within-class scatter matrix:

SW=Σ1+Σ2S_W = \Sigma_1 + \Sigma_2SW=Σ1+Σ2

2. Between-class scatter matrix:

SB=(μ1−μ2)(μ1−μ2)TS_B = (\mu_1 - \mu_2)(\mu_1 - \mu_2)^TSB=(μ1−μ2)(μ1−μ2)T

3. Optimal projection vector:


w=SW−1(μ1−μ2)w = S_W^{-1} (\mu_1 - \mu_2)w=SW−1(μ1−μ2)

 Projecting data onto www maximizes class separation.

Classification Rule

1. Project feature vectors xxx onto www:

y=wTxy = w^T xy=wTx

2. Set a threshold (usually midway between projected class means):

y0=wTμ1+wTμ22y_0 = \frac{w^T \mu_1 + w^T \mu_2}{2}y0=2wTμ1+wTμ2

3. Decision:

 If y>y0y > y_0y>y0 → Class C1C_1C1


 If y<y0y < y_0y<y0 → Class C2C_2C2

Advantages

1. Reduces Dimensionality
o Projects high-dimensional data to 1D or few dimensions.
2. Maximizes Class Separability
o Ensures classes are well-separated along projection.
3. Simple Computation
o Only requires mean vectors and covariance matrices.
4. Widely Used
o Foundation for LDA and other linear classifiers.

Disadvantages

1. Linear Assumption
o Assumes classes are linearly separable after projection.
2. Sensitive to Outliers
o Outliers can distort scatter matrices and projection vector.
3. Two-Class Limitation
o Direct formulation is for two classes; multiple classes require generalization.
Applications

 Face Recognition – Linear projection of facial features to distinguish people.


 Medical Diagnosis – Separating healthy vs diseased patients.
 Pattern Recognition – Handwritten digit or letter recognition.
 Finance – Credit risk classification.

Example (Simplified)

Suppose two classes with features:

 Class 1: mean μ1=[2,3]\mu_1 = [2, 3]μ1=[2,3]


 Class 2: mean μ2=[5,7]\mu_2 = [5, 7]μ2=[5,7]

1. Compute SWS_WSW using covariance matrices.


2. Compute w=SW−1(μ1−μ2)w = S_W^{-1} (\mu_1 - \mu_2)w=SW−1(μ1−μ2)
3. Project data points: y=wTxy = w^T xy=wTx
4. Classify using threshold y0=(wTμ1+wTμ2)/2y_0 = (w^T \mu_1 + w^T \mu_2)/2y0
=(wTμ1+wTμ2)/2

Data points fall on a 1D line along www, making classification easier.

Conclusion

Fisher’s Linear Discriminant is a powerful technique to maximize class separation while


reducing dimensionality.

 It forms the basis for Linear Discriminant Analysis (LDA) and is widely used in
pattern recognition, machine learning, and computer vision.

I can also create a diagram showing projection of two classes onto Fisher’s line for visual
understanding.

Do you want me to do that?

You said:
Feed-forward network mappings.
ChatGPT said:

Here’s a detailed long answer on Feed-Forward Network Mappings, suitable for exams or
assignments:

Feed-Forward Network Mappings


Definition

A Feed-Forward Neural Network (FFNN) is a type of artificial neural network where


connections between neurons do not form cycles.

 Information flows in one direction: from input layer → hidden layer(s) → output layer.
 FFNNs are used to map input vectors to output vectors, performing tasks like
classification, regression, and function approximation.

Network Structure

1. Input Layer
o Receives the raw input features x=[x1,x2,...,xn]x = [x_1, x_2, ..., x_n]x=[x1,x2
,...,xn].
o No computation occurs here; neurons just pass inputs forward.
2. Hidden Layer(s)
o Performs weighted summation followed by an activation function (e.g.,
sigmoid, ReLU, tanh).
o Captures non-linear relationships between inputs and outputs.

hj=f(∑i=1nwijxi+bj)h_j = f\Big(\sum_{i=1}^{n} w_{ij} x_i + b_j\Big)hj=f(i=1∑nwijxi+bj)

Where:

 wijw_{ij}wij = weight from input iii to hidden neuron jjj


 bjb_jbj = bias of hidden neuron jjj
 f()f()f() = activation function

3. Output Layer
o Computes final network output based on hidden layer outputs:

yk=g(∑j=1mvjkhj+ck)y_k = g\Big(\sum_{j=1}^{m} v_{jk} h_j + c_k\Big)yk=g(j=1∑mvjkhj+ck)

Where:

 vjkv_{jk}vjk = weight from hidden neuron jjj to output neuron kkk


 ckc_kck = bias of output neuron kkk
 g()g()g() = output activation function (linear for regression, softmax for classification)

Feed-Forward Mapping

The network implements a mapping function:

y=F(x;θ)y = F(x; \theta)y=F(x;θ)

Where:

 xxx = input vector


 yyy = output vector
 θ\thetaθ = set of all network parameters (weights and biases)

Essentially, the network learns a non-linear function that maps inputs to outputs.

Characteristics of Feed-Forward Mapping

1. One-Way Flow
o Data moves forward; no loops or cycles.
2. Universal Approximation
o A feed-forward network with at least one hidden layer can approximate any
continuous function to arbitrary accuracy (Universal Approximation Theorem).
3. Layered Structure
o Multiple layers allow learning complex hierarchical features.
4. Non-Linearity
o Activation functions introduce non-linear transformations, enabling complex
mappings.

Advantages

1. Can model complex, non-linear relationships.


2. Simple structure → easier to train than recurrent networks.
3. Universal approximator for continuous functions.
4. Flexible → can handle classification, regression, and function approximation.
Disadvantages

1. Cannot handle sequential data (no memory of previous inputs).


2. May require large datasets for good performance.
3. Susceptible to overfitting if network is too large.
4. Training can be computationally intensive, especially with deep networks.

Applications

 Pattern Recognition: Handwriting and speech recognition.


 Regression: Predicting stock prices or sensor outputs.
 Function Approximation: Modeling complex relationships in engineering and physics.
 Classification: Image recognition, spam detection, medical diagnosis.

Example

Suppose we want to predict house prices based on features like area (x1x_1x1), bedrooms
(x2x_2x2), and location (x3x_3x3):

1. Input layer: neurons for x1,x2,x3x_1, x_2, x_3x1,x2,x3


2. Hidden layer: 5 neurons with ReLU activation
3. Output layer: 1 neuron (price) with linear activation

 Feed-forward mapping:

y=F(x1,x2,x3;θ)≈Predicted Pricey = F(x_1, x_2, x_3; \theta) \approx \text{Predicted Price}y=F(x1,x2,x3


;θ)≈Predicted Price

Weights and biases (θ\thetaθ) are learned using training data via backpropagation.

Conclusion

Feed-forward network mappings are fundamental building blocks in neural networks.

 They provide a flexible way to learn non-linear input-output relationships.


 While simple in concept, they are extremely powerful when combined with multiple
layers and sufficient training data.

You might also like