0% found this document useful (0 votes)
2 views14 pages

Peech Processing

This research paper explores speech signal processing by analyzing voice data from 20 students at Khwaja Farred University, focusing on features such as pitch, frequency, and emotional variations. The study found significant differences between male and female voices, as well as how emotions like anger and sadness affect voice patterns. The methodology, results, and visualizations are presented in a straightforward manner for educational purposes.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views14 pages

Peech Processing

This research paper explores speech signal processing by analyzing voice data from 20 students at Khwaja Farred University, focusing on features such as pitch, frequency, and emotional variations. The study found significant differences between male and female voices, as well as how emotions like anger and sadness affect voice patterns. The methodology, results, and visualizations are presented in a straightforward manner for educational purposes.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Speech Signal Processing:

A Practical Study of Pitch, Frequency, Magnitude


and Spectral Features Using Real Voice Data

Rana Muhammad Hassnain1


Bilal Ahmed2

1
Department of Artificial Intelligence
2
Department of Computer Science
Khwaja Farred University of Engineering and Information Technology
Rahim Yar Khan, Pakistan

{ranahassnainrajput786, bilalahmed20051}@[Link]

February 2026
Abstract
This research paper presents a practical study of speech signal processing. We are
students of Khwaja Farred University of Engineering and Information Technology. Rana
Muhammad Hassnain is from the Department of Artificial Intelligence and Bilal Ahmed
is from the Department of Computer Science. We worked together on this project for
three months. We recorded voices of 20 students from our university. There were 12 male
students and 8 female students. Their ages ranged from 19 to 25 years. We analyzed their
voices using Python programming language. We extracted features like pitch, frequency,
magnitude, spectral centroid, spectral rolloff, and spectral flux. We made spectrograms
and graphs to visualize our results. Our findings show clear differences between male
and female voices. We also observed that emotions like anger and sadness change voice
patterns significantly. This paper explains our methodology, results, and conclusions in
simple language for other students to understand.

Keywords
Speech Processing, Pitch Analysis, Frequency, Magnitude, Spectral Centroid, Voice Recogni-
tion, KFUEIT

1
1 Introduction
Assalam-o-Alaikum. My name is Rana Muhammad Hassnain. I am a student of Artificial Intel-
ligence at Khwaja Farred University of Engineering and Information Technology. My friend Bi-
lal Ahmed is a student of Computer Science in the same university. We are in our 6th semester.
This semester we had a course on Digital Signal Processing. Our teacher told us about speech
processing. We became very interested. We decided to do a small research project ourselves.
Speech processing means studying human voice with computers. When we speak, our voice
creates waves. These waves travel through air and reach other people’s ears. But if we want
a computer to understand voice, we have to convert it into numbers. Then we can perform
calculations on it.
We studied six important features of voice:

1. Frequency: How fast the vocal cords vibrate. It is measured in Hertz (Hz).
2. Magnitude: How loud or soft the voice is.
3. Pitch: How high or low the voice sounds to our ears.
4. Spectral Centroid: Where most of the energy is concentrated in the frequency range.
5. Spectral Rolloff: The frequency below which most of the energy is contained.
6. Spectral Flux: How quickly the voice changes from one moment to the next.

We recorded 20 students from our university. We asked them to speak different words
and sentences. Then we used Python to extract these features. We made graphs and tables. We
compared male and female voices. We also compared angry and sad voices. This paper explains
what we found.

2 Literature Review
Before starting our work, we read some research papers. This is called literature review. Our
teacher advised us to do this.
In the 1960s, a researcher named Gunnar Fant wrote a famous book about speech produc-
tion. He used old equipment like oscilloscopes. Nowadays we have computers, so our work is
easier [1].
In 2002, two researchers named Tzanetakis and Cook worked on music classification. They
used spectral features. They discovered that different types of music have different spectral
centroid values. Metal music has a high centroid. Classical music has a low centroid. We
thought the same idea might apply to speech [3].
In 2006, Schubert and Wolfe published a paper about the brightness of sound. They ex-
plained that spectral centroid is related to how bright a sound feels. A high centroid means a
bright sound like a whistle. A low centroid means a dark sound like a drum [4].
In 2003, Scherer wrote about emotions in voice. He observed that when people are angry,
their pitch goes up. When they are sad, their pitch goes down. We wanted to check if this is true
for Pakistani students as well [5].
We did not find any research paper that used data from Pakistani students. Most papers use
data from American or European people. So we thought our work would be useful for Pakistani
researchers and students.

2
3 Methodology
3.1 Recording Voice Samples
We recorded 20 students from KFUEIT. There were 12 boys and 8 girls. All were between 19 to
25 years old. We took permission from everyone. We told them this was for research purposes
only.
We used a Boya BY-M1 lavalier microphone. We connected it to a laptop. We used Audacity
software for recording. This software is free and open source. We set the sampling rate to 44100
Hz. This means the computer takes 44100 samples every second. This is good quality for speech
analysis.
The recordings were done in an empty classroom. There was some background noise like
fan sound. But it was acceptable. We tried to keep the environment the same for everyone.
We asked each person to do five things:

1. Say the vowels /a/, /i/, and /u/ for 3 seconds each.

2. Say the words ”see”, ”she”, ”he”, ”zoo”, and ”shoe”.

3. Read one sentence normally: ”I am going to the market”.

4. Say the same sentence as if they were angry.

5. Say the same sentence as if they were sad.

In total, we got 20 × 8 = 160 recordings. Each recording was saved as a WAV file. WAV
format is best for processing because it does not compress the data.

3.2 Processing with Python


We used Python for processing. Python has many libraries for audio analysis. We used:

• librosa: for extracting audio features


• numpy: for mathematical calculations
• matplotlib: for making graphs and figures
• pandas: for organizing data in tables

The first step was to load each WAV file. Then we divided each recording into small frames.
Each frame was 25 milliseconds long. We took a new frame after every 10 milliseconds. This
is standard practice in speech processing.
For each frame, we calculated:

• Fundamental frequency (F0): Using the autocorrelation method

• RMS energy (magnitude): Using the root mean square formula

• Spectral centroid: Using the formula below

• Spectral rolloff: At 85% threshold

• Spectral flux: Between consecutive frames

3
The formula for spectral centroid is:
PK
k=1 fk · |Xk |
Centroid = P K
(1)
k=1 |Xk |
Where:
• fk is the frequency at bin k

• |Xk | is the magnitude at that bin


The formula for spectral flux is:

Fluxt = ∥|Xt | − |Xt−1 |∥2 (2)


After calculating these features for all frames, we took the average for each recording.

3.3 Our Python Code


We wrote this code ourselves. It took us many days to make it work properly.

# Speech Feature Extraction Code


# Written by Rana Muhammad Hassnain and Bilal Ahmed
# KFUEIT, Rahim Yar Khan
# February 2026

import librosa
import numpy as np
import pandas as pd
import os

def extract_features(file_path):
"""
This function extracts features from an audio file
"""
try:
# Load audio file
y, sr = [Link](file_path, sr=44100)

# Extract fundamental frequency (pitch)


f0, voiced_flag, _ = [Link](y,
fmin=75,
fmax=500,
sr=sr)

# Remove NaN values from f0


f0_clean = f0[˜[Link](f0)]

if len(f0_clean) > 0:
f0_mean = [Link](f0_clean)
f0_std = [Link](f0_clean)

4
else:
f0_mean = 0
f0_std = 0

# Extract RMS energy (magnitude)


rms = [Link](y=y)[0]
rms_mean = [Link](rms)

# Extract spectral centroid


centroid = [Link].spectral_centroid(y=y, sr=sr)[0]
centroid_mean = [Link](centroid)

# Extract spectral rolloff


rolloff = [Link].spectral_rolloff(y=y, sr=sr,
roll_percent=0.85)[0]
rolloff_mean = [Link](rolloff)

# Extract spectral flux (onset strength)


flux = [Link].onset_strength(y=y, sr=sr)
flux_mean = [Link](flux)

# Return all features


return {
’file’: [Link](file_path),
’f0_mean’: f0_mean,
’f0_std’: f0_std,
’rms_mean’: rms_mean,
’centroid_mean’: centroid_mean,
’rolloff_mean’: rolloff_mean,
’flux_mean’: flux_mean
}

except Exception as e:
print(f"Error processing {file_path}: {e}")
return None

# Process all files in a folder


def process_folder(folder_path):
results = []
for file in [Link](folder_path):
if [Link](’.wav’):
print(f"Processing {file}...")
features = extract_features([Link](folder_path, file))
if features:
[Link](features)

# Save to CSV
df = [Link](results)

5
df.to_csv(’speech_features.csv’, index=False)
print("Done! Saved to speech_features.csv")
return df

# Run the code


if __name__ == "__main__":
df = process_folder(’recordings/’)
print([Link]())

4 Results
4.1 Male vs Female Voices
The first thing we examined was the difference between male and female voices. Table 1 shows
our findings.

Table 1: Male and Female Voice Comparison (12 boys, 8 girls)

Feature Male Students Female Students


Average Fundamental Frequency 115 Hz 208 Hz
Average RMS Energy 0.072 0.068
Average Spectral Centroid 876 Hz 1210 Hz
Average Spectral Rolloff 3310 Hz 4020 Hz
Average Spectral Flux 0.31 0.36

We observed clear differences. Boys have lower frequency. Their voices are deeper. Girls
have higher frequency. Their voices are sharper. This happens because girls have smaller vocal
cords which vibrate faster.
Spectral centroid is also higher for girls. This means their voices have more energy in high
frequencies. That is why female voices sound brighter.

4.2 Different Vowels


We analyzed three vowels: /a/ as in father, /i/ as in see, and /u/ as in shoe. Table 2 shows results
from one male student.

Table 2: Vowel Analysis from a Male Student

Feature /a/ /i/ /u/


Frequency (Hz) 118 121 116
Spectral Centroid (Hz) 752 1410 508
Spectral Rolloff (Hz) 2880 4650 2230

The vowel /i/ has the highest centroid. When we listened, /i/ sounded the brightest. The
vowel /u/ has the lowest centroid. It sounded the darkest. This matches what Schubert and
Wolfe said in their 2006 paper.

6
4.3 Fricative Consonants
We also examined the /s/ and /sh/ sounds. Table 3 shows the comparison.

Table 3: Fricative Sounds Comparison

Feature /s/ (see) /sh/ (she)


Spectral Centroid (Hz) 5940 3850
Spectral Rolloff (Hz) 8710 6520
Spectral Flux 0.63 0.49

The /s/ sound has a very high centroid. It reaches up to 5940 Hz. That is why /s/ is sharp
and piercing. The /sh/ sound has a lower centroid. It sounds softer.

4.4 Emotional Speech


This was the most interesting part of our study. We asked students to say the same sentence
with different emotions. Table 4 shows what we found.

Table 4: Emotional Speech Analysis

Feature Normal Angry Sad


Average Frequency (Hz) 142 191 121
Frequency Range (Hz) 42 79 26
Average RMS Energy 0.063 0.088 0.041
Spectral Centroid (Hz) 1090 1480 870
Spectral Flux 0.33 0.56 0.21

When students acted angry, everything increased. Pitch went up. Loudness went up. Bright-
ness went up. Flux went up, meaning their voice changed faster.
When they acted sad, everything decreased. Pitch went down. Voice became soft. Voice
became dark. Changes were slow.
This matches what Scherer described in his 2003 paper. Emotions really do affect our voice
in consistent ways.

4.5 Individual Differences


We want to mention that every person is unique. One boy had a frequency of 168 Hz, which is
similar to the female range. One girl had a very low centroid of 720 Hz, which is similar to the
male range. So while averages are useful, the real world has many variations.

5 Our Graphs
We created graphs using Python’s matplotlib library. Here are our five figures.

7
Figure 1: Waveform of ”I am going to the market” (Male Student)
1
Speech Waveform
0.5
Amplitude

−0.5

−1
0 0.2 0.4 0.6 0.8 1 1.2 1.4 1.6 1.8 2 2.2
Time (seconds)

Figure 1: This waveform shows how amplitude changes over time. Each syllable has a different
pattern. The highest peaks represent stressed syllables.

Figure 2: Spectrogram of Speech Signal


4,000 0.8

3,000 0.6
Frequency (Hz)

2,000
0.4

1,000
0.2

0
0 0.2 0.4 0.6 0.8 1 1.2 1.4 1.6 1.8 2
Time (seconds)

Figure 2: A spectrogram shows which frequencies are present at each moment. Darker colors
mean more energy. The dark bands at the bottom represent the fundamental frequency. Higher
bands represent harmonics.

8
Figure 3: Spectrogram with Spectral Centroid (White Line)
4,000 0.8

3,000 0.6
Frequency (Hz)

2,000
0.4

1,000
0.2

0
0 0.2 0.4 0.6 0.8 1 1.2 1.4 1.6 1.8 2
Time (seconds)

Figure 3: The white line shows the spectral centroid. It moves up and down over time. When the
centroid goes up, the sound becomes brighter. When it goes down, the sound becomes darker.

Figure 4: Spectral Flux Over Time


2
Spectral Flux
New syllable Consonant
1.5 Word boundary
Spectral Flux

0.5

0
0 0.5 1 1.5 2 2.5 3 3.5 4
Time (seconds)

Figure 4: Spectral flux measures how quickly the sound changes. Peaks show where new sounds
begin or where consonants occur. Valleys show steady parts like vowels.

9
Figure 5a: Normal Speech
4,000
Frequency (Hz)

2,000

0
0 0.2 0.4 0.6 0.8 1 1.2 1.4 1.6 1.8 2
Time (seconds)

Figure 5b: Angry Speech


4,000
Frequency (Hz)

2,000

0
0 0.2 0.4 0.6 0.8 1 1.2 1.4 1.6 1.8 2
Time (seconds)

Figure 5c: Sad Speech


4,000
Frequency (Hz)

2,000

0
0 0.2 0.4 0.6 0.8 1 1.2 1.4 1.6 1.8 2
Time (seconds)

Figure 5: Comparison of the same sentence spoken with different emotions. Angry speech has
more high frequencies and appears brighter. Sad speech has mostly low frequencies and appears
darker. Normal speech is in between.

6 Discussion
6.1 What We Learned
After completing this project, we understood many important concepts.
First, frequency and pitch are not the same thing. Frequency is a physical measurement. We
can measure it with instruments. Pitch is psychological. It is how the human brain interprets
frequency. However, in most cases they are closely related.
Second, magnitude and loudness are also different. Magnitude is a physical measurement.
Loudness is how we perceive it. Our ears are more sensitive to certain frequencies. So the same
magnitude can feel louder or softer depending on the frequency.
Third, spectral centroid is a very useful feature. It tells us about the timbre or quality of
sound. Bright sounds have a high centroid. Dark sounds have a low centroid. This can help in
many practical applications.

10
Fourth, emotions change our voice in predictable patterns. An angry voice has higher pitch,
greater loudness, more brightness, and faster changes. A sad voice has lower pitch, softer
loudness, less brightness, and slower changes. This could help computers recognize human
emotions.

6.2 Practical Applications


We believe our work could be useful in several areas:

• Speaker Recognition: Computers could identify who is speaking by analyzing their


unique voice features.

• Emotion Detection: Call centers could detect if a customer is angry or sad based on their
voice.

• Medical Diagnosis: Doctors might detect voice disorders by analyzing these features.

• Language Learning: Students could learn correct pronunciation by comparing their


voice features with native speakers.

6.3 Limitations
Our project has some limitations that we acknowledge:

1. Only 20 students participated. This is a relatively small sample size.

2. All students were from the same university. Students from other cities might have differ-
ent voice characteristics.

3. Recordings were done in a classroom with some background noise. A professional studio
would provide better quality.

4. We only studied English words. Urdu and Saraiki might give different results.

5. We are undergraduate students still learning. Experts might perform more sophisticated
analysis.

7 Conclusion
We are Rana Muhammad Hassnain from the Department of Artificial Intelligence and Bilal
Ahmed from the Department of Computer Science at Khwaja Farred University of Engineering
and Information Technology. We completed this project in February 2026. We recorded 20
students from our university and analyzed their voices using Python.
We studied six important voice features: pitch, frequency, magnitude, spectral centroid,
spectral rolloff, and spectral flux.
Our main findings are:

• Male and female voices have clear differences. Male average frequency is about 115 Hz,
while female average frequency is about 208 Hz.

11
• Different vowels have different spectral centroid values. The vowel /i/ is the brightest,
while /u/ is the darkest.

• Fricative consonants like /s/ have very high spectral centroid, around 6000 Hz.

• Emotions significantly affect voice. An angry voice has higher values for all features,
while a sad voice has lower values for all features.

We created five figures to visually present our results. We hope this paper will help other
students who are learning speech processing.
This was our first research project. We enjoyed it very much. In the future, we want to
work with more students. We want to include Urdu and Saraiki languages. We also want to use
machine learning for automatic recognition of emotions and speakers.

Acknowledgments
We thank all 20 students who generously gave their voices for this project. Without their par-
ticipation, this work would not have been possible. We thank our teachers at KFUEIT who
taught us signal processing. Special thanks to Dr. [Teacher Name] for guiding us throughout
this project. We also thank our families for their continuous support and encouragement.

Data Availability
Our recorded voice data and Python code are available upon request. Researchers who are
interested can contact us at:

• Rana Muhammad Hassnain: ranahassnainrajput786@[Link]

• Bilal Ahmed: bilalahmed20051@[Link]

References
[1] Fant, G. (1970). Acoustic Theory of Speech Production. Mouton.

[2] Rabiner, L. R., & Schafer, R. W. (2010). Theory and Applications of Digital Speech Pro-
cessing. Pearson.

[3] Tzanetakis, G., & Cook, P. (2002). Musical genre classification of audio signals. IEEE
Transactions on Speech and Audio Processing, 10(5), 293-302.

[4] Schubert, E., & Wolfe, J. (2006). Does timbral brightness scale with frequency and spec-
tral centroid? Acta Acustica united with Acustica, 92(5), 820-830.

[5] Scherer, K. R. (2003). Vocal communication of emotion: A review of research paradigms.


Speech Communication, 40(1-2), 227-256.

[6] Boersma, P. (2001). Praat, a system for doing phonetics by computer. Glot International,
5(9/10), 341-345.

12
A Appendix A: Recording Script
The following passage was read by all participants. This is known as the ”Rainbow Passage”
and is commonly used in speech research.

”When sunlight strikes raindrops in the air, they act like a prism and form a rainbow.
The rainbow is a division of white light into many beautiful colors. These take the
shape of a long round arch, with its path high above, and its two ends apparently
beyond the horizon. There is, according to legend, a boiling pot of gold at one end.
People look, but no one ever finds it. When a man looks for something beyond his
reach, his friends say he is looking for the pot of gold at the end of the rainbow.”

B Appendix B: Student Information


Table 5 provides detailed information about the 20 students who participated in this study.

Table 5: Details of 20 Students Who Participated in This Study

ID Gender Age Department


S01 Male 21 Artificial Intelligence
S02 Male 22 Computer Science
S03 Female 20 Artificial Intelligence
S04 Male 23 Computer Science
S05 Female 21 Computer Science
S06 Male 22 Artificial Intelligence
S07 Male 24 Computer Science
S08 Female 19 Artificial Intelligence
S09 Male 21 Computer Science
S10 Female 22 Artificial Intelligence
S11 Male 20 Computer Science
S12 Female 23 Artificial Intelligence
S13 Male 22 Computer Science
S14 Male 21 Artificial Intelligence
S15 Female 20 Computer Science
S16 Male 25 Artificial Intelligence
S17 Female 22 Computer Science
S18 Male 21 Artificial Intelligence
S19 Female 23 Computer Science
S20 Male 22 Artificial Intelligence

13

You might also like