0% found this document useful (0 votes)
8 views5 pages

Text Data Cleaning in Python

The document provides a step-by-step guide on processing text data, including techniques such as lowercasing, punctuation removal, stop words removal, URL and HTML tag removal, stemming, lemmatization, tokenization, and text normalization. It also explains the 'bins' parameter used in creating histograms with Python's Matplotlib library, detailing how it splits data into intervals. Examples of Python code are included for each processing step and the histogram creation.

Uploaded by

srushti.more
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views5 pages

Text Data Cleaning in Python

The document provides a step-by-step guide on processing text data, including techniques such as lowercasing, punctuation removal, stop words removal, URL and HTML tag removal, stemming, lemmatization, tokenization, and text normalization. It also explains the 'bins' parameter used in creating histograms with Python's Matplotlib library, detailing how it splits data into intervals. Examples of Python code are included for each processing step and the histogram creation.

Uploaded by

srushti.more
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

🔹 Dummy Dataset Setup

python
CopyEdit
data = ["The quick brown fox jumps over the lazy dog! Visit
[Link]
"<html><body>This is <b>HTML</b> text.</body></html>",
"He was running and had been troubled by the geese."] #
Sample text data

We'll process each sentence in this list.

✅ 1. Lowercasing
python
CopyEdit
lowercased = [[Link]() for sentence in data]
print(lowercased)

🔎 Explanation: Converts all text to lowercase to ensure uniformity, e.g., "The" and "the" are
treated the same.

✅ 2. Removing Punctuation & Special Characters


python
CopyEdit
import re
no_punct = [[Link](r'[^\w\s]', '', sentence) for sentence in
lowercased]
print(no_punct)

🔎 Explanation: Removes punctuation marks and special characters using regex [^\w\s]
which keeps only words and spaces.

✅ 3. Stop Words Removal


python
CopyEdit
import nltk
from [Link] import stopwords
[Link]('stopwords')

stop_words = set([Link]('english'))
no_stopwords = [' '.join([word for word in [Link]() if word
not in stop_words]) for sentence in no_punct]
print(no_stopwords)

🔎 Explanation: Filters out common words like “the”, “is”, etc., which carry less meaningful
information.

✅ 4. Removing URLs
python
CopyEdit
no_urls = [[Link](r'http\S+|www\S+', '', sentence) for sentence in
data]
print(no_urls)

🔎 Explanation: Removes URLs using regex that matches any text starting with http or www.

✅ 5. Removing HTML Tags


python
CopyEdit
from bs4 import BeautifulSoup

no_html = [BeautifulSoup(sentence, "[Link]").get_text() for


sentence in data]
print(no_html)

🔎 Explanation: Uses BeautifulSoup to clean out HTML tags, leaving just the visible text.
✅ 6a. Stemming
python
CopyEdit
from [Link] import PorterStemmer
stemmer = PorterStemmer()

stemmed = [' '.join([[Link](word) for word in [Link]()])


for sentence in no_stopwords]
print(stemmed)

🔎 Explanation: Converts words to their root form (e.g., “connected” → “connect”) by chopping
suffixes.

✅ 6b. Lemmatization
python
CopyEdit
from [Link] import WordNetLemmatizer
[Link]('wordnet')
lemmatizer = WordNetLemmatizer()

lemmatized = [' '.join([[Link](word) for word in


[Link]()]) for sentence in no_stopwords]
print(lemmatized)

🔎 Explanation: Similar to stemming but smarter; uses vocabulary and grammar rules (e.g.,
“geese” → “goose”).

✅ 7. Tokenization
python
CopyEdit
from [Link] import word_tokenize
[Link]('punkt')

tokens = [word_tokenize(sentence) for sentence in data]


print(tokens)
🔎 Explanation: Splits sentences into individual words or tokens.

✅ 8. Text Normalization
(e.g., converting numbers to text, expanding contractions, etc. Here we just normalize
whitespace)

python
CopyEdit
normalized = [' '.join([Link]()) for sentence in data]
print(normalized)

🔎 Explanation: Cleans up extra spaces, tabs, and inconsistent formatting.

💡 Summary Output Example (Cleaned Steps Combined)


python
CopyEdit
final_cleaned = [' '.join([[Link](word) for word in
[Link](r'[^\w\s]', '', [Link]()).split()
if word not in stop_words]) for sentence in data]
print(final_cleaned)

Let me know if you'd like to apply this to a real dataset or wrap it into a reusable function!

You said:
what is bin parameter''

ChatGPT said:

The bins parameter (not "bin") is most commonly used in Python's data visualization and
analysis libraries like Matplotlib, Seaborn, or Pandas when creating histograms.

✅ Definition:
The bins parameter determines how the data range is split into intervals (called bins) for a
histogram.

✅ Example: Using Matplotlib


python
CopyEdit
import [Link] as plt

data = [10, 20, 21, 22, 22, 25, 30, 35, 40, 45, 50, 60]

[Link](data, bins=5)
[Link]("Histogram Example")
[Link]("Value")
[Link]("Frequency")
[Link]()

✅ Explanation:
●​ bins=5 → The range of data will be divided into 5 equal-width intervals.​

●​ Each bar in the histogram represents the number of data points falling into that bin.​

●​ You can also pass a list of bin edges:​

python
CopyEdit
[Link](data, bins=[0, 20, 40, 60, 80])

This gives more control over the bin ranges.

You might also like