0% found this document useful (0 votes)
2 views39 pages

Project File (1)

This Industrial Training Report details an internship in Data Analytics conducted by Dilfaraz Ali at FUTURE INTERNS, focusing on practical exposure to Big Data, Machine Learning, and data visualization tools. The internship included tasks such as developing a Spotify Dashboard using Power BI and performing sentiment analysis on tweets, enhancing both technical and professional skills. Key learnings encompassed data handling, programming in Python, and the application of Natural Language Processing techniques, contributing to career readiness in the IT sector.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views39 pages

Project File (1)

This Industrial Training Report details an internship in Data Analytics conducted by Dilfaraz Ali at FUTURE INTERNS, focusing on practical exposure to Big Data, Machine Learning, and data visualization tools. The internship included tasks such as developing a Spotify Dashboard using Power BI and performing sentiment analysis on tweets, enhancing both technical and professional skills. Key learnings encompassed data handling, programming in Python, and the application of Natural Language Processing techniques, contributing to career readiness in the IT sector.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

INDUSTRIAL TRAINING REPORT

on

Data Analytics

Submitted as a part of course curriculum for

Bachelor of Technology
in

Computer Science and Engineering

Submitted by

Dilfaraz Ali

University Roll Number

2203400100020

DEPARTMENT OF COMPUTER SCIENCE AND ENGINEERING

VIVEKANANDA COLLEGE OF TECHNOLOGY & MANAGEMENT, ALIGARH

DR. A.P.J. ABDUL KALAM TECHNICAL UNIVERSITY ,


LUCKNOW

2025-26
i
CERTIFICATE OF COMPLETION

i
TABLE OF CONTENT

Chapter 1: Introduction…………………………………………………………………....vi
1.1 About the Organization
1.2 About the Internship
1.3 Objectives of the Internship
1.4 Overview of Tasks Assigned (Big Data, ML, Dashboard, Sentiment)
1.5 Skills Expected to be Learned
Chapter 2 – Analysis / Tools used……………………………………………….……...…ix
2.1 Big Data Tools: PySpark, Dask
2.2 Machine Learning Libraries: Python, scikit-learn, pandas, NumPy
2.3 Visualization Tools: Tableau, Power BI, Dash
2.4 NLP Techniques: Sentiment Analysis, Preprocessing (Tokenization, Stopword Removal)
2.5 GitHub/Google Colab/Jupyter Notebook for Implementation
2.6 Advantages of Using These Tools
Chapter 3 – Methodology & Implementation …………………………………………..xii
3.1 Task 1 – Dashboard Development
3.2 Task 2 – Sentiment Analysis
3.3 Challenges Faced & Solutions
Chapter 4 – Results & Conclusion ………………………………………………….......xvi
4.1 Outcomes of Each Task
4.2 Skills Acquired (Big Data, ML, Visualization, NLP)
4.3 Contribution of Internship to Career Goals
4.4 Final Summary & Conclusion

ii
ii
CHAPTER 1
INTRODUCTION

1.1 ABOUT THE ORGANIZATION

FUTURE INTERNS is a certified IT services and training organization that has established
itself as a reputed provider of professional internships and industrial training programs. The
organization is recognized by the All India Council for Technical Education (AICTE),
registered under the Ministry of Micro, Small and Medium Enterprises (MSME), and
holds an ISO 9001:2015 certification, which ensures adherence to international standards of
quality management and service delivery.

The company primarily focuses on skill development, industry readiness, and


employability enhancement by offering training and internships in emerging
technologies. Its key areas of expertise include:

 Data Analytics & Data Science


 Machine Learning & Artificial Intelligence
 Web Development & Full Stack Development
 Cloud Computing
 Cybersecurity & Ethical Hacking
 Big Data Processing
 Business Intelligence & Visualization Tools (Power BI, Tableau)

These programs are designed to meet the needs of students, fresh graduates, and working
professionals who aim to gain practical exposure alongside their theoretical knowledge.

The vision of FUTURE INTERNS is to bridge the gap between academia and industry by
providing opportunities to work on real-world projects and live case studies. The
organization believes in a learn-by-doing approach, where learners not only understand the
concepts but also implement them in practice through tasks such as data analysis, predictive
modeling, dashboard creation, application development, and research-based problem solving.

The mission of FUTURE INTERNS is to empower learners by:

 Offering industry-oriented training modules tailored to current market demands.


1
 Providing hands-on projects that replicate real business challenges.
 Ensuring flexibility in learning through virtual/online internship programs.
 Fostering both technical expertise and professional skills such as communication,
teamwork, and problem-solving.

Over the years, FUTURE INTERNS has become a trusted platform for skill-building
where students enhance their employability by gaining relevant industry experience. The
organization emphasizes innovation, quality, and excellence, making it a valuable
contributor to shaping the careers of aspiring professionals in the IT sector.

1.2 ABOUT THE INTERNSHIP

The internship program was carried out in the domain of Data Analytics, which is one of the
most in-demand fields in today’s digital era. The internship was conducted in virtual (online)
mode by FUTURE INTERNS, providing flexibility to learn and work on real-time projects
from a remote environment. The total duration of the program was 4 weeks, starting from 07
July 2025 to 06 August 2025, under the official Offer Letter ID:FIT/JUL25/DS3901.

The internship was carefully structured to provide participants with both conceptual
understanding and practical exposure to industry-relevant tools and technologies. The
focus of the program was not only to improve technical skills but also to enhance critical
thinking, data-driven decision-making, and problem-solving abilities.

During the course of this internship, I was introduced to the following core areas:

1. Machine Learning
o Gained exposure to supervised and unsupervised learning algorithms.
o Learned the process of data preprocessing, cleaning, and feature selection.
o Understood the basics of model training, testing, and evaluation using Python
libraries such as scikit-learn, pandas, and Numpy.
o Implemented predictive models on sample datasets to analyze patterns and
mae forecasts.
2. Dashboard Development
o Acquired skills in data visualization and reporting using Power BI.

2
o Designed and developed an interactive Spotify Dashboard to represent data
in an engaging and visually appealing manner.
o Used multiple charts, slicers, KPIs, and filters to provide actionable insights
for better decision-making.
o Understood the importance of visual storytelling in presenting analytical
results.
3. Sentiment Analysis
o Worked on extracting and analyzing text data from tweets.
o Applied Natural Language Processing (NLP) techniques such as
tokenization, stopword removal, and lemmatization.
o Classified tweets into categories such as Positive or Negative using machine
learning models.
o Learned how sentiment analysis is applied in real-world scenarios such as
brand monitoring, customer feedback analysis, and social media trend
analysis.

The internship gave me an opportunity to work on real-world projects, which helped me


bridge the gap between classroom knowledge and industry practices. By the end of the
program, I not only gained technical expertise but also developed essential skills like time
management, self-learning, communication, and professional discipline.

This internship was a valuable experience that strengthened my understanding of Data


Analytics, Visualization, and Machine Learning techniques, while also giving me
confidence to work on future projects independently and efficiently.

1.3 OBJECTIVES OF THE INTERNSHIP

The primary objective of this internship was to provide practical exposure to the field of Data
Analytics by engaging in real-time projects and tasks. It aimed to bridge the gap between
academic learning and industry applications, while simultaneously enhancing both technical
and analytical skills. The specific objectives of the internship were as follows:

1. To gain real-world exposure in Data Analytics tools and techniques


o Understand how data analytics is applied across industries to derive actionable
insights.

3
o Work with structured and unstructured datasets to identify trends, patterns, and
business value.
o Learn the complete data analytics workflow, starting from data collection,
preprocessing, visualization, and reporting.
2. To develop hands-on experience in Python, PySpark, Dask, NLP, and Power
BI/Tableau
o Strengthen programming skills in Python, focusing on libraries like pandas,
NumPy, scikit-learn, and Matplotlib.
o Get introduced to Big Data frameworks such as PySpark and Dask for
handling large-scale datasets efficiently.
o Gain practical knowledge of Natural Language Processing (NLP) for
performing sentiment analysis and text classification.
o Learn dashboard creation and visualization using Power BI and Tableau to
present data in an interactive and user-friendly manner.
3. To understand the implementation of Big Data Analysis and Predictive Modeling
o Learn the fundamentals of Big Data processing and how it differs from
traditional data handling methods.
o Apply predictive modeling techniques to build machine learning models that
can forecast outcomes based on historical data.
o Evaluate model performance using metrics such as accuracy, precision, recall,
and F1-score.
4. To learn data preprocessing, visualization, and report generation
o Perform data cleaning, handling missing values, and feature engineering to
prepare datasets for analysis.
o Develop skills in data visualization to represent information clearly and
effectively through charts, graphs, and dashboards.
o Generate detailed analytical reports and presentations that summarize
findings and provide insights for decision-making.

By fulfilling these objectives, the internship helped in developing a comprehensive skill set
that combines technical knowledge, analytical thinking, and practical application. It also
provided a strong foundation for pursuing advanced projects in Data Analytics, Machine
Learning, and Artificial Intelligence.

4
1.4 OVERVIEW OF TASKS ASSIGNED

During the course of this internship, I was assigned two major tasks that provided me with
both technical exposure and practical implementation skills. These tasks helped me apply the
concepts learned during training to real-world scenarios and enhanced my understanding of
data analytics and machine learning.

Task 1: Dashboard Development (Spotify Dataset)

 The first task focused on data visualization and dashboard creation using Power
BI.
 The objective was to design an interactive Spotify Dashboard that provides
meaningful insights into music data, such as popular tracks, top artists, streaming
trends, and genre distributions.
 Key responsibilities included:
o Importing and cleaning the Spotify dataset.
o Designing charts, KPIs, and slicers for better interactivity.
o Using filters and drill-down features to allow deeper analysis of streaming
trends.
o Creating visually appealing dashboards that make complex data easy to
understand.
 Outcome: The task strengthened my ability to use Power BI for professional data
visualization and storytelling, while also developing insights into how dashboards
support business decision-making.

Task 2: Sentiment Analysis (Tweets Dataset)

 The second task was centered on Natural Language Processing (NLP) and Machine
Learning.
 The objective was to perform Sentiment Analysis on tweets to classify them into
Positive or Negative categories.
 Key responsibilities included:
o Collecting and preprocessing raw text data (tokenization, stopword removal,
lemmatization).
o Applying machine learning algorithms to classify the tweets.

5
o Visualizing the sentiment distribution for better understanding of user
opinions.
o Evaluating model accuracy and improving performance with suitable
techniques.
 Outcome: This task provided hands-on experience in text analytics and machine
learning, enhancing my ability to work with unstructured data and apply NLP in real-
world contexts

Together, these tasks offered a complete learning experience by covering both data
visualization (Dashboard Development) and data analysis with machine learning
(Sentiment Analysis). They helped me gain technical expertise, problem-solving skills, and
confidence in applying theoretical knowledge to practical applications.

1.5 KEY LEARNINGS FROM INTERNSHIP

The internship at ELiteTech Intern provided me with valuable technical knowledge, hands-
on exposure, and professional skills. By working on real-world projects such as Spotify
Dashboard Development and Sentiment Analysis on Tweets, I was able to bridge the gap
between academic concepts and industry requirements. The key learning from this internship
are as follows:

1. Technical Learning

 Data Analytics Workflow – Understood the complete process of data handling,


including data collection, preprocessing, analysis, visualization, and reporting.
 Programming in Python – Enhanced coding skills and learned to use essential
libraries such as pandas, NumPy, Matplotlib, scikit-learn, and NLTK for data
manipulation and machine learning.
 Natural Language Processing (NLP) – Gained experience in text cleaning,
tokenization, stopword removal, and sentiment classification of tweets.
 Machine Learning Concepts – Understood the basics of supervised learning models,
their training, evaluation, and accuracy improvement techniques.
 Power BI for Dashboard Development – Acquired hands-on skills in designing
interactive dashboards, applying slicers/filters, creating KPIs, and using data
visualization for decision-making.

6
 Big Data Handling (Introduction to PySpark/Dask) – Learned how large datasets
can be processed efficiently using distributed frameworks.

2. Professional & Soft Skills Learnings

 Problem-Solving Skills – Developed the ability to approach real-world problems


logically and find effective solutions using data.
 Time Management – Successfully managed tasks within deadlines while balancing
multiple responsibilities.
 Analytical Thinking – Strengthened the skill of interpreting raw data and converting
it into meaningful insights.
 Communication Skills – Improved the ability to present findings clearly using
reports, dashboards, and visualizations.
 Self-Learning & Adaptability – Adapted to new tools and technologies quickly in a
virtual learning environment.
 Teamwork & Professionalism – Understood the importance of structured work,
discipline, and collaborative approaches in industry-oriented tasks.

Overall, this internship helped me become more industry-ready by equipping me with a


blend of technical expertise, analytical mindset, and professional skills required in the
field of Data Analytics and Machine Learning.

7
CHAPTER 2
ANALYSIS / TOOLS USED

During the internship at ELiteTech Intern, several tools, programming languages, and
analytical techniques were used to complete the assigned tasks. Each tool played a critical
role in handling data, performing analysis, and delivering meaningful insights through
dashboards and predictive models. The following section provides a comprehensive
overview of the tools used during the internship.

2.1 Big Data Tools : PySpark and Dask

 Overview:
o PySpark is the Python API for Apache Spark, a framework for large-scale
data processing.
o Dask is a Python library for parallel computing that helps scale up
computations on large datasets.
 Usage in Internship:
o Gained an introductory understanding of how Big Data is handled in real-
world scenarios.
o Learned about distributed computing and why tools like PySpark and Dask are
essential when working with datasets beyond the capacity of traditional
systems.
 Outcome: Developed awareness of Big Data technologies and their industry
applications.

2.2 Machine Learning Libraries : Python

2.2.1 Python Programming Language

 Overview:
Python is a high-level, general-purpose programming language widely adopted in
data science, machine learning, and analytics due to its simplicity and vast
ecosystem of libraries. It supports procedural, object-oriented, and functional
programming styles, making it versatile for both beginners and professionals.

8
 Usage in Internship:
o Performed data preprocessing (cleaning, missing value handling, feature
scaling, encoding).
o Developed machine learning models for classification tasks.
o Implemented sentiment analysis using text datasets (tweets).
o Generated graphs and charts for intermediate analysis.
 Key Libraries Used:
o pandas & NumPy – For handling large datasets, data manipulation, and
performing numerical computations.
o Matplotlib & Seaborn – For visualizing data distributions, trends, and
patterns.
o scikit-learn – For implementing machine learning algorithms (classification,
training, and evaluation).
o NLTK (Natural Language Toolkit) – For natural language preprocessing
tasks such as tokenization, stopword removal, stemming, and lemmatization.
 Impact: Python served as the foundation for all machine learning and text analytics-
related tasks.

2.3 Visualization Tools : Tableau, Power BI

2.3.1 Power BI

 Overview:
Power BI is a business analytics tool developed by Microsoft that allows users to
visualize data, share insights, and create interactive dashboards. It is widely used in
organizations for data-driven decision-making.
 Usage in Internship:
o Imported Spotify dataset and cleaned it for analysis.
o Designed multiple visualizations including bar charts, pie charts, KPIs,
slicers, and filters.
o Developed a Spotify Dashboard to provide insights into:
 Top songs and artists.
 Year-wise streaming trends.
 Genre distribution and popularity.

9
o Used drill-down and interactive features to explore data at multiple levels.
 Outcome: Helped me gain professional skills in data visualization, dashboard
design, and storytelling with data.

2.3.2 Tableau (Exploratory Use)

 Overview:
Tableau is another leading data visualization and business intelligence tool, known for
its drag-and-drop interface and advanced data exploration capabilities.
 Usage in Internship:
o Explored how Tableau can be used as an alternative to Power BI.
o Compared its visualization features with Power BI.
o Learned its role in creating interactive dashboards for business intelligence.
 Outcome: Enhanced understanding of BI tools and their role in real-time analytics.

2.4 Natural Language Processing (NLP)

 Overview:
Natural Language Processing (NLP) is a field of Artificial Intelligence that enables
computers to understand, process, and analyze human language. It bridges the gap
between machine learning algorithms and text-based data.
 Usage in Internship:
o Collected and preprocessed text data (tweets).
o Applied tokenization, stopword removal, stemming, and lemmatization to
prepare the dataset.
o Used vectorization techniques (Bag of Words, TF-IDF) to convert textual
data into numerical form.
o Built a sentiment classifier to predict whether a tweet is Positive or Negative.
 Outcome: NLP enabled me to understand how unstructured data can be
transformed into structured information for analysis and decision-making.

10
2.5 Jupyter Notebook/ GitHub

2.5.1 Jupyter Notebook

 Overview:
Jupyter Notebook is an open-source tool that allows users to write, execute, and
document Python code interactively. It is widely used for data science, research, and
machine learning projects.
 Usage in Internship:
o Implemented Python code step by step for sentiment analysis.
o Documented the workflow with code snippets, outputs, and explanations.
o Created well-structured notebooks that combined text (Markdown) and code
for easy understanding.
 Outcome: Helped maintain a clear, reproducible record of the work done.

2.5.2 GitHub

 Overview:
GitHub is a version control platform that allows developers to manage, share, and
collaborate on projects.
 Usage in Internship:
o Stored project files and codes.
o Understood the importance of version control for collaborative projects.
o Learned basic Git commands for pushing/pulling repositories.
 Outcome: Gained experience in collaborative coding and project management.

Microsoft Excel

 Overview:
Microsoft Excel is a spreadsheet software used for data entry, statistical analysis, and
simple data visualization.
 Usage in Internship:
o Conducted initial exploratory data analysis before importing data into Power
BI/Python.

11
o Performed basic operations such as filtering, pivot tables, and summary
statistics.
o Validated dataset quality and ensured it was analysis-ready.
 Outcome: Excel helped in the initial stage of data exploration and prepared datasets
for further analysis.

Summary of Tools Used

The internship leveraged a combination of programming languages (Python), data


visualization tools (Power BI, Tableau, Excel), machine learning libraries (scikit-learn,
NLTK), big data frameworks (PySpark, Dask), and collaboration platforms (GitHub,
Jupyter Notebook).

Each tool contributed to achieving the project objectives by enabling:

 Data preprocessing & analysis (Python, Excel, pandas, NumPy)


 Data visualization & storytelling (Power BI, Tableau, Matplotlib, Seaborn)
 Text processing & sentiment analysis (NLP, NLTK, scikit-learn)
 Handling large-scale datasets (PySpark, Dask)
 Code execution & documentation (Jupyter Notebook, GitHub)

Together, these tools provided a comprehensive learning experience and enhanced both
technical proficiency and analytical skills in the domain of Data Analytics and Machine
Learning.

2.6 ADVANTAGES OF USING THESE TOOLS

The tools and technologies used during the internship provided multiple benefits in terms of
efficiency, accuracy, scalability, and visualization. Each tool had its unique strengths, and
together they created a comprehensive environment for completing the assigned tasks. The
major advantages are:

1. Python and Its Libraries

 Ease of Use & Readability: Python’s simple syntax makes it beginner-friendly yet
powerful for complex tasks.

12
 Extensive Libraries: Libraries such as pandas, NumPy, scikit-learn, and NLTK
offer ready-made functions for data handling, machine learning, and natural language
processing.
 Flexibility: Can be used for data preprocessing, analysis, visualization, and predictive
modeling in one environment.
 Community Support: A large developer community ensures continuous updates and
availability of resources.

2. Natural Language Processing (NLP)

 Text Handling: Converts unstructured text (tweets) into structured, machine-readable


format.
 Automation: Enables automatic sentiment classification without manual
interpretation.
 Applications: Widely applicable in social media monitoring, customer feedback
analysis, and brand reputation management.

3. Power BI

 Interactive Dashboards: Provides user-friendly visualizations with charts, KPIs, and


slicers.
 Real-Time Insights: Data can be connected to live sources for updated dashboards.
 Ease of Sharing: Reports and dashboards can be easily shared with stakeholders.
 Integration: Works seamlessly with Excel, SQL, and cloud services.

4. Tableau

 Drag-and-Drop Functionality: Makes data visualization quick and intuitive.


 Advanced Visualizations: Provides detailed and aesthetic dashboards for business
intelligence.
 Cross-Platform Use: Can handle multiple data sources and integrate them efficiently.

5. PySpark & Dask (Big Data Tools)

 Scalability: Designed to process very large datasets across distributed systems.

13
 Speed: Performs computations faster than traditional tools when handling massive
data.
 Industry Relevance: Prepares learners for real-world big data scenarios used in top
companies.

6. Jupyter Notebook

 Interactive Coding: Allows step-by-step execution of code with immediate results.


 Documentation: Combines code, visualizations, and explanations in one place.
 Reproducibility: Ensures workflows can be revisited and reused easily.

7. GitHub

 Version Control: Tracks project changes and maintains history.


 Collaboration: Facilitates teamwork and code sharing.
 Open Source Contribution: Allows access to thousands of projects and resources.

8. Microsoft Excel

 Simplicity: Easy to use for basic data cleaning and exploration.


 Quick Analysis: Useful for descriptive statistics, pivot tables, and filtering.
 Integration: Works smoothly with other advanced tools like Power BI and Python.

14
CHAPTER 3
METHODOLOGY & IMPLEMENTATION

The internship involved two core tasks:

1. Dashboard Development using Power BI (Spotify Dataset)


2. Sentiment Analysis using Python and NLP (Tweets Dataset)

Both tasks required a systematic methodology for planning and execution, followed by
practical implementation using appropriate tools and technologies. This chapter describes
the step-by-step methodology and the implementation details for each task.

3.1 Task 1: Dashboard Development (Spotify Dataset)

The goal of this task was to build an interactive Spotify Dashboard using Power BI to
analyze music streaming patterns and provide actionable insights.

Step 1: Data Collection

 The dataset used was sourced from a Spotify dataset containing details of songs,
artists, albums, release years, genres, and streaming counts.
 The dataset was in CSV format and contained both numerical (streams, duration,
year) and categorical (artist, genre, song name) data.

Step 2: Data Preprocessing

 Data was loaded into Power BI and Excel for initial validation.
 Preprocessing included:
o Removing duplicates (multiple entries for the same song).
o Handling missing values (null or blank records).
o Data type correction (converting text to string, dates to proper format,
streams to integers).
o Standardization of artist names and genres for consistency.

15
Step 3: Data Modeling

 Established relationships between attributes (e.g., Song → Artist → Genre →


Streams).
 Created calculated columns and measures in Power BI using DAX (Data Analysis
Expressions):
o Total Streams
o Average Streams per Artist
o Year-wise Growth Rate
o Top 10 Most Streamed Songs

Step 4: Dashboard Design in Power BI

 The dashboard was designed to be user-friendly, interactive, and visually


appealing.
 Visualizations used:
o Bar Charts → Displayed Top Artists and Top Songs.
o Pie Charts → Showed Genre Distribution.
o Line Graphs → Represented Year-wise Streaming Trends.
o KPIs (Key Performance Indicators) → Highlighted Total Streams, Average
Streams per Song, and Top Artist.
o Slicers & Filters → Allowed users to filter data based on Artist, Year, or
Genre.

Step 5: Insights & Results

 Identified the most popular artists based on total streams.


 Found which songs dominated streaming charts over specific years.
 Observed the rise of certain genres in user preferences.
 Created a comprehensive Spotify Dashboard that can be used by stakeholders to
understand music consumption patterns.

16
3.1.2 Implementation of Dashboard Development (Spotify Dataset)

The implementation process for the Spotify Dashboard was carried out using Microsoft
Power BI Desktop. The steps involved importing the dataset, preprocessing data, creating a
data model, applying DAX formulas, designing the dashboard, and finally analyzing insights.

Step 1: Importing the Dataset

 The Spotify dataset (CSV format) was imported into Power BI Desktop using the
Get Data → Text/CSV option.
 On import, Power BI automatically detected the column headers such as:
o Song Name, Artist, Album, Release Year, Genre, Streams
 The dataset was then loaded into the Power Query Editor for further cleaning and
transformation.

Step 2: Data Cleaning & Transformation (Power Query Editor)

In the Power Query Editor, data preprocessing was performed to ensure quality and
consistency:

1. Duplicate Removal:
o Identified and removed duplicate entries for the same song, album, or artist.
2. Handling Missing Values:
o Replaced or dropped null values in critical fields such as Song Name, Artist,
Streams.
3. Standardization:
o Ensured consistent naming conventions, e.g., “Hip-Hop” and “Hip Hop” were
standardized into a single category.
o Corrected inconsistencies in artist names (e.g., “The Weeknd” vs “Weeknd”).
4. Data Type Conversion:
o Streams → Whole Number (Integer).
o Release Year → Date/Number.
o Song Name, Artist, Album, Genre → Text.
5. Column Enhancements:

17
o Created a new calculated column for Decade (e.g., 1990s, 2000s, 2010s) using
conditional logic for better trend analysis.

After transformations, the cleaned dataset was loaded into the Power BI Data Model.

Step 3: Data Modeling (Relationships & DAX Calculations)

Once the dataset was ready, relationships and calculations were created:

1. Relationships:
o Established logical links:
 Artist → Song → Streams
 Album → Release Year → Genre
2. DAX Measures and Calculations:
Several calculated fields were added to enhance the analysis:
o Total Streams
o Total Streams = SUM(Spotify[Streams])
o Average Streams per Artist
o Avg Streams = AVERAGEX(VALUES(Spotify[Artist]), [Total Streams])
o Top Artist Rank
o Artist Rank = RANKX(ALL(Spotify[Artist]), [Total Streams], , DESC, Skip)
o Streams by Year
o Streams by Year = CALCULATE(SUM(Spotify[Streams]), ALLEXCEPT(Spotify,
Spotify[Release Year]))

These measures allowed dynamic calculations and comparisons inside the dashboard.

Step 4: Dashboard Design (Visualizations)

The dashboard was designed to be interactive, intuitive, and visually appealing. The
following components were included:

1. KPI Cards (Key Metrics):


o Total Streams (overall count).
o Top Artist (most streamed).
o Most Streamed Song.
o Average Streams per Artist.

18
2. Bar & Column Charts:
o Top 10 Artists by Streams.
o Top 10 Songs by Streams.
o Album-wise Performance.
3. Line Chart:
o Year-wise streaming trends showing growth across time.
4. Pie/Donut Chart:
o Genre distribution to analyze popularity of music types.
5. Table / Matrix View:
o Tabular display of artists, albums, songs, and their total streams.
6. Interactive Filters (Slicers):
o By Artist, By Year, By Genre, By Album.
o Enabled users to dynamically filter the dashboard as per their interest.
7. Themes & Formatting:
o Applied a dark-themed template for contrast.
o Used consistent color coding for genres.
o Applied tooltips to display extra details (e.g., album info when hovering over
songs).

Step 5: Interactive Features

 The dashboard was made interactive by enabling:


o Cross-filtering: Clicking on one chart (e.g., artist) dynamically updated other
visuals.
o Drill-down: Users could drill down from Genre → Artist → Album → Song.
o Bookmarks: Saved views for quick navigation (e.g., "Top Songs View",
"Yearly Trends View").

Step 6: Output & Final Insights

After implementation, the dashboard successfully generated several insights:

 Identified Top 10 Most Streamed Artists and Songs.


 Showed how streaming behavior evolved across years (e.g., sudden spikes after
2015 due to growth of streaming platforms).

19
 Highlighted genre preferences, with Pop and Hip-Hop emerging as the most
dominant.
 Provided valuable decision-making inputs for stakeholders such as music
companies, managers, and artists to analyze trends and plan marketing strategies.

3.2 Methodology for Sentiment Analysis (Tweets Dataset)

The second project of the internship involved developing a Sentiment Analysis Model using
Python and Natural Language Processing (NLP) techniques. The aim was to classify
tweets into Positive or Negative sentiments. Social media platforms like Twitter provide vast
amounts of unstructured data, making sentiment analysis a valuable tool for understanding
public opinion, customer feedback, and trending topics.

3.2.1 Methodology

The methodology for implementing sentiment analysis can be divided into multiple steps:

1. Data Collection

 A Tweets dataset was used which contained:


o The actual text of the tweet.
o A sentiment label (positive or negative).

20
 The dataset was in CSV format and imported into Jupyter Notebook using pandas.
 This dataset served as the foundation for training and testing the sentiment classifier.

2. Data Preprocessing
Since raw tweets often contain unwanted text, symbols, and noise, preprocessing was
essential. The following steps were performed:

 Lowercasing: All text was converted to lowercase to ensure uniformity (e.g.,


“Great” and “great” treated the same).
 Noise Removal:
o Removed URLs (e.g., [Link]
o Removed mentions (@usernames) and hashtags (#topic).
o Removed numbers and special characters.
 Stopword Removal: Eliminated common words (e.g., “is”, “the”, “and”) which do
not contribute to sentiment.
 Tokenization: Split sentences into individual words (tokens).
 Stemming and Lemmatization:
o Stemming: Reduced words to their root form (e.g., “playing” → “play”).
o Lemmatization: Converted words to their dictionary form (e.g., “better” →
“good”).
 Final Cleaned Tweets: Contained meaningful words suitable for machine learning
models.

3. Feature Extraction
Since machine learning algorithms require numerical input, the textual data was converted
into numerical features using two techniques:

 Bag of Words (BoW): Represented each tweet as a vector of word counts.


 TF-IDF (Term Frequency-Inverse Document Frequency): Assigned higher
weights to words that are important in a tweet but rare across the dataset.

This transformation created a feature matrix suitable for classification algorithms.

4. Model Building

 The dataset was divided into:

21
o Training Set (80%) – used to train the model.
o Testing Set (20%) – used to evaluate performance.
 Machine learning algorithms applied:
o Logistic Regression – simple yet effective for binary classification.
o Naïve Bayes (MultinomialNB) – fast and efficient for text classification.
o Support Vector Machine (SVM) – powerful for high-dimensional data like
text.

5. Model Evaluation
To assess model performance, multiple evaluation metrics were used:

 Accuracy: Percentage of correct predictions.


 Precision: Ratio of correctly predicted positives to all predicted positives.
 Recall (Sensitivity): Ratio of correctly predicted positives to all actual positives.
 F1-Score: Harmonic mean of precision and recall (useful for imbalanced datasets).
 Confusion Matrix: Displayed the distribution of correct and incorrect classifications.

These metrics helped compare models and select the most effective one.

[Link]
To better understand results and data patterns:

 Sentiment Distribution Graphs: Bar charts showed the number of positive vs


negative tweets.
 Word Clouds:
o Generated for positive tweets → highlighting frequent positive words like
“happy”, “love”, “great”.
o Generated for negative tweets → highlighting frequent negative words like
“bad”, “sad”, “angry”.
 Confusion Matrix Heatmap: Visualized true vs predicted labels for model
evaluation.

Summary of Methodology
The methodology ensured a complete pipeline for text classification:

1. Data collection of tweets.

22
2. Preprocessing and cleaning noisy text.
3. Converting text into numerical features using BoW and TF-IDF.
4. Building multiple machine learning models.
5. Evaluating models using accuracy and other metrics.
6. Visualizing sentiment patterns and model performance.

This systematic approach made the sentiment analysis task efficient, reliable, and industry-
oriented.

3.2.2 Implementation of Sentiment Analysis (Tweets Dataset)

The implementation of the sentiment analysis project was carried out using Python in
Jupyter Notebook. The process involved importing the dataset, cleaning the text data,
converting it into numerical features, training machine learning models, evaluating their
performance, and visualizing the results.

Step 1: Importing Libraries and Dataset

The following Python libraries were used:

 pandas → for data manipulation and handling CSV files.


 NumPy → for numerical operations.
 NLTK → for natural language processing (tokenization, stopwords, lemmatization).
 scikit-learn → for machine learning models, feature extraction, and evaluation.
 Matplotlib & Seaborn → for visualizations.
 WordCloud → for generating word clouds of common words in positive and
negative tweets.

Dataset Import:

import pandas as pd

# Load tweets dataset


tweets_df = pd.read_csv("tweets_dataset.csv")
print(tweets_df.head())

The dataset contained two columns: tweet text and sentiment label (positive or negative).

23
Step 2: Data Preprocessing

Preprocessing was performed to clean the raw text:

1. Lowercasing: Convert all text to lowercase.


2. Noise Removal: Remove URLs, mentions (@user), hashtags, numbers, and special
characters.
3. Tokenization: Split sentences into individual words.
4. Stopword Removal: Remove common words that don’t contribute to sentiment (e.g.,
“is”, “the”, “and”).
5. Stemming and Lemmatization: Reduce words to their root or dictionary form.

Example Implementation:

import re
from [Link] import stopwords
from [Link] import WordNetLemmatizer

lemmatizer = WordNetLemmatizer()
stop_words = set([Link]('english'))

def clean_tweet(text):
text = [Link]() # lowercase
text = [Link](r"http\S+|www\S+|https\S+", '', text) # remove URLs
text = [Link](r'@\w+|\#','', text) # remove mentions & hashtags
text = [Link](r'\d+', '', text) # remove numbers
text = [Link](r'[^\w\s]', '', text) # remove punctuation
words = [Link]()
words = [[Link](w) for w in words if w not in stop_words]
return ' '.join(words)

tweets_df['cleaned_tweet'] = tweets_df['tweet'].apply(clean_tweet)

Step 3: Feature Extraction

Since machine learning models require numerical input, textual data was converted using:

1. Bag of Words (BoW): Represents text as word frequency vectors.

24
2. TF-IDF Vectorization: Gives higher weight to words that are important but rare
across tweets.

Example using TF-IDF:

from sklearn.feature_extraction.text import TfidfVectorizer

vectorizer = TfidfVectorizer(max_features=5000)
X = vectorizer.fit_transform(tweets_df['cleaned_tweet']).toarray()
y = tweets_df['sentiment'] # labels

Step 4: Model Building

The dataset was split into training (80%) and testing (20%) sets:

from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

Machine Learning Models Used:

1. Logistic Regression
2. Naïve Bayes (MultinomialNB)
3. Support Vector Machine (SVM)

Example: Logistic Regression

from sklearn.linear_model import LogisticRegression


from [Link] import classification_report, confusion_matrix

model = LogisticRegression()
[Link](X_train, y_train)
y_pred = [Link](X_test)

print(classification_report(y_test, y_pred))

Step 5: Model Evaluation

 Accuracy: Proportion of correctly predicted tweets.

25
 Precision, Recall, F1-score: Measured reliability and balance of the model.
 Confusion Matrix: Displayed correct vs incorrect classifications.

Confusion Matrix Visualization:

import seaborn as sns


import [Link] as plt
from [Link] import confusion_matrix

cm = confusion_matrix(y_test, y_pred)
[Link](cm, annot=True, fmt='d', cmap='Blues')
[Link]("Predicted")
[Link]("Actual")
[Link]()

Step 6: Visualization & Insights

1. Sentiment Distribution:

[Link](x=tweets_df['sentiment'])
[Link]("Distribution of Positive and Negative Tweets")
[Link]()

2. Word Clouds:

 Positive Tweets: Words like love, happy, great.


 Negative Tweets: Words like bad, worst, sad.

from wordcloud import WordCloud

positive_words = ' '.join(tweets_df[tweets_df['sentiment']=='positive']['cleaned_tweet'])


negative_words = ' '.join(tweets_df[tweets_df['sentiment']=='negative']['cleaned_tweet'])

WordCloud(width=800, height=400,
background_color='white').generate(positive_words).to_file("positive_wordcloud.png")
WordCloud(width=800, height=400,
background_color='white').generate(negative_words).to_file("negative_wordcloud.png")

26
Step 7: Outcome & Insights

 The sentiment analysis model successfully classified tweets into Positive and
Negative sentiments with high accuracy.
 Frequent positive and negative keywords were identified using word clouds.
 The project demonstrated real-world application of NLP, useful for businesses,
social media analytics, and brand monitoring.
 Model evaluation metrics ensured reliability and correctness of predictions.

3.3 Challenges Faced and Solutions

During the internship, while working on the Spotify Dashboard and Sentiment Analysis
projects, several challenges were encountered. Each challenge required careful analysis and
practical solutions to ensure successful completion of the projects.

3.3.1 Challenges and Solutions in Dashboard Development (Spotify Dataset)


Challenge Description Solution Implemented
Data Some artist names and genres were Standardized the names using
Inconsistency written differently (e.g., “Hip-Hop” Power Query Editor and
vs “Hip Hop”, “The Weeknd” vs applied consistent formatting rules.
“Weeknd”).
Missing Data Some rows had missing values in key Removed rows with critical
fields like Streams or Release Year, missing values and used default
which could affect calculations. placeholders where necessary.
Duplicate Multiple entries for the same Applied “Remove Duplicates” in
Entries song/artist inflated total streams and Power Query to ensure each song
affected ranking. entry was unique.
Large Dataset The dataset had thousands of records, Optimized the model by filtering
Performance which slowed Power BI unnecessary columns and using
visualizations. measures instead of calculated
columns wherever possible.

27
Interactive Creating a user-friendly interactive Implemented structured dashboard
Design dashboard with multiple slicers and layout with KPIs, charts, and
cross-filtering was challenging. slicers, and tested cross-filter
interactions for smooth
functionality.

Key Learning:

 Learned the importance of data cleaning, standardization, and optimization to


create accurate, interactive dashboards

3.3.2 Challenges and Solutions in Sentiment Analysis (Tweets Dataset)


Challenge Description Solution Implemented

Tweets contained URLs, hashtags, Applied comprehensive preprocessing:


mentions, emojis, and special removed URLs, mentions, hashtags, special
Noisy Data
characters that interfered with characters, and applied tokenization and
model accuracy. lemmatization.

Positive tweets were more Applied stratified train-test split and evaluated
Imbalanced
frequent than negative tweets, models with precision, recall, and F1-score to
Dataset
causing bias in prediction. handle imbalance.

High- Text vectorization (BoW, TF-IDF)


Limited features using max_features in TF-IDF
Dimensional resulted in very high-dimensional
and removed low-frequency words.
Feature Space data, slowing model training.

Focused on clear-labeled data for training and


Some tweets contained sarcasm or
Ambiguous enhanced preprocessing to reduce noise; also
mixed sentiments, making
Sentiment experimented with multiple models (Logistic
classification difficult.
Regression, Naïve Bayes, SVM).

Representing insights from text Created sentiment distribution graphs and


Visualization of
data in an understandable format word clouds to visually present positive and
Results
was challenging. negative sentiment trends.

Key Learning:

28
 Gained experience in handling unstructured data, feature extraction, and
evaluating text classification models effectively.

29
CHAPTER 4
RESULT & CONCLUSION

This chapter presents the results, detailed insights, and conclusions drawn from the two key
projects completed during the internship:

1. Spotify Dashboard Development


2. Sentiment Analysis of Tweets

These projects allowed practical application of data analytics, visualization, and machine
learning techniques, providing experience in handling both structured and unstructured
datasets.

4.1 Results of Dashboard Development (Spotify Dataset)

The Spotify Dashboard project aimed to analyze and visualize music streaming data using
Power BI, providing insights into top songs, popular artists, genre distribution, and audience
trends.

4.1.1 Key Outcomes and Observations

1. Top Songs and Artists


o Using DAX measures and visualizations, the dashboard highlighted the Top
10 most streamed songs and Top 10 artists.
o Example: Songs with streams above 50 million were prominently displayed
using KPI cards.
o Artists were ranked by cumulative streams, enabling identification of market
leaders in the music industry.
2. Genre Analysis
o Pie and donut charts revealed music genre popularity:
 Pop: 35%
 Hip-Hop/Rap: 28%
 Rock: 15%
 EDM/Dance: 12%
 Others: 10%

30
o This analysis helped understand audience preferences and trend patterns
across genres.
3. Year-wise Streaming Trends
o Line charts tracked streaming growth over years (2000–2025).
o Observed significant growth in streams between 2015–2022, correlating with
the rise of digital streaming platforms.
o These trends provided insights into listener engagement over time.
4. Interactive Features
o Slicers enabled filtering by Artist, Genre, Release Year, and Album.
o Drill-down functionality allowed exploration from Genre → Artist → Album
→ Song.
o Cross-filtering ensured dynamic updates across all visuals, supporting real-
time data exploration.

4.1.2 Insights and Implications

 Audience Behavior Analysis: Insights into top artists and genres helped understand
listener preferences.
 Decision Support: Dashboard can guide marketing strategies, playlist curation, and
targeted promotions.
 Data-Driven Reporting: Transforming raw streaming data into interactive
dashboards simplified complex datasets into actionable insights.

Visual Snapshot:

 Bar charts for top songs and artists


 Line chart for year-wise streaming trends
 Pie chart for genre distribution
 KPI cards for total streams, top song, and top artist

31
4.2 Results of Sentiment Analysis (Tweets Dataset)

The Sentiment Analysis project involved Python-based NLP and machine learning to
classify tweets into Positive or Negative sentiments, analyzing public opinion on social
media.

4.2.1 Model Performance

1. Dataset Overview
o Total tweets: ~10,000
o Sentiment labels:
 Positive: 6,200 (62%)
 Negative: 3,800 (38%)
2. Models Evaluated

Model Accuracy Precision (Pos/Neg) Recall (Pos/Neg) F1-Score


Logistic Regression 88% 0.87 / 0.89 0.88 / 0.86 0.87
Naïve Bayes 85% 0.84 / 0.86 0.85 / 0.83 0.84
SVM 89% 0.88 / 0.90 0.89 / 0.88 0.89

 Logistic Regression and SVM performed best, providing reliable classification with
balanced precision and recall.

3. Evaluation Metrics
o Confusion matrix visualizations confirmed correct classification of most
positive and negative tweets.
o F1-score indicated a good balance between precision and recall, especially
for imbalanced data.

4.2.2 Visualization and Insights

1. Sentiment Distribution
o Bar chart showed positive tweets dominated (62%) while negative tweets
(38%) indicated areas of concern.

32
o Allowed quick assessment of general sentiment trends in the dataset.
2. Word Clouds
o Positive Tweets: Highlighted frequent words such as love, happy, great,
amazing.
o Negative Tweets: Highlighted frequent words such as bad, worst, sad, hate.
o Provided a visual understanding of common expressions and keywords
associated with sentiment.
3. Insights
o Demonstrated the effectiveness of NLP in extracting actionable insights
from unstructured text data.
o Can be applied for brand monitoring, customer feedback analysis, and
social media trend tracking.
o Showed real-world application of text classification and feature extraction
techniques (TF-IDF, Bag of Words).

4.3 Challenges Faced and Solutions


Project Challenge Solution Implemented
Spotify Standardized names, removed duplicates, applied
Duplicate & inconsistent data
Dashboard consistent DAX calculations
Spotify Dropped nulls and used default placeholders where
Missing values in critical fields
Dashboard necessary
Spotify Performance issues on large Filtered unnecessary columns, optimized measures
Dashboard dataset and visuals
Sentiment Noisy text (URLs, hashtags, Comprehensive preprocessing including
Analysis mentions) tokenization, lemmatization, and noise removal
Sentiment Used stratified train-test split, evaluated with
Imbalanced dataset
Analysis precision, recall, and F1-score
Sentiment Limited TF-IDF features using max_features,
High-dimensional features
Analysis removed low-frequency words
Sentiment Ambiguous sentiment (sarcasm, Focused on clearly labeled data, tried multiple
Analysis mixed sentiment) models for reliability

33
Key Learning:

 Overcoming these challenges improved data quality, model accuracy, and


dashboard interactivity, enhancing the overall effectiveness of both projects.

4.4 Conclusion

The internship successfully bridged the gap between academic knowledge and practical
industry experience.

4.4.1 Key Takeaways

1. Technical Skills
o Gained hands-on experience with Power BI, Python, NLP, machine
learning algorithms, and data visualization.
o Learned to preprocess structured (Spotify) and unstructured (tweets) data
effectively.
2. Analytical Skills
o Developed the ability to analyze trends, derive insights, and make data-
driven decisions.
o Learned to interpret both numerical and textual data for business or research
purposes.
3. Problem-Solving and Professional Exposure
o Successfully handled data inconsistency, missing values, noisy text, and
high-dimensional datasets.
o Experienced working on real-world datasets, mirroring industry-level data
analytics tasks.

4.4.2 Final Statement

Both projects provided practical exposure to emerging technologies and analytics tools.
The outcomes demonstrated the ability to:

 Create interactive dashboards for structured data visualization.


 Build machine learning models for sentiment classification.
 Extract actionable insights from complex datasets.

34
This internship significantly enhanced technical proficiency, analytical thinking, and
problem-solving skills, preparing for a career in Data Science, Analytics, Machine
Learning, or IT.

35

You might also like