0% found this document useful (0 votes)
8 views64 pages

Python Data Science Internship Report

Python internship report based on summer internship of python for personal use don't use it for illegal things

Uploaded by

ym9599
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
8 views64 pages

Python Data Science Internship Report

Python internship report based on summer internship of python for personal use don't use it for illegal things

Uploaded by

ym9599
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

PYTHON FOR DATA SCIENCE

(INTERNSHIP -1)
an Internship Report submitted to
SRM institute of science and technology,tiruchirapalli
in partial fullfillment of the requirements for the award of the

BSC COMPUTER SCIENCE


by
[Link]
(Reg No:RA2431005050030)

under the Guidance of


[Link]
(Assistant Professor Department of Computer Science )

DEPARTMENT OF COMPUTER SCIENCE


SRM INSTITUTE OF SCIENCE AND TECHNOLOGY
TIRUCHIRAPPALLI

September 2025
Completion Certificate
Learning Objectives/Internship Objectives:

Internships are generally thought of to be reserved for college students looking to gain experience
in a particular field. However, a wide array of people can benefit from Training Internships in
order to receive real world experience and develop their skills.

An objective for this position should emphasize the skills you already possess in the area and you
interest in learning more

Internships are utilized in a number of different career fields, including architecture, engineering,
healthcare, economics, advertising and many more.

Some Internships are used to allow individuals to perform scientific research while others are
specifically designed to allow people to gain first-hand experience working.

Utilizing Internships is a great way to build your resume and develop skills that can be
emphasized in your resume for future jobs. When you are applying for Training internship, make
you stand apart from the rest of the applicants so that you have an approved chance of landing
the position.
INDEX

S. No CONTENTS page no

1. Introduction
1.1 Modules

2. Analysis
2.1 Requirement analysis
2.2 Feasibility study
3. Software requirements specifications
3.1 System configuration
3.2 Software requirements
3.3 Hardware requirements
4. Technology
4.1
4.2 PYTHON

5. Testing
5.1 Introduction
5.2 Types of testing
5.3 Test Cases
Overview of internship activities:

DATE DAY NAME OF THE TOPIC/MODULE COMPLETED

Day 1 Thursday Introduction to Data analytics, Role of data analytics in AI


developments
Day 2 Introduction to POWERBI / Case study on development of
interactive dashboard for live room parameter dataset
Day 3 Introduction to ThingSpeak Cloud for Data collection /
Configuration, Data downloading, Public sharing / Cloud reading
Day 4 Case study work 1 Data analytics with Predictive analysis model /
Virtual sensor
Day 5 Case study work 2: Creation of EDA plots
Day 6 Introduction to Data augmentation
Day 7 Case study work: 3: Vgg16 based Socket image classification (Data
augmentation)
Day 8 Case study work 4 : Stress detection and analysis
Day 9 Case study work 5 : Word cloud for Big Data
Day 10 Case study work 6 : Phishing attack detection
Day 11 Case study work 7: Fake news detection
Day 12 Case study work 8 : Kidney disease analysis
Day 13 Case study work 9: Groundnut oil adulteration analysis
Day 14 Case study work 10: Nature inspired algorithms
Day 15 Doubts clarifications / Materials and feedback session

2. ANALYSIS
REQUIREMENT ANALYSIS:
The purpose of this project is to upload books in online more efficiently and
effectively to satisfy every client on authors disk.
To Develop a College Internal Research Papers Publishing Journals website with can use
publish Papers, Articles, Conferences took in College is a quarterly, open access,
multidisciplinary, peer reviewed, online and fully refereed international journal. We primarily
aim to bring out the research talent and the works done by scientists, academia, engineers,
practitioners, scholars, post graduate students of engineering, science and other related subjects
so that fellow researchers can get benefit from the research done. This journal aims to cover the
scientific research in a broader sense and not just publishing a niche area of research facilitating
researchers from various verticals to publish their papers. Each Research paper is evaluated in
depth by IJAC reviewer panel that ensures the novelty in each research manuscript being
published.

Feasibility Study

Preliminary investigation examine project feasibility, the likelihood the system


will be useful to the organization. The main objective of the feasibility study is to test the
Technical, Operational and Economical feasibility for adding new modules and debugging old
running system. All system is feasible if they are unlimited resource and infinite time. There are
aspects in the feasibility study portion of the preliminary investigation:
✔ Technical Feasibility
✔ Operation Feasibility
✔ Economic Feasibility

Technical Feasibility:
The technical issue usually raised during the feasibility stage of the
investigation includes the following:
✔ Does then necessary technology exist to do what is suggested?
✔ Does the proposed equipment have the technical capacity to hold the data required to use the
new system?
✔ Will the proposed system provide adequate response to inquiries, regardless of the number or
location of users?
✔ Can the system be upgraded if developed?
✔ Are there technical guarantees of accuracy, reliability, ease of access and data security?

Earlier no system existed to cater to the needs of „Secure Infrastructure Implementation System‟.
The current system developed is technically feasible. It is a web based user interface for audit
workflow at NIC-CSD. Thus it provides an easy access to the users. The database‟s purpose is
to create, establish and maintain a workflow among various entities in order to facilities all
concerned users in their various capacities or roles. Permission to the users would be granted
based on the roles specified. Therefore, it provides the technical guarantee of accuracy, reliability
and security. The software and hardware requirement for the development of this project are not
many and are already available in-house at NIC or are available as free as open source. The work
for the project is done with the current equipment and existing software technology. Necessary
bandwidth exists for providing a fast feedback to the users irrespective of the number of users
using the system.
Operational Feasibility
Proposed projects are beneficial only if they can be turned out into information system. That will
meet the organization‟s operating requirements. Operational feasibility aspects of the project are
to be taken as an importantpart of the project implementation. Some of the important issues raised
are to test the operational feasibility of a project includes the following:

✔ Is there sufficient support for the management from the users?


✔ Will the system be used and work properly if it being developed and implement?
✔ Will there be any resistance from the user that will undermine the possible applications
benefits?
✔ This system is targeted to be in accordance with the above-mentioned issues. Beforehand, the
management is issues and user requirements have
✔ been taken into [Link] there is to question of resistance from the users that can
undermine the possible applications benefits.

The well-planned designed would ensure the optimal utilization of the computer resource and
would help in the improvement of performance status.

Economic Feasibility
A system can be developed technically and the will be used if installed must still be a good
investment for the organization. In the economic feasibility, the development cost in creating the
system is evaluated against the ultimate benefits derived from the new system. Financial benefits
must equal or exceed the costs.
The system is economically feasible. It does not require any addition hardware or software.
Since the interface for this system is developed using the existing resources and technologies
available at NIC, there is nominal expenditure and economic feasibility for certain.

3. SOFTWARE REQUIREMENTS SPECIFICATIONS 3.1System configurations


The software requirement specification can produce at the culmination of the analysis task. The
function and performance allocated to software as part of system engineering are refined by
established a complete information description, a detailed functional description, a
representation of system behaviour, and indication of performance and design constrain,
appropriate validate criteria, and other information pertinent to requirements.

Software requirements:
Operating System : Windows 7, 10,11
,,Front End : Python
Back End : MySQL
Hardware Requirements:
Processor : Intel core i3
Memory : 4GB RAM
Hard Disk : 500GB

Data Analysis

Data Analysis or Data Analytics is studying, cleaning, modeling, and transforming data to find
useful information, suggest conclusions, and support decision-making. This Data Analytics
Tutorial will cover all the basic to advanced concepts of Excel data analysis like data
visualization, data preprocessing, time series, data analysis tools, etc.

Data Analysis Process

Data Analysis is developed by the statistician John Tukey in the 1970s. It is a procedure for
analyzing data, methods for interpreting the results of such systems, and modes of planning the
group of data to make its analysis easier, more accurate, or more factual.
Therefore, data analysis is a process for getting large, unstructured data from different sources
and converting it into information that is gone through the below process:

 Data Requirements Specification


 Data Collection
 Data Processing
 Data Cleaning
 Data Analysis
 Communication

Need for Data Analysis

Data analytics is significant for business optimization performance. An organization can also use
data analytics to make better business decisions and support analyzing customer trends and
fulfillment, which can lead to unknown and better products and services. Executing it into the
business model indicates businesses can help reduce costs by recognizing more efficient modes
of doing business.

Applications of Data Analysis

Better decision-making: The Key advantage of data analysis is better decision-making in the long
term. Rather than depending only on knowledge, businesses are increasingly looking at data
before deciding.

Identification of potential risks: Companies in today’s world succeed in high-risk conditions, but
those environments require critical risk management processes, and extensive data has
contributed to developing new risk management solutions. Data can enhance the effectiveness of
actual simulations to predict future risks and create better planning.

Increase the efficiency of work: Data analysis allows you to analyze a large set of data and present
it in a structured way to help reach your organization’s objectives. Possibilities and progress
within the organization are reflected, and activities can increase work efficiency and productivity.
It enables a culture of efficiency and collaboration by allowing managers to share detailed data
with employees.

Delivering relevant products: Products are the oil for every organization, and often the most
important asset of organizations. The role of the product management team is to determine trends
that drive strategic creation, and activity plans for unique functions and services.

Track customer behavioral changes: Consumers have a lot to choose from in products available
in the markets. Organizations have to pay attention to consumer demands and expectations, So
to analyze the behavior of the customer data analysis is very important.

Prerequisites for Data Analysis

To strong skill for Data Analysis we needs to learn this resources to have a best practice in this
domains.

Python For Data Analysis

SQL For Data Analysis

Python Data Visulization

Data Analysis Datasets

Data Analysis Libraries

Pandas Tutorial

Learn Pandas to unlock powerful tools for data analysis in Python. This essential library offers
versatile data structures like DataFrames, enabling efficient data manipulation, analysis, and
visualization. Mastering Pandas will significantly enhance your ability to handle and extract
insights from complex datasets, making it an indispensable skill for any data analyst or scientist.
CSV files are the Comma Separated Files. To access data from the CSV file, we require a function
read_csv() from Pandas that retrieves data in the form of the data frame.

Syntax of Pandas read_csv

Here is the Pandas read CSV syntax with its parameters.

Syntax: pd.read_csv(filepath_or_buffer, sep=’ ,’ , header=’infer’, index_col=None,


usecols=None, engine=None, skiprows=None, nrows=None)

Parameters:

 filepath_or_buffer: Location of the csv file. It accepts any string path or URL of the file.

 sep: It stands for separator, default is ‘, ‘.

 header: It accepts int, a list of int, row numbers to use as the column names, and the start
of the data. If no names are passed, i.e., header=None, then, it will display the first
column as 0, the second as 1, and so on.

 usecols: Retrieves only selected columns from the CSV file.

 nrows: Number of rows to be displayed from the dataset.

 index_col: If None, there are no index numbers displayed along with records.

 skiprows: Skips passed rows in the new data frame.

Code

# Import pandas
import pandas as pd
# reading csv file
df = pd.read_csv("[Link]")
print([Link]())
Exploratory Data Analysis

Exploratory Data Analysis (EDA) is also crucial step in the data analysis process that involves
summarizing the main characteristics of a dataset, often with visual methods. The goal of EDA
is to understand the data’s underlying structure, detect patterns and anomalies, test hypotheses,
and check assumptions. EDA is essential for making informed decisions about data
preprocessing, feature engineering, and modeling.

In Python, exploratory data analysis, or EDA, is a crucial step in the data analysis process that
involves studying, exploring, and visualizing information to derive important insights. To find
patterns, trends, and relationships in the data, it makes use of statistical tools and visualizations.
This helps to formulate hypotheses and direct additional investigations.

Python provides strong EDA tools with its diverse library ecosystem, which includes Seaborn,
Matplotlib, and Pandas. An essential phase in the data science pipeline, this procedure improves
data comprehension and provides information for further modeling decisions.

Exploratory Data Analysis(EDA) is the main step in the process of various data analysis. It helps
data to visualize the patterns, characteristics, and relationships between variables. Python
provides various libraries used for EDA such as NumPy, Pandas, Matplotlib, Seaborn, and Plotly.

EDA is a phenomenon under data analysis used for gaining a better understanding of data aspects
like:
main features of data

variables and relationships that hold between them

Identifying which variables are important for our problem

We shall look at various exploratory data analysis methods like:

Reading dataset

Analyzing the data

Checking for the duplicates

Missing Values Calculation

Exploratory Data Analysis

Univariate Analysis

Bivariate Analysis

Multivariate Analysis

Numpy Tutorial

Learn NumPy to master numerical computing in Python. This foundational library provides
support for arrays, matrices, and high-level mathematical functions, making data manipulation
and computation highly efficient. Understanding NumPy is crucial for performing advanced data
analysis and scientific computing, and it serves as a cornerstone for many other data science
libraries.

Data Preprocessing:

Data preparation is a critical step in any data analysis or machine learning project. It involves a
variety of tasks aimed at transforming raw data into a clean and usable format. Properly prepared
data ensures more accurate and reliable analysis results, leading to better decision-making and
more effective predictive models. This guide will cover key aspects of data preparation, including
data formatting, data cleaning, outlier detection, data transformation, and data sampling.

When referring to data preparation and cleaning, preprocessing is done before raw data is
entered into an analytical tool or machine learning model. Missing value handling, feature
scaling, categorical variable encoding, and outlier removal are all part of it. To improve the
performance and interpretability of the model, it is important to make sure the data is in the
right format. Data-driven jobs are more successful overall when preprocessing is used to reduce
noise, standardize data, and optimize it for effective analysis.
GOOGLE COLLAB

Google Colab, or Colaboratory, is a powerful and user-friendly cloud-based platform designed


for writing and executing Python code in an interactive Jupyter Notebook environment. It is
particularly valuable for data analytics learning due to its integration with Google’s cloud
infrastructure, which provides a range of features that enhance the learning and data analysis
experience.

Key Features and Benefits:


Interactive Notebooks: Google Colab offers an interactive notebook interface where users can
write and run Python code, visualize data, and document their analyses in a single environment.
This makes it easier to develop and share data analysis workflows and results.

Access to Computational Resources: Users can take advantage of Google’s cloud infrastructure,
which includes access to powerful computational resources like GPUs (Graphics Processing
Units) and TPUs (Tensor Processing Units). These resources are especially beneficial for
handling large datasets and training complex machine learning models, which can otherwise be
resource-intensive on local machines.

Pre-installed Libraries and Tools: Colab comes with a wide range of pre-installed libraries and
tools essential for data analytics and machine learning, including popular ones like NumPy,
Pandas, Matplotlib, Scikit-learn, and TensorFlow. This reduces the setup time and allows
learners to focus on their data analysis tasks.

Seamless Integration with Google Drive: Integration with Google Drive allows users to easily
store, access, and manage their notebooks and datasets. This integration also simplifies sharing
and collaboration, as users can directly open and save their notebooks in Google Drive.

Real-time Collaboration: Colab supports real-time collaboration, enabling multiple users to


work on the same notebook simultaneously. This feature is useful for group projects, peer
reviews, and collaborative learning, where users can see each other’s changes and comments in
real time.
Ease of Use: Google Colab’s interface is designed to be user-friendly, making it accessible for
learners of all levels. The environment supports various features such as code completion,
inline documentation, and error highlighting, which facilitate a smoother coding experience.

Cost-effectiveness: The platform is free to use, which makes it an attractive option for learners
and educators who may not have access to high-performance computing resources. While
Colab offers paid plans with enhanced features, the free tier provides substantial capabilities for
most educational and analytical tasks.

In summary, Google Colab is a versatile and accessible tool for data analytics learning,
providing an interactive and collaborative environment with robust computational resources and
a range of pre-installed libraries. Its integration with Google Drive and support for real-time
collaboration further enhance its utility as a platform for developing and sharing data analysis
projects.

LIBRARY INSTALLATIONS

Google Colab provides access to a range of pre-installed libraries that offer significant benefits
for data analysis, machine learning, and scientific computing. Here’s an overview of the
advantages of using these libraries in Google Colab:

Ease of Setup: Many essential libraries are pre-installed in Google Colab, which means users
don’t need to spend time installing and configuring them. This simplifies the setup process and
allows users to focus on their analysis and modeling tasks immediately.

Comprehensive Ecosystem: Colab includes a broad set of libraries that cover various aspects of
data science and machine learning. Key libraries include:

NumPy: Provides support for large, multi-dimensional arrays and matrices, along with a
collection of mathematical functions to operate on these arrays.
Pandas: Offers data structures and data analysis tools for handling and analyzing structured data,
such as data frames and series.

Matplotlib and Seaborn: Used for data visualization, allowing users to create a wide range of
plots and charts to visualize data distributions and relationships.

Scikit-learn: Contains tools for data mining and data analysis, including algorithms for
classification, regression, clustering, and dimensionality reduction.

TensorFlow and PyTorch: Popular frameworks for deep learning that provide extensive tools and
libraries for building and training neural networks.

Accelerated Computation: Libraries such as TensorFlow and PyTorch leverage the cloud-based
GPUs and TPUs available in Colab to accelerate the training of machine learning models. This
significantly speeds up computations compared to running on a CPU, especially for complex
models and large datasets.

Up-to-date Versions: Google Colab often updates its libraries to the latest versions, providing
users with access to the newest features, improvements, and bug fixes. This ensures that users
can take advantage of the latest advancements in the libraries without needing to manually update
them.

Seamless Integration: The libraries in Colab are well-integrated with each other and with Colab’s
environment, which enhances productivity. For example, Pandas and Matplotlib work together
smoothly for data analysis and visualization, while TensorFlow and Keras offer high-level APIs
for building and training deep learning models.

Collaboration and Sharing: Pre-installed libraries in Colab facilitate easy sharing and
collaboration on notebooks. Users can share their notebooks with colleagues or the community,
and others can reproduce the analysis or extend it using the same libraries, ensuring consistency
and reproducibility.

Educational Resource: For learners and educators, having a well-curated set of libraries available
in Colab provides a comprehensive learning environment. It enables students to experiment with
various tools and techniques without the overhead of setting up their own development
environments.

POWER BI

Power BI is a powerful business analytics tool developed by Microsoft that enables users to
visualize data, create interactive reports, and share insights across organizations. It integrates
with various data sources and provides a comprehensive suite of features designed to facilitate
data analysis and decision-making.

Key Features of Power BI:

Interactive Dashboards and Reports:

Visualizations: Power BI offers a wide range of visualization options, including charts, graphs,
maps, and custom visuals, allowing users to present data in an easily digestible format.

Real-Time Data: Users can create dashboards that reflect real-time data updates, which is
crucial for monitoring key performance indicators (KPIs) and making timely decisions.

Data Integration:

Connectors: Power BI connects to various data sources, including databases (SQL Server,
Oracle), cloud services (Azure, Google Analytics), spreadsheets (Excel), and online services
(Salesforce). This flexibility allows users to consolidate data from multiple sources into a single
view.

Data Transformation: The Power Query Editor enables users to clean, transform, and shape data
before analysis, ensuring data quality and consistency.

Advanced Analytics:
DAX (Data Analysis Expressions): Power BI uses DAX for creating custom calculations and
measures within reports, providing powerful analytical capabilities.

Integration with R and Python: Users can integrate R and Python scripts to perform advanced
analytics and create custom visualizations, enhancing the analytical depth of reports.

Collaboration and Sharing:

Power BI Service: The online service allows users to publish and share reports and dashboards
with others, facilitating collaboration and ensuring that stakeholders have access to the latest
insights.

Access Control: Granular access control and security features ensure that only authorized users
can view or interact with specific data and reports.

Mobile Access:

Mobile App: Power BI provides a mobile app for iOS and Android, enabling users to access
reports and dashboards on the go, which is essential for remote work and decision-making.

Artificial Intelligence (AI) Integration:

AI Insights: Power BI incorporates AI capabilities, such as automated insights and natural


language queries, allowing users to explore data using conversational language and uncover
hidden patterns.

Benefits of Power BI in Data Analytics:

Enhanced Data Visualization and Interpretation:


Power BI’s rich set of visualization tools helps users present complex data in a clear and
intuitive manner, making it easier to identify trends, patterns, and anomalies. This leads to more
informed decision-making and better data-driven strategies.

Seamless Data Integration and Consolidation:

By connecting to various data sources and integrating data into a unified platform, Power BI
eliminates data silos and provides a comprehensive view of organizational performance. This
consolidated view supports more accurate analysis and reporting.

Improved Data Accessibility and Collaboration:

The ability to share interactive reports and dashboards through Power BI Service enhances
collaboration among team members and stakeholders. Real-time updates and mobile access
ensure that users stay informed and can make decisions based on the most current data.

Cost-Effective Solution:

Power BI offers a range of pricing options, including a free version with core features and paid
versions with advanced capabilities. This flexibility makes it accessible to organizations of all
sizes, from small businesses to large enterprises.

Empowered Self-Service Analytics:

Power BI’s user-friendly interface and drag-and-drop functionality empower non-technical


users to perform their own data analysis and create custom reports without relying heavily on
IT departments. This self-service approach accelerates insights and reduces bottlenecks.

Scalability and Flexibility:


Power BI’s cloud-based architecture allows for scalability, accommodating growing data
volumes and user needs. Its integration with other Microsoft tools and services enhances its
versatility and adaptability to various organizational requirements.

DATA COLLECTION USING THING SPEAK

ThingSpeak is an Internet of Things (IoT) platform and cloud service that allows developers to
build and deploy IoT applications and devices. It is a platform that provides an easy-to-use
interface and tools to collect, store, analyze, and visualize data from IoT devices in real-time.
Internet of Things (IoT) describes an emerging trend where a large number of embedded devices
(things) are connected to the Internet. These connected devices communicate with people and
other things and often provide sensor data to cloud storage and cloud computing resources where
the data is processed and analyzed to gain important insights. Cheap cloud computing power and
increased device connectivity is enabling this [Link] solutions are built for many vertical
applications such as environmental monitoring and control, health monitoring, vehicle fleet
monitoring, industrial monitoring and control, and home automation.

At a high level, many IoT systems can be described using the diagram below:
Some key features of ThingSpeak include:

On the left, we have the smart devices (the “things” in IoT) that live at the edge of the network.
These devices collect data and include things like wearable devices, wireless temperatures
sensors, heart rate monitors, and hydraulic pressure sensors, and machines on the factory floor.

In the middle, we have the cloud where data from many sources is aggregated and analyzed in
real time, often by an IoT analytics platform designed for this purpose.

The right side of the diagram depicts the algorithm development associated with the IoT
application. Here an engineer or data scientist tries to gain insight into the collected data by
performing historical analysis on the data. In this case, the data is pulled from the IoT platform
into a desktop software environment to enable the engineer or scientist to prototype algorithms
that may eventually execute in the cloud or on the smart device itself.
An IoT system includes all these elements. ThingSpeak fits in the cloud part of the diagram and
provides a platform to quickly collect and analyze data from internet connected sensors.

ThingSpeak Key Features

ThingSpeak allows you to aggregate, visualize and analyze live data streams in the cloud. Some
of the key capabilities of ThingSpeak include the ability to:

Easily configure devices to send data to ThingSpeak using popular IoT protocols.

Visualize your sensor data in real-time.

Aggregate data on-demand from third-party sources.

Use the power of MATLAB to make sense of your IoT data.

Run your IoT analytics automatically based on schedules or events.

Prototype and build IoT systems without setting up servers or developing web software.

Automatically act on your data and communicate using third-party services like Twilio® or
Twitter®.

To learn how you can collect, analyze and act on your IoT data with ThingSpeak, explore the
topics below:

Data collection: ThingSpeak provides APIs and tools to collect data from various IoT devices
such as sensors, cameras, and other data sources.

Data storage: ThingSpeak stores data in channels, which are essentially time-series databases
that allow users to store and organize data from multiple devices.
Data analytics: ThingSpeak allows developers to analyze and manipulate data using built-in
analytics tools such as MATLAB analytics, data visualizations, and machine learning models.

Integration with other services: ThingSpeak can be integrated with other IoT services and
platforms such as Arduino, Particle, and other cloud services.

IoT device management: ThingSpeak offers a device management portal that allows developers
to monitor and manage IoT devices and their connections.

One of the main advantages of ThingSpeak is its simplicity and ease-of-use. The platform
offers a user-friendly interface and requires minimal coding skills, making it accessible to
developers with varying levels of experience. Additionally, ThingSpeak is a free cloud service
that offers a range of features, making it an attractive option for developers looking to build and
deploy IoT applications quickly and cost-effectively.

Introduction to Python:

Python is a general-purpose interpreted, interactive, object-oriented, and high-level


programming language.

Characteristics of Python

Following are important characteristics of Python Programming −


 It supports functional and structured programming methods as well as OOP.
 It can be used as a scripting language or can be compiled to byte-code for building large
applications.
 It provides very high-level dynamic data types and supports dynamic type checking.
 It supports automatic garbage collection.
 It can be easily integrated with C, C++, COM, ActiveX, CORBA, and Java.

Program 1:

Print(“"Hello, Python!")

Python has five standard data types

 Numbers
 String
 List
 Tuple
 Dictionary

Python Numbers

Number data types store numeric values. Number objects are created when you assign a value
to them. For example −
var1 = 1
var2 = 10
You can also delete the reference to a number object by using the del statement. The syntax of
the del statement is −
del var1[,var2[,var3[....,varN]]]]
You can delete a single object or multiple objects by using the del statement. For example −
del var
del var_a, var_b

Python supports four different numerical types −

 int (signed integers)


 long (long integers, they can also be represented in octal and hexadecimal)
 float (floating point real values)
 complex (complex numbers)
Python Strings
Strings in Python are identified as a contiguous set of characters represented in the quotation
marks. Python allows for either pairs of single or double quotes. Subsets of strings can be taken
using the slice operator ([ ] and [:] ) with indexes starting at 0 in the beginning of the string and
working their way from -1 at the end.
The plus (+) sign is the string concatenation operator and the asterisk (*) is the repetition
operator. For example −

str = 'Hello World!'


print str # Prints complete string
print str[0] # Prints first character of the string
print str[2:5] # Prints characters starting from 3rd to 5th
print str[2:] # Prints string starting from 3rd character
print str * 2 # Prints string two times
print str + "TEST" # Prints concatenated string
Python Lists
Lists are the most versatile of Python's compound data types. A list contains items separated by
commas and enclosed within square brackets ([]). To some extent, lists are similar to arrays in
C. One difference between them is that all the items belonging to a list can be of different data
type.
list = [ 'abcd', 786 , 2.23, 'john', 70.2 ]
tinylist = [123, 'john']

print list # Prints complete list


print list[0] # Prints first element of the list
print list[1:3] # Prints elements starting from 2nd till 3rd
print list[2:] # Prints elements starting from 3rd element
print tinylist * 2 # Prints list two times
print list + tinylist # Prints concatenated lists
Python Tuples
A tuple is another sequence data type that is similar to the list. A tuple consists of a number of
values separated by commas. Unlike lists, however, tuples are enclosed within parentheses.
The main differences between lists and tuples are: Lists are enclosed in brackets ( [ ] ) and their
elements and size can be changed, while tuples are enclosed in parentheses ( ( ) ) and cannot be
updated. Tuples can be thought of as read-only lists. For example

tuple = ( 'abcd', 786 , 2.23, 'john', 70.2 )


tinytuple = (123, 'john')

print tuple # Prints the complete tuple


print tuple[0] # Prints first element of the tuple
print tuple[1:3] # Prints elements of the tuple starting from 2nd till 3rd
print tuple[2:] # Prints elements of the tuple starting from 3rd element
print tinytuple * 2 # Prints the contents of the tuple twice
print tuple + tinytuple # Prints concatenated tuples
Python Dictionary
Python's dictionaries are kind of hash table type. They work like associative arrays or hashes
found in Perl and consist of key-value pairs. A dictionary key can be almost any Python type,
but are usually numbers or strings. Values, on the other hand, can be any arbitrary Python object.
Dictionaries are enclosed by curly braces ({ }) and values can be assigned and accessed using
square braces ([]). For example −#!/usr/bin/python

dict = {}
dict['one'] = "This is one"
dict[2] = "This is two"

tinydict = {'name': 'john','code':6734, 'dept': 'sales'}

print dict['one'] # Prints value for 'one' key


print dict[2] # Prints value for 2 key
print tinydict # Prints complete dictionary
print [Link]() # Prints all the keys
print [Link]() # Prints all the values

Conditional Exceptions

Decision making is anticipation of conditions occurring while execution of the program and
specifying actions taken according to the conditions.
Decision structures evaluate multiple expressions which produce TRUE or FALSE as outcome.
You need to determine which action to take and which statements to execute if outcome is TRUE
or FALSE otherwise.
Following is the general form of a typical decision making structure found in most of the
programming languages −
Python programming language assumes any non-zero and non-null values as TRUE, and if it
is either zero or null, then it is assumed as FALSE value.
Python programming language provides following types of decision making statements. Click
the following links to check their detail.

[Link]. Statement & Description

1 if statements

An if statement consists of a boolean expression followed by one or more


statements.

2 if...else statements

An if statement can be followed by an optional else statement, which


executes when the boolean expression is FALSE.

3 nested if statements

You can use one if or else if statement inside another if or else


if statement(s).
Functions
A function is a block of organized, reusable code that is used to perform a single, related action.
Functions provide better modularity for your application and a high degree of code reusing.
As you already know, Python gives you many built-in functions like print(), etc. but you can
also create your own functions. These functions are called user-defined functions.
Defining a Function
You can define functions to provide the required functionality. Here are simple rules to define a
function in Python.
 Function blocks begin with the keyword def followed by the function name and
parentheses ( ( ) ).
 Any input parameters or arguments should be placed within these parentheses. You can
also define parameters inside these parentheses.
 The first statement of a function can be an optional statement - the documentation string
of the function or docstring.
 The code block within every function starts with a colon (:) and is indented.
 The statement return [expression] exits a function, optionally passing back an expression
to the caller. A return statement with no arguments is the same as return None.
Syntax
def functionname( parameters ):
"function_docstring"
function_suite
return [expression]
By default, parameters have a positional behavior and you need to inform them in the same order
that they were defined.
Example
The following function takes a string as input parameter and prints it on standard screen.

def printme( str ):


"This prints a passed string into this function"
print str
return

Loop Control Statements


Loop control statements change execution from its normal sequence. When execution leaves a
scope, all automatic objects that were created in that scope are destroyed.
Python supports the following control statements. Click the following links to check their detail.
Let us go through the loop control statements briefly

[Link]. Control Statement & Description

1 break statement

Terminates the loop statement and transfers execution to the statement


immediately following the loop.

2 continue statement

Causes the loop to skip the remainder of its body and immediately retest its
condition prior to reiterating.

3 pass statement

The pass statement in Python is used when a statement is required syntactically


but you do not want any command or code to execute.

Strings are amongst the most popular types in Python. We can create them simply by enclosing
characters in quotes. Python treats single quotes the same as double quotes. Creating strings is
as simple as assigning a value to a variable. For example −

var1 = 'Hello World!'


var2 = "Python Programming"

Lambda Function

 A lambda function is also called as an anonymous function as it is a function that is


defined without a name.
 A lambda function behaves similar to a standard function except it is defined in one-
line.
 It is defined using a lambda key-word.
Lambda functions can have any number of arguments but only one expression. The
expression is evaluated and returned.
 Lambda functions can be used wherever function objects are required.
Syntax of Lambda Function –

add = lambda a, b : a + b

add(2, 3)

NUMPY

1. What is Python Numpy¶

The NumPy is a python package that stands for ‘Numerical Python’.

Numpy is the core library for scientific computing, which contains a powerful n-dimensional
array object, provides tools for integrating C, C++, etc.

It contains a powerful N-dimensional array object.

Python NumPy Array v/s List

We use python numpy array instead of a list because of the below three reasons:

Arrays occupy less memory

work faster

are convenient to use

Advantages of numpy arrays over lists¶

Python numpy arrays are more compact as compared to lists.

Functions to Create Array:

 An array of random numbers(random(),randn(),randint())


The random() function returns random numbers in the half-open interval [0.0, 1.0). The
half-open interval includes 0 but excludes 1. The required number of random numbers is
passed through the ‘size’ parameter.
The randn() creates an array of the given shape with random variables from a uniform
distribution between (0, 1)
The randint() returns random integers from low (inclusive) to high (exclusive)

Arithmetic Functions in Numpy

Sum(),min(),power()

1. Pandas

Pandas contain data structures and data manipulation tools designed for data cleaning and
analysis.

While pandas adopt much code from NumPy, the difference is that Pandas is designed for
tabular, heterogeneous data. NumPy, by difference, is best suited for working with
homogeneous numerical array data.

The name Pandas is derived from the term 'panel data' (an econometrics term for
multidimensional structured data sets).

1. You can import it as 'pd'


import pandas as pd

Pandas Series

Pandas Series is a one-dimensional labelled array capable of holding any data type. However, a
series is a sequence of similar data types, similar to an array, list, or column in a table.

It will assign a labelled index to each item in the [Link]. By default, each item will receive
an index label from 0 to N, where N is the length of the Series minus one.

4. Pandas DataFrames¶

A DataFrame is a tabular representation of data containing an ordered collection of columns,


each of which can be a different type (such as numeric, string, boolean).

The DataFrame has both row and column index; it can be thought of as a dict of Series all
sharing the same index. In a data frame, the data is stored as one or more two-dimensional
blocks rather than a list, dict, or some other collection of one-dimensional arrays.
While a DataFrame is physically two-dimensional, it can be used to represent higher
dimensional data in a tabular format using hierarchical indexing.

Concat in pandas: joining tables

IMPORTING DATA

Reading CSV

Saving IN python

MATPLOT and SEABORN DATA ANALYSIS:

 Univariate Analysis:
Numerical column – Histogram,Kdeplot,Distplot,Boxplot,Lineplot
Categorical column – Bar Graph,pie
 Bi Variate Analysis:
Numerical – Numerical – Scatter plot,Heat map,pairplot
Numerical – Categorical – bar graph,sworm plot,boxplot
Categorical- Categorical- Bar with cross tab

What Is Regression?
Regression searches for relationships among variables. For example, you can observe several
employees of some company and try to understand how their salaries depend on their features,
such as experience, education level, role, city of employment, and so on.

This is a regression problem where data related to each employee represents one observation.
The presumption is that the experience, education, role, and city are the independent features,
while the salary depends on them.

Similarly, you can try to establish the mathematical dependence of housing prices on area,
number of bedrooms, distance to the city center, and so on.

Linear Regression
Linear regression is probably one of the most important and widely used regression techniques.
It’s among the simplest regression methods. One of its main advantages is the ease of interpreting
results.

Example:

import numpy as np

import [Link] as plt


def estimate_coef(x, y):

# number of observations/points

n = [Link](x)

# mean of x and y vector

m_x = [Link](x)

m_y = [Link](y)

# calculating cross-deviation and deviation about x

SS_xy = [Link](y*x) - n*m_y*m_x

SS_xx = [Link](x*x) - n*m_x*m_x

# calculating regression coefficients

b_1 = SS_xy / SS_xx

b_0 = m_y - b_1*m_x

return (b_0, b_1)

def plot_regression_line(x, y, b):

# plotting the actual points as scatter plot

[Link](x, y, color = "m",

marker = "o", s = 30)

# predicted response vector

y_pred = b[0] + b[1]*x


# plotting the regression line

[Link](x, y_pred, color = "g")

# putting labels

[Link]('x')

[Link]('y')

# function to show plot

[Link]()

def main():

# observations / data

x = [Link]([0, 1, 2, 3, 4, 5, 6, 7, 8, 9])

y = [Link]([1, 3, 2, 5, 7, 8, 8, 9, 10, 12])

# estimating coefficients

b = estimate_coef(x, y)

print("Estimated coefficients:\nb_0 = {} \

\nb_1 = {}".format(b[0], b[1]))

# plotting regression line

plot_regression_line(x, y, b)

if __name__ == "__main__":

main()
Multiple Linear Regression
Multiple or multivariate linear regression is a case of linear regression with two or more
independent variables.

If there are just two independent variables, then the estimated regression function is, It represents
a regression plane in a three-dimensional space. The goal of regression is to determine the values
of the weights 𝑏₀, 𝑏₁, and 𝑏₂ such that this plane is as close as possible to the actual responses,
while yielding the minimal SSR.

import numpy as np

import matplotlib as mpl

from mpl_toolkits.mplot3d import Axes3D

import [Link] as plt

def generate_dataset(n):

x = []

y = []

random_x1 = [Link]()

random_x2 = [Link]()

for i in range(n):

x1 = i

x2 = i/2 + [Link]()*n

[Link]([1, x1, x2])

[Link](random_x1 * x1 + random_x2 * x2 + 1)

return [Link](x), [Link](y)

x, y = generate_dataset(200)

[Link]['[Link]'] = 12

fig = [Link]()

ax = fig.add_subplot(projection ='3d')
[Link](x[:, 1], x[:, 2], y, label ='y', s = 5)

[Link]()

ax.view_init(45, 0)

[Link]()

ROLE OF MACHINE LEARNING / AI

Supervised Learning

Learning models are divided into three conventional types based on the way of Input and target
is being analyzed. Supervised learning model is a technique in which the machine learning
analysis is with the specific target is predictable. Based on the data provided to the labels the
prediction process takes place. The label information provider to the machine learning algorithm
keep focus on achieving the target matches score. Since the labels are non-factors to identify the
target data supervisor learning approaches at times a high accuracy. Most of the practical
applications utilized supervised learning process for fast and efficient outcome the supervisor
learning model is represented by the expression given below.

Y=f(x)

Here f(x) is based on the mapping variable associated with it. Since supervised learning model
depends on labels f(x) is always focused towards known selection of input data where the input
variable x and the output variable Y is controlled by the certain mapping function derived by the
algorithm. The mapping function is being supervised by the known data of certain pattern from
the given input data.
Figure 1.1 Supervised learning model

The output y is derived as the output function used as a derivative value of the whole process.
The method is qualified into supervised learning process as algorithm keep on learning and
understanding the pattern present with the input and interpret the results according to it.
Supervised learning models are normally used in two problems. Regression problem and
classification problem other two different problems is being solved by using the supervisor
learning approaches. Aggression problems deal with the data that is unstructured in nature. Most
of the real time data are unstructured and unbalanced in nature. Learning algorithm is used to
produce continuously wearing data to analyse unique pattern present in IT. Certain patterns are
relatively helpful to make the analysis and Decision Process easier.

Steps involved in supervisor learning process


First thing to determine the type of the data set. The collected data set need to be labelled based
on the training data. Further split the dataset into training data, testing data and validation data.
The machine learning systems contains 70% of training data 15% of testing data and 15% of
validation data. This dataset are trained with respect to the features extracted from the future
extraction techniques adapted to the machine learning algorithm. Feature extraction is a vital
process in machine learning algorithm that support the decision making process. Execute the
algorithm on training data set. The model created using the machine learning algorithm need to
be tested with the training data set. If the performance of the model is adaptively higher
comparing with other existing systems then the model need to be considered for evaluation. For
the testing data set need to be evaluated using the training model and further performance analysis
can be vibrator through accuracy Precision recall and fonts code sensitivity.

Advantages of supervised learning


Supervised learning algorithms can give partial information about the outcome to be
experienced. Since the supervised learning algorithm provides labeling of classes to be
validator the supervised model solve various real time issues with respect to the labelled
data such as anomaly detection fraud evaluation I am deduction filtering Malware
detection etc.

With the help of supervisor learning process the model can predict the output of the given
pattern of data effectively. The disadvantages of supervisor learning process is always
relay on the handling work for Complex task given by the Real Time systems. Supervised
learning process cannot predict different output if the labeling is not properly assigned in
supervised learning process in a data need to be given for classification else the
classification process itself deviated from the complete analysis.

Applications of supervisor learning model


In case of fryer experience with the supervised learning algorithms of existing
frameworks the algorithm used for pattern pattern matching is effectively utilized. If there
is no proper input provided with the system then the classification process is getting
deviated not recommended for Complex inputs

1.3.6 Unsupervised learning


Unsupervised learning model is a kind of machine learning algorithm in which the
labeling are not given to analyze the real pattern present in the input data. The instructor
the data is given to the machine learning algorithm and for the analysis module need to
conclude the unique pattern present with the input data set. And like other pattern
recognition process and classification algorithms the kind of supervisor learning is
effectively used to for real time applications and complex problem solving systems.

Figure 1. 2 Unsupervised Learning Model

Unsupervised learning models are not frequently used for regression problems and
classification issues. Unsupervised learning model required to structure the data into a
common format from the unstructured format. The combination need to procure unique
pattern present in the data. On the other hand real time applications does not depend upon
the data set available in such cases and supervised learning model cannot produce Highly
Effective outcome hands learning model need to be tune.

Clustering is the process of converging similar kind of objects present in the group of
input data. Clustering is a group of input features and their similarities present between
the data set. Association is a process involved in clustering process where the relativity
between the data are randomly studied. Real time example such as purchasing similar
product from in customers provides different feedbacks and different purchasing plans
that produces the uniqueness present in the transaction details. Learning algorithms such
as K means of frequently used in many applications

Some of the main reasons to define the importance of on supervised learning process are
it is helpful to analyze the structure data and finding out the unique structure present in
the data set. Unsupervised learning process is similar to the top human brain where it
consideres so many constraints on making valid decisions. And supervised the learning
process is closely related to artificial intelligence since there is no guidance for the system
to make decision. The system need to take the pattern related decision accurately based
on the learning process involved in it.

The steps involved in a and supervised learning process consider the input data set and
divide the data set into training data testing data validation data etc. The real world data
are not always depend on the problem to be solved sometimes the data set need to be
analyzed to predict the problem present in it.

The advantages of unsupervised learning process considered Complex task comparing


with the supervised learning process since that is no libeling of data is available in the
classification process need more learning procedures. Unsupervised learning is preferable
in the cases where similar pattern of occurrences of frequently repeated throat the data.

The disadvantages of unsupervised learning process are the time taken to learn the
complete data set for Complex inputs. Most of the real time data sets are unbalanced in
nature. It takes adequate time to learn the pattern present in it. Many cases if the similar
pattern is not available within the data set then the answer learning process takes more
relativity extraction process hence the propagation time increases Applications of
unsupervised learning model is completely Complex in computations since there is no
labeled data are available to predict.

1.3.7 Reinforcement learning


Most of the real time applications are dynamic in nature. From the dynamic environment
the data collector are completely variable and keep on getting we read based on the
changes occurring in the environment. In order to achieve pattern reorganization with
respect to the target a particular feedback is required to continuously very the analysis
process based on the dynamically changing environment. Learning process is a dynamic
method where the system process is getting keep on engaged and change the weight of
the analysis based on the feedback coming up from the output. Rainfalls to learning
process is also used for decoration as well as [Link] randomly distributed data
is clustered with reference to the unique inputs, to form the classification shown in
Figure1.3.

Figure 1.3 Reinforcement Learning


Applications of Reinforcement Learning
The applications of rainforest to learning includes the following things that need to be
discussed such as maximum performance of the system is implemented to reach certain
content. In manufacturing industries reinforced learning algorithm for used to automate
the manual process hens the changes occurring in the system need to be updated by the
rainfalls to learning cycle. In inventory management in order to reduce the transmit time
packing and unfaking poster learning process is used. Using Power Analysis systems the
load and distribution capability are keep on changing with respect to the dynamic
environment there the usage of free and force to learning is highly applicable.
In the robotic applications where robotic navigation robotic soccer and many robotic
controls are continuously understood by the rain forced learning technique.
Reinforced learning systems are adaptive in factory process and machine control
telecommunication self-healing networks automated vehicles helicopters etc. In most of
the gaming applications where the changing options completely very the outcome of the
game free and forced learning algorithms utilised.

In chemical industry in order to optimise the amount of chemical process need to be read
according to the changes in the environment and demanded chemical combinations based
on Rain first learning method specific combinations are created.
In manufacturing industries various automobile companies utilise deep learning
reinforced learning process to study the changing environment and further adopt the
robotic system to the Dynamic input full stop the controls are completely we read based
on the dynamic changes occurring at the input.
In financial sectors the prediction of future financial growth and fall are highly
demandable. The enforced learning process completely learn the history of financial
growth sudden rice and sudden fall conditions and determine the outcome predicted using
reinforce to learning method.

Types of machine learning based classification algorithms

Classification algorithms are used to divide the input data into two different forms of
outcomes.
Classifications are divided into linear method and nonlinear method.
Linear method have the inputs that very with respect to Linear changes in the input
example Logistic regression support vector machine etc.

Nonlinear models involved in handling the unstructured data set sometimes it is also
called as an supervised model. Some of the nonlinear machine learning models are K
nearest neighbour algorithm kernel SVM Nav base algorithm decision tree classification
random forest algorithm etc

1.3.8 Logistic Regression in Machine Learning


The Logistic regression in machine learning algorithm is relatively used algorithm bar
supervised technique is involved. Reduced to predict categrical based outcome the depend
on the input at some cases based on the labelled data. The label the data is touching but
the dependent variable given as a short hint. Logistic regression predict output of the
dependent variable and the classification output may belongs to logical 1 and logical zero.
To different outcomes hence it is also called as binary classification technique.
Figure 1.4 Logistic regression curves

Logistic regression is completely similar to than that of solving regulation problems


where Logistic regression is used to solve various classification techniques involved in
IT of the real time data sets are not enough to evaluate using Logistic regression. Hence
fitting of Logistic regression for real time data set is not completely feasible fifth of based
on the pattern present in the input data set the binary classification are nonlinear
classification need to be decided. Logistic regression is significant are its aggregate the
given data set into binary format and further divide the data set into two types of groups.
Sometimes logical regression is used to transform the input data and for the divide the
data into useful data and useful data. The Logistic regression curve depends on the two
dress where the escrow cross that restored value and provides the Useful information
beyond the threshold level are below the sold level.

1.3.9 Support Vector Machine Algorithm

Support vector machine(SVM) is one of the popular supervised learning algorithm this
used for classification problem as well as regression problem. It is ultimately used for
classification problems in machine learning algorithms. The goal of the SVM model is to
provide a best feasible boundary line that divides the input data into two different category
of classes. In many cases multi class SVM also utilized where the hyperplain has one are
more threshold levels. Extreme vectors by while creating the hyperplane. These extreme
cases handle the supporter machine algorithm in frequently involved data set. Figure
shows the basic structure of support factor machine algorithm will stop

Figure 1.5 Support vector machine curves

1.3.10 K-Nearest Neighbor (KNN) Algorithm


one of the simple form of machine learning algorithm in which supervised learning
technique is initially used. Depends on the input data group of training data and labels
provided with the model. K-nearest neighbour algorithm find out the similarity between
the training data and the testing data and evaluate the relative label associated with the
training data. This kind of algorithm are normally used for classification process in which
K nearest neighbour algorithm also utilized insert and regression problems. Depend on
any of the input features patterns. It is also called as lazy Lon algorithm since the training
data always relay on the input label provided to eat. Kane and algorithm after the training
phase stores various kinds of relativity data into the memory such that it can compare the
data with the future inputs

1.3.11 Kernel SVM

Analyser specific function used in support vector machine algorithm in order to optimise
the problem solving process. It provides short procedures for Complex calculations. The
amazing thing associated with kernel is that it can go to higher level of analysis process
based on smoothening the input data. The kernel support the term machine goes to infinite
number of dimensions using the associated kernels. Sometimes the support rectang
machine device the given input data as per the problem s scenario between the
hyperplains.

A kernel helps to divide the input data into the hyperplane such that it leaves only the
relativity between the data set of. Kernels are also helpful in nonlinear problems where
multiple classification is feasible. Support rectang machines are normally used in binary
classification technique in which the hyperplains divide the input data set into positive
plane are negative plane. Where is most of the real time data are unbalanced and
unstructured in nature. Karnal SVM device the input data collector from the real time data
set and further divide the hyperplane between multiple classes of data full stop are fitting
is one of the issue that can happen in the support formation algorithm where the number
of features associated with the input data are not linear in nature. This problem need to be
rectified by choosing right kernel with the support that our machine .

Supporter machine depends on the radial basis function that acts as a activation
function for the algorithm. Smaller the data set then the chances of higher accuracy.
The simple kernels users linear and polynomial data exchange is encouraged here.
1.3.12 Naïve Bayes Classifier Algorithm

Naive Bayes Classifier Algorithm

Maybe algorithm is a kind of supervised learning algorithm in which the basic


process is depend upon the base theorem for solving multiple classification issues. It
is mainly used in text related classification process where it includes high
dimensional training data set. Is considered as one of the robust classifier for text
analysis and helpful for handling the nonlinear data set. It is a probably stick
classifier where the outcome depends upon the probability of occurrence between
the input data and the training data. Some popular algorithms utilised maybe as a
supporting classify your to determine spam filtration sentimental analysis Mall were
detection and classification

1.3.13 Decision tree


Decision tree acts as a group of various levels of smaller decisions all located together to
make high level of classification and regression result. It probably depends upon the
pattern of the input data in which the tree structured classified classifies the internal nodes
based on the leaf node under tree note. In decision tree classify your decision node and
the leaf node are completely used to make decisions the final outcome of the decision tree
algorithm is nothing but the accumulation of smaller decisions all together associated with
the branches. The decision tree algorithm performance is measured based on the graphical
approach and the accuracy of the system sensitivity level and decision making capability.

1.3.14 Random Forest Algorithm


Figure 1.6 Evaluating a Classification model:

Cross-Entropy Loss:

Cross entropylos it is used to available the performance of the classifier which is being
used for the presented analysis. Most of the machine learning algorithm depend on the
statistical analysis part where sums out of a statistical measure need to be validate to
consider the predicted results. The cross entropy loss is determined to check the outcome
and the expected result for the higher accuracy of the model for binary classification cross
entropy need to be used.

For Binary classification, cross-entropy can be calculated as:

?(ylog(p)+(1?y)log(1?p))

Where y= Actual output, p= predicted output.

Confusion Matrix:
Analysis model in machine learning and deep learning algorithm is based on the expected
result versus The predicted result. The performance of the system need to be unless to the
any kind of statistical measures such as error rate true positive rate false posturate to
negative rate false negative rate etc. The number of correctly classified parameters need
to be com part with the incorrect parameters.

1.3.15 Challenges of Machine Learning

One of the prime challenges in machine learning algorithm are it is a subfield of artificial
intelligence hence the number of features used in machine learning algorithm need to be
improved to apply the machine learning models into artificial intelligence frameworks.
The features election is highly sensitive hence improved version of features selection
process is recommended. That wants but admachine learning Technology as improved
various effective results towards Technology such that few of them are not related to the
results of single machine learning algorithms are recommended.

Machine learning algorithms are recent trend in which the scientific computing need to
be done. Artificial intelligent frameworks are highly depend upon the machine learning
algorithm and deep learning algorithm. In spite of super intelligence model and strong
artificial intelligence model the capability of machine learning algorithm and hybriding
the machine learning algorithm together to form high accuracy is recommended. Various
medical analysis data analysis singularity formulation are implemented with the help of
autonomous validation of data using artificial intelligence. The impact of Artificial
Intelligence and its growth is enabled to apply the patient monitoring systems with recent
innovations. Medical data are highly sensitive information collect the directly from the
patience. This information are process effectively in order to identify the required amount
of clarity in the form of data pattern values etc. The impact of artificial intelligence
enabled various applications get involved with machine learning algorithms. The
presented research work is focused on studying the various benefits along with machine
learning algorithm and deep learning algorithm for the innovative idea of deep learning
process with lightweight architecture is created.
Artificial intelligence algorithm are applied in various areas of automotive industry
where the human intervention need to be reduced. It doesn't mean that human
intervention is completely ignored by the evaluation of artificial intelligence on the
other hand the enormous growth of artificial intelligence in medical industry provide
capable solutions in automated way where manual interventions are reduced. Manual
errors are highly reduced and hence the evaluation of recent technology in patient
monitoring systems are highly recommended. Medical emergency are very sensitive that
need to be treated in a fast and precise way. Automatic detection prevention and
predictive analysis models are helpful to analyse the patient data and for the detect the
abnormalities in the yearly stages. The sensitive cases are handled by the technological
growth hence these kind of systems are highly recommended.

CASE STUDY WORKS

Case study work 1: Data analytics with Predictive analysis model / Virtual sensor

Predictive analytics with virtual sensors combines data analytics techniques and sensor
technology to forecast future outcomes based on historical data and real-time sensor inputs.
Virtual sensors are algorithm-based models that simulate physical sensor readings using data
from other sensors or sources. These models can predict values for conditions that are difficult or
costly to measure directly. By applying predictive analytics to virtual sensor data, businesses can
anticipate equipment failures, optimize maintenance schedules, enhance process efficiency, and
make informed decisions, leading to improved operational performance and reduced downtime.
This approach is widely used in industries such as manufacturing, agriculture, and energy

Case study work 2: Creation of EDA plots

The creation of Exploratory Data Analysis (EDA) plots is a vital project in data science, aimed
at uncovering insights, patterns, and relationships within datasets before advanced modeling. The
project involves the generation of various plots to understand data distributions, identify outliers,
detect correlations, and highlight missing data. Key plots such as histograms, box plots, scatter
plots, pair plots, and correlation heatmaps are used to visualize both univariate and multivariate
aspects of the data. Histograms provide a view of data distribution, while box plots offer insights
into data variability and potential outliers. Scatter plots and pair plots are essential for identifying
relationships between variables, and heatmaps are used to detect correlations. The project
emphasizes the importance of transforming raw data into visual formats, enabling better decision-
making and preparation for machine learning models. Additionally, the project may include data
preprocessing techniques, such as normalization and handling missing values, to ensure the
accuracy of insights. The outcome of EDA enhances the understanding of data structure and
guides subsequent analytical processes, leading to more effective data-driven strategies.

Case study work: 3: Vgg16 based Socket image classification (Data augmentation)
This project focuses on developing a socket image classification system using a VGG16-based
deep learning model with data augmentation techniques. The VGG16 model, a pre-trained
convolutional neural network known for its depth and simplicity, serves as the backbone for
classifying various types of electrical sockets. The system aims to distinguish between different
socket types, configurations, or conditions (such as damaged versus intact sockets) based on
image inputs.

Data augmentation plays a crucial role in enhancing the performance and generalization of the
model. By applying transformations such as rotation, flipping, scaling, brightness adjustment,
and cropping, the dataset is artificially expanded, thus improving the model's ability to recognize
patterns in varied real-world conditions. This technique also helps mitigate overfitting, ensuring
the model can accurately classify unseen images.

The project involves several key stages, including data collection (acquiring a labeled dataset of
socket images), preprocessing (resizing, normalization, and augmentation), and fine-tuning the
VGG16 model on the specific classification task. The pre-trained layers of VGG16, designed to
extract high-level image features, are leveraged, while the top layers are retrained to focus on
socket-specific features.

By integrating transfer learning and data augmentation, the project aims to achieve high accuracy
in socket classification. The outcome is a robust model capable of identifying different socket
types, which can be applied in manufacturing, quality control, and safety monitoring domains.

Case study work 4 : Stress detection and analysis

The "Stress Detection and Analysis" project aims to develop a comprehensive system for
monitoring and analyzing human stress levels using physiological and behavioral data. This
project integrates various data sources such as heart rate, skin conductance, facial expressions,
voice modulation, and physical activity to detect stress in real-time. Advanced machine learning
algorithms are applied to analyze these signals, identifying patterns indicative of stress. The
system may utilize wearable sensors or smartphone applications for continuous data collection,
ensuring non-intrusive and real-time monitoring.

The core objective of the project is to build a model that can accurately classify stress levels and
provide actionable insights into its triggers. This involves developing features from raw
physiological signals and applying algorithms like Support Vector Machines (SVM), Random
Forest, or deep learning models for classification. The project also emphasizes data
preprocessing, including filtering noise from sensor inputs, normalizing data, and handling
missing values to ensure high accuracy in predictions.

Additionally, the analysis phase includes identifying stress trends over time, correlating stress
levels with external factors (e.g., work environment, physical exertion), and offering
recommendations for stress management. The system could be expanded to provide personalized
stress relief suggestions, such as relaxation techniques, based on detected stress patterns. The
outcomes of the project are aimed at contributing to mental health research, improving workplace
well-being, and empowering individuals with tools for better stress management. The system’s
capability to continuously monitor stress and provide insights makes it valuable for applications
in healthcare, corporate wellness programs, and personal health management.

Case study work 5 : Word cloud for Big Data

The Word Cloud for Big Data project aims to visualize the most significant and frequently
occurring terms within large datasets, offering an intuitive overview of key topics, trends, and
patterns. In the context of big data, where traditional text analysis can be challenging due to sheer
volume and complexity, word clouds provide a simplified yet powerful visual representation.
This project involves extracting text from diverse big data sources such as social media feeds,
customer reviews, research papers, or log files, and processing it using Natural Language
Processing (NLP) techniques.
Key stages of the project include data collection, text cleaning (removal of stop words,
punctuation, and irrelevant terms), and term frequency analysis. The size of each word in the
cloud corresponds to its frequency or relevance in the dataset, with larger words representing
higher occurrence or importance. Advanced techniques such as stemming, lemmatization, and
filtering of domain-specific keywords can be employed to improve the accuracy and relevance
of the word cloud.

The primary goal of the project is to offer stakeholders a quick and insightful view of the most
pertinent themes within the dataset, facilitating decision-making, content summarization, and
trend detection. This project is particularly useful in applications like sentiment analysis, market
research, and trend monitoring, where rapid insights are essential. Furthermore, by leveraging big
data platforms and tools (such as Hadoop, Spark, or cloud-based data processing services), the
project ensures scalability and efficiency in processing vast amounts of text.

Case study work 6 : Phishing attack detection

The project on "Phishing Attack Detection Using a Machine Learning Hybrid Learning Model"
aims to develop an advanced system capable of accurately identifying phishing attacks in real-
time by leveraging the strengths of multiple machine learning algorithms. Phishing attacks, which
involve deceptive tactics to trick users into divulging sensitive information, are a major
cybersecurity threat. The hybrid learning model integrates both supervised and unsupervised
learning techniques to enhance detection capabilities. The model uses features extracted from
email content, URLs, and web pages, including keyword analysis, domain age, SSL certificates,
and IP address characteristics.

In this project, a combination of algorithms such as Random Forest, Support Vector Machine
(SVM), and Neural Networks are employed, with ensemble methods to improve prediction
accuracy. The supervised component focuses on classifying known phishing attacks, while the
unsupervised part identifies novel, unknown threats by recognizing anomalous patterns in the
data. Data preprocessing techniques, such as feature selection and dimensionality reduction, are
applied to improve model efficiency and reduce computational complexity.

The hybrid model is evaluated using performance metrics like accuracy, precision, recall, and
F1-score, compared against individual models to assess improvements in phishing detection. The
project also explores real-time implementation possibilities through the deployment of the hybrid
model in email filtering systems, web security protocols, and browser plugins. The ultimate goal
is to reduce false positives and improve the early detection of phishing attempts, providing robust
defense mechanisms in dynamic cybersecurity environments.

Case study work 7: Fake news detection

The "Fake News Detection Using Machine Learning" project aims to build a system that can
automatically classify news articles as either legitimate or fake, using advanced machine learning
techniques. In today's digital age, the rapid spread of misinformation poses significant challenges,
making fake news detection a critical application in media and communications. This project
involves collecting a large dataset of labeled news articles, including both fake and genuine news,
and then applying text processing techniques to extract meaningful features from the content,
such as word frequencies, n-grams, and sentiment analysis. The project focuses on the
implementation of machine learning algorithms, such as Logistic Regression, Support Vector
Machines (SVM), Decision Trees, and deep learning models like LSTM (Long Short-Term
Memory) networks. Natural Language Processing (NLP) techniques play a key role in
transforming raw text into structured data that machine learning models can process. Feature
engineering, such as the use of TF-IDF (Term Frequency-Inverse Document Frequency) and
word embeddings like Word2Vec or BERT, helps capture the linguistic and contextual nuances
of the news content.

The model's performance is evaluated using standard metrics such as accuracy, precision, recall,
and F1-score to ensure the system’s reliability in distinguishing fake news from authentic articles.
Techniques such as cross-validation and hyperparameter tuning are employed to optimize model
performance. Additionally, the project addresses the ethical considerations of biased datasets,
ensuring fairness and robustness. Upon completion, the project delivers a machine learning model
capable of real-time fake news detection, which can be integrated into social media platforms,
news aggregators, and other information dissemination systems to mitigate the spread of
misinformation and enhance content credibility.

Case study work 8 : Kidney disease analysis


The project focuses on kidney disease analysis using hypertunable machine learning models to
improve diagnostic accuracy and predictive capabilities. Chronic kidney disease (CKD) is a
growing global health issue, and early detection is crucial for effective treatment and
management. This project leverages machine learning techniques to develop a robust
classification model for detecting CKD, utilizing clinical datasets containing patient information
such as blood pressure, glomerular filtration rate (GFR), creatinine levels, and other biomarkers.

Hypertunable machine learning models, including algorithms like Random Forest, Support
Vector Machine (SVM), XGBoost, and neural networks, are employed to optimize model
performance. Hyperparameter tuning techniques such as Grid Search, Random Search, and
Bayesian Optimization are used to fine-tune model parameters, ensuring the best predictive
performance. The tuning process focuses on maximizing metrics such as accuracy, precision,
recall, and F1-score to balance false positives and false negatives effectively.

The project also includes data preprocessing steps such as handling missing values, normalizing
numerical features, and encoding categorical variables. EDA (Exploratory Data Analysis) is
conducted to identify data trends, outliers, and correlations among features. Feature selection
techniques, like Principal Component Analysis (PCA) and Recursive Feature Elimination (RFE),
are applied to reduce dimensionality and enhance model interpretability.

The hypertuned model is evaluated using cross-validation techniques, and the results are
compared with baseline models to assess improvements. The final model is designed to assist
healthcare professionals in diagnosing kidney disease more accurately and at an earlier stage,
potentially reducing the need for invasive procedures and improving patient outcomes. This
project has significant implications for the future of predictive healthcare analytics, with the
potential to be adapted for various other diseases.

Case study work 9: Groundnut oil adulteration analysis

The project on "Groundnut Oil Adulteration Analysis Using Machine Learning" focuses on
leveraging advanced machine learning techniques to detect and quantify adulteration in
groundnut oil, a common issue affecting quality and safety in the food industry. The objective is
to develop a robust, automated system capable of identifying adulterants—such as other
vegetable oils, synthetic substances, or contaminants—by analyzing various features of the oil.
The project begins with the collection of a diverse dataset comprising samples of pure groundnut
oil and adulterated variants. Key features for analysis include chemical composition, physical
properties, and spectroscopic data, which are extracted using techniques such as Gas
Chromatography (GC), Fourier Transform Infrared Spectroscopy (FTIR), or Near-Infrared
Spectroscopy (NIRS). Machine learning algorithms, such as Support Vector Machines (SVM),
Random Forest, and Convolutional Neural Networks (CNN), are then applied to this dataset to
build predictive models capable of distinguishing between pure and adulterated samples. These
models are trained on labeled data, where the presence and type of adulterants are known, and
validated using cross-validation techniques to ensure accuracy and robustness.

The project involves preprocessing steps such as normalization, feature selection, and
dimensionality reduction to improve model performance. Performance metrics, including
accuracy, precision, recall, and F1 score, are used to evaluate the effectiveness of the models.
The final system aims to provide a real-time, cost-effective solution for quality control in the food
industry, ensuring consumer safety and product integrity. Overall, this project demonstrates the
application of machine learning in enhancing food safety by providing a reliable method for
detecting adulteration in groundnut oil, thereby supporting industry compliance and protecting
consumer health.

Case study work 10: Nature inspired algorithms

The project on "Nature-Inspired Algorithms" explores the application of computational


techniques that mimic natural processes to solve complex optimization and problem-solving
tasks. These algorithms draw inspiration from biological, ecological, and physical systems,
leveraging their inherent strategies and behaviors to develop efficient solutions for various
challenges. Key nature-inspired algorithms include Genetic Algorithms (GAs), which simulate
evolutionary processes such as selection, crossover, and mutation; Particle Swarm Optimization
(PSO), inspired by the social behavior of bird flocks and fish schools; Ant Colony Optimization
(ACO), which mimics the foraging behavior of ants to find optimal paths; and Artificial Bee
Colony (ABC) algorithms, based on the foraging behavior of honeybees. These algorithms are
applied to a range of problems, from engineering design and logistics to machine learning and
data mining. The project involves designing and implementing these algorithms to address
specific optimization tasks, such as minimizing cost, maximizing efficiency, or finding optimal
configurations. It includes the development of computational models, parameter tuning, and
performance evaluation against benchmark problems to assess their effectiveness and efficiency.
By leveraging nature-inspired algorithms, the project aims to provide innovative solutions that
are both computationally efficient and adaptable to diverse problem domains. The project
highlights the potential of bio-inspired methods to tackle complex real-world challenges, offering
insights into how natural systems can inform and enhance computational problem-solving
strategies.

TESTING

INTRODUCTION:
Testing is a process used to help identi
fy the correctness, completeness and quality of developed computer software. With that in mind,
testing can never completely establish the correctness of computer software. There are many
approaches to software testing from using tools to automated testing, but effective testing of
complex products is essentially a process of investigation, not merely a matter of creating and
following rote procedure.

One definition of testing is "the process of questioning a product in order to evaluate it", where
the "questions" are things the tester tries to do with the product, and the product answers with its
behaviour in reaction to the probing of the tester. Although most of the intellectual processes of
testing are nearly identical to that of review or inspection, the word testing is connoted to mean
the dynamic analysis of the product putting the product through its paces.
The quality of the application can and normally does vary widely from system to system but some
of the common quality attributes include reliability, stability, portability, maintainability and
usability. Refer to the ISO standard ISO 9126 for a more complete list of attributes and criteria.

1. Testing is a process of executing a program with the intent of finding an error.


2. A good test case is one that has a high probability of finding an as yet undiscovered error.
3. A successful test is one that uncovers an as yet undiscovered
error.

Testing should systematically uncover different classes of errors in a minimum amount of time
and with a minimum amount of effort. A secondary benefit of testing is that it demonstrates that
the software appears to be working as stated in the specifications. The data collected through
testing can also provide an indication of the software's reliability and quality. But, testing cannot
show the absence of defect -- it can only show that software defects are present.

TYPE OF TESTING:

Manual Testing

Manual testing includes testing a software manually, i.e., without using any automated tool or
any script. In this type, the tester takes over the role of an end-user and tests the software to
identify any unexpected behaviour or bug. There are different stages for manual testing such as
unit testing, integration testing, system testing, and user acceptance testing.

Testers use test plans, test cases, or test scenarios to test a software to ensure the completeness of
testing. Manual testing also includes exploratory testing, as testers explore the software to identify
errors in it.

White-Box Testing:
White-box testing is the detailed investigation of internal logic and structure of the code. White-
box testing is also called glass testing or open-box testing. In order to perform white- box testing
on an application, a tester needs to know the internal workings of the code.
The tester needs to have a look inside the source code and find out which unit/chunk of the code
is behaving inappropriately.

The following table lists the advantages and disadvantages of white-box testing.

Advantages Disadvantages

As the tester has knowledge of the source Due to the fact that a skilled tester is needed to
code, it becomes very easy to find out which perform white-box testing, the costs are
type of data can help in testing the application increased.
effectively. Sometimes it is impossible to look into every
It helps in optimizing the code. nook and corner to find out hidden errors that
Extra lines of code can be removed which can may create problems, as many paths will go
bring in hidden defects. untested.
Due to the tester's knowledge about the code, It is difficult to maintain white-box testing, as it
maximum coverage is attained during test requires specialized tools like code analysers
scenario writing. and debugging tools.

Unit testing is a software development process in which the smallest testable parts of an
application, called units, are individually and independently scrutinized for proper operation.
Unit testing can be done manually but is often automated.

Test cases

The purpose of a test case is to describe how you intend to empirically verify that the software
being developed conforms to the specifications. In other words, you need to be able to show that
it can correctly carry out its intended functions. The test case should be written with enough
clarity and detail that it could be given to an independent tester and have the tests properly carried
out.

CONCLUSION
● The training covered the basic knowledge required for understanding artificial intelligence
frameworks.
● The detailed knowledge on IoT Thing Speak cloud is studied.
● The training provides in-depth knowledge on involvement of hardware to learn about
environment and creating a wireless sensor cloud is formulated.
● Various case studies relavent to data analytics, Image processing, Cloud configurations, AI and
Neural computing are studied.
● Further the training motivated us to enhance the knowledge on AI frameworks in cloud
configuration etc.

Common questions

Powered by AI

Supervised learning requires labeled data for training to predict outcomes based on input features and is effective for tasks like classification and regression. In contrast, unsupervised learning does not need labeled data and seeks to discover underlying patterns or groupings within the data. While supervised learning is good for anomaly detection and fraud evaluation, unsupervised learning excels in clustering and associative tasks, revealing intrinsic data patterns .

Google Colab offers several advantages for data analytics learning: it provides access to powerful computational resources like GPUs and TPUs, allowing efficient handling of large datasets and complex models. It also supports real-time collaboration and integrates seamlessly with Google Drive, facilitating shared developments and instant updates among team members .

Unsupervised learning for real-time applications demands significant computational resources and time due to the absence of labeled data. It requires extensive data learning to identify patterns accurately, which can be complex and time-consuming, particularly for unbalanced datasets. The lack of initial guidance makes it challenging to ensure timely and efficient data processing, often necessitating extended analysis time .

Logistic regression is unique in its approach to binary classification by estimating the probability of class membership, using a logistic function to transform input values. However, its assumptions, such as linearity between dependent and independent variables, limit its applicability in more complex, non-linear real-time applications, where more flexible algorithms, like decision trees or SVM, might be better suited .

Preprocessing improves model performance and interpretability by converting data into a clean, noise-reduced format. This involves handling missing values, feature scaling, and encoding categorical variables. By ensuring data is standardized, models can better learn from the data, leading to enhanced accuracy and generalizability of predictions. Proper preprocessing also ensures that the model's outputs are understandable and actionable .

Clustering in unsupervised learning helps by grouping data with similar characteristics, revealing inherent structure and commonalities in datasets. Association explores relationships and dependencies between grouped items. These processes allow for the discovery of patterns and associations in data without prior knowledge of label structures, facilitating comprehensive understanding and analysis of complex data sets .

ThingSpeak acts as a cloud service for real-time IoT data collection by providing an easy-to-use interface that collects, stores, and processes data from IoT devices. This enables real-time analytics and visualization. Potential use cases include environmental monitoring, health monitoring, industrial control, and vehicle fleet monitoring, where real-time data processing and insights are crucial .

Feature extraction in supervised learning is crucial as it transforms raw data into a set of features that encapsulate the essential information required for the predictive task. By focusing on the most informative attributes, it enhances the model's decision-making efficiency, enabling quicker and more accurate predictions. Feature extraction also reduces dimensionality, simplifying the model and potentially improving performance .

NumPy provides optimized support for arrays and matrices along with high-level mathematical functions which significantly enhance data manipulation and computation efficiency. It is foundational for data science because many other libraries use NumPy’s array object as their basis, allowing for compatibility and performance optimization in scientific computing tasks .

Power BI enhances data visualization by offering rich visualization tools that make complex data comprehensible. This clarity helps identify trends, patterns, and anomalies, leading to informed decision-making and effective data-driven strategies. By facilitating a clear visual representation, stakeholders can quickly grasp analytics insights, enhancing overall organizational performance .

You might also like