0% found this document useful (0 votes)
3 views751 pages

What Is Python

python

Uploaded by

Aastha Kumar
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views751 pages

What Is Python

python

Uploaded by

Aastha Kumar
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

What is Python?

Python is a popular programming language. It was created by Guido van


Rossum, and released in 1991.

It is used for:

 web development (server-side),


 software development,
 mathematics,
 system scripting.

What can Python do?


 Python can be used on a server to create web applications.
 Python can be used alongside software to create workflows.
 Python can connect to database systems. It can also read and modify
files.
 Python can be used to handle big data and perform complex
mathematics.
 Python can be used for rapid prototyping, or for production-ready
software development.

Why Python?

Python works on different platforms (Windows, Mac, Linux, Raspberry Pi, etc).

Python has a simple syntax similar to the English language.

Python has syntax that allows developers to write programs with fewer lines than some other
programming languages.

Python runs on an interpreter system, meaning that code can be executed as soon as it is written. This
means that prototyping can be very quick.

Python can be treated in a procedural way, an object-oriented way or a functional way.

Python Syntax compared to other programming


languages
 Python was designed for readability, and has some similarities to the
English language with influence from mathematics.
 Python uses new lines to complete a command, as opposed to other
programming languages which often use semicolons or parentheses.
 Python relies on indentation, using whitespace, to define scope; such
as the scope of loops, functions and classes. Other programming
languages often use curly-brackets for this purpose.
Execute Python Syntax
As we learned in the previous page, Python syntax can be executed by writing
directly in the Command Line:

Or by creating a python file on the server, using the .py file extension, and
running it in the Command Line:

C:\Users\Your Name>python [Link]

Python Indentation
Indentation refers to the spaces at the beginning of a code line.

Where in other programming languages the indentation in code is for


readability only, the indentation in Python is very important.

Python uses indentation to indicate a block of code.

Python will give you an error if you skip the indentation:


The number of spaces is up to you as a programmer, the most common use is four,
but it has to be at least one.
You have to use the same number of spaces in the same block of code, otherwise
Python will give you an error:

Python Variables
In Python, a variable is created when you assign a value to it

Python has no command for declaring a variable.

Comments
Python has commenting capability for the purpose of in-code documentation.

Comments start with a #, and Python will render the rest of the line as a
comment:
Statements
A computer program is a list of "instructions" to be "executed" by a
computer.

In a programming language, these programming instructions are


called statements.

In Python, a statement usually ends when the line ends. You do not need to use a
semicolon (;) like in many other programming languages (for example, Java or C).

Many Statements
Most Python programs contain many statements.

The statements are executed one by one, in the same order as they are
written:

From the example above, we have three statements:

1. print("Hello World!")
2. print("Have a good day.")
3. print("Learning Python is fun!")

The first statement is executed first (print "Hello World!").


Then the second statement is executed (print "Have a good day.").
And at last, the third statement is executed (print "Learning Python is fun!").

Semicolons (Optional, Rarely


Used)
Semicolons are optional in Python. You can write multiple statements on one
line by separating them with ; but this is rarely used because it makes it
hard to read:

However, if you put two statements on the same line without a separator (newline
or ;), Python will give an error:

Print Text

You have already learned that you can use the print() function to display text or
output values:
You can use the print() function as many times as you want. Each call prints text on a new line by
default:

Double Quotes

Text in Python must be inside quotes. You can use either " double quotes or ' single quotes:

If you forget to put the text inside quotes, Python will give an error:

Print Without a New Line


By default, the print() function ends with a new line.

If you want to print multiple words on the same line, you can use
the end parameter:

Note that we add a space after end=" " for better readability.

Print Numbers
You can also use the print() function to display numbers:

However, unlike text, we don't put numbers inside quotes:

You can also do math inside the print() function:

Mix Text and Numbers

You can combine text and numbers in one output by separating them with a comma:

Python Comments

Comments can be used to explain Python code.

Comments can be used to make the code more readable.

Comments can be used to prevent execution when testing code.

Creating a Comment
Comments starts with a #, and Python will ignore them:

Comments can be placed at the end of a line, and Python will ignore the rest of the
line:
A comment does not have to be text that explains the code, it can also be used to prevent Python from
executing code:

Multiline Comments

Python does not really have a syntax for multiline comments.

To add a multiline comment, you could insert a # for each line:

Or, not quite as intended, you can use a multiline string.

Since Python will ignore string literals that are not assigned to a variable, you can add a multiline string
(triple quotes) in your code, and place your comment inside it:

As long as the string is not assigned to a variable, Python will read the code, but then ignore it, and you
have made a multiline comment.

Variables
Variables are containers for storing data values.

Creating Variables
Python has no command for declaring a variable.

A variable is created the moment you first assign a value to it.

Variables do not need to be declared with any particular type, and can even
change type after they have been set.

Example
x = 4 # x is of type int
x = "Sally" # x is now of type str
print(x)

Casting
If you want to specify the data type of a variable, this can be done with
casting.
Example
x = str(3) # x will be '3'
y = int(3) # y will be 3
z = float(3) # z will be 3.0

Get the Type


You can get the data type of a variable with the type() function.

Single or Double Quotes?

String variables can be declared either by using single or double quotes:

Case-Sensitive
Variable names are case-sensitive.

Example
This will create two variables:

a = 4
A = "Sally"
#A will not overwrite a

Variable Names
A variable can have a short name (like x and y) or a more descriptive name
(age, carname, total_volume).

Rules for Python variables:

 A variable name must start with a letter or the underscore character


 A variable name cannot start with a number
 A variable name can only contain alpha-numeric characters and
underscores (A-z, 0-9, and _ )
 Variable names are case-sensitive (age, Age and AGE are three
different variables)
 A variable name cannot be any of the Python keywords.

 Example
 Legal variable names:
 myvar = "John"
my_var = "John"
_my_var = "John"
myVar = "John"
MYVAR = "John"
myvar2 = "John"

 Example
 Illegal variable names:

 2myvar = "John"
my-var = "John"
my var = "John"

 Many Values to Multiple


Variables
 Python allows you to assign values to multiple variables in one line:

 Example
 x, y, z = "Orange", "Banana", "Cherry"
print(x)
print(y)
print(z)

One Value to Multiple Variables


And you can assign the same value to multiple variables in one line:

Example
x = y = z = "Orange"
print(x)
print(y)
print(z)
Unpack a Collection
If you have a collection of values in a list, tuple etc. Python allows you to
extract the values into variables. This is called unpacking.

Example
Unpack a list:

fruits = ["apple", "banana", "cherry"]


x, y, z = fruits
print(x)
print(y)
print(z)

Output Variables
The print() function is often used to output variables.

Example
x = "Python is awesome"
print(x)

In the print() function, you output multiple variables, separated by a


comma:

Example
x = "Python"
y = "is"
z = "awesome"
print(x, y, z)

You can also use the + operator to output multiple variables:

Example
x = "Python "
y = "is "
z = "awesome"
print(x + y + z)

For numbers, the + character works as a mathematical operator:


Example
x = 5
y = 10
print(x + y)

In the print() function, when you try to combine a string and a number with
the + operator, Python will give you an error:

Example
x = 5
y = "John"
print(x + y)

The best way to output multiple variables in the print() function is to


separate them with commas, which even support different data types:

Example
x = 5
y = "John"
print(x, y)

Python - Global Variables


Global Variables
Variables that are created outside of a function (as in all of the examples in
the previous pages) are known as global variables.

Global variables can be used by everyone, both inside of functions and


outside.

Example
Create a variable outside of a function, and use it inside the function

x = "awesome"

def myfunc():
print("Python is " + x)

myfunc()
If you create a variable with the same name inside a function, this variable will be
local, and can only be used inside the function. The global variable with the same
name will remain as it was, global and with the original value.

Example
Create a variable inside a function, with the same name as the global
variable

x = "awesome"

def myfunc():
x = "fantastic"
print("Python is " + x)

myfunc()

print("Python is " + x)

The global Keyword


Normally, when you create a variable inside a function, that variable is local,
and can only be used inside that function.

To create a global variable inside a function, you can use


the global keyword.

Example
If you use the global keyword, the variable belongs to the global scope:

def myfunc():
global x
x = "fantastic"

myfunc()

print("Python is " + x)

Also, use the global keyword if you want to change a global variable inside
a function.

Example
To change the value of a global variable inside a function, refer to the
variable by using the global keyword:

x = "awesome"

def myfunc():
global x
x = "fantastic"

myfunc()

print("Python is " + x)

Python Data Types


Built-in Data Types
In programming, data type is an important concept.

Variables can store data of different types, and different types can do
different things.

Python has the following data types built-in by default, in these categories:

Text Type: str

Numeric Types: int, float, complex

Sequence Types: list, tuple, range

Mapping Type: dict

Set Types: set, frozen set

Boolean Type: bool

Binary Types: bytes, byte array, memory view

None Type: NoneType


Getting the Data Type
You can get the data type of any object by using the type() function:

Example
Print the data type of the variable x:

x = 5
print(type(x))

Setting the Data Type


In Python, the data type is set when you assign a value to a variable:

Example

x = "Hello World"

x = 20

x = 20.5

x = 1j

x = ["apple", "banana", "cherry"]


x = ("apple", "banana", "cherry")

x = range(6)

x = {"name" : "John", "age" : 36}

x = {"apple", "banana", "cherry"}

x = frozen set({"apple", "banana", "cherry"})

x = True

x = b"Hello"

x = byte array(5)

x = memory view(bytes(5))

x = None
Setting the Specific Data Type
If you want to specify the data type, you can use the following constructor
functions:

Example

x = str("Hello World")

x = int(20)

x = float(20.5)

x = complex(1j)

x = list(("apple", "banana", "cherry"))

x = tuple(("apple", "banana", "cherry"))

x = range(6)

x = dict(name="John", age=36)
x = set(("apple", "banana", "cherry"))

x = frozen set(("apple", "banana", "cherry"))

x = bool(5)

x = bytes(5)

x = byte array(5)

x = memory view(bytes(5))

Python Numbers
There are three numeric types in Python:

 int
 float
 complex

Variables of numeric types are created when you assign a value to them:

Example
x = 1 # int
y = 2.8 # float
z = 1j # complex

To verify the type of any object in Python, use the type() function:
Example
print(type(x))
print(type(y))
print(type(z))

Int
Int, or integer, is a whole number, positive or negative, without decimals, of
unlimited length.

Example
Integers:

x = 1
y = 35656222554887711
z = -3255522

print(type(x))
print(type(y))
print(type(z))

Float
Float, or "floating point number" is a number, positive or negative,
containing one or more decimals.

Example
Floats:

x = 1.10
y = 1.0
z = -35.59

print(type(x))
print(type(y))
print(type(z))

Float can also be scientific numbers with an "e" to indicate the power of 10.

Example
Floats:

x = 35e3
y = 12E4
z = -87.7e100

print(type(x))
print(type(y))
print(type(z))

Complex
Complex numbers are written with a "j" as the imaginary part:

Example
Complex:

x = 3+5j
y = 5j
z = -5j

print(type(x))
print(type(y))
print(type(z))

Type Conversion
You can convert from one type to another with the int(), float(),
and complex() methods:

Example
Convert from one type to another:
x = 1 # int
y = 2.8 # float
z = 1j # complex

#convert from int to float:


a = float(x)
#convert from float to int:
b = int(y)

#convert from int to complex:


c = complex(x)

print(a)
print(b)
print(c)

print(type(a))
print(type(b))
print(type(c))

Random Number
Python does not have a random() function to make a random number, but
Python has a built-in module called random that can be used to make random
numbers:

Example
Import the random module, and display a random number from 1 to 9:

import random

print([Link](1, 10))

Python Casting
Specify a Variable Type
There may be times when you want to specify a type on to a variable. This
can be done with casting. Python is an object-orientated language, and as
such it uses classes to define data types, including its primitive types.

Casting in python is therefore done using constructor functions:

 int() - constructs an integer number from an integer literal, a float


literal (by removing all decimals), or a string literal (providing the
string represents a whole number)
 float() - constructs a float number from an integer literal, a float
literal or a string literal (providing the string represents a float or an
integer)
 str() - constructs a string from a wide variety of data types, including
strings, integer literals and float literals

 Example
 x = int(1) # x will be 1
y = int(2.8) # y will be 2
z = int("3") # z will be 3

 Example
 Floats:
 x = float(1) # x will be 1.0
y = float(2.8) # y will be 2.8
z = float("3") # z will be 3.0
w = float("4.2") # w will be 4.2

 Example
 Strings:

 x = str("s1") # x will be 's1'


y = str(2) # y will be '2'
z = str(3.0) # z will be '3.0'

Python Strings
Strings
Strings in python are surrounded by either single quotation marks, or double
quotation marks.

'hello' is the same as "hello".

You can display a string literal with the print() function:

Example
print("Hello")
print('Hello')
Quotes Inside Quotes
You can use quotes inside a string, as long as they don't match the quotes
surrounding the string:

Example
print("It's alright")
print("He is called 'Johnny'")
print('He is called "Johnny"')

Assign String to a Variable


Assigning a string to a variable is done with the variable name followed by
an equal sign and the string:

Example
a = "Hello"
print(a)

Multiline Strings
You can assign a multiline string to a variable by using three quotes:

Example
You can use three double quotes:

a = """Lorem ipsum dolor sit amet,


consectetur adipiscing elit,
sed do eiusmod tempor incididunt
ut labore et dolore magna aliqua."""
print(a)

Or three single quotes:

Example
a = '''Lorem ipsum dolor sit amet,
consectetur adipiscing elit,
sed do eiusmod tempor incididunt
ut labore et dolore magna aliqua.'''
print(a)

Strings are Arrays


Like many other popular programming languages, strings in Python are
arrays of unicode characters.

However, Python does not have a character data type, a single character is
simply a string with a length of 1.

Square brackets can be used to access elements of the string.

Example
Get the character at position 1 (remember that the first character has the
position 0):

a = "Hello, World!"
print(a[1])

Looping Through a String


Since strings are arrays, we can loop through the characters in a string, with
a for loop.

Example
Loop through the letters in the word "banana":

for x in "banana":
print(x)

String Length
To get the length of a string, use the len() function.

Example
The len() function returns the length of a string:

a = "Hello, World!"
print(len(a))
Check String
To check if a certain phrase or character is present in a string, we can use
the keyword in.

Example
Check if "free" is present in the following text:

txt = "The best things in life are free!"


print("free" in txt)

Use it in an if statement:

Example
Print only if "free" is present:

txt = "The best things in life are free!"


if "free" in txt:
print("Yes, 'free' is present.")

Check if NOT
To check if a certain phrase or character is NOT present in a string, we can
use the keyword not in.

Example
Check if "expensive" is NOT present in the following text:

txt = "The best things in life are free!"


print("expensive" not in txt)

Use it in an if statement:

Example
print only if "expensive" is NOT present:

txt = "The best things in life are free!"


if "expensive" not in txt:
print("No, 'expensive' is NOT present.")
Python - Slicing Strings
Slicing
You can return a range of characters by using the slice syntax.

Specify the start index and the end index, separated by a colon, to return a
part of the string.

Example
Get the characters from position 2 to position 5 (not included):

b = "Hello, World!"
print(b[2:5])

Slice from the Start


By leaving out the start index, the range will start at the first character:

Example
Get the characters from the start to position 5 (not included):

b = "Hello, World!"
print(b[:5])

Slice to the End


By leaving out the end index, the range will go to the end:

Example
Get the characters from position 2, and all the way to the end:

b = "Hello, World!"
print(b[2:])

Negative Indexing
Use negative indexes to start the slice from the end of the string:
Example
Get the characters:

From: "o" in "World!" (position -5)

To, but not included: "d" in "World!" (position -2):

b = "Hello, World!"
print(b[-5:-2])

Python - Modify Strings


Python has a set of built-in methods that you can use on strings.

Upper Case
Example
The upper() method returns the string in upper case:

a = "Hello, World!"
print([Link]())

Lower Case
Example
The lower() method returns the string in lower case:

a = "Hello, World!"
print([Link]())

Remove Whitespace
Whitespace is the space before and/or after the actual text, and very often
you want to remove this space.

Example
The strip() method removes any whitespace from the beginning or the
end:

a = " Hello, World! "


print([Link]()) # returns "Hello, World!"

Replace String
Example
The replace() method replaces a string with another string:

a = "Hello, World!"
print([Link]("H", "J"))

Split String
The split() method returns a list where the text between the specified
separator becomes the list items.

Example
The split() method splits the string into substrings if it finds instances of
the separator:

a = "Hello, World!"
print([Link](",")) # returns ['Hello', ' World!']

Python - String
Concatenation
String Concatenation
To concatenate, or combine, two strings you can use the + operator.

Example
Merge variable a with variable b into variable c:

a = "Hello"
b = "World"
c = a + b
print(c)

Example
To add a space between them, add a " ":

a = "Hello"
b = "World"
c = a + " " + b
print(c)

Python - Format - Strings


String Format
As we learned in the Python Variables chapter, we cannot combine strings
and numbers like this:

Example
age = 36
#This will produce an error:
txt = "My name is John, I am " + age
print(txt)

But we can combine strings and numbers by using f-strings or


the format() method!

F-Strings
F-String was introduced in Python 3.6, and is now the preferred way of
formatting strings.

To specify a string as an f-string, simply put an f in front of the string literal,


and add curly brackets {} as placeholders for variables and other operations.

Example
Create an f-string:
age = 36
txt = f"My name is John, I am {age}"
print(txt)

Placeholders and Modifiers


A placeholder can contain variables, operations, functions, and modifiers to
format the value.

Example
Add a placeholder for the price variable:

price = 59
txt = f"The price is {price} dollars"
print(txt)

A placeholder can include a modifier to format the value.

A modifier is included by adding a colon: followed by a legal formatting type,


like .2f which means fixed point number with 2 decimals:

Example
Display the price with 2 decimals:

price = 59
txt = f"The price is {price:.2f} dollars"
print(txt)

A placeholder can contain Python code, like math operations:

Example
Perform a math operation in the placeholder, and return the result:

txt = f"The price is {20 * 59} dollars"


print(txt)

Python - Escape Characters


Escape Character
To insert characters that are illegal in a string, use an escape character.

An escape character is a backslash \ followed by the character you want to


insert.

An example of an illegal character is a double quote inside a string that is


surrounded by double quotes:

Example
You will get an error if you use double quotes inside a string that is
surrounded by double quotes:

txt = "We are the so-called "Vikings" from the north."

To fix this problem, use the escape character \":

Example
The escape character allows you to use double quotes when you normally
would not be allowed:

txt = "We are the so-called \"Vikings\" from the north."

Escape Characters
Other escape characters used in Python:

Code Result

\' Single Quote

\\ Backslash

\n New Line
\r Carriage Return

\t Tab

\b Backspace

\f Form Feed

\ooo Octal value

\xhh Hex value

Python - String Methods


String Methods
Python has a set of built-in methods that you can use on strings.

Method Description

capitalize() Converts the first character to upper case

casefold() Converts string into lower case


center() Returns a centered string

count() Returns the number of times a specified value occurs in a string

encode() Returns an encoded version of the string

endswith() Returns true if the string ends with the specified value

expandtabs() Sets the tab size of the string

find() Searches the string for a specified value and returns the position of w

format() Formats specified values in a string

format_map() Formats specified values in a string

index() Searches the string for a specified value and returns the position of w

isalnum() Returns True if all characters in the string are alphanumeric

isalpha() Returns True if all characters in the string are in the alphabet
isascii() Returns True if all characters in the string are ascii characters

isdecimal() Returns True if all characters in the string are decimals

isdigit() Returns True if all characters in the string are digits

isidentifier() Returns True if the string is an identifier

islower() Returns True if all characters in the string are lower case

isnumeric() Returns True if all characters in the string are numeric

isprintable() Returns True if all characters in the string are printable

isspace() Returns True if all characters in the string are whitespaces

istitle() Returns True if the string follows the rules of a title

isupper() Returns True if all characters in the string are upper case

join() Joins the elements of an iterable to the end of the string


ljust() Returns a left justified version of the string

lower() Converts a string into lower case

lstrip() Returns a left trim version of the string

maketrans() Returns a translation table to be used in translations

partition() Returns a tuple where the string is parted into three parts

replace() Returns a string where a specified value is replaced with a specified v

rfind() Searches the string for a specified value and returns the last position

rindex() Searches the string for a specified value and returns the last position

rjust() Returns a right justified version of the string

rpartition() Returns a tuple where the string is parted into three parts

rsplit() Splits the string at the specified separator, and returns a list
rstrip() Returns a right trim version of the string

split() Splits the string at the specified separator, and returns a list

split lines() Splits the string at line breaks and returns a list

starts with() Returns true if the string starts with the specified value

strip() Returns a trimmed version of the string

swap case() Swaps cases, lower case becomes upper case and vice versa

title() Converts the first character of each word to upper case

translate() Returns a translated string

upper() Converts a string into upper case

zfill() Fills the string with a specified number of 0 values at the beginning

Python Booleans
booleans represent one of two values: True or False.

Boolean Values
In programming you often need to know if an expression is True or False.

You can evaluate any expression in Python, and get one of two
answers, True or False.

When you compare two values, the expression is evaluated and Python
returns the Boolean answer:

Example
print(10 > 9)
print(10 == 9)
print(10 < 9

When you run a condition in an if statement, Python returns True or False:

Example
Print a message based on whether the condition is True or False:

a = 200
b = 33

if b > a:
print("b is greater than a")
else:
print("b is not greater than a")

Evaluate Values and Variables


The bool() function allows you to evaluate any value, and give
you True or False in return,

Example
Evaluate a string and a number:
print(bool("Hello"))
print(bool(15))

Example
Evaluate two variables:

x = "Hello"
y = 15

print(bool(x))
print(bool(y))

Most Values are True


Almost any value is evaluated to True if it has some sort of content.

Any string is True, except empty strings.

Any number is True, except 0.

Any list, tuple, set, and dictionary are True, except empty ones.

Example
The following will return True:

bool("abc")
bool(123)
bool(["apple", "cherry", "banana"])

Some Values are False


In fact, there are not many values that evaluate to False, except empty
values, such as (), [], {}, "", the number 0, and the value None. And of
course the value False evaluates to False.

Example
The following will return False:

bool(False)
bool(None)
bool(0)
bool("")
bool(())
bool([])
bool({})

One more value, or object in this case, evaluates to False, and that is if you
have an object that is made from a class with a __len__ function that
returns 0 or False:

Example
class my class():
def __len__(self):
return 0

myobj = myclass()
print(bool(myobj))

Functions can Return a Boolean


You can create functions that returns a Boolean Value:

Example
Print the answer of a function:

def myFunction() :
return True

print(myFunction())

You can execute code based on the Boolean answer of a function:

Example
Print "YES!" if the function returns True, otherwise print "NO!":

def myFunction() :
return True

if myFunction():
print("YES!")
else:
print("NO!")
Python also has many built-in functions that return a boolean value, like
the isinstance() function, which can be used to determine if an object is of
a certain data type:

Example
Check if an object is an integer or not:

x = 200
print(isinstance(x, int))

Python Operators
Python Operators
Operators are used to perform operations on variables and values.

In the example below, we use the + operator to add together two values:

Example
print (10 + 5)

Although the + operator is often used to add together two values, like in the
example above, it can also be used to add together a variable and a value,
or two variables:

Example
sum1 = 100 + 50 # 150 (100 + 50)
sum2 = sum1 + 250 # 400 (150 + 250)
sum3 = sum2 + sum2 # 800 (400 + 400)

Python divides the operators in the following groups:

 Arithmetic operators
 Assignment operators
 Comparison operators
 Logical operators
 Identity operators
 Membership operators
 Bitwise operators
Python Arithmetic Operators
Arithmetic Operators
Arithmetic operators are used with numeric values to perform common
mathematical operations:

Operator Name Example

+ Addition x+y

- Subtraction x-y

* Multiplication x*y

/ Division x/y

% Modulus x%y

** Exponentiation x ** y

// Floor division x // y
Examples
Here is an example using different arithmetic operators:

Example
x = 15
y = 4

print(x + y)
print(x - y)
print(x * y)
print(x / y)
print(x % y)
print(x ** y)
print(x // y)

Division in Python
Python has two division operators:

 / - Division (returns a float)


 // - Floor division (returns an integer)
Example
Division always returns a float:

x = 12
y = 5

print(x / y)

Example
Floor division always returns an integer.

It rounds DOWN to the nearest integer:

x = 12
y = 5

print(x // y)
Python Assignment
Operators
Assignment Operators
Assignment operators are used to assign values to variables:

Operator Example Same As

= x=5 x=5

+= x += 3 x=x+3

-= x -= 3 x=x-3

*= x *= 3 x=x*3

/= x /= 3 x=x/3

%= x %= 3 x=x%3

//= x //= 3 x = x // 3
**= x **= 3 x = x ** 3

&= x &= 3 x=x&3

|= x |= 3 x=x|3

^= x ^= 3 x=x^3

>>= x >>= 3 x = x >> 3

<<= x <<= 3 x = x << 3

:= print(x := 3) x=3
print(x)

The Walrus Operator


Python 3.8 introduced the := operator, known as the "walrus operator". It
assigns values to variables as part of a larger expression:

Example
The count variable is assigned in the if statement, and given the value 5:

numbers = [1, 2, 3, 4, 5]

if (count := len(numbers)) > 3:


print(f"List has {count} elements")
Python Comparison
Operators
Comparison Operators
Comparison operators are used to compare two values:

Operator Name Example

== Equal x == y

!= Not equal x != y

> Greater than x>y

< Less than x<y

>= Greater than or equal to x >= y

<= Less than or equal to x <= y

Examples
Comparison operators return True or False based on the comparison:

Example
x = 5
y = 3

print(x == y)
print(x != y)
print(x > y)
print(x < y)
print(x >= y)
print(x <= y)

Chaining Comparison Operators


Python allows you to chain comparison operators:

Example
x = 5

print(1 < x < 10)

print(1 < x and x < 10)

Python Logical Operators


Logical Operators
Logical operators are used to combine conditional statements:

Operator Description Example

and Returns True if both statements are true x < 5 and x

or Returns True if one of the statements is x < 5 or x <


true
not Reverse the result, returns False if the not(x < 5 an
result is true

Examples
Example
Test if a number is greater than 0 and less than 10:

x = 5

print(x > 0 and x < 10)

Example
Test if a number is less than 5 or greater than 10:

x = 5

print(x < 5 or x > 10)


Example
Reverse the result with not:

x = 5

print(not(x > 3 and x < 10))

Python Identity Operators


Identity Operators
Identity operators are used to compare the objects, not if they are equal, but
if they are actually the same object, with the same memory location:

Operator Description Example


is Returns True if both variables are the same x is y
object

is not Returns True if both variables are not the x is not y


same object

Examples
Example
The is operator returns True if both variables point to the same object:

x = ["apple", "banana"]
y = ["apple", "banana"]
z = x

print(x is z)
print(x is y)
print(x == y)

Example
The is not operator returns True if both variables do not point to the same
object:

x = ["apple", "banana"]
y = ["apple", "banana"]

print(x is not y)

Difference Between is and ==


 is - Checks if both variables point to the same object in memory
 == - Checks if the values of both variables are equal
Example
x = [1, 2, 3]
y = [1, 2, 3]
print(x == y)
print(x is y)

Python Membership
Operators
Membership Operators
Membership operators are used to test if a sequence is presented in an
object:

Operator Description Exam

in Returns True if a sequence with the specified value is x in y


present in the object

not in Returns True if a sequence with the specified value is x not


not present in the object

Examples
Example
Check if "banana" is present in a list:

fruits = ["apple", "banana", "cherry"]

print("banana" in fruits)

Example
Check if "pineapple" is NOT present in a list:
fruits = ["apple", "banana", "cherry"]

print("pineapple" not in fruits)

Membership in Strings
The membership operators also work with strings:

Example
text = "Hello World"

print("H" in text)
print("hello" in text)
print("z" not in text)

Python Bitwise Operators


Bitwise Operators
Bitwise operators are used to compare (binary) numbers:

Operat Name Description


or

& AND Sets each bit to 1 if both bits are 1

| OR Sets each bit to 1 if one of two bits is 1

^ XOR Sets each bit to 1 if only one of two bits is 1


~ NOT Inverts all the bits

<< Zero fill left Shift left by pushing zeros in from the right and let the leftmos
shift bits fall off

>> Signed right Shift right by pushing copies of the leftmost bit in from the lef
shift and let the rightmost bits fall off

Examples
Example
The & operator compares each bit and set it to 1 if both are 1, otherwise it is
set to 0:

print (6 & 3)

The binary representation of 6 is 0110


The binary representation of 3 is 0011

Then the & operator compares the bits and returns 0010, which is 2 in
decimal.

Example
The | operator compares each bit and set it to 1 if one or both is 1, otherwise
it is set to 0:

print (6 | 3)

The binary representation of 6 is 0110


The binary representation of 3 is 0011

Then the | operator compares the bits and returns 0111, which is 7 in
decimal.

Example
The ^ operator compares each bit and set it to 1 if only one is 1, otherwise
(if both are 1 or both are 0) it is set to 0:

print (6 ^ 3)

The binary representation of 6 is 0110


The binary representation of 3 is 0011

Then the ^ operator compares the bits and returns 0101, which is 5 in
decimal.

Python Operator Precedence


Operator Precedence
Operator precedence describes the order in which operations are performed.

Example
Parentheses has the highest precedence, meaning that expressions inside
parentheses must be evaluated first:

print((6 + 3) - (6 + 3))

Example
Multiplication * has higher precedence than addition +, and therefore
multiplications are evaluated before additions:

print(100 + 5 * 3)

Precedence Order
The precedence order is described in the table below, starting with the
highest precedence at the top:

Operator Description
() Parentheses

** Exponentiation

+x -x ~x Unary plus, unary minus, and bitwise NOT

* / // % Multiplication, division, floor division, and modulus

+ - Addition and subtraction

<< >> Bitwise left and right shifts

& Bitwise AND

^ Bitwise XOR

| Bitwise OR

== != > >= < <= is is Comparisons, identity, and membership operators


not in not in

not Logical NOT


and AND

or OR

Left-to-Right Evaluation
If two operators have the same precedence, the expression is evaluated
from left to right.

Example
Addition + and subtraction - has the same precedence, and therefore we
evaluate the expression from left to right:

print (5 + 4 - 7 + 3)

Python Lists
mylist = ["apple", "banana", "cherry"]

List
Lists are used to store multiple items in a single variable.

Lists are one of 4 built-in data types in Python used to store collections of
data, the other 3 are Tuple, Set, and Dictionary, all with different qualities
and usage.

Lists are created using square brackets:

Example
Create a List:

this list = ["apple", "banana", "cherry"]


print(this list)
List Items
List items are ordered, changeable, and allow duplicate values.

List items are indexed, the first item has index [0], the second item has
index [1] etc.

Ordered
When we say that lists are ordered, it means that the items have a defined
order, and that order will not change.

If you add new items to a list, the new items will be placed at the end of the
list.

Changeable
The list is changeable, meaning that we can change, add, and remove items
in a list after it has been created.

Allow Duplicates
Since lists are indexed, lists can have items with the same value:

Example
Lists allow duplicate values:

thislist = ["apple", "banana", "cherry", "apple", "cherry"]


print(thislist)

List Length
To determine how many items a list has, use the len() function:
Example
Print the number of items in the list:

thislist = ["apple", "banana", "cherry"]


print(len(thislist))

List Items - Data Types


List items can be of any data type:

Example
String, int and boolean data types:

list1 = ["apple", "banana", "cherry"]


list2 = [1, 5, 7, 9, 3]
list3 = [True, False, False]

A list can contain different data types:

Example
A list with strings, integers and boolean values:

list1 = ["abc", 34, True, 40, "male"]

type()
From Python's perspective, lists are defined as objects with the data type
'list':

<class 'list'>
Example
What is the data type of a list?

mylist = ["apple", "banana", "cherry"]


print(type(mylist))

The list() Constructor


It is also possible to use the list() constructor when creating a new list.

Example
Using the list() constructor to make a List:

thislist = list(("apple", "banana", "cherry")) # note the double


round-brackets
print(thislist)

Python Collections (Arrays)


There are four collection data types in the Python programming language:

 List is a collection which is ordered and changeable. Allows duplicate


members.
 Tuple is a collection which is ordered and unchangeable. Allows
duplicate members.
 Set is a collection which is unordered, unchangeable*, and unindexed.
No duplicate members.
 Dictionary is a collection which is ordered** and changeable. No
duplicate members.
When choosing a collection type, it is useful to understand the properties of that
type. Choosing the right type for a particular data set could mean retention of
meaning, and, it could mean an increase in efficiency or security.

Python - Access List Items


Access Items
List items are indexed and you can access them by referring to the index
number:

Example
Print the second item of the list:

thislist = ["apple", "banana", "cherry"]


print(thislist[1])
Negative Indexing
Negative indexing means start from the end

-1 refers to the last item, -2 refers to the second last item etc.

Example
Print the last item of the list:

thislist = ["apple", "banana", "cherry"]


print(thislist[-1])

Range of Indexes
You can specify a range of indexes by specifying where to start and where to
end the range.

When specifying a range, the return value will be a new list with the
specified items.

Example
Return the third, fourth, and fifth item:

thislist =
["apple", "banana", "cherry", "orange", "kiwi", "melon", "mango"]
print(thislist[2:5])

By leaving out the start value, the range will start at the first item:

Example
This example returns the items from the beginning to, but NOT including,
"kiwi":

thislist =
["apple", "banana", "cherry", "orange", "kiwi", "melon", "mango"]
print(thislist[:4])

By leaving out the end value, the range will go on to the end of the list:

Example
This example returns the items from "cherry" to the end:

thislist =
["apple", "banana", "cherry", "orange", "kiwi", "melon", "mango"]
print(thislist[2:])

Range of Negative Indexes


Specify negative indexes if you want to start the search from the end of the
list:

Example
This example returns the items from "orange" (-4) to, but NOT including
"mango" (-1):

thislist =
["apple", "banana", "cherry", "orange", "kiwi", "melon", "mango"]
print(thislist[-4:-1])

Check if Item Exists


To determine if a specified item is present in a list use the in keyword:

Example
Check if "apple" is present in the list:

thislist = ["apple", "banana", "cherry"]


if "apple" in thislist:
print("Yes, 'apple' is in the fruits list")

Python - Change List Items


Change Item Value
To change the value of a specific item, refer to the index number:

Example
Change the second item:
thislist = ["apple", "banana", "cherry"]
thislist[1] = "blackcurrant"
print(thislist)

Change a Range of Item Values


To change the value of items within a specific range, define a list with the
new values, and refer to the range of index numbers where you want to
insert the new values:

Example
Change the values "banana" and "cherry" with the values "blackcurrant" and
"watermelon":

thislist = ["apple", "banana", "cherry", "orange", "kiwi", "mango"]


thislist[1:3] = ["blackcurrant", "watermelon"]
print(thislist)

If you insert more items than you replace, the new items will be inserted
where you specified, and the remaining items will move accordingly:

Example
Change the second value by replacing it with two new values:

thislist = ["apple", "banana", "cherry"]


thislist[1:2] = ["blackcurrant", "watermelon"]
print(thislist)

If you insert less items than you replace, the new items will be inserted
where you specified, and the remaining items will move accordingly:

Example
Change the second and third value by replacing it with one value:

thislist = ["apple", "banana", "cherry"]


thislist[1:3] = ["watermelon"]
print(thislist)

Insert Items
To insert a new list item, without replacing any of the existing values, we can
use the insert() method.

The insert() method inserts an item at the specified index:

Example
Insert "watermelon" as the third item:

thislist = ["apple", "banana", "cherry"]


[Link](2, "watermelon")
print(thislist)

Python - Add List Items


Append Items
To add an item to the end of the list, use the append() method:

Example
Using the append() method to append an item:

thislist = ["apple", "banana", "cherry"]


[Link]("orange")
print(thislist)

Insert Items
To insert a list item at a specified index, use the insert() method.

The insert() method inserts an item at the specified index:

Example
Insert an item as the second position:

thislist = ["apple", "banana", "cherry"]


[Link](1, "orange")
print(thislist)
Extend List
To append elements from another list to the current list, use
the extend() method.

Example
Add the elements of tropical to thislist:

thislist = ["apple", "banana", "cherry"]


tropical = ["mango", "pineapple", "papaya"]
[Link](tropical)
print(thislist)

The elements will be added to the end of the list.

Add Any Iterable


The extend() method does not have to append lists, you can add any
iterable object (tuples, sets, dictionaries etc.).

Example
Add elements of a tuple to a list:

thislist = ["apple", "banana", "cherry"]


thistuple = ("kiwi", "orange")
[Link](thistuple)
print(thislist)

Python - Remove List Items


Remove Specified Item
The remove() method removes the specified item.

Example
Remove "banana":
thislist = ["apple", "banana", "cherry"]
[Link]("banana")
print(thislist)

If there are more than one item with the specified value,
the remove() method removes the first occurrence:

Example
Remove the first occurrence of "banana":

thislist = ["apple", "banana", "cherry", "banana", "kiwi"]


[Link]("banana")
print(thislist)

Remove Specified Index


The pop() method removes the specified index.

Example
Remove the second item:

thislist = ["apple", "banana", "cherry"]


[Link](1)
print(thislist)

If you do not specify the index, the pop() method removes the last item.

Example
Remove the last item:

thislist = ["apple", "banana", "cherry"]


[Link]()
print(thislist)

The del keyword also removes the specified index:

Example
Remove the first item:
thislist = ["apple", "banana", "cherry"]
del thislist[0]
print(thislist)

The del keyword can also delete the list completely.

Example
Delete the entire list:

thislist = ["apple", "banana", "cherry"]


del thislist

Clear the List


The clear() method empties the list.

The list still remains, but it has no content.

Example
Clear the list content:

thislist = ["apple", "banana", "cherry"]


[Link]()
print(thislist)

Python - Loop Lists


Loop Through a List
You can loop through the list items by using a for loop:

Example
Print all items in the list, one by one:

thislist = ["apple", "banana", "cherry"]


for x in thislist:
print(x)

Learn more about for loops in our Python For Loops Chapter.
Loop Through the Index Numbers
You can also loop through the list items by referring to their index number.

Use the range() and len() functions to create a suitable iterable.

Example
Print all items by referring to their index number:

thislist = ["apple", "banana", "cherry"]


for i in range(len(thislist)):
print(thislist[i])

The iterable created in the example above is [0, 1, 2].

Using a While Loop


You can loop through the list items by using a while loop.

Use the len() function to determine the length of the list, then start at 0 and
loop your way through the list items by referring to their indexes.

Remember to increase the index by 1 after each iteration.

Example
Print all items, using a while loop to go through all the index numbers

thislist = ["apple", "banana", "cherry"]


i = 0
while i < len(thislist):
print(thislist[i])
i = i + 1

Learn more about while loops in our Python While Loops Chapter.

Looping Using List Comprehension


List Comprehension offers the shortest syntax for looping through lists:

Example
A short hand for loop that will print all items in a list:

thislist = ["apple", "banana", "cherry"]


[print(x) for x in thislist]

Python - List Comprehension


List Comprehension
List comprehension offers a shorter syntax when you want to create a new
list based on the values of an existing list.

Example:

Based on a list of fruits, you want a new list, containing only the fruits with
the letter "a" in the name.

Without list comprehension you will have to write a for statement with a
conditional test inside:

Example
fruits = ["apple", "banana", "cherry", "kiwi", "mango"]
newlist = []

for x in fruits:
if "a" in x:
[Link](x)

print(newlist)

With list comprehension you can do all that with only one line of code:

Example
fruits = ["apple", "banana", "cherry", "kiwi", "mango"]

newlist = [x for x in fruits if "a" in x]

print(newlist)
The Syntax
newlist = [expression for item in iterable if condition == True]

The return value is a new list, leaving the old list unchanged.

Condition
The condition is like a filter that only accepts the items that evaluate to True.

Example
Only accept items that are not "apple":

newlist = [x for x in fruits if x != "apple"]

The condition if x != "apple" will return True for all elements other than
"apple", making the new list contain all fruits except "apple".

The condition is optional and can be omitted:

Example
With no if statement:

newlist = [x for x in fruits]

Iterable
The iterable can be any iterable object, like a list, tuple, set etc.

Example
You can use the range() function to create an iterable:

newlist = [x for x in range(10)]

Same example, but with a condition:


Example
Accept only numbers lower than 5:

newlist = [x for x in range(10) if x < 5]

Expression
The expression is the current item in the iteration, but it is also the outcome,
which you can manipulate before it ends up like a list item in the new list:

Example
Set the values in the new list to upper case:

newlist = [[Link]() for x in fruits]

You can set the outcome to whatever you like:

Example
Set all values in the new list to 'hello':

newlist = ['hello' for x in fruits]

The expression can also contain conditions, not like a filter, but as a way to
manipulate the outcome:

Example
Return "orange" instead of "banana":

newlist = [x if x != "banana" else "orange" for x in fruits]

The expression in the example above says:

"Return the item if it is not banana, if it is banana return orange".

Python - Sort Lists


Sort List Alphanumerically
List objects have a sort() method that will sort the list alphanumerically,
ascending, by default:

Example
Sort the list alphabetically:

thislist = ["orange", "mango", "kiwi", "pineapple", "banana"]


[Link]()
print(thislist)

Example
Sort the list numerically:

thislist = [100, 50, 65, 82, 23]


[Link]()
print(thislist)

Sort Descending
To sort descending, use the keyword argument reverse = True:

Example
Sort the list descending:

thislist = ["orange", "mango", "kiwi", "pineapple", "banana"]


[Link](reverse = True)
print(thislist)

Example
Sort the list descending:

thislist = [100, 50, 65, 82, 23]


[Link](reverse = True)
print(thislist)
Customize Sort Function
You can also customize your own function by using the keyword
argument key = function.

The function will return a number that will be used to sort the list (the lowest
number first):

Example
Sort the list based on how close the number is to 50:

def myfunc(n):
return abs(n - 50)

thislist = [100, 50, 65, 82, 23]


[Link](key = myfunc)
print(thislist)

Case Insensitive Sort


By default the sort() method is case sensitive, resulting in all capital letters
being sorted before lower case letters:

Example
Case sensitive sorting can give an unexpected result:

thislist = ["banana", "Orange", "Kiwi", "cherry"]


[Link]()
print(thislist)

Luckily we can use built-in functions as key functions when sorting a list.

So if you want a case-insensitive sort function, use [Link] as a key


function:

Example
Perform a case-insensitive sort of the list:
thislist = ["banana", "Orange", "Kiwi", "cherry"]
[Link](key = [Link])
print(thislist)

Reverse Order
What if you want to reverse the order of a list, regardless of the alphabet?

The reverse() method reverses the current sorting order of the elements.

Example
Reverse the order of the list items:

thislist = ["banana", "Orange", "Kiwi", "cherry"]


[Link]()
print(thislist)

Python - Copy Lists


Copy a List
You cannot copy a list simply by typing list2 = list1, because: list2 will
only be a reference to list1, and changes made in list1 will automatically
also be made in list2.

Use the copy() method


You can use the built-in List method copy() to copy a list.

Example
Make a copy of a list with the copy() method:

thislist = ["apple", "banana", "cherry"]


mylist = [Link]()
print(mylist)
Use the list() method
Another way to make a copy is to use the built-in method list().

Example
Make a copy of a list with the list() method:

thislist = ["apple", "banana", "cherry"]


mylist = list(thislist)
print(mylist)

Use the slice Operator


You can also make a copy of a list by using the : (slice) operator.

Example
Make a copy of a list with the : operator:

thislist = ["apple", "banana", "cherry"]


mylist = thislist[:]
print(mylist)

Python - Join Lists


Join Two Lists
There are several ways to join, or concatenate, two or more lists in Python.

One of the easiest ways are by using the + operator.

Example
Join two list:
list1 = ["a", "b", "c"]
list2 = [1, 2, 3]

list3 = list1 + list2


print(list3)

Another way to join two lists is by appending all the items from list2 into
list1, one by one:

Example
Append list2 into list1:

list1 = ["a", "b" , "c"]


list2 = [1, 2, 3]

for x in list2:
[Link](x)

print(list1)

Or you can use the extend() method, where the purpose is to add elements
from one list to another list:

Example
Use the extend() method to add list2 at the end of list1:

list1 = ["a", "b" , "c"]


list2 = [1, 2, 3]

[Link](list2)
print(list1)

Python - List Methods


List Methods
Python has a set of built-in methods that you can use on lists.
Method Description

append() Adds an element at the end of the list

clear() Removes all the elements from the list

copy() Returns a copy of the list

count() Returns the number of elements with the specified value

extend() Add the elements of a list (or any iterable), to the end of the current lis

index() Returns the index of the first element with the specified value

insert() Adds an element at the specified position

pop() Removes the element at the specified position

remove() Removes the item with the specified value

reverse() Reverses the order of the list


sort() Sorts the list

Python Tuples
mytuple = ("apple", "banana", "cherry")

Tuple
Tuples are used to store multiple items in a single variable.

Tuple is one of 4 built-in data types in Python used to store collections of


data, the other 3 are List, Set, and Dictionary, all with different qualities and
usage.

A tuple is a collection which is ordered and unchangeable.

Tuples are written with round brackets.

Example
Create a Tuple:
thistuple = ("apple", "banana", "cherry")
print(thistuple)

Tuple Items
Tuple items are ordered, unchangeable, and allow duplicate values.

Tuple items are indexed, the first item has index [0], the second item has
index [1] etc.
Ordered
When we say that tuples are ordered, it means that the items have a defined
order, and that order will not change.

Unchangeable
Tuples are unchangeable, meaning that we cannot change, add or remove
items after the tuple has been created.

Allow Duplicates
Since tuples are indexed, they can have items with the same value:

Example
Tuples allow duplicate values:

thistuple = ("apple", "banana", "cherry", "apple", "cherry")


print(thistuple)

Tuple Length
To determine how many items a tuple has, use the len() function:

Example
Print the number of items in the tuple:

thistuple = ("apple", "banana", "cherry")


print(len(thistuple))
Create Tuple With One Item
To create a tuple with only one item, you have to add a comma after the
item, otherwise Python will not recognize it as a tuple.

Example
One item tuple, remember the comma:

thistuple = ("apple",)
print(type(thistuple))

#NOT a tuple
thistuple = ("apple")
print(type(thistuple))

Tuple Items - Data Types


Tuple items can be of any data type:

Example
String, int and boolean data types:

tuple1 = ("apple", "banana", "cherry")


tuple2 = (1, 5, 7, 9, 3)
tuple3 = (True, False, False)

A tuple can contain different data types:

Example
A tuple with strings, integers and boolean values:

tuple1 = ("abc", 34, True, 40, "male")

type()
From Python's perspective, tuples are defined as objects with the data type
'tuple':

<class 'tuple'>
Example
What is the data type of a tuple?

mytuple = ("apple", "banana", "cherry")


print(type(mytuple))

The tuple() Constructor


It is also possible to use the tuple() constructor to make a tuple.

Example
Using the tuple() method to make a tuple:

thistuple = tuple(("apple", "banana", "cherry")) # note the double


round-brackets
print(thistuple)

Python Collections (Arrays)


There are four collection data types in the Python programming language:

 List is a collection which is ordered and changeable. Allows duplicate


members.
 Tuple is a collection which is ordered and unchangeable. Allows
duplicate members.
 Set is a collection which is unordered, unchangeable*, and unindexed.
No duplicate members.
 Dictionary is a collection which is ordered** and changeable. No
duplicate members

 When choosing a collection type, it is useful to understand the


properties of that type. Choosing the right type for a particular data set
could mean retention of meaning, and, it could mean an increase in
efficiency or security.

Python - Access Tuple Items


Access Tuple Items
You can access tuple items by referring to the index number, inside square
brackets:

Example
Print the second item in the tuple:

thistuple = ("apple", "banana", "cherry")


print(thistuple[1]

Negative Indexing
Negative indexing means start from the end.

-1 refers to the last item, -2 refers to the second last item etc.

Example
Print the last item of the tuple:

thistuple = ("apple", "banana", "cherry")


print(thistuple[-1])

Range of Indexes
You can specify a range of indexes by specifying where to start and where to
end the range.

When specifying a range, the return value will be a new tuple with the
specified items.

Example
Return the third, fourth, and fifth item:
thistuple =
("apple", "banana", "cherry", "orange", "kiwi", "melon", "mango")
print(thistuple[2:5])

By leaving out the start value, the range will start at the first item:

Example
This example returns the items from the beginning to, but NOT included,
"kiwi":

thistuple =
("apple", "banana", "cherry", "orange", "kiwi", "melon", "mango")
print(thistuple[:4])

By leaving out the end value, the range will go on to the end of the tuple:

Example
This example returns the items from "cherry" and to the end:

thistuple =
("apple", "banana", "cherry", "orange", "kiwi", "melon", "mango")
print(thistuple[2:])

Range of Negative Indexes


Specify negative indexes if you want to start the search from the end of the
tuple:

Example
This example returns the items from index -4 (included) to index -1
(excluded)

thistuple =
("apple", "banana", "cherry", "orange", "kiwi", "melon", "mango")
print(thistuple[-4:-1])

Check if Item Exists


To determine if a specified item is present in a tuple use the in keyword:

Example
Check if "apple" is present in the tuple:

thistuple = ("apple", "banana", "cherry")


if "apple" in thistuple:
print("Yes, 'apple' is in the fruits tuple")

Python - Update Tuples


Tuples are unchangeable, meaning that you cannot change, add, or remove
items once the tuple is created.

But there are some workarounds.

Change Tuple Values


Once a tuple is created, you cannot change its values. Tuples
are unchangeable, or immutable as it also is called.

But there is a workaround. You can convert the tuple into a list, change the
list, and convert the list back into a tuple.

Example
Convert the tuple into a list to be able to change it:

x = ("apple", "banana", "cherry")


y = list(x)
y[1] = "kiwi"
x = tuple(y)

print(x)

Add Items
Since tuples are immutable, they do not have a built-in append() method, but
there are other ways to add items to a tuple.

1. Convert into a list: Just like the workaround for changing a tuple, you
can convert it into a list, add your item(s), and convert it back into a tuple.
Example
Convert the tuple into a list, add "orange", and convert it back into a tuple:

thistuple = ("apple", "banana", "cherry")


y = list(thistuple)
[Link]("orange")
thistuple = tuple(y)

2. Add tuple to a tuple. You are allowed to add tuples to tuples, so if you
want to add one item, (or many), create a new tuple with the item(s), and
add it to the existing tuple:

Example
Create a new tuple with the value "orange", and add that tuple:

thistuple = ("apple", "banana", "cherry")


y = ("orange",)
thistuple += y

print(thistuple)

Remove Items
Note: You cannot remove items in a tuple.

Tuples are unchangeable, so you cannot remove items from it, but you can
use the same workaround as we used for changing and adding tuple items:

Example
Convert the tuple into a list, remove "apple", and convert it back into a tuple:

thistuple = ("apple", "banana", "cherry")


y = list(thistuple)
[Link]("apple")
thistuple = tuple(y)

Or you can delete the tuple completely:

Example
The del keyword can delete the tuple completely:

thistuple = ("apple", "banana", "cherry")


del thistuple
print(thistuple) #this will raise an error because the tuple no
longer exists

Python - Unpack Tuples


Unpacking a Tuple
When we create a tuple, we normally assign values to it. This is called
"packing" a tuple:

Example
Packing a tuple:

fruits = ("apple", "banana", "cherry")

But, in Python, we are also allowed to extract the values back into variables.
This is called "unpacking":

Example
Unpacking a tuple:

fruits = ("apple", "banana", "cherry")

(green, yellow, red) = fruits

print(green)
print(yellow)
print(red)

Using Asterisk*
If the number of variables is less than the number of values, you can add
an * to the variable name and the values will be assigned to the variable as
a list:

Example
Assign the rest of the values as a list called "red":

fruits = ("apple", "banana", "cherry", "strawberry", "raspberry")

(green, yellow, *red) = fruits

print(green)
print(yellow)
print(red)

If the asterisk is added to another variable name than the last, Python will
assign values to the variable until the number of values left matches the
number of variables left.

Example
Add a list of values the "tropic" variable:

fruits = ("apple", "mango", "papaya", "pineapple", "cherry")

(green, *tropic, red) = fruits

print(green)
print(tropic)
print(red)

Python - Loop Tuples


Loop Through a Tuple
You can loop through the tuple items by using a for loop.

Example
Iterate through the items and print the values:

thistuple = ("apple", "banana", "cherry")


for x in thistuple:
print(x)

Loop Through the Index Numbers


You can also loop through the tuple items by referring to their index number.
Use the range() and len() functions to create a suitable iterable.

Example
Print all items by referring to their index number:

thistuple = ("apple", "banana", "cherry")


for i in range(len(thistuple)):
print(thistuple[i])

Using a While Loop


You can loop through the tuple items by using a while loop.

Use the len() function to determine the length of the tuple, then start at 0
and loop your way through the tuple items by referring to their indexes.

Remember to increase the index by 1 after each iteration.

Example
Print all items, using a while loop to go through all the index numbers:

thistuple = ("apple", "banana", "cherry")


i = 0
while i < len(thistuple):
print(thistuple[i])
i = i + 1

Python - Join Tuples


Join Two Tuples
To join two or more tuples you can use the + operator:

Example
Join two tuples:
tuple1 = ("a", "b" , "c")
tuple2 = (1, 2, 3)

tuple3 = tuple1 + tuple2


print(tuple3)

Multiply Tuples
If you want to multiply the content of a tuple a given number of times, you
can use the * operator:

Example
Multiply the fruits tuple by 2:

fruits = ("apple", "banana", "cherry")


mytuple = fruits * 2

print(mytuple)

Python - Tuple Methods


Tuple Methods
Python has two built-in methods that you can use on tuples.

Method Description

count() Returns the number of times a specified value occurs in a tuple

index() Searches the tuple for a specified value and returns the position

Python Sets
myset = {"apple", "banana", "cherry"}

Set
Sets are used to store multiple items in a single variable.

Set is one of 4 built-in data types in Python used to store collections of data,
the other 3 are List, Tuple, and Dictionary, all with different qualities and
usage.

A set is a collection which is unordered, unchangeable*, and unindexed.

Sets are written with curly brackets.

ExampleGet your own Python Server


Create a Set:

thisset = {"apple", "banana", "cherry"}


print(thisset)
Note: Sets are unordered, so you cannot be sure in which order the items
will appear.

Set Items
Set items are unordered, unchangeable, and do not allow duplicate values.

Unordered
Unordered means that the items in a set do not have a defined order.

Set items can appear in a different order every time you use them, and
cannot be referred to by index or key.
Unchangeable
Set items are unchangeable, meaning that we cannot change the items after
the set has been created.

Once a set is created, you cannot change its items, but you can remove
items and add new items.

Duplicates Not Allowed


Sets cannot have two items with the same value.

Example
Duplicate values will be ignored:

thisset = {"apple", "banana", "cherry", "apple"}

print(thisset)
Note: The values True and 1 are considered the same value in sets, and are
treated as duplicates:

Example
True and 1 is considered the same value:

thisset = {"apple", "banana", "cherry", True, 1, 2}

print(thisset)
Note: The values False and 0 are considered the same value in sets, and
are treated as duplicates:

Example
False and 0 is considered the same value:

thisset = {"apple", "banana", "cherry", False, True, 0}

print(thisset)
Get the Length of a Set
To determine how many items a set has, use the len() function.

Example
Get the number of items in a set:

thisset = {"apple", "banana", "cherry"}

print(len(thisset))

Set Items - Data Types


Set items can be of any data type:

Example
String, int and boolean data types:

set1 = {"apple", "banana", "cherry"}


set2 = {1, 5, 7, 9, 3}
set3 = {True, False, False}

A set can contain different data types:

Example
A set with strings, integers and boolean values:

set1 = {"abc", 34, True, 40, "male"}

type()
From Python's perspective, sets are defined as objects with the data type
'set':

<class 'set'>
Example
What is the data type of a set?

myset = {"apple", "banana", "cherry"}


print(type(myset))

The set() Constructor


It is also possible to use the set() constructor to make a set.

Example
Using the set() constructor to make a set:

thisset = set(("apple", "banana", "cherry")) # note the double round-


brackets
print(thisset)

Python Collections (Arrays)


There are four collection data types in the Python programming language:

 List is a collection which is ordered and changeable. Allows duplicate


members.
 Tuple is a collection which is ordered and unchangeable. Allows
duplicate members.
 Set is a collection which is unordered, unchangeable*, and unindexed.
No duplicate members.
 Dictionary is a collection which is ordered** and changeable. No
duplicate members.
*Set items are unchangeable, but you can remove items and add new items.

**As of Python version 3.7, dictionaries are ordered. In Python 3.6 and
earlier, dictionaries are unordered.

When choosing a collection type, it is useful to understand the properties of


that type. Choosing the right type for a particular data set could mean
retention of meaning, and, it could mean an increase in efficiency or
security.

Python - Access Set Items


Access Items
You cannot access items in a set by referring to an index or a key.

But you can loop through the set items using a for loop, or ask if a specified
value is present in a set, by using the in keyword.

Example
Loop through the set, and print the values:

thisset = {"apple", "banana", "cherry"}

for x in thisset:
print(x)

Example
Check if "banana" is present in the set:

thisset = {"apple", "banana", "cherry"}

print("banana" in thisset)

Example
Check if "banana" is NOT present in the set:

thisset = {"apple", "banana", "cherry"}

print("banana" not in thisset)

Change Items
Once a set is created, you cannot change its items, but you can add new
items.

Python - Add Set Items


Add Items
Once a set is created, you cannot change its items, but you can add new
items.

To add one item to a set use the add() method.

Example
Add an item to a set, using the add() method:

thisset = {"apple", "banana", "cherry"}

[Link]("orange")

print(thisset)

Add Sets
To add items from another set into the current set, use
the update() method.

Example
Add elements from tropical into thisset:

thisset = {"apple", "banana", "cherry"}


tropical = {"pineapple", "mango", "papaya"}

[Link](tropical)

print(thisset)
Add Any Iterable
The object in the update() method does not have to be a set, it can be any
iterable object (tuples, lists, dictionaries etc.).

Example
Add elements of a list to a set:

thisset = {"apple", "banana", "cherry"}


mylist = ["kiwi", "orange"]

[Link](mylist)

print(thisset)

Python - Remove Set Items


Remove Item
To remove an item in a set, use the remove(), or the discard() method.

Example
Remove "banana" by using the remove() method:

thisset = {"apple", "banana", "cherry"}

[Link]("banana")

print(thisset)
Note: If the item to remove does not exist, remove() will raise an error.

Example
Remove "banana" by using the discard() method:

thisset = {"apple", "banana", "cherry"}

[Link]("banana")
print(thisset)
Note: If the item to remove does not exist, discard() will NOT raise an
error.

You can also use the pop() method to remove an item, but this method will
remove a random item, so you cannot be sure what item that gets removed.

The return value of the pop() method is the removed item.

Example
Remove a random item by using the pop() method:

thisset = {"apple", "banana", "cherry"}

x = [Link]()

print(x)

print(thisset)
Note: Sets are unordered, so when using the pop() method, you do not
know which item that gets removed.

Example
The clear() method empties the set:

thisset = {"apple", "banana", "cherry"}

[Link]()

print(thisset)

Example
The del keyword will delete the set completely:

thisset = {"apple", "banana", "cherry"}

del thisset

print(thisset)
Python - Loop Sets
Loop Items
You can loop through the set items by using a for loop:

Example
Loop through the set, and print the values:

thisset = {"apple", "banana", "cherry"}

for x in thisset:
print(x)

Python - Join Sets


Join Sets
There are several ways to join two or more sets in Python.

The union() and update() methods joins all items from both sets.

The intersection() method keeps ONLY the duplicates.

The difference() method keeps the items from the first set that are not in
the other set(s).

The symmetric_difference() method keeps all items EXCEPT the


duplicates.

Union
The union() method returns a new set with all items from both sets.

Example
Join set1 and set2 into a new set:
set1 = {"a", "b", "c"}
set2 = {1, 2, 3}

set3 = [Link](set2)
print(set3)

You can use the | operator instead of the union() method, and you will get
the same result.

Example
Use | to join two sets:

set1 = {"a", "b", "c"}


set2 = {1, 2, 3}

set3 = set1 | set2


print(set3)

Join Multiple Sets


All the joining methods and operators can be used to join multiple sets.

When using a method, just add more sets in the parentheses, separated by
commas:

Example
Join multiple sets with the union() method:

set1 = {"a", "b", "c"}


set2 = {1, 2, 3}
set3 = {"John", "Elena"}
set4 = {"apple", "bananas", "cherry"}

myset = [Link](set2, set3, set4)


print(myset)

When using the | operator, separate the sets with more | operators:

Example
Use | to join two sets:

set1 = {"a", "b", "c"}


set2 = {1, 2, 3}
set3 = {"John", "Elena"}
set4 = {"apple", "bananas", "cherry"}

myset = set1 | set2 | set3 |set4


print(myset)

Join a Set and a Tuple


The union() method allows you to join a set with other data types, like lists
or tuples.

The result will be a set.

Example
Join a set with a tuple:

x = {"a", "b", "c"}


y = (1, 2, 3)

z = [Link](y)
print(z)
Note: The | operator only allows you to join sets with sets, and not with
other data types like you can with the union() method.

Update
The update() method inserts all items from one set into another.

The update() changes the original set, and does not return a new set.

Example
The update() method inserts the items in set2 into set1:
set1 = {"a", "b" , "c"}
set2 = {1, 2, 3}

[Link](set2)
print(set1)
Note: Both union() and update() will exclude any duplicate items

Intersection
Keep ONLY the duplicates

The intersection() method will return a new set, that only contains the
items that are present in both sets.

Example
Join set1 and set2, but keep only the duplicates:

set1 = {"apple", "banana", "cherry"}


set2 = {"google", "microsoft", "apple"}

set3 = [Link](set2)
print(set3)

You can use the & operator instead of the intersection() method, and you
will get the same result.

Example
Use & to join two sets:

set1 = {"apple", "banana", "cherry"}


set2 = {"google", "microsoft", "apple"}

set3 = set1 & set2


print(set3)
Note: The & operator only allows you to join sets with sets, and not with
other data types like you can with the intersection() method.

The intersection_update() method will also keep ONLY the duplicates, but
it will change the original set instead of returning a new set.

Example
Keep the items that exist in both set1, and set2:

set1 = {"apple", "banana", "cherry"}


set2 = {"google", "microsoft", "apple"}

set1.intersection_update(set2)

print(set1)

The values True and 1 are considered the same value. The same goes
for False and 0.

Example
Join sets that contains the values True, False, 1, and 0, and see what is
considered as duplicates:

set1 = {"apple", 1, "banana", 0, "cherry"}


set2 = {False, "google", 1, "apple", 2, True}

set3 = [Link](set2)

print(set3)

Difference
The difference() method will return a new set that will contain only the
items from the first set that are not present in the other set.

Example
Keep all items from set1 that are not in set2:

set1 = {"apple", "banana", "cherry"}


set2 = {"google", "microsoft", "apple"}

set3 = [Link](set2)

print(set3)

You can use the - operator instead of the difference() method, and you
will get the same result.
Example
Use - to join two sets:

set1 = {"apple", "banana", "cherry"}


set2 = {"google", "microsoft", "apple"}

set3 = set1 - set2


print(set3)
Note: The - operator only allows you to join sets with sets, and not with
other data types like you can with the difference() method.

The difference_update() method will keep the items from the first set that
are not in the other set, but it will change the original set instead of returning
a new set.

Example
Use the difference_update() method to keep only the items from the first
set that are not present in the other set:

set1 = {"apple", "banana", "cherry"}


set2 = {"google", "microsoft", "apple"}

set1.difference_update(set2)

print(set1)

Symmetric Differences
The symmetric_difference() method will keep only the elements that are
NOT present in both sets.

Example
Keep the items that are not present in both sets:

set1 = {"apple", "banana", "cherry"}


set2 = {"google", "microsoft", "apple"}

set3 = set1.symmetric_difference(set2)
print(set3)

You can use the ^ operator instead of the symmetric_difference() method,


and you will get the same result.

Example
Use ^ to join two sets:

set1 = {"apple", "banana", "cherry"}


set2 = {"google", "microsoft", "apple"}

set3 = set1 ^ set2


print(set3)
Note: The ^ operator only allows you to join sets with sets, and not with
other data types like you can with the symmetric_difference() method.

The symmetric_difference_update() method will also keep all but the


duplicates, but it will change the original set instead of returning a new set.

Example
Use the symmetric_difference_update() method to keep the items that
are not present in both sets:

set1 = {"apple", "banana", "cherry"}


set2 = {"google", "microsoft", "apple"}

set1.symmetric_difference_update(set2)

print(set1)

Python frozenset
Python frozenset
frozenset is an immutable version of a set.

Like sets, it contains unique, unordered, unchangeable elements.

Unlike sets, elements cannot be added or removed from a frozenset.


Creating a frozenset
Use the frozenset() constructor to create a frozenset from any iterable.

Example
Create a frozenset and check its type:

x = frozenset({"apple", "banana", "cherry"})


print(x)
print(type(x))

Frozenset Methods
Being immutable means you cannot add or remove elements. However,
frozensets support all non-mutating operations of sets.

Method Shortcut Description

copy() Returns a shallow copy

difference() - Returns a new frozenset with the difference

intersection() & Returns a new frozenset with the intersection

isdisjoint() Returns True if there is NO intersection betwee


issubset() <= / < Returns True if this frozenset is a (proper) subs

issuperset() >= / > Returns True if this frozenset is a (proper) supe

symmetric_difference() ^ Returns a new frozenset with the symmetric d

union() | Returns a new frozenset containing the union

Python - Set Methods


Set Methods
Python has a set of built-in methods that you can use on sets.

Method Shortcu Description


t

add() Adds an element to the set

clear() Removes all the elements from the set

copy() Returns a copy of the set


difference() - Returns a set containing the difference bet

difference_update() -= Removes the items in this set that are also

discard() Remove the specified item

intersection() & Returns a set, that is the intersection of tw

intersection_update() &= Removes the items in this set that are not

isdisjoint() Returns True if NO items of this set is pres

issubset() <= Returns True if all items of this set is prese

< Returns True if all items of this set is prese

issuperset() >= Returns True if all items of another set is p

> Returns True if all items of another, smalle

pop() Removes an element from the set


remove() Removes the specified element

symmetric_difference() ^ Returns a set with the symmetric differenc

symmetric_difference_update() ^= Inserts the symmetric differences from this

union() | Return a set containing the union of sets

update() |= Update the set with the union of this set an

Python Dictionaries
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}

Dictionary
Dictionaries are used to store data values in key:value pairs.

A dictionary is a collection which is ordered*, changeable and do not allow


duplicates.

As of Python version 3.7, dictionaries are ordered. In Python 3.6 and earlier,
dictionaries are unordered.

Dictionaries are written with curly brackets, and have keys and values:
Example
Create and print a dictionary:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
print(thisdict)

Dictionary Items
Dictionary items are ordered, changeable, and do not allow duplicates.

Dictionary items are presented in key:value pairs, and can be referred to by


using the key name.

Example
Print the "brand" value of the dictionary:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
print(thisdict["brand"])

Ordered or Unordered?
As of Python version 3.7, dictionaries are ordered. In Python 3.6 and earlier,
dictionaries are unordered.

When we say that dictionaries are ordered, it means that the items have a
defined order, and that order will not change.
Unordered means that the items do not have a defined order, you cannot
refer to an item by using an index.

Changeable
Dictionaries are changeable, meaning that we can change, add or remove
items after the dictionary has been created.

Duplicates Not Allowed


Dictionaries cannot have two items with the same key:

Example
Duplicate values will overwrite existing values:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964,
"year": 2020
}
print(thisdict)

Dictionary Length
To determine how many items a dictionary has, use the len() function:

Example
Print the number of items in the dictionary:

print(len(thisdict))
Dictionary Items - Data Types
The values in dictionary items can be of any data type:

Example
String, int, boolean, and list data types:

thisdict = {
"brand": "Ford",
"electric": False,
"year": 1964,
"colors": ["red", "white", "blue"]
}

type()
From Python's perspective, dictionaries are defined as objects with the data
type 'dict':

<class 'dict'>
Example
Print the data type of a dictionary:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
print(type(thisdict))

The dict() Constructor


It is also possible to use the dict() constructor to make a dictionary.

Example
Using the dict() method to make a dictionary:

thisdict = dict(name = "John", age = 36, country = "Norway")


print(thisdict)

Python Collections (Arrays)


There are four collection data types in the Python programming language:

 List is a collection which is ordered and changeable. Allows duplicate


members.
 Tuple is a collection which is ordered and unchangeable. Allows
duplicate members.
 Set is a collection which is unordered, unchangeable*, and unindexed.
No duplicate members.
 Dictionary is a collection which is ordered** and changeable. No
duplicate members.
*Set items are unchangeable, but you can remove and/or add items
whenever you like.

**As of Python version 3.7, dictionaries are ordered. In Python 3.6 and
earlier, dictionaries are unordered.

When choosing a collection type, it is useful to understand the properties of


that type. Choosing the right type for a particular data set could mean
retention of meaning, and, it could mean an increase in efficiency or
security.

Python - Access Dictionary


Items
Accessing Items
You can access the items of a dictionary by referring to its key name, inside
square brackets:

Example
Get the value of the "model" key:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
x = thisdict["model"]

There is also a method called get() that will give you the same result:

Example
Get the value of the "model" key:

x = [Link]("model")

Get Keys
The keys() method will return a list of all the keys in the dictionary.

Example
Get a list of the keys:

x = [Link]()

The list of the keys is a view of the dictionary, meaning that any changes
done to the dictionary will be reflected in the keys list.

Example
Add a new item to the original dictionary, and see that the keys list gets
updated as well:

car = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
x = [Link]()

print(x) #before the change

car["color"] = "white"

print(x) #after the change

Get Values
The values() method will return a list of all the values in the dictionary.

Example
Get a list of the values:

x = [Link]()

The list of the values is a view of the dictionary, meaning that any changes
done to the dictionary will be reflected in the values list.

Example
Make a change in the original dictionary, and see that the values list gets
updated as well:

car = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}

x = [Link]()

print(x) #before the change

car["year"] = 2020

print(x) #after the change

Example
Add a new item to the original dictionary, and see that the values list gets
updated as well:
car = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}

x = [Link]()

print(x) #before the change

car["color"] = "red"

print(x) #after the change

Get Items
The items() method will return each item in a dictionary, as tuples in a list.

Example
Get a list of the key:value pairs

x = [Link]()

The returned list is a view of the items of the dictionary, meaning that any
changes done to the dictionary will be reflected in the items list.

Example
Make a change in the original dictionary, and see that the items list gets
updated as well:

car = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}

x = [Link]()

print(x) #before the change

car["year"] = 2020
print(x) #after the change

Example
Add a new item to the original dictionary, and see that the items list gets
updated as well:

car = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}

x = [Link]()

print(x) #before the change

car["color"] = "red"

print(x) #after the change

Check if Key Exists


To determine if a specified key is present in a dictionary use the in keyword:

Example
Check if "model" is present in the dictionary:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
if "model" in thisdict:
print("Yes, 'model' is one of the keys in the thisdict dictionary")

Python - Change Dictionary


Items
Change Values
You can change the value of a specific item by referring to its key name:

Example
Change the "year" to 2018:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
thisdict["year"] = 2018

Update Dictionary
The update() method will update the dictionary with the items from the
given argument.

The argument must be a dictionary, or an iterable object with key:value


pairs.

Example
Update the "year" of the car by using the update() method:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
[Link]({"year": 2020})

Python - Add Dictionary


Items
Adding Items
Adding an item to the dictionary is done by using a new index key and
assigning a value to it:

Example
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
thisdict["color"] = "red"
print(thisdict)

Update Dictionary
The update() method will update the dictionary with the items from a given
argument. If the item does not exist, the item will be added.

The argument must be a dictionary, or an iterable object with key:value


pairs.

Example
Add a color item to the dictionary by using the update() method:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
[Link]({"color": "red"})

Python - Remove Dictionary


Items
Removing Items
There are several methods to remove items from a dictionary:

Example
The pop() method removes the item with the specified key name:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
[Link]("model")
print(thisdict)

Example
The popitem() method removes the last inserted item (in versions before
3.7, a random item is removed instead):

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
[Link]()
print(thisdict)

Example
The del keyword removes the item with the specified key name:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
del thisdict["model"]
print(thisdict)

Example
The del keyword can also delete the dictionary completely:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
del thisdict
print(thisdict) #this will cause an error because "thisdict" no
longer exists.

Example
The clear() method empties the dictionary:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
[Link]()
print(thisdict)

Python - Loop Dictionaries


Loop Through a Dictionary
You can loop through a dictionary by using a for loop.

When looping through a dictionary, the return value are the keys of the
dictionary, but there are methods to return the values as well.

Example
Print all key names in the dictionary, one by one:

for x in thisdict:
print(x)

Example
Print all values in the dictionary, one by one:

for x in thisdict:
print(thisdict[x])

Example
You can also use the values() method to return values of a dictionary:

for x in [Link]():
print(x)

Example
You can use the keys() method to return the keys of a dictionary:

for x in [Link]():
print(x)

Example
Loop through both keys and values, by using the items() method:

for x, y in [Link]():
print(x, y)

Python - Copy Dictionaries


Copy a Dictionary
You cannot copy a dictionary simply by typing dict2 = dict1,
because: dict2 will only be a reference to dict1, and changes made
in dict1 will automatically also be made in dict2.

There are ways to make a copy, one way is to use the built-in Dictionary
method copy().

Example
Make a copy of a dictionary with the copy() method:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
mydict = [Link]()
print(mydict)
Another way to make a copy is to use the built-in function dict().

Example
Make a copy of a dictionary with the dict() function:

thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
mydict = dict(thisdict)
print(mydict)

Python - Nested Dictionaries


Nested Dictionaries
A dictionary can contain dictionaries, this is called nested dictionaries.

Example
Create a dictionary that contain three dictionaries:

myfamily = {
"child1" : {
"name" : "Emil",
"year" : 2004
},
"child2" : {
"name" : "Tobias",
"year" : 2007
},
"child3" : {
"name" : "Linus",
"year" : 2011
}
}

Or, if you want to add three dictionaries into a new dictionary:


Example
Create three dictionaries, then create one dictionary that will contain the
other three dictionaries:

child1 = {
"name" : "Emil",
"year" : 2004
}
child2 = {
"name" : "Tobias",
"year" : 2007
}
child3 = {
"name" : "Linus",
"year" : 2011
}

myfamily = {
"child1" : child1,
"child2" : child2,
"child3" : child3
}

Access Items in Nested Dictionaries


To access items from a nested dictionary, you use the name of the
dictionaries, starting with the outer dictionary:

Example
Print the name of child 2:

print(myfamily["child2"]["name"])

Loop Through Nested Dictionaries


You can loop through a dictionary by using the items() method like this:
Example
Loop through the keys and values of all nested dictionaries:

for x, obj in [Link]():


print(x)

for y in obj:
print(y + ':', obj[y])

Python Dictionary Methods


Dictionary Methods
Python has a set of built-in methods that you can use on dictionaries.

Method Description

clear() Removes all the elements from the dictionary

copy() Returns a copy of the dictionary

fromkeys() Returns a dictionary with the specified keys and value

get() Returns the value of the specified key

items() Returns a list containing a tuple for each key value pair

keys() Returns a list containing the dictionary's keys


pop() Removes the element with the specified key

popitem() Removes the last inserted key-value pair

setdefault() Returns the value of the specified key. If the key does not exist: insert the

update() Updates the dictionary with the specified key-value pairs

values() Returns a list of all the values in the dictionary

Python If Statement
Python Conditions and If statements
Python supports the usual logical conditions from mathematics:

 Equals: a == b
 Not Equals: a != b
 Less than: a < b
 Less than or equal to: a <= b
 Greater than: a > b
 Greater than or equal to: a >= b

These conditions can be used in several ways, most commonly in "if


statements" and loops.

An "if statement" is written by using the if keyword.

Example
If statement:
a = 33
b = 200
if b > a:
print("b is greater than a")

In this example we use two variables, a and b, which are used as part of the
if statement to test whether b is greater than a. As a is 33, and b is 200, we
know that 200 is greater than 33, and so we print to screen that "b is greater
than a".

How If Statements Work


The if statement evaluates a condition (an expression that results
in True or False). If the condition is true, the code block inside the if
statement is executed. If the condition is false, the code block is skipped.

Example
Checking if a number is positive:

number = 15
if number > 0:
print("The number is positive")

Indentation
Python relies on indentation (whitespace at the beginning of a line) to define
scope in the code. Other programming languages often use curly-brackets
for this purpose.

Example
If statement, without indentation (will raise an error):

a = 33
b = 200
if b > a:
print("b is greater than a") # you will get an error
Note: You can use spaces or tabs for indentation, but you must use the
same amount of indentation for all statements within the same code block.

Multiple Statements in If Block


You can have multiple statements inside an if block. All statements must be
indented at the same level.

Example
Multiple statements in an if block:

age = 20
if age >= 18:
print("You are an adult")
print("You can vote")
print("You have full legal rights")

Using Variables in Conditions


Boolean variables can be used directly in if statements without comparison
operators.

Example
Using a boolean variable:

is_logged_in = True
if is_logged_in:
print("Welcome back!")

Python can evaluate many types of values as True or False in an if


statement.

Zero (0), empty strings (""), None, and empty collections are treated
as False. Everything else is treated as True.

This includes positive numbers (5), negative numbers (-3), and any non-
empty string (even "False" is treated as True because it's a non-empty
string).
Python Elif Statement
The Elif Keyword
The elif keyword is Python's way of saying "if the previous conditions were
not true, then try this condition".

The elif keyword allows you to check multiple expressions for True and
execute a block of code as soon as one of the conditions evaluates to True.

ExampleGet your own Python Server


a = 33
b = 33
if b > a:
print("b is greater than a")
elif a == b:
print("a and b are equal")

In this example a is equal to b, so the first condition is not true, but


the elif condition is true, so we print to screen that "a and b are equal".

Multiple Elif Statements


You can have as many elif statements as you need. Python will check each
condition in order and execute the first one that is true.

Example
Testing multiple conditions:

score = 75

if score >= 90:


print("Grade: A")
elif score >= 80:
print("Grade: B")
elif score >= 70:
print("Grade: C")
elif score >= 60:
print("Grade: D")
In this example, the program checks each condition in order. Since score is
75, it prints "Grade: C" (the first condition that evaluates to true).

How Elif Works


When you use elif, Python evaluates the conditions from top to bottom. As
soon as it finds a condition that is true, it executes that block and skips all
remaining conditions.

Important: Only the first true condition will be executed. Even if multiple
conditions are true, Python stops after executing the first matching block.

Example
Categorizing age groups:

age = 25

if age < 13:


print("You are a child")
elif age < 20:
print("You are a teenager")
elif age < 65:
print("You are an adult")
elif age >= 65:
print("You are a senior")

When to Use Elif


Use elif when you have multiple mutually exclusive conditions to check.
This is more efficient than using multiple separate if statements because
Python stops checking once it finds a true condition.

Example
Day of the week checker:

day = 3

if day == 1:
print("Monday")
elif day == 2:
print("Tuesday")
elif day == 3:
print("Wednesday")
elif day == 4:
print("Thursday")
elif day == 5:
print("Friday")
elif day == 6:
print("Saturday")
elif day == 7:
print("Sunday")

Python Else Statement


The Else Keyword
The else keyword catches anything which isn't caught by the preceding
conditions.

The else statement is executed when the if condition (and


any elif conditions) evaluate to False.

Example
a=200
b = 33
if b > a:
print("b is greater than a")
elif a == b:
print("a and b are equal")
else:
print("a is greater than b")

In this example a is greater than b, so the first condition is not true, also
the elif condition is not true, so we go to the else condition and print to
screen that "a is greater than b".

Else Without Elif


You can also have an else without the elif:
Example
a = 200
b = 33
if b > a:
print("b is greater than a")
else:
print("b is not greater than a")

This creates a simple two-way choice: if the condition is true, execute one
block; otherwise, execute the else block.

How Else Works


The else statement provides a default action when none of the previous
conditions are true. Think of it as a "catch-all" for any scenario not covered
by your if and elif statements.

Note: The else statement must come last. You cannot have an elif after
an else.

Example
Checking even or odd numbers:

number = 7

if number % 2 == 0:
print("The number is even")
else:
print("The number is odd")

Complete If-Elif-Else Chain


You can combine if, elif, and else to create a comprehensive decision-
making structure.

Example
Temperature classifier:

temperature = 22

if temperature > 30:


print("It's hot outside!")
elif temperature > 20:
print("It's warm outside")
elif temperature > 10:
print("It's cool outside")
else:
print("It's cold outside!")

Else as Fallback
The else statement acts as a fallback that executes when none of the
preceding conditions are true. This makes it useful for error handling,
validation, and providing default values.

Example
Validating user input:

username = "Emil"

if len(username) > 0:
print(f"Welcome, {username}!")
else:
print("Error: Username cannot be empty")

Python Shorthand If
Short Hand If
If you have only one statement to execute, you can put it on the same line
as the if statement.

Example
One-line if statement:

a = 5
b = 2
if a > b: print("a is greater than b")
Note: You still need the colon : after the condition.

Short Hand If ... Else


If you have one statement for if and one for else, you can put them on the
same line using a conditional expression:

Example
One-line if/else that prints a value:

a = 2
b = 330
print("A") if a > b else print("B")
This is called a conditional expression (sometimes known as a "ternary
operator").

Assign a Value With If ... Else


You can also use a one-line if/else to choose a value and assign it to a
variable:

Example
a = 10
b = 20
bigger = a if a > b else b
print("Bigger is", bigger)

The syntax follows this pattern:

variable = value_if_true if condition else value_if_false

Multiple Conditions on One Line


You can chain conditional expressions, but keep it short so it stays readable:

Example
One line, three outcomes:

a = 330
b = 330
print("A") if a > b else print("=") if a == b else print("B")

Practical Examples
Ternary operators are particularly useful for simple assignments and return
statements.

Example
Finding the maximum of two numbers:

x = 15
y = 20
max_value = x if x > y else y
print("Maximum value:", max_value)

Example
Setting a default value:

username = ""
display_name = username if username else "Guest"
print("Welcome,", display_name)

When to Use Shorthand If


Shorthand if statements and ternary operators should be used when:

 The condition and actions are simple


 It improves code readability
 You want to make a quick assignment based on a condition
Important: While shorthand if statements can make code more concise,
avoid overusing them for complex conditions. For readability, use regular if-
else statements when dealing with multiple lines of code or complex logic.
Python Logical Operators
Python Logical Operators
Logical operators are used to combine conditional statements. Python has
three logical operators:

 and - Returns True if both statements are true


 or - Returns True if one of the statements is true
 not - Reverses the result, returns False if the result is true

The and Operator


The and keyword is a logical operator, and is used to combine conditional
statements. Both conditions must be true for the entire expression to be
true.

Example
Test if a is greater than b, AND if c is greater than a:

a = 200
b = 33
c = 500
if a > b and c > a:
print("Both conditions are True")

The or Operator
The or keyword is a logical operator, and is used to combine conditional
statements. At least one condition must be true for the entire expression to
be true.

Example
Test if a is greater than b, OR if a is greater than c:

a = 200
b = 33
c = 500
if a > b or a > c:
print("At least one of the conditions is True")

The not Operator


The not keyword is a logical operator, and is used to reverse the result of the
conditional statement.

Example
Test if a is NOT greater than b:

a = 33
b = 200
if not a > b:
print("a is NOT greater than b")

Combining Multiple Operators


You can combine multiple logical operators in a single expression. Python
evaluates not first, then and, then or.

Example
Combining and, or, and not:

age = 25
is_student = False
has_discount_code = True

if (age < 18 or age > 65) and not is_student or has_discount_code:


print("Discount applies!")

Truth Tables
Understanding how logical operators work with different values:
and Operator Truth Table

Condition 1 Condition 2 Resul

True True True

True False False

False True False

False False False

or Operator Truth Table

Condition 1 Condition 2 Resul

True True True

True False True

False True True

False False False


Using Parentheses for Clarity
When combining multiple logical operators, use parentheses to make your
intentions clear and control the order of evaluation.

Example
Using parentheses for complex conditions:

temperature = 25
is_raining = False
is_weekend = True

if (temperature > 20 and not is_raining) or is_weekend:


print("Great day for outdoor activities!")

More Examples
Example
User authentication check:

username = "Tobias"
password = "secret123"
is_verified = True

if username and password and is_verified:


print("Login successful")
else:
print("Login failed")

Example
Range checking with logical operators:

score = 85

if score >= 0 and score <= 100:


print("Valid score")
else:
print("Invalid score")

Python Nested If
Nested If Statements
You can have if statements inside if statements. This is
called nested if statements.

Example
x = 41

if x > 10:
print("Above ten,")
if x > 20:
print("and also above 20!")
else:
print("but not above 20.")

In this example, the inner if statement only runs if the outer condition (x >
10) is true.

How Nested If Works


Each level of nesting creates a deeper level of decision-making. The code
evaluates from the outermost condition inward.

Example
Checking multiple conditions with nesting:

age = 25
has_license = True

if age >= 18:


if has_license:
print("You can drive")
else:
print("You need a license")
else:
print("You are too young to drive")

Multiple Levels of Nesting


You can nest as many levels deep as needed, but keep in mind that too
many levels can make code harder to read.

Example
Three levels of nesting:

score = 85
attendance = 90
submitted = True

if score >= 60:


if attendance >= 80:
if submitted:
print("Pass with good standing")
else:
print("Pass but missing assignment")
else:
print("Pass but low attendance")
else:
print("Fail")

Nested If vs Logical Operators


Sometimes nested if statements can be simplified using logical operators
like and. The choice depends on your logic.

Example
This nested if:

temperature = 25
is_sunny = True
if temperature > 20:
if is_sunny:
print("Perfect beach weather!")

Example
Could also be written with and:

temperature = 25
is_sunny = True

if temperature > 20 and is_sunny:


print("Perfect beach weather!")

Both approaches produce the same result. Use nested if statements when
the inner logic is complex or depends on the outer condition. Use and when
both conditions are simple and equally important.

More Examples
Example
Login validation with nested checks:

username = "Emil"
password = "python123"
is_active = True

if username:
if password:
if is_active:
print("Login successful")
else:
print("Account is not active")
else:
print("Password required")
else:
print("Username required")

Example
Grade calculation with nested logic:
score = 92
extra_credit = 5

if score >= 90:


if extra_credit > 0:
print("A+ grade")
else:
print("A grade")
elif score >= 80:
print("B grade")
else:
print("C grade or below")

Python Pass Statement


The pass Statement
if statements cannot be empty, but if you for some reason have
an if statement with no content, put in the pass statement to avoid getting
an error.

Example
a = 33
b = 200

if b > a:
pass

The pass statement is a null operation - nothing happens when it executes. It


serves as a placeholder.

Why Use pass?


The pass statement is useful in several situations:

 When you're creating code structure but haven't implemented the


logic yet
 When a statement is required syntactically but no action is needed
 As a placeholder for future code during development
 In empty functions or classes that you plan to implement later
pass in Development
During development, you might want to sketch out your program structure
before implementing the details. The pass statement allows you to do this
without syntax errors.

Example
Placeholder for future implementation:

age = 16

if age < 18:


pass # TODO: Add underage logic later
else:
print("Access granted")

pass vs Comments
A comment is ignored by Python, but pass is an actual statement that gets
executed (though it does nothing). You need pass where Python expects a
statement, not just a comment.

Example
This will cause an error (empty code block):

score = 85

if score > 90:


# This is excellent
# This will raise an IndentationError

Example
This works correctly with pass:

score = 85

if score > 90:


pass # This is excellent
print("Score processed")
pass with Multiple Conditions
You can use pass in any branch of an if-elif-else statement.

Example
Using pass in different branches:

value = 50

if value < 0:
print("Negative value")
elif value == 0:
pass # Zero case - no action needed
else:
print("Positive value")

pass in Other Contexts


While we focus on pass with if statements here, it's also commonly used with
loops, functions, and classes.

Example
Using pass with functions:

def calculate_discount(price):
pass # TODO: Implement discount logic

# Function exists but doesn't do anything yet

Python Match
The match statement is used to perform different actions based on
different conditions.

The Python Match Statement


Instead of writing many if..else statements, you can use
the match statement.

The match statement selects one of many code blocks to be executed.

Syntax
match expression:
case x:
code block
case y:
code block
case z:
code block

This is how it works:

 The match expression is evaluated once.


 The value of the expression is compared with the values of each case.
 If there is a match, the associated block of code is executed.

The example below uses the weekday number to print the weekday name:

Example
day = 4
match day:
case 1:
print("Monday")
case 2:
print("Tuesday")
case 3:
print("Wednesday")
case 4:
print("Thursday")
case 5:
print("Friday")
case 6:
print("Saturday")
case 7:
print("Sunday")

Default Value
Use the underscore character _ as the last case value if you want a code
block to execute when there are not other matches:
Example
day = 4
match day:
case 6:
print("Today is Saturday")
case 7:
print("Today is Sunday")
case _:
print("Looking forward to the Weekend")

The value _ will always match, so it is important to place it as


the last case to make it behave as a default case.

Combine Values
Use the pipe character | as an or operator in the case evaluation to check
for more than one value match in one case:

Example
day = 4
match day:
case 1 | 2 | 3 | 4 | 5:
print("Today is a weekday")
case 6 | 7:
print("I love weekends!")

If Statements as Guards
You can add if statements in the case evaluation as an extra condition-
check:

Example
month = 5
day = 4
match day:
case 1 | 2 | 3 | 4 | 5 if month == 4:
print("A weekday in April")
case 1 | 2 | 3 | 4 | 5 if month == 5:
print("A weekday in May")
case _:
print("No match")

Python While Loops


Python Loops
Python has two primitive loop commands:

 while loops
 for loops

The while Loop


With the while loop we can execute a set of statements as long as a
condition is true.

ExampleGet your own Python Server


Print i as long as i is less than 6:

i = 1
while i < 6:
print(i)
i += 1

Note: remember to increment i, or else the loop will continue forever.

The while loop requires relevant variables to be ready, in this example we


need to define an indexing variable, i, which we set to 1.

The break Statement


With the break statement we can stop the loop even if the while condition is
true:

Example
Exit the loop when i is 3:

i = 1
while i < 6:
print(i)
if i == 3:
break
i += 1

The continue Statement


With the continue statement we can stop the current iteration, and continue
with the next:

Example
Continue to the next iteration if i is 3:

i = 0
while i < 6:
i += 1
if i == 3:
continue
print(i)

The else Statement


With the else statement we can run a block of code once when the condition
no longer is true:

Example
Print a message once the condition is false:
i = 1
while i < 6:
print(i)
i += 1
else:
print("i is no longer less than 6")

Note: The else block will NOT be executed if the loop is stopped by
a break statement.

Python For Loops


Python For Loops
A for loop is used for iterating over a sequence (that is either a list, a tuple,
a dictionary, a set, or a string).

This is less like the for keyword in other programming languages, and works
more like an iterator method as found in other object-orientated
programming languages.

With the for loop we can execute a set of statements, once for each item in
a list, tuple, set etc.

ExampleGet your own Python Server


Print each fruit in a fruit list:

fruits = ["apple", "banana", "cherry"]


for x in fruits:
print(x)

The for loop does not require an indexing variable to set beforehand.

Looping Through a String


Even strings are iterable objects, they contain a sequence of characters:

Example
Loop through the letters in the word "banana":

for x in "banana":
print(x)

The break Statement


With the break statement we can stop the loop before it has looped through
all the items:

Example
Exit the loop when x is "banana":

fruits = ["apple", "banana", "cherry"]


for x in fruits:
print(x)
if x == "banana":
break

Example
Exit the loop when x is "banana", but this time the break comes before the
print:

fruits = ["apple", "banana", "cherry"]


for x in fruits:
if x == "banana":
break
print(x)

The continue Statement


With the continue statement we can stop the current iteration of the loop,
and continue with the next:

Example
Do not print banana:
fruits = ["apple", "banana", "cherry"]
for x in fruits:
if x == "banana":
continue
print(x)

The range() Function


To loop through a set of code a specified number of times, we can use
the range() function,

The range() function returns a sequence of numbers, starting from 0 by


default, and increments by 1 (by default), and ends at a specified number.

Example
Using the range() function:

for x in range(6):
print(x)
Note that range(6) is not the values of 0 to 6, but the values 0 to 5.

The range() function defaults to 0 as a starting value, however it is possible


to specify the starting value by adding a parameter: range(2, 6), which
means values from 2 to 6 (but not including 6):

Example
Using the start parameter:

for x in range(2, 6):


print(x)

The range() function defaults to increment the sequence by 1, however it is


possible to specify the increment value by adding a third parameter: range(2,
30, 3):

Example
Increment the sequence with 3 (default is 1):
for x in range(2, 30, 3):
print(x)

Else in For Loop


The else keyword in a for loop specifies a block of code to be executed
when the loop is finished:

Example
Print all numbers from 0 to 5, and print a message when the loop has ended:

for x in range(6):
print(x)
else:
print("Finally finished!")
Note: The else block will NOT be executed if the loop is stopped by
a break statement.

Example
Break the loop when x is 3, and see what happens with the else block:

for x in range(6):
if x == 3: break
print(x)
else:
print("Finally finished!")

Nested Loops
A nested loop is a loop inside a loop.

The "inner loop" will be executed one time for each iteration of the "outer
loop":

Example
Print each adjective for every fruit:
adj = ["red", "big", "tasty"]
fruits = ["apple", "banana", "cherry"]

for x in adj:
for y in fruits:
print(x, y)

The pass Statement


for loops cannot be empty, but if you for some reason have a for loop with
no content, put in the pass statement to avoid getting an error.

Example
for x in [0, 1, 2]:
pass

Python Functions
Python Functions
A function is a block of code which only runs when it is called.

A function can return data as a result.

A function helps avoiding code repetition.

Creating a Function
In Python, a function is defined using the def keyword, followed by a function
name and parentheses:

Example
def my_function():
print("Hello from a function")
This creates a function named my_function that prints "Hello from a
function" when called.

The code inside the function must be indented. Python uses indentation to
define code blocks.

Calling a Function
To call a function, write its name followed by parentheses:

Example
def my_function():
print("Hello from a function")

my_function()

You can call the same function multiple times:

Example
def my_function():
print("Hello from a function")

my_function()
my_function()
my_function()

Function Names
Function names follow the same rules as variable names in Python:

 A function name must start with a letter or underscore


 A function name can only contain letters, numbers, and underscores
 Function names are case-sensitive (myFunction and myfunction are
different)
Example
Valid function names:
calculate_sum()
_private_function()
myFunction2()
It's good practice to use descriptive names that explain what the function
does.

Why Use Functions?


Imagine you need to convert temperatures from Fahrenheit to Celsius
several times in your program. Without functions, you would have to write
the same calculation code repeatedly:

Example
Without functions - repetitive code:

temp1 = 77
celsius1 = (temp1 - 32) * 5 / 9
print(celsius1)

temp2 = 95
celsius2 = (temp2 - 32) * 5 / 9
print(celsius2)

temp3 = 50
celsius3 = (temp3 - 32) * 5 / 9
print(celsius3)

With functions, you write the code once and reuse it:

Example
With functions - reusable code:

def fahrenheit_to_celsius(fahrenheit):
return (fahrenheit - 32) * 5 / 9

print(fahrenheit_to_celsius(77))
print(fahrenheit_to_celsius(95))
print(fahrenheit_to_celsius(50))
Return Values
Functions can send data back to the code that called them using
the return statement.

When a function reaches a return statement, it stops executing and sends


the result back:

Example
A function that returns a value:

def get_greeting():
return "Hello from a function"

message = get_greeting()
print(message)

You can use the returned value directly:

Example
Using the return value directly:

def get_greeting():
return "Hello from a function"

print(get_greeting())
If a function doesn't have a return statement, it returns None by default.

The pass Statement


Function definitions cannot be empty. If you need to create a function
placeholder without any code, use the pass statement:

Example
def my_function():
pass
The pass statement is often used when developing, allowing you to define
the structure first and implement details later.

Python Function Arguments


Arguments
Information can be passed into functions as arguments.

Arguments are specified after the function name, inside the parentheses.
You can add as many arguments as you want, just separate them with a
comma.

The following example has a function with one argument (fname). When the
function is called, we pass along a first name, which is used inside the
function to print the full name:

ExampleGet your own Python Server


A function with one argument:

def my_function(fname):
print(fname + " Refsnes")

my_function("Emil")
my_function("Tobias")
my_function("Linus")

Parameters vs Arguments
The terms parameter and argument can be used for the same thing:
information that are passed into a function.

From a function's perspective:

A parameter is the variable listed inside the parentheses in the function


definition.

An argument is the actual value that is sent to the function when it is


called.
Example
def my_function(name): # name is a parameter
print("Hello", name)

my_function("Emil") # "Emil" is an argument

Number of Arguments
By default, a function must be called with the correct number of arguments.

If your function expects 2 arguments, you must call it with exactly 2


arguments.

Example
This function expects 2 arguments, and gets 2 arguments::

def my_function(fname, lname):


print(fname + " " + lname)

my_function("Emil", "Refsnes")

If you try to call the function with the wrong number of arguments, you will
get an error:

Example
This function expects 2 arguments, but gets only 1:

def my_function(fname, lname):


print(fname + " " + lname)

my_function("Emil")

Default Parameter Values


You can assign default values to parameters. If the function is called without
an argument, it uses the default value:

Example
def my_function(name = "friend"):
print("Hello", name)

my_function("Emil")
my_function("Tobias")
my_function()
my_function("Linus")

Example
Default value for country parameter:

def my_function(country = "Norway"):


print("I am from", country)

my_function("Sweden")
my_function("India")
my_function()
my_function("Brazil")

Keyword Arguments
You can send arguments with the key = value syntax.

Example
def my_function(animal, name):
print("I have a", animal)
print("My", animal + "'s name is", name)

my_function(animal = "dog", name = "Buddy")

This way, with keyword arguments, the order of the arguments does not
matter.

Example
def my_function(animal, name):
print("I have a", animal)
print("My", animal + "'s name is", name)

my_function(name = "Buddy", animal = "dog")


The phrase Keyword Arguments is often shortened to kwargs in Python
documentation.
Positional Arguments
When you call a function with arguments without using keywords, they are
called positional arguments.

Positional arguments must be in the correct order:

Example
def my_function(animal, name):
print("I have a", animal)
print("My", animal + "'s name is", name)

my_function("dog", "Buddy")

The order matters with positional arguments:

Example
Switching the order changes the result:

def my_function(animal, name):


print("I have a", animal)
print("My", animal + "'s name is", name)

my_function("Buddy", "dog")

Mixing Positional and Keyword


Arguments
You can mix positional and keyword arguments in a function call.

However, positional arguments must come before keyword arguments:

Example
def my_function(animal, name, age):
print("I have a", age, "year old", animal, "named", name)

my_function("dog", name = "Buddy", age = 5)


Passing Different Data Types
You can send any data type as an argument to a function (string, number,
list, dictionary, etc.).

The data type will be preserved inside the function:

Example
Sending a list as an argument:

def my_function(fruits):
for fruit in fruits:
print(fruit)

my_fruits = ["apple", "banana", "cherry"]


my_function(my_fruits)

Example
Sending a dictionary as an argument:

def my_function(person):
print("Name:", person["name"])
print("Age:", person["age"])

my_person = {"name": "Emil", "age": 25}


my_function(my_person)

Return Values
Functions can return values using the return statement:

Example
def my_function(x, y):
return x + y

result = my_function(5, 3)
print(result)
Returning Different Data Types
Functions can return any data type, including lists, tuples, dictionaries, and
more.

Example
A function that returns a list:

def my_function():
return ["apple", "banana", "cherry"]

fruits = my_function()
print(fruits[0])
print(fruits[1])
print(fruits[2])

Example
A function that returns a tuple:

def my_function():
return (10, 20)

x, y = my_function()
print("x:", x)
print("y:", y)

Positional-Only Arguments
You can specify that a function can have ONLY positional arguments.

To specify positional-only arguments, add , / after the arguments:

Example
def my_function(name, /):
print("Hello", name)

my_function("Emil")
Without the , / you are actually allowed to use keyword arguments even if
the function expects positional arguments:

Example
def my_function(name):
print("Hello", name)

my_function(name = "Emil")

With , /, you will get an error if you try to use keyword arguments:

Example
def my_function(name, /):
print("Hello", name)

my_function(name = "Emil")

Keyword-Only Arguments
To specify that a function can have only keyword arguments,
add *, before the arguments:

Example
def my_function(*, name):
print("Hello", name)

my_function(name = "Emil")

Without *,, you are allowed to use positional arguments even if the function
expects keyword arguments:

Example
def my_function(name):
print("Hello", name)

my_function("Emil")

With *,, you will get an error if you try to use positional arguments:
Example
def my_function(*, name):
print("Hello", name)

my_function("Emil")

Combining Positional-Only and


Keyword-Only
You can combine both argument types in the same function.

Arguments before / are positional-only, and arguments after * are keyword-


only:

Example
def my_function(a, b, /, *, c, d):
return a + b + c + d

result = my_function(5, 10, c = 15, d = 20)


print(result)

Python *args and **kwargs


*args and **kwargs
By default, a function must be called with the correct number of arguments.

However, sometimes you may not know how many arguments that will be
passed into your function.

*args and **kwargs allow functions to accept a unknown number of


arguments.

Arbitrary Arguments - *args


If you do not know how many arguments will be passed into your function,
add a * before the parameter name.

This way, the function will receive a tuple of arguments and can access the
items accordingly:

Example
Using *args to accept any number of arguments:

def my_function(*kids):
print("The youngest child is " + kids[2])

my_function("Emil", "Tobias", "Linus")


Arbitrary Arguments are often shortened to *args in Python documentation.

What is *args?
The *args parameter allows a function to accept any number of positional
arguments.

Inside the function, args becomes a tuple containing all the passed
arguments:

Example
Accessing individual arguments from *args:

def my_function(*args):
print("Type:", type(args))
print("First argument:", args[0])
print("Second argument:", args[1])
print("All arguments:", args)

my_function("Emil", "Tobias", "Linus")

Using *args with Regular Arguments


You can combine regular parameters with *args.
Regular parameters must come before *args:

Example
def my_function(greeting, *names):
for name in names:
print(greeting, name)

my_function("Hello", "Emil", "Tobias", "Linus")

In this example, "Hello" is assigned to greeting, and the rest are collected
in names.

Practical Example with *args


*args is useful when you want to create flexible functions:

Example
A function that calculates the sum of any number of values:

def my_function(*numbers):
total = 0
for num in numbers:
total += num
return total

print(my_function(1, 2, 3))
print(my_function(10, 20, 30, 40))
print(my_function(5))

Example
Finding the maximum value:

def my_function(*numbers):
if len(numbers) == 0:
return None
max_num = numbers[0]
for num in numbers:
if num > max_num:
max_num = num
return max_num
print(my_function(3, 7, 2, 9, 1))

Arbitrary Keyword Arguments -


**kwargs
If you do not know how many keyword arguments will be passed into your
function, add two asterisks ** before the parameter name.

This way, the function will receive a dictionary of arguments and can access
the items accordingly:

Example
Using **kwargs to accept any number of keyword arguments:

def my_function(**kid):
print("His last name is " + kid["lname"])

my_function(fname = "Tobias", lname = "Refsnes")


Arbitrary Keyword Arguments are often shortened to **kwargs in Python
documentation.

What is **kwargs?
The **kwargs parameter allows a function to accept any number of keyword
arguments.

Inside the function, kwargs becomes a dictionary containing all the keyword
arguments:

Example
Accessing values from **kwargs:

def my_function(**myvar):
print("Type:", type(myvar))
print("Name:", myvar["name"])
print("Age:", myvar["age"])
print("All data:", myvar)

my_function(name = "Tobias", age = 30, city = "Bergen")

Using **kwargs with Regular


Arguments
You can combine regular parameters with **kwargs.

Regular parameters must come before **kwargs:

Example
def my_function(username, **details):
print("Username:", username)
print("Additional details:")
for key, value in [Link]():
print(" ", key + ":", value)

my_function("emil123", age = 25, city = "Oslo", hobby = "coding")

Combining *args and **kwargs


You can use both *args and **kwargs in the same function.

The order must be:

1. regular parameters
2. *args
3. **kwargs
Example
def my_function(title, *args, **kwargs):
print("Title:", title)
print("Positional arguments:", args)
print("Keyword arguments:", kwargs)

my_function("User Info", "Emil", "Tobias", age = 25, city = "Oslo")


Unpacking Arguments
The * and ** operators can also be used when calling functions to unpack
(expand) a list or dictionary into separate arguments.

Unpacking Lists with *


If you have values stored in a list, you can use * to unpack them into
individual arguments:

Example
Using * to unpack a list into arguments:

def my_function(a, b, c):


return a + b + c

numbers = [1, 2, 3]
result = my_function(*numbers) # Same as: my_function(1, 2, 3)
print(result)

Unpacking Dictionaries with **


If you have keyword arguments stored in a dictionary, you can use ** to
unpack them:

Example
Using ** to unpack a dictionary into keyword arguments:

def my_function(fname, lname):


print("Hello", fname, lname)

person = {"fname": "Emil", "lname": "Refsnes"}


my_function(**person) # Same as: my_function(fname="Emil",
lname="Refsnes")
Remember: Use * and ** in function definitions to collect arguments, and
use them in function calls to unpack arguments

Python Scope
Scope
A variable is only available from inside the region it is created. This is
called scope.

Local Scope
A variable created inside a function belongs to the local scope of that
function, and can only be used inside that function.

Example
A variable created inside a function is available inside that function:

def myfunc():
x = 300
print(x)

myfunc()

Function Inside Function


As explained in the example above, the variable x is not available outside the
function, but it is available for any function inside the function:

Example
The local variable can be accessed from a function within the function:

def myfunc():
x = 300
def myinnerfunc():
print(x)
myinnerfunc()

myfunc()

Global Scope
A variable created in the main body of the Python code is a global variable
and belongs to the global scope.

Global variables are available from within any scope, global and local.

Example
A variable created outside of a function is global and can be used by anyone:

x = 300

def myfunc():
print(x)

myfunc()

print(x)

Naming Variables
If you operate with the same variable name inside and outside of a function,
Python will treat them as two separate variables, one available in the global
scope (outside the function) and one available in the local scope (inside the
function):

Example
The function will print the local x, and then the code will print the global x:

x = 300

def myfunc():
x = 200
print(x)

myfunc()

print(x)

Global Keyword
If you need to create a global variable, but are stuck in the local scope, you
can use the global keyword.
The global keyword makes the variable global.

Example
If you use the global keyword, the variable belongs to the global scope:

def myfunc():
global x
x = 300

myfunc()

print(x)

Also, use the global keyword if you want to make a change to a global
variable inside a function.

Example
To change the value of a global variable inside a function, refer to the
variable by using the global keyword:

x = 300

def myfunc():
global x
x = 200

myfunc()

print(x)

Nonlocal Keyword
The nonlocal keyword is used to work with variables inside nested
functions.

The nonlocal keyword makes the variable belong to the outer function.

Example
If you use the nonlocal keyword, the variable will belong to the outer
function:
def myfunc1():
x = "Jane"
def myfunc2():
nonlocal x
x = "hello"
myfunc2()
return x

print(myfunc1())

The LEGB Rule


Python follows the LEGB rule when looking up variable names, and searches
for them in this order:

1. Local - Inside the current function


2. Enclosing - Inside enclosing functions (from inner to outer)
3. Global - At the top level of the module
4. Built-in - In Python's built-in namespace
Example
Understanding the LEGB rule:

x = "global"

def outer():
x = "enclosing"
def inner():
x = "local"
print("Inner:", x)
inner()
print("Outer:", x)

outer()
print("Global:", x)

Python Decorators
Decorators let you add extra behavior to a function, without changing
the function's code.
A decorator is a function that takes another function as input and
returns a new function.

Basic Decorator
Define the decorator first, then apply it with @decorator_name above the
function.

Exampl
A basic decorator that uppercases the return value of the decorated function.

def changecase(func):
def myinner():
return func().upper()
return myinner

@changecase
def myfunction():
return "Hello Sally"

print(myfunction())

By placing @changecase directly above the function definition, the


function myfunction is being "decorated" with the changecase function.

The function changecase is the decorator.

The function myfunction is the function that gets decorated.

Multiple Decorator Calls


A decorator can be called multiple times. Just place the decorator above the
function you want to decorate.

Example
Using the @changecase decorator on two functions:
def changecase(func):
def myinner():
return func().upper()
return myinner

@changecase
def myfunction():
return "Hello Sally"

@changecase
def otherfunction():
return "I am speed!"

print(myfunction())
print(otherfunction())

Arguments in the Decorated Function


Functions that require arguments can also be decorated, just make sure you
pass the arguments to the wrapper function:

Example
Functions with arguments can also be decorated:

def changecase(func):
def myinner(x):
return func(x).upper()
return myinner

@changecase
def myfunction(nam):
return "Hello " + nam

print(myfunction("John"))

*args and **kwargs


Sometimes the decorator function has no control over the arguments passed
from decorated function, to solve this problem, add (*args, **kwargs) to
the wrapper function, this way the wrapper function can accept any number,
and any type of arguments, and pass them to the decorated function.
Example
Secure the function with *args and **kwargs arguments:

def changecase(func):
def myinner(*args, **kwargs):
return func(*args, **kwargs).upper()
return myinner

@changecase
def myfunction(nam):
return "Hello " + nam

print(myfunction("John"))

Decorator With Arguments


Decorators can accept their own arguments by adding another wrapper
level.

Example
A decorator factory that takes an argument and transforms the casing based
on the argument value.

def changecase(n):
def changecase(func):
def myinner():
if n == 1:
a = func().lower()
else:
a = func().upper()
return a
return myinner
return changecase

@changecase(1)
def myfunction():
return "Hello Linus"

print(myfunction())

Multiple Decorators
You can use multiple decorators on one function.

This is done by placing the decorator calls on top of each other.

Decorators are called in the reverse order, starting with the one closest to
the function.

Example
One decorator for upper case, and one for adding a greeting:

def changecase(func):
def myinner():
return func().upper()
return myinner

def addgreeting(func):
def myinner():
return "Hello " + func() + " Have a good day!"
return myinner

@changecase
@addgreeting
def myfunction():
return "Tobias"

print(myfunction())

Preserving Function Metadata


Functions in Python has metadata that can be accessed using
the __name__ and __doc__ attributes.

Example
Normally, a function's name can be returned with the __name__ attribute:

def myfunction():
return "Have a great day!"

print(myfunction.__name__)
But, when a function is decorated, the metadata of the original function is
lost.

Example
Try returning the name from a decorated function and you will not get the
same result:

def changecase(func):
def myinner():
return func().upper()
return myinner

@changecase
def myfunction():
return "Have a great day!"

print(myfunction.__name__)

To fix this, Python has a built-in function called [Link] that can be
used to preserve the original function's name and docstring.

Example
Import [Link] to preserve the original function name and docstring.

import functools

def changecase(func):
@[Link](func)
def myinner():
return func().upper()
return myinner

@changecase
def myfunction():
return "Have a great day!"

print(myfunction.__name__)

Python Lambda
Lambda Functions
A lambda function is a small anonymous function.
A lambda function can take any number of arguments, but can only have one
expression.

Syntax
lambda arguments : expression

The expression is executed and the result is returned:

Example
Add 10 to argument a, and return the result:

x = lambda a : a + 10
print(x(5))

Lambda functions can take any number of arguments:

Example
Multiply argument a with argument b and return the result:

x = lambda a, b : a * b
print(x(5, 6))

Example
Summarize argument a, b, and c and return the result:

x = lambda a, b, c : a + b + c
print(x(5, 6, 2))

Why Use Lambda Functions?


The power of lambda is better shown when you use them as an anonymous
function inside another function.

Say you have a function definition that takes one argument, and that
argument will be multiplied with an unknown number:

def myfunc(n):
return lambda a : a * n
Use that function definition to make a function that always doubles the
number you send in:

Example
def myfunc(n):
return lambda a : a * n

mydoubler = myfunc(2)

print(mydoubler(11))

Or, use the same function definition to make a function that


always triples the number you send in:

Example
def myfunc(n):
return lambda a : a * n

mytripler = myfunc(3)

print(mytripler(11))

Or, use the same function definition to make both functions, in the same
program:

Example
def myfunc(n):
return lambda a : a * n

mydoubler = myfunc(2)
mytripler = myfunc(3)

print(mydoubler(11))
print(mytripler(11))
Use lambda functions when an anonymous function is required for a short
period of time.

Lambda with Built-in Functions


Lambda functions are commonly used with built-in functions
like map(), filter(), and sorted().

Using Lambda with map()


The map() function applies a function to every item in an iterable:

Example
Double all numbers in a list:

numbers = [1, 2, 3, 4, 5]
doubled = list(map(lambda x: x * 2, numbers))
print(doubled)

Using Lambda with filter()


The filter() function creates a list of items for which a function returns True:

Example
Filter out odd numbers from a list:

numbers = [1, 2, 3, 4, 5, 6, 7, 8]
odd_numbers = list(filter(lambda x: x % 2 != 0, numbers))
print(odd_numbers)

Using Lambda with sorted()


The sorted() function can use a lambda as a key for custom sorting:

Example
Sort a list of tuples by the second element:

students = [("Emil", 25), ("Tobias", 22), ("Linus", 28)]


sorted_students = sorted(students, key=lambda x: x[1])
print(sorted_students)

Example
Sort strings by length:
words = ["apple", "pie", "banana", "cherry"]
sorted_words = sorted(words, key=lambda x: len(x))
print(sorted_words)

Python Recursion
Recursion
Recursion is when a function calls itself.

Recursion is a common mathematical and programming concept. It means


that a function calls itself. This has the benefit of meaning that you can loop
through data to reach a result.

The developer should be very careful with recursion as it can be quite easy
to slip into writing a function which never terminates, or one that uses
excess amounts of memory or processor power. However, when written
correctly recursion can be a very efficient and mathematically-elegant
approach to programming.

ExampleGet your own Python Server


A simple recursive function that counts down from 5:

def countdown(n):
if n <= 0:
print("Done!")
else:
print(n)
countdown(n - 1)

countdown(5)

Base Case and Recursive Case


Every recursive function must have two parts:

 A base case - A condition that stops the recursion


 A recursive case - The function calling itself with a modified
argument
Without a base case, the function would call itself forever, causing a stack
overflow error.

Example
Identifying base case and recursive case:

def factorial(n):
# Base case
if n == 0 or n == 1:
return 1
# Recursive case
else:
return n * factorial(n - 1)

print(factorial(5))
The base case is crucial. Always make sure your recursive function has a
condition that will eventually be met.

Fibonacci Sequence
The Fibonacci sequence is a classic example where each number is the sum
of the two preceding ones. The sequence starts with 0 and 1:

0, 1, 1, 2, 3, 5, 8, 13, ...

The sequence continues indefinitely, with each number being the sum of the
two preceding ones.

We can use recursion to find a specific number in the sequence:

Example
Find the 7th number in the Fibonacci sequence:

def fibonacci(n):
if n <= 1:
return n
else:
return fibonacci(n - 1) + fibonacci(n - 2)

print(fibonacci(7))
Recursion with Lists
Recursion can be used to process lists by handling one element at a time:

Example
Calculate the sum of all elements in a list:

def sum_list(numbers):
if len(numbers) == 0:
return 0
else:
return numbers[0] + sum_list(numbers[1:])

my_list = [1, 2, 3, 4, 5]
print(sum_list(my_list))

Example
Find the maximum value in a list:

def find_max(numbers):
if len(numbers) == 1:
return numbers[0]
else:
max_of_rest = find_max(numbers[1:])
return numbers[0] if numbers[0] > max_of_rest else max_of_rest

my_list = [3, 7, 2, 9, 1]
print(find_max(my_list))

Recursion Depth Limit


Python has a limit on how deep recursion can go. The default limit is usually
around 1000 recursive calls.

Example
Check the recursion limit:

import sys
print([Link]())
If you need deeper recursion, you can increase the limit, but be careful as
this can cause crashes:

Example
import sys
[Link](2000)
print([Link]())
Increasing the recursion limit should be done with caution. For very deep
recursion, consider using iteration instead.

Python Generators
Generators
Generators are functions that can pause and resume their execution.

When a generator function is called, it returns a generator object, which is an


iterator.

The code inside the function is not executed yet, it is only compiled. The
function only executes when you iterate over the generator.

ExampleGet your own Python Server


A simple generator function:

def my_generator():
yield 1
yield 2
yield 3

for value in my_generator():


print(value)
Generators allow you to iterate over data without storing the entire dataset
in memory.

Instead of using return, generators use the yield keyword.

The yield Keyword


The yield keyword is what makes a function a generator.

When yield is encountered, the function's state is saved, and the value is
returned. The next time the generator is called, it continues from where it
left off.

Example
Generator that yields numbers:

def count_up_to(n):
count = 1
while count <= n:
yield count
count += 1

for num in count_up_to(5):


print(num)
Unlike return, which terminates the function, yield pauses it and can be
called multiple times.

Generators Saves Memory


Generators are memory-efficient because they generate values on-the-fly
instead of storing everything in memory.

For large datasets, generators save memory:

Example
Generator for large sequences:

def large_sequence(n):
for i in range(n):
yield i

# This doesn't create a million numbers in memory


gen = large_sequence(1000000)
print(next(gen))
print(next(gen))
print(next(gen))
Using next() with Generators
You can manually iterate through a generator using the next() function:

Example
def simple_gen():
yield "Emil"
yield "Tobias"
yield "Linus"

gen = simple_gen()
print(next(gen))
print(next(gen))
print(next(gen))

When there are no more values to yield, the generator raises


a StopIteration exception:

Example
def simple_gen():
yield 1
yield 2

gen = simple_gen()
print(next(gen))
print(next(gen))
print(next(gen)) # This will raise StopIteration

Generator Expressions
Similar to list comprehensions, you can create generators using generator
expressions with parentheses instead of square brackets:

Example
List comprehension vs generator expression:

# List comprehension - creates a list


list_comp = [x * x for x in range(5)]
print(list_comp)

# Generator expression - creates a generator


gen_exp = (x * x for x in range(5))
print(gen_exp)
print(list(gen_exp))

Example
Using a generator expression with sum:

# Calculate sum of squares without creating a list


total = sum(x * x for x in range(10))
print(total)

Fibonacci Sequence Generator


Generators can be used to create the Fibonacci sequence.

It can continue generating values indefinitely, without running out of


memory:

Example
Generate 100 Fibonacci numbers:

def fibonacci():
a, b = 0, 1
while True:
yield a
a, b = b, a + b

# Get first 100 Fibonacci numbers


gen = fibonacci()
for _ in range(100):
print(next(gen))

Generator Methods
Generators have special methods for advanced control:

send() Method
The send() method allows you to send a value to the generator:

Example
def echo_generator():
while True:
received = yield
print("Received:", received)

gen = echo_generator()
next(gen) # Prime the generator
[Link]("Hello")
[Link]("World")

close() Method
The close() method stops the generator:

Example
def my_gen():
try:
yield 1
yield 2
yield 3
finally:
print("Generator closed")

gen = my_gen()
print(next(gen))
[Link]()

Python range
Python range
The built-in range() function returns an immutable sequence of numbers,
commonly used for looping a specific number of times.

This set of numbers has its own data type called range.

Note: Immutable means that it cannot be modified after it is created.


Creating ranges
The range() function can be called with 1, 2, or 3 arguments, using this
syntax:

range(start, stop, step)

Call range() With One Argument


If the range function is called with only one argument, the argument
represents the stop value.

The start argument is optional, and if not provided, it defaults to 0.

range(10) returns a sequence of each number from 0 to 9. (The start


argument, 0 is inclusive, and the stop argument, 10 is exclusive).

ExampleGet your own Python Server


Create a range of numbers from 0 to 9:

x = range(10)

Call range() With Two Arguments


If the range function is called with two arguments, the first argument
represents the start value, and the second argument represents
the stop value.

range(3, 10) returns a sequence of each number from 3 to 9:

Example
Create a range of numbers from 3 to 9:
x = range(3, 10)

Call range() With Three Arguments


If the range function is called with three arguments, the third argument
represents the step value.

The step value means the difference between each number in the sequence.
It is optional, and if not provided, it defaults to 1.

range(3, 10, 2) returns a sequence of each number from 3 to 9, with a


step of 2:

Example
Create a range of numbers from 3 to 9:

x = range(3, 10, 2)

Using ranges
Ranges are often used in for loops to iterate over a sequence of numbers.

Example
Iterate over each value in a range:

for i in range(10):
print(i)

Using List to Display Ranges


The range object is a data type that represents an immutable sequence of
numbers, and it is not directly displayable.
Therefore, ranges are often converted to lists for display.

Example
Convert different ranges to lists:

print(list(range(5)))
print(list(range(1, 6)))
print(list(range(5, 20, 3)))

Slicing Ranges
Like other sequences, ranges can be sliced to extract a subsequence.

Example
Extract a subsequence from a range:

r = range(10)
print(r[2])
print(r[:3])

Note: The first print statement returns the value at index 2, and the second
print statement returns a new range object, from index 0 to 3.

Membership Testing
Ranges support membership testing with the in operator.

Example
Test if the numbers 6 and 7 are present in a range:

r = range(0, 10, 2)
print(6 in r)
print(7 in r)
The return value is True when the number is present in the range,
and False when it is not.

Length
Ranges support the len() function to get the number of elements in the
range.

Example
Get the length of a range:

r = range(0, 10, 2)
print(len(r))

Python Arrays
Note: Python does not have built-in support for Arrays, but Python Lists can
be used instead.

Arrays
Note: This page shows you how to use LISTS as ARRAYS, however, to work
with arrays in Python you will have to import a library, like the NumPy library.

Arrays are used to store multiple values in one single variable:

ExampleGet your own Python Server


Create an array containing car names:

cars = ["Ford", "Volvo", "BMW"]

What is an Array?
An array is a special variable, which can hold more than one value at a time.

If you have a list of items (a list of car names, for example), storing the cars
in single variables could look like this:

car1 = "Ford"
car2 = "Volvo"
car3 = "BMW"

However, what if you want to loop through the cars and find a specific one?
And what if you had not 3 cars, but 300?

The solution is an array!

An array can hold many values under a single name, and you can access the
values by referring to an index number.

Access the Elements of an Array


You refer to an array element by referring to the index number.

Example
Get the value of the first array item:

x = cars[0]

Example
Modify the value of the first array item:

cars[0] = "Toyota"

The Length of an Array


Use the len() method to return the length of an array (the number of
elements in an array).

Example
Return the number of elements in the cars array:

x = len(cars)
Note: The length of an array is always one more than the highest array
index.

Looping Array Elements


You can use the for in loop to loop through all the elements of an array.

Example
Print each item in the cars array:

for x in cars:
print(x)

Adding Array Elements


You can use the append() method to add an element to an array.

Example
Add one more element to the cars array:

[Link]("Honda")

Removing Array Elements


You can use the pop() method to remove an element from the array.

Example
Delete the second element of the cars array:
[Link](1)

You can also use the remove() method to remove an element from the
array.

Example
Delete the element that has the value "Volvo":

[Link]("Volvo")
Note: The list's remove() method only removes the first occurrence of the
specified value.

Array Methods
Python has a set of built-in methods that you can use on lists/arrays.

Method Description

append() Adds an element at the end of the list

clear() Removes all the elements from the list

copy() Returns a copy of the list

count() Returns the number of elements with the specified value

extend() Add the elements of a list (or any iterable), to the end of the current lis
index() Returns the index of the first element with the specified value

insert() Adds an element at the specified position

pop() Removes the element at the specified position

remove() Removes the first item with the specified value

reverse() Reverses the order of the list

sort() Sorts the list

Note: Python does not have built-in support for Arrays, but Python Lists can
be used instead.

Python Iterators
Python Iterators
An iterator is an object that contains a countable number of values.

An iterator is an object that can be iterated upon, meaning that you can
traverse through all the values.

Technically, in Python, an iterator is an object which implements the iterator


protocol, which consist of the methods __iter__() and __next__().

Iterator vs Iterable
Lists, tuples, dictionaries, and sets are all iterable objects. They are
iterable containers which you can get an iterator from.

All these objects have a iter() method which is used to get an iterator:

ExampleGet your own Python Server


Return an iterator from a tuple, and print each value:

mytuple = ("apple", "banana", "cherry")


myit = iter(mytuple)

print(next(myit))
print(next(myit))
print(next(myit))

Even strings are iterable objects, and can return an iterator:

Example
Strings are also iterable objects, containing a sequence of characters:

mystr = "banana"
myit = iter(mystr)

print(next(myit))
print(next(myit))
print(next(myit))
print(next(myit))
print(next(myit))
print(next(myit))

Looping Through an Iterator


We can also use a for loop to iterate through an iterable object:

Example
Iterate the values of a tuple:
mytuple = ("apple", "banana", "cherry")

for x in mytuple:
print(x)

Example
Iterate the characters of a string:

mystr = "banana"

for x in mystr:
print(x)

The for loop actually creates an iterator object and executes


the next() method for each loop.

Create an Iterator
To create an object/class as an iterator you have to implement the
methods __iter__() and __next__() to your object.

As you will learn in the Python Classes/Objects chapter, all classes have a
function called __init__(), which allows you to do some initializing when
the object is being created.

The __iter__() method acts similar, you can do operations (initializing etc.),
but must always return the iterator object itself.

The __next__() method also allows you to do operations, and must return
the next item in the sequence.

Example
Create an iterator that returns numbers, starting with 1, and each sequence
will increase by one (returning 1,2,3,4,5 etc.):

class MyNumbers:
def __iter__(self):
self.a = 1
return self

def __next__(self):
x = self.a
self.a += 1
return x

myclass = MyNumbers()
myiter = iter(myclass)

print(next(myiter))
print(next(myiter))
print(next(myiter))
print(next(myiter))
print(next(myiter))

StopIteration
The example above would continue forever if you had enough next()
statements, or if it was used in a for loop.

To prevent the iteration from going on forever, we can use


the StopIteration statement.

In the __next__() method, we can add a terminating condition to raise an


error if the iteration is done a specified number of times:

Example
Stop after 20 iterations:

class MyNumbers:
def __iter__(self):
self.a = 1
return self

def __next__(self):
if self.a <= 20:
x = self.a
self.a += 1
return x
else:
raise StopIteration

myclass = MyNumbers()
myiter = iter(myclass)

for x in myiter:
print(x)
Python Modules
What is a Module?
Consider a module to be the same as a code library.

A file containing a set of functions you want to include in your application.

Create a Module
To create a module just save the code you want in a file with the file
extension .py:

ExampleGet your own Python Server


Save this code in a file named [Link]

def greeting(name):
print("Hello, " + name)

Use a Module
Now we can use the module we just created, by using the import statement:

Example
Import the module named mymodule, and call the greeting function:

import mymodule

[Link]("Jonathan")
Note: When using a function from a module, use the
syntax: module_name.function_name.

Variables in Module
The module can contain functions, as already described, but also variables of
all types (arrays, dictionaries, objects etc):

Example
Save this code in the file [Link]

person1 = {
"name": "John",
"age": 36,
"country": "Norway"
}

Example
Import the module named mymodule, and access the person1 dictionary:

import mymodule

a = mymodule.person1["age"]
print(a)

Naming a Module
You can name the module file whatever you like, but it must have the file
extension .py

Re-naming a Module
You can create an alias when you import a module, by using the as keyword:

Example
Create an alias for mymodule called mx:

import mymodule as mx

a = mx.person1["age"]
print(a)

Built-in Modules
There are several built-in modules in Python, which you can import whenever
you like.

Example
Import and use the platform module:

import platform

x = [Link]()
print(x)

Using the dir() Function


There is a built-in function to list all the function names (or variable names)
in a module. The dir() function:

Example
List all the defined names belonging to the platform module:

import platform

x = dir(platform)
print(x)
Note: The dir() function can be used on all modules, also the ones you
create yourself.

Import From Module


You can choose to import only parts from a module, by using
the from keyword.

Example
The module named mymodule has one function and one dictionary:
def greeting(name):
print("Hello, " + name)

person1 = {
"name": "John",
"age": 36,
"country": "Norway"
}

Example
Import only the person1 dictionary from the module:

from mymodule import person1

print (person1["age"])

Note: When importing using the from keyword, do not use the module name
when referring to elements in the module.
Example: person1["age"], not mymodule.person1["age"]

Python Datetime

Python Dates
A date in Python is not a data type of its own, but we can import a module
named datetime to work with dates as date objects.

ExampleGet your own Python Server


Import the datetime module and display the current date:

import datetime

x = [Link]()
print(x)
Date Output
When we execute the code from the example above the result will be:

2026-05-24 14:22:42.674664

The date contains year, month, day, hour, minute, second, and microsecond.

The datetime module has many methods to return information about the
date object.

Here are a few examples, you will learn more about them later in this
chapter:

Example
Return the year and name of weekday:

import datetime

x = [Link]()

print([Link])
print([Link]("%A"))

Creating Date Objects


To create a date, we can use the datetime() class (constructor) of
the datetime module.

The datetime() class requires three parameters to create a date: year,


month, day.

Example
Create a date object:

import datetime

x = [Link](2020, 5, 17)
print(x)

The datetime() class also takes parameters for time and timezone (hour,
minute, second, microsecond, tzone), but they are optional, and has a
default value of 0, (None for timezone).

The strftime() Method


The datetime object has a method for formatting date objects into readable
strings.

The method is called strftime(), and takes one parameter, format, to


specify the format of the returned string:

Example
Display the name of the month:

import datetime

x = [Link](2018, 6, 1)

print([Link]("%B"))

A reference of all the legal format codes:

Directive Description Examp

%a Weekday, short version Wed

%A Weekday, full version Wednes

%w Weekday as a number 0-6, 0 is Sunday 3


%d Day of month 01-31 31

%b Month name, short version Dec

%B Month name, full version Decemb

%m Month as a number 01-12 12

%y Year, short version, without century 18

%Y Year, full version 2018

%H Hour 00-23 17

%I Hour 00-12 05

%p AM/PM PM

%M Minute 00-59 41

%S Second 00-59 08
%f Microsecond 000000-999999 548513

%z UTC offset +0100

%Z Timezone CST

%j Day number of year 001-366 365

%U Week number of year, Sunday as the first day of 52


week, 00-53

%W Week number of year, Monday as the first day of 52


week, 00-53

%c Local version of date and time Mon De

%C Century 20

%x Local version of date 12/31/1

%X Local version of time 17:41:0

%% A % character %
%G ISO 8601 year 2018

%u ISO 8601 weekday (1-7) 1

%V ISO 8601 weeknumber (01-53) 01

Python Math
Python has a set of built-in math functions, including an extensive
math module, that allows you to perform mathematical tasks on
numbers.

Built-in Math Functions


The min() and max() functions can be used to find the lowest or highest
value in an iterable:

ExampleGet your own Python Server


x = min(5, 10, 25)
y = max(5, 10, 25)

print(x)
print(y)

The abs() function returns the absolute (positive) value of the specified
number:

Example
x = abs(-7.25)

print(x)

The pow(x, y) function returns the value of x to the power of y (x y).


Example
Return the value of 4 to the power of 3 (same as 4 * 4 * 4):

x = pow(4, 3)

print(x)

The Math Module


Python has also a built-in module called math, which extends the list of
mathematical functions.

To use it, you must import the math module:

import math

When you have imported the math module, you can start using methods and
constants of the module.

The [Link]() method for example, returns the square root of a number:

Example
import math

x = [Link](64)

print(x)

The [Link]() method rounds a number upwards to its nearest integer,


and the [Link]() method rounds a number downwards to its nearest
integer, and returns the result:

Example
import math

x = [Link](1.4)
y = [Link](1.4)

print(x) # returns 2
print(y) # returns 1
The [Link] constant, returns the value of PI (3.14...):

Example
import math

x = [Link]

print(x)

Python JSON
JSON is a syntax for storing and exchanging data.

JSON is text, written with JavaScript object notation.

JSON in Python
Python has a built-in package called json, which can be used to work with
JSON data.

ExampleGet your own Python Server


Import the json module:

import json

Parse JSON - Convert from JSON to


Python
If you have a JSON string, you can parse it by using
the [Link]() method.

The result will be a Python dictionary.

Example
Convert from JSON to Python:

import json

# some JSON:
x = '{ "name":"John", "age":30, "city":"New York"}'

# parse x:
y = [Link](x)

# the result is a Python dictionary:


print(y["age"])

Convert from Python to JSON


If you have a Python object, you can convert it into a JSON string by using
the [Link]() method.

Example
Convert from Python to JSON:

import json

# a Python object (dict):


x = {
"name": "John",
"age": 30,
"city": "New York"
}

# convert into JSON:


y = [Link](x)

# the result is a JSON string:


print(y)

You can convert Python objects of the following types, into JSON strings:

 dict
 list
 tuple
 string
 int
 float
 True
 False
 None
Example
Convert Python objects into JSON strings, and print the values:

import json

print([Link]({"name": "John", "age": 30}))


print([Link](["apple", "bananas"]))
print([Link](("apple", "bananas")))
print([Link]("hello"))
print([Link](42))
print([Link](31.76))
print([Link](True))
print([Link](False))
print([Link](None))

When you convert from Python to JSON, Python objects are converted into
the JSON (JavaScript) equivalent:

Python JSON

dict Object

list Array

tuple Array
str String

int Number

float Number

True true

False false

None null

Example
Convert a Python object containing all the legal data types:

import json

x = {
"name": "John",
"age": 30,
"married": True,
"divorced": False,
"children": ("Ann","Billy"),
"pets": None,
"cars": [
{"model": "BMW 230", "mpg": 27.5},
{"model": "Ford Edge", "mpg": 24.1}
]
}
print([Link](x))

Format the Result


The example above prints a JSON string, but it is not very easy to read, with
no indentations and line breaks.

The [Link]() method has parameters to make it easier to read the result:

Example
Use the indent parameter to define the numbers of indents:

[Link](x, indent=4)

You can also define the separators, default value is (", ", ": "), which means
using a comma and a space to separate each object, and a colon and a
space to separate keys from values:

Example
Use the separators parameter to change the default separator:

[Link](x, indent=4, separators=(". ", " = "))

Order the Result


The [Link]() method has parameters to order the keys in the result:

Example
Use the sort_keys parameter to specify if the result should be sorted or not:

[Link](x, indent=4, sort_keys=True)

Python RegEx
A RegEx, or Regular Expression, is a sequence of characters that forms
a search pattern.

RegEx can be used to check if a string contains the specified search


pattern.

RegEx Module
Python has a built-in package called re, which can be used to work with
Regular Expressions.

Import the re module:

import re

RegEx in Python
When you have imported the re module, you can start using regular
expressions:

ExampleGet your own Python Server


Search the string to see if it starts with "The" and ends with "Spain":

import re

txt = "The rain in Spain"


x = [Link]("^The.*Spain$", txt)

RegEx Functions
The re module offers a set of functions that allows us to search a string for a
match:
Function Description

findall Returns a list containing all matches

search Returns a Match object if there is a match anywhere in the string

split Returns a list where the string has been split at each match

sub Replaces one or many matches with a string

Metacharacters
Metacharacters are characters with a special meaning:

Character Description

[] A set of characters

\ Signals a special sequence (can also be used to escape special characters

. Any character (except newline character)


^ Starts with

$ Ends with

* Zero or more occurrences

+ One or more occurrences

? Zero or one occurrences

{} Exactly the specified number of occurrences

| Either or

() Capture and group

Flags
You can add flags to the pattern when using regular expressions.

Flag Shorthand Description


[Link] re.A Returns only ASCII matches

[Link] Returns debug information

[Link] re.S Makes the . character match all characters (including new

[Link] re.I Case-insensitive matching


SE

[Link] re.M Returns matches at the start/end of each line

[Link] Specifies that no flag is set for this pattern

[Link] re.U Returns Unicode matches. This is default from Python 3.


flag to return only Unicode matches

[Link] re.X Allows whitespaces and comments inside patterns. Make


readable

Special Sequences
A special sequence is a \ followed by one of the characters in the list below,
and has a special meaning:
Character Description

\A Returns a match if the specified characters are at the beginning of the str

\b Returns a match where the specified characters are at the beginning or at


of a word
(the "r" in the beginning is making sure that the string is being treated as
string")

\B Returns a match where the specified characters are present, but NOT at th
beginning (or at the end) of a word
(the "r" in the beginning is making sure that the string is being treated as
string")

\d Returns a match where the string contains digits (numbers from 0-9)

\D Returns a match where the string DOES NOT contain digits

\s Returns a match where the string contains a white space character

\S Returns a match where the string DOES NOT contain a white space charac

\w Returns a match where the string contains any word characters (characte
to Z, digits from 0-9, and the underscore _ character)
\W Returns a match where the string DOES NOT contain any word characters

\Z Returns a match if the specified characters are at the end of the string

Sets
A set is a set of characters inside a pair of square brackets [] with a special
meaning:

Set Description

[arn] Returns a match where one of the specified characters (a, r, or n) is prese

[a-n] Returns a match for any lower case character, alphabetically between a a

[^arn] Returns a match for any character EXCEPT a, r, and n

[0123] Returns a match where any of the specified digits (0, 1, 2, or 3) are presen

[0-9] Returns a match for any digit between 0 and 9

[0-5][0-9] Returns a match for any two-digit numbers from 00 and 59


[a-zA-Z] Returns a match for any character alphabetically between a and z, lower c

[+] In sets, +, *, ., |, (), $,{} has no special meaning, so [+] means: return a
any + character in the string

The findall() Function


The findall() function returns a list containing all matches.

Example
Print a list of all matches:

import re

txt = "The rain in Spain"


x = [Link]("ai", txt)
print(x)

The list contains the matches in the order they are found.

If no matches are found, an empty list is returned:

Example
Return an empty list if no match was found:

import re

txt = "The rain in Spain"


x = [Link]("Portugal", txt)
print(x)
The search() Function
The search() function searches the string for a match, and returns a Match
object if there is a match.

If there is more than one match, only the first occurrence of the match will
be returned:

Example
Search for the first white-space character in the string:

import re

txt = "The rain in Spain"


x = [Link]("\s", txt)

print("The first white-space character is located in position:",


[Link]())

If no matches are found, the value None is returned:

Example
Make a search that returns no match:

import re

txt = "The rain in Spain"


x = [Link]("Portugal", txt)
print(x)

The split() Function


The split() function returns a list where the string has been split at each
match:

Example
Split at each white-space character:

import re

txt = "The rain in Spain"


x = [Link]("\s", txt)
print(x)

You can control the number of occurrences by specifying


the maxsplit parameter:

Example
Split the string only at the first occurrence:

import re

txt = "The rain in Spain"


x = [Link]("\s", txt, 1)
print(x)

The sub() Function


The sub() function replaces the matches with the text of your choice:

Example
Replace every white-space character with the number 9:

import re

txt = "The rain in Spain"


x = [Link]("\s", "9", txt)
print(x)
You can control the number of replacements by specifying
the count parameter:

Example
Replace the first 2 occurrences:

import re

txt = "The rain in Spain"


x = [Link]("\s", "9", txt, 2)
print(x)

Match Object
A Match Object is an object containing information about the search and the
result.

Note: If there is no match, the value None will be returned, instead of the
Match Object.

Example
Do a search that will return a Match Object:

import re

txt = "The rain in Spain"


x = [Link]("ai", txt)
print(x) #this will print an object

The Match object has properties and methods used to retrieve information
about the search, and the result:

.span() returns a tuple containing the start-, and end positions of the
match.
.string returns the string passed into the function
.group() returns the part of the string where there was a match

Example
Print the position (start- and end-position) of the first match occurrence.

The regular expression looks for any words that starts with an upper case
"S":

import re

txt = "The rain in Spain"


x = [Link](r"\bS\w+", txt)
print([Link]())

Example
Print the string passed into the function:

import re

txt = "The rain in Spain"


x = [Link](r"\bS\w+", txt)
print([Link])

Example
Print the part of the string where there was a match.

The regular expression looks for any words that starts with an upper case
"S":

import re

txt = "The rain in Spain"


x = [Link](r"\bS\w+", txt)
print([Link]())
Note: If there is no match, the value None will be returned, instead of the
Match Object.

Python PIP

What is PIP?
PIP is a package manager for Python packages, or modules if you like.
Note: If you have Python version 3.4 or later, PIP is included by default.

What is a Package?
A package contains all the files you need for a module.

Modules are Python code libraries you can include in your project.

Check if PIP is Installed


Navigate your command line to the location of Python's script directory, and
type the following:

ExampleGet your own Python Server


Check PIP version:

C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>pip --version

Install PIP
If you do not have PIP installed, you can download and install it from this
page: [Link]

Download a Package
Downloading a package is very easy.

Open the command line interface and tell PIP to download the package you
want.
Navigate your command line to the location of Python's script directory, and
type the following:

Example
Download a package named "camelcase":

C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>pip install camelcase

Now you have downloaded and installed your first package!

REMOVE ADS

Using a Package
Once the package is installed, it is ready to use.

Import the "camelcase" package into your project.

Example
Import and use "camelcase":

import camelcase

c = [Link]()

txt = "hello world"

print([Link](txt))

Find Packages
Find more packages at [Link]
Remove a Package
Use the uninstall command to remove a package:

Example
Uninstall the package named "camelcase":

C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>pip uninstall camelcase

The PIP Package Manager will ask you to confirm that you want to remove
the camelcase package:

Uninstalling camelcase-02.1:
Would remove:
c:\users\Your Name\appdata\local\programs\python\python36-32\lib\
site-packages\[Link]-info
c:\users\Your Name\appdata\local\programs\python\python36-32\lib\
site-packages\camelcase\*
Proceed (y/n)?

Press y and the package will be removed.

List Packages
Use the list command to list all the packages installed on your system:

Example
List installed packages:

C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>pip list

Result:

Package Version
-----------------------
camelcase 0.2
mysql-connector 2.1.6
pip 18.1
pymongo 3.6.1
setuptools 39.0.1

Python Try Except

The try block lets you test a block of code for errors.

The except block lets you handle the error.

The else block lets you execute code when there is no error.

The finally block lets you execute code, regardless of the result of
the try- and except blocks.

Exception Handling
When an error occurs, or exception as we call it, Python will normally stop
and generate an error message.

These exceptions can be handled using the try statement:

ExampleGet your own Python Server


The try block will generate an exception, because x is not defined:

try:
print(x)
except:
print("An exception occurred")

Since the try block raises an error, the except block will be executed.

Without the try block, the program will crash and raise an error:

Example
This statement will raise an error, because x is not defined:
print(x)

Many Exceptions
You can define as many exception blocks as you want, e.g. if you want to
execute a special block of code for a special kind of error:

Example
Print one message if the try block raises a NameError and another for other
errors:

try:
print(x)
except NameError:
print("Variable x is not defined")
except:
print("Something else went wrong")

Else
You can use the else keyword to define a block of code to be executed if no
errors were raised:

Example
In this example, the try block does not generate any error:

try:
print("Hello")
except:
print("Something went wrong")
else:
print("Nothing went wrong")

Finally
The finally block, if specified, will be executed regardless if the try block
raises an error or not.

Example
try:
print(x)
except:
print("Something went wrong")
finally:
print("The 'try except' is finished")

This can be useful to close objects and clean up resources:

Example
Try to open and write to a file that is not writable:

try:
f = open("[Link]")
try:
[Link]("Lorum Ipsum")
except:
print("Something went wrong when writing to the file")
finally:
[Link]()
except:
print("Something went wrong when opening the file")

The program can continue, without leaving the file object open.

Raise an exception
As a Python developer you can choose to throw an exception if a condition
occurs.

To throw (or raise) an exception, use the raise keyword.

Example
Raise an error and stop the program if x is lower than 0:

x = -1
if x < 0:
raise Exception("Sorry, no numbers below zero")

The raise keyword is used to raise an exception.

You can define what kind of error to raise, and the text to print to the user.

Example
Raise a TypeError if x is not an integer:

x = "hello"

if not type(x) is int:


raise TypeError("Only integers are allowed")

Python String Formatting

F-String was introduced in Python 3.6, and is now the preferred way of
formatting strings.

Before Python 3.6 we had to use the format() method.

F-Strings
F-string allows you to format selected parts of a string.

To specify a string as an f-string, simply put an f in front of the string literal,


like this:

ExampleGet your own Python Server


Create an f-string:

txt = f"The price is 49 dollars"


print(txt)
Placeholders and Modifiers
To format values in an f-string, add placeholders {}, a placeholder can
contain variables, operations, functions, and modifiers to format the value.

Example
Add a placeholder for the price variable:

price = 59
txt = f"The price is {price} dollars"
print(txt)

A placeholder can also include a modifier to format the value.

A modifier is included by adding a colon : followed by a legal formatting


type, like .2f which means fixed point number with 2 decimals:

Example
Display the price with 2 decimals:

price = 59
txt = f"The price is {price:.2f} dollars"
print(txt)

You can also format a value directly without keeping it in a variable:

Example
Display the value 95 with 2 decimals:

txt = f"The price is {95:.2f} dollars"


print(txt)

Perform Operations in F-Strings


You can perform Python operations inside the placeholders.
You can do math operations:

Example
Perform a math operation in the placeholder, and return the result:

txt = f"The price is {20 * 59} dollars"


print(txt)

You can perform math operations on variables:

Example
Add taxes before displaying the price:

price = 59
tax = 0.25
txt = f"The price is {price + (price * tax)} dollars"
print(txt)

You can perform if...else statements inside the placeholders:

Example
Return "Expensive" if the price is over 50, otherwise return "Cheap":

price = 49
txt = f"It is very {'Expensive' if price>50 else 'Cheap'}"

print(txt)

Execute Functions in F-Strings


You can execute functions inside the placeholder:

Example
Use the string method upper()to convert a value into upper case letters:

fruit = "apples"
txt = f"I love {[Link]()}"
print(txt)
The function does not have to be a built-in Python method, you can create
your own functions and use them:

Example
Create a function that converts feet into meters:

def myconverter(x):
return x * 0.3048

txt = f"The plane is flying at a {myconverter(30000)} meter altitude"


print(txt)

More Modifiers
At the beginning of this chapter we explained how to use the .2f modifier to
format a number into a fixed point number with 2 decimals.

There are several other modifiers that can be used to format values:

Example
Use a comma as a thousand separator:

price = 59000
txt = f"The price is {price:,} dollars"
print(txt)

Here is a list of all the formatting types.

Formatting Types

:< Left aligns the result (within the available space)

:> Right aligns the result (within the available space)


:^ Center aligns the result (within the available space)

:= Places the sign to the left most position

:+ Use a plus sign to indicate if the result is positive or negative

:- Use a minus sign for negative values only

: Use a space to insert an extra space before positive numbers (and a minus sign

:, Use a comma as a thousand separator

:_ Use a underscore as a thousand separator

:b Binary format

:c Converts the value into the corresponding Unicode character

:d Decimal format

:e Scientific format, with a lower case e


:E Scientific format, with an upper case E

:f Fix point number format

:F Fix point number format, in uppercase format (show inf and nan as INF and NA

:g General format

:G General format (using a upper case E for scientific notations)

:o Octal format

:x Hex format, lower case

:X Hex format, upper case

:n Number format

:% Percentage format

String format()
Before Python 3.6 we used the format() method to format strings.

The format() method can still be used, but f-strings are faster and the
preferred way to format strings.

The next examples in this page demonstrates how to format strings with
the format() method.

The format() method also uses curly brackets as placeholders {}, but the
syntax is slightly different:

Example
Add a placeholder where you want to display the price:

price = 49
txt = "The price is {} dollars"
print([Link](price))

You can add parameters inside the curly brackets to specify how to convert
the value:

Example
Format the price to be displayed as a number with two decimals:

txt = "The price is {:.2f} dollars"

Check out all formatting types in our String format() Reference.

Multiple Values
If you want to use more values, just add more values to
the format() method:

print([Link](price, itemno, count))

And add more placeholders:

Example
quantity = 3
itemno = 567
price = 49
myorder = "I want {} pieces of item number {} for {:.2f} dollars."
print([Link](quantity, itemno, price))

Index Numbers
You can use index numbers (a number inside the curly brackets {0}) to be
sure the values are placed in the correct placeholders:

Example
quantity = 3
itemno = 567
price = 49
myorder = "I want {0} pieces of item number {1} for {2:.2f} dollars."
print([Link](quantity, itemno, price))

Also, if you want to refer to the same value more than once, use the index
number:

Example
age = 36
name = "John"
txt = "His name is {1}. {1} is {0} years old."
print([Link](age, name))

Named Indexes
You can also use named indexes by entering a name inside the curly
brackets {carname}, but then you must use names when you pass the
parameter values [Link](carname = "Ford"):

Example
myorder = "I have a {carname}, it is a {model}."
print([Link](carname = "Ford", model = "Mustang"))
Python None

Python None
None is a special constant in Python that represents the absence of a value.

Its data type is NoneType, and None is the only instance of a NoneType object.

NoneType
Variables can be assigned None to indicate "no value" or "not set".

ExampleGet your own Python Server


Assign and display a None value:

x = None
print(x)

Use type() to see the type of a None value.

Example
Assign and print the data type of a None value:

x = None
print(type(x))

Comparing to None
To compare a value to None, use the identity operator is or is not

Example
Use the identity operator is for comparisons with None:

result = None
if result is None:
print("No result yet")
else:
print("Result is ready")

Example
Similar example, but using is not instead:

result = None
if result is not None:
print("Result is ready")
else:
print("No result yet")

True or False
None evaluates to False in a boolean context.

Example
Check truthiness:

print(bool(None))

Functions returning None


Functions that do not explicitly return a value return None by default.

Example
A function without a return statement returns None:

def myfunc():
x = 5

x = myfunc()
print(x)
Python User Input

User Input
Python allows for user input.

That means we are able to ask the user for input.

The following example asks for your name, and when you enter a name, it
gets printed on the screen:

ExampleGet your own Python Server


Ask for user input:

print("Enter your name:")


name = input()
print(f"Hello {name}")

Python stops executing when it comes to the input() function, and


continues when the user has given some input.

Using prompt
In the example above, the user had to input their name on a new line. The
Python input() function has a prompt parameter, which acts as a message
you can put in front of the user input, on the same line:

Example
Add a message in front of the user input:

name = input("Enter your name:")


print(f"Hello {name}")
Multiple Inputs
You can add as many inputs as you want, Python will stop executing at each
of them, waiting for user input:

Example
Multiple inputs:

name = input("Enter your name:")


print(f"Hello {name}")
fav1 = input("What is your favorite animal:")
fav2 = input("What is your favorite color:")
fav3 = input("What is your favorite number:")
print(f"Do you want a {fav2} {fav1} with {fav3} legs?")

Input Number
The input from the user is treated as a string. Even if, in the example above,
you can input a number, the Python interpreter will still treat it as a string.

You can convert the input into a number with the float() function:

Example
To find the square root, the input has to be converted into a number:

x = input("Enter a number:")

#find the square root of the number:


y = [Link](float(x))

print(f"The square root of {x} is {y}")

Validate Input
It is a good practice to validate any input from the user. In the example
above, an error will occur if the user inputs something other than a number.
To avoid getting an error, we can test the input, and if it is not a number, the
user could get a message like "Wrong input, please try again", and allowed
to make a new input:

Example
Keep asking until you get a number:

y = True
while y == True:
x = input("Enter a number:")
try:
x = float(x);
y = False
except:
print("Wrong input, please try again.")

print("Thank you!")

Python Virtual Environment


What is a Virtual Environment?
A virtual environment in Python is an isolated environment on your
computer, where you can run and test your Python projects.

It allows you to manage project-specific dependencies without interfering


with other projects or the original Python installation.

Think of a virtual environment as a separate container for each Python


project. Each container:

 Has its own Python interpreter


 Has its own set of installed packages
 Is isolated from other virtual environments
 Can have different versions of the same package

Using virtual environments is important because:

 It prevents package version conflicts between projects


 Makes projects more portable and reproducible
 Keeps your system Python installation clean
 Allows testing with different Python versions

Creating a Virtual Environment


Python has the built-in venv module for creating virtual environments.

To create a virtual environment on your computer, open the command


prompt, and navigate to the folder where you want to create your project,
then type this command:

ExampleGet your own Python Server


Run this command to create a virtual environment named myfirstproject:

Windows
macOS/Linux
C:\Users\Your Name> python -m venv myfirstproject

This will set up a virtual environment, and create a folder named


"myfirstproject" with subfolders and files, like this:

Result
The file/folder structure will look like this:

myfirstproject
Include
Lib
Scripts
.gitignore
[Link]

Activate Virtual Environment


To use the virtual environment, you have to activate it with this command:

Example
Activate the virtual environment:

Windows
macOS/Linux
C:\Users\Your Name> myfirstproject\Scripts\activate

After activation, your prompt will change to show that you are now working
in the active environment:

Result
The command line will look like this when the virtual environment is active:

Windows
macOS/Linux
(myfirstproject) C:\Users\Your Name>

Install Packages
Once your virtual environment is activated, you can install packages in it,
using pip.

We will install a package called 'cowsay':

Example
Install 'cowsay' in the virtual environment:

Windows
macOS/Linux
(myfirstproject) C:\Users\Your Name> pip install cowsay

Result
'cowsay' is installed only in the virtual environment:

Collecting cowsay
Downloading [Link] (5.6 kB)
Downloading [Link] (25 kB)
Installing collected packages: cowsay
Successfully installed cowsay-6.1
[notice] A new release of pip is available: 25.0.1 -> 25.1.1
[notice] To update, run: [Link] -m pip install --upgrade pip

Using Package
Now that the 'cowsay' module is installed in your virtual environment, lets
use it to display a talking cow.

Create a file called [Link] on your computer. You can place it wherever you
want, but I will place it in the same location as the myfirstproject folder -
not in the folder, but in the same location.

Open the file and insert these three lines in it:

Example
Insert two lines in [Link]:

[Link]
import cowsay

[Link]("Good Mooooorning!")

Then, try to execute the file while you are in the virtual environment:

Example
Execute [Link] in the virtual environment:

Windows
macOS/Linux
(myfirstproject) C:\Users\Your Name> python [Link]

As a result a cow will appear in you terminal:

Result
The purpose of the 'cowsay' module is to draw a cow that says whatever
input you give it:

_________________
| Good Mooooorning! |
=================
\
\
^__^
(oo)\_______
(__)\ )\/\
||----w |
|| ||

Deactivate Virtual Environment


To deactivate the virtual environment use this command:

Example
Deactivate the virtual environment:

Windows
macOS/Linux
(myfirstproject) C:\Users\Your Name> deactivate

As a result, you are now back in the normal command line interface:

Result
Normal command line interface:

Windows
macOS/Linux
C:\Users\Your Name>

If you try to execute the [Link] file outside of the virtual environment, you
will get an error because 'cowsay' is missing. It was only installed in the
virtual environment:

Example
Execute [Link] outside of the virtual environment:

Windows
macOS/Linux
C:\Users\Your Name> python [Link]

Result
Error because 'cowsay' is missing:

Traceback (most recent call last):


File "C:\Users\Your Name\[Link]", line 1, in <module>
import cowsay
ModuleNotFoundError: No module named 'cowsay'
Note: The virtual environment myfirstproject still exists, it is just not
activated. If you activate the virtual environment again, you can execute
the [Link] file, and the diagram will be displayed.

Delete Virtual Environment


Another nice thing about working with a virtual environment is that when
you, for some reason want to delete it, there are no other projects depend on
it, and only the modules and files in the specified virtual environment are
deleted.

To delete a virtual environment, you can simply delete its folder with all its
content. Either directly in the file system, or use the command line interface
like this:

Example
Delete myfirstproject from the command line interface:

Windows
macOS/Linux
C:\Users\Your Name> rmdir /s /q myfirstproject

Python OOP
What is OOP?
OOP stands for Object-Oriented Programming.

Python is an object-oriented language, allowing you to structure your code


using classes and objects for better organization and reusability.
Advantages of OOP
 Provides a clear structure to programs
 Makes code easier to maintain, reuse, and debug
 Helps keep your code DRY (Don't Repeat Yourself)
 Allows you to build reusable applications with less code

Tip: The DRY principle means you should avoid writing the same code more
than once. Move repeated code into functions or classes and reuse it.

What are Classes and Objects?


Classes and objects are the two core concepts in object-oriented
programming.

A class defines what an object should look like, and an object is created
based on that class. For example:

Class Objects

Fruit Apple, Banana, Mango

Car Volvo, Audi, Toyota

When you create an object from a class, it inherits all the variables and
functions defined inside that class.

In the next chapters, you will learn about:

 Classes and objects


 The __init__() method
 The self parameter
 Properties and methods
 Inheritance and polymorphism
 Encapsulation and inner classes
Python Classes and Objects
Python Classes/Objects
Python is an object oriented programming language.

Almost everything in Python is an object, with its properties and methods.

A Class is like an object constructor, or a "blueprint" for creating objects.

Create a Class
To create a class, use the keyword class:

ExampleGet your own Python Server


Create a class named MyClass, with a property named x:

class MyClass:
x = 5

Create Object
Now we can use the class named MyClass to create objects:

Example
Create an object named p1, and print the value of x:

p1 = MyClass()
print(p1.x)

Delete Objects
You can delete objects by using the del keyword:

Example
Delete the p1 object:

del p1

Multiple Objects
You can create multiple objects from the same class:

Example
Create three objects from the MyClass class:

p1 = MyClass()
p2 = MyClass()
p3 = MyClass()

print(p1.x)
print(p2.x)
print(p3.x)
Note: Each object is independent and has its own copy of the class
properties.

The pass Statement


class definitions cannot be empty, but if you for some reason have
a class definition with no content, put in the pass statement to avoid
getting an error.

Example
class Person:
pass

Python __init__() Method


The __init__() Method
All classes have a built-in method called __init__(), which is always
executed when the class is being initiated.

The __init__() method is used to assign values to object properties, or to


perform operations that are necessary when the object is being created.

ExampleGet your own Python Server


Create a class named Person, use the __init__() method to assign values
for name and age:

class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age

p1 = Person("Emil", 36)

print([Link])
print([Link])
Note: The __init__() method is called automatically every time the class is
being used to create a new object.

Why Use __init__()?


Without the __init__() method, you would need to set properties manually
for each object:

Example
Create a class without __init__():

class Person:
pass

p1 = Person()
[Link] = "Tobias"
[Link] = 25

print([Link])
print([Link])
Using __init__() makes it easier to create objects with initial values:

Example
With __init__(), you can set initial values when creating the object:

class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age

p1 = Person("Linus", 28)

print([Link])
print([Link])

Default Values in __init__()


You can also set default values for parameters in the __init__() method:

Example
Set a default value for the age parameter:

class Person:
def __init__(self, name, age=18):
[Link] = name
[Link] = age

p1 = Person("Emil")
p2 = Person("Tobias", 25)

print([Link], [Link])
print([Link], [Link])

Multiple Parameters
The __init__() method can have as many parameters as you need:

Example
Create a Person class with multiple parameters:
class Person:
def __init__(self, name, age, city, country):
[Link] = name
[Link] = age
[Link] = city
[Link] = country

p1 = Person("Linus", 30, "Oslo", "Norway")

print([Link])
print([Link])
print([Link])
print([Link])

Python self Parameter

The self Parameter


The self parameter is a reference to the current instance of the class.

It is used to access properties and methods that belong to the class.

ExampleGet your own Python Server


Use self to access class properties:

class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age

def greet(self):
print("Hello, my name is " + [Link])

p1 = Person("Emil", 25)
[Link]()
Note: The self parameter must be the first parameter of any method in the
class.
Why Use self?
Without self, Python would not know which object's properties you want to
access:

Example
The self parameter links the method to the specific object:

class Person:
def __init__(self, name):
[Link] = name

def printname(self):
print([Link])

p1 = Person("Tobias")
p2 = Person("Linus")

[Link]()
[Link]()

REMOVE ADS

self Does Not Have to Be Named "self"


It does not have to be named self, you can call it whatever you like, but it
has to be the first parameter of any method in the class:

Example
Use the words myobject and abc instead of self:

class Person:
def __init__(myobject, name, age):
[Link] = name
[Link] = age

def greet(abc):
print("Hello, my name is " + [Link])
p1 = Person("Emil", 36)
[Link]()
Note: While you can use a different name, it is strongly recommended to
use self as it is the convention in Python and makes your code more
readable to others.

Accessing Properties with self


You can access any property of the class using self:

Example
Access multiple properties using self:

class Car:
def __init__(self, brand, model, year):
[Link] = brand
[Link] = model
[Link] = year

def display_info(self):
print(f"{[Link]} {[Link]} {[Link]}")

car1 = Car("Toyota", "Corolla", 2020)


car1.display_info()
REMOVE ADS

Calling Methods with self


You can also call other methods within the class using self:

Example
Call one method from another method using self:

class Person:
def __init__(self, name):
[Link] = name
def greet(self):
return "Hello, " + [Link]

def welcome(self):
message = [Link]()
print(message + "! Welcome to our website.")

p1 = Person("Tobias")
[Link]()

Python Class Properties


Class Properties
Properties are variables that belong to a class. They store data for each
object created from the class.

ExampleGet your own Python Server


Create a class with properties:

class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age

p1 = Person("Emil", 36)

print([Link])
print([Link])

Access Properties
You can access object properties using dot notation:

Example
Access the properties of an object:
class Car:
def __init__(self, brand, model):
[Link] = brand
[Link] = model

car1 = Car("Toyota", "Corolla")

print([Link])
print([Link])

REMOVE ADS

Modify Properties
You can modify the value of properties on objects:

Example
Change the age property:

class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age

p1 = Person("Tobias", 25)
print([Link])

[Link] = 26
print([Link])

Delete Properties
You can delete properties from objects using the del keyword:

Example
Delete the age property:

class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age

p1 = Person("Linus", 30)

del [Link]

print([Link]) # This works


# print([Link]) # This would cause an error
REMOVE ADS

Class Properties vs Object Properties


Properties defined inside __init__() belong to each object (instance
properties).

Properties defined outside methods belong to the class itself (class


properties) and are shared by all objects:

Example
Class property vs instance property:

class Person:
species = "Human" # Class property

def __init__(self, name):


[Link] = name # Instance property

p1 = Person("Emil")
p2 = Person("Tobias")

print([Link])
print([Link])
print([Link])
print([Link])
Modifying Class Properties
When you modify a class property, it affects all objects:

Example
Change a class property:

class Person:
lastname = ""

def __init__(self, name):


[Link] = name

p1 = Person("Linus")
p2 = Person("Emil")

[Link] = "Refsnes"

print([Link])
print([Link])

Add New Properties


You can add new properties to existing objects:

Example
Add a new property to an object:

class Person:
def __init__(self, name):
[Link] = name

p1 = Person("Tobias")

[Link] = 25
[Link] = "Oslo"

print([Link])
print([Link])
print([Link])
Note: Adding properties this way only adds them to that specific object, not
to all objects of the class.

Python Class Methods

Class Methods
Methods are functions that belong to a class. They define the behavior of
objects created from the class.

ExampleGet your own Python Server


Create a method in a class:

class Person:
def __init__(self, name):
[Link] = name

def greet(self):
print("Hello, my name is " + [Link])

p1 = Person("Emil")
[Link]()
Note: All methods must have self as the first parameter.

Methods with Parameters


Methods can accept parameters just like regular functions:

Example
Create a method with parameters:

class Calculator:
def add(self, a, b):
return a + b

def multiply(self, a, b):


return a * b

calc = Calculator()
print([Link](5, 3))
print([Link](4, 7))

REMOVE ADS

Methods Accessing Properties


Methods can access and modify object properties using self:

Example
A method that accesses object properties:

class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age

def get_info(self):
return f"{[Link]} is {[Link]} years old"

p1 = Person("Tobias", 28)
print(p1.get_info())

Methods Modifying Properties


Methods can modify the properties of an object:

Example
A method that changes a property value:

class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age

def celebrate_birthday(self):
[Link] += 1
print(f"Happy birthday! You are now {[Link]}")

p1 = Person("Linus", 25)
p1.celebrate_birthday()
p1.celebrate_birthday()

The __str__() Method


The __str__() method is a special method that controls what is returned
when the object is printed:

Example
Without the __str__() method:

class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age

p1 = Person("Emil", 36)
print(p1)

Example
With the __str__() method:

class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age

def __str__(self):
return f"{[Link]} ({[Link]})"

p1 = Person("Tobias", 36)
print(p1)
Multiple Methods
A class can have multiple methods that work together:

Example
Create multiple methods in a class:

class Playlist:
def __init__(self, name):
[Link] = name
[Link] = []

def add_song(self, song):


[Link](song)
print(f"Added: {song}")

def remove_song(self, song):


if song in [Link]:
[Link](song)
print(f"Removed: {song}")

def show_songs(self):
print(f"Playlist '{[Link]}':")
for song in [Link]:
print(f"- {song}")

my_playlist = Playlist("Favorites")
my_playlist.add_song("Bohemian Rhapsody")
my_playlist.add_song("Stairway to Heaven")
my_playlist.show_songs()

Delete Methods
You can delete methods from a class using the del keyword:

Example
Delete a method from a class:

class Person:
def __init__(self, name):
[Link] = name

def greet(self):
print("Hello!")

p1 = Person("Emil")

del [Link]

[Link]() # This will cause an error

Python Inheritance
Python Inheritance
Inheritance allows us to define a class that inherits all the methods and
properties from another class.

Parent class is the class being inherited from, also called base class.

Child class is the class that inherits from another class, also called derived
class.

Create a Parent Class


Any class can be a parent class, so the syntax is the same as creating any
other class:

ExampleGet your own Python Server


Create a class named Person, with firstname and lastname properties, and
a printname method:

class Person:
def __init__(self, fname, lname):
[Link] = fname
[Link] = lname

def printname(self):
print([Link], [Link])
#Use the Person class to create an object, and then execute the
printname method:

x = Person("John", "Doe")
[Link]()

Create a Child Class


To create a class that inherits the functionality from another class, send the
parent class as a parameter when creating the child class:

Example
Create a class named Student, which will inherit the properties and methods
from the Person class:

class Student(Person):
pass
Note: Use the pass keyword when you do not want to add any other
properties or methods to the class.

Now the Student class has the same properties and methods as the Person
class.

Example
Use the Student class to create an object, and then execute
the printname method:

x = Student("Mike", "Olsen")
[Link]()

REMOVE ADS

Add the __init__() Function


So far we have created a child class that inherits the properties and methods
from its parent.

We want to add the __init__() function to the child class (instead of


the pass keyword).

Note: The __init__() function is called automatically every time the class is
being used to create a new object.

Example
Add the __init__() function to the Student class:

class Student(Person):
def __init__(self, fname, lname):
#add properties etc.

When you add the __init__() function, the child class will no longer inherit
the parent's __init__() function.

Note: The child's __init__() function overrides the inheritance of the


parent's __init__() function.

To keep the inheritance of the parent's __init__() function, add a call to the
parent's __init__() function:

Example
class Student(Person):
def __init__(self, fname, lname):
Person.__init__(self, fname, lname)

Now we have successfully added the __init__() function, and kept the
inheritance of the parent class, and we are ready to add functionality in
the __init__() function.

REMOVE ADS

Use the super() Function


Python also has a super() function that will make the child class inherit all
the methods and properties from its parent:
Example
class Student(Person):
def __init__(self, fname, lname):
super().__init__(fname, lname)

By using the super() function, you do not have to use the name of the
parent element, it will automatically inherit the methods and properties from
its parent.

Add Properties
Example
Add a property called graduationyear to the Student class:

class Student(Person):
def __init__(self, fname, lname):
super().__init__(fname, lname)
[Link] = 2019

In the example below, the year 2019 should be a variable, and passed into
the Student class when creating student objects. To do so, add another
parameter in the __init__() function:

Example
Add a year parameter, and pass the correct year when creating objects:

class Student(Person):
def __init__(self, fname, lname, year):
super().__init__(fname, lname)
[Link] = year

x = Student("Mike", "Olsen", 2019)

Add Methods
Example
Add a method called welcome to the Student class:

class Student(Person):
def __init__(self, fname, lname, year):
super().__init__(fname, lname)
[Link] = year

def welcome(self):
print("Welcome", [Link], [Link], "to the class
of", [Link])

If you add a method in the child class with the same name as a function in
the parent class, the inheritance of the parent method will be overridden.

Python Polymorphism

The word "polymorphism" means "many forms", and in programming it


refers to methods/functions/operators with the same name that can be
executed on many objects or classes.

Function Polymorphism
An example of a Python function that can be used on different objects is
the len() function.

String
For strings len() returns the number of characters:

ExampleGet your own Python Server


x = "Hello World!"

print(len(x))

Tuple
For tuples len() returns the number of items in the tuple:
Example
mytuple = ("apple", "banana", "cherry")

print(len(mytuple))

Dictionary
For dictionaries len() returns the number of key/value pairs in the
dictionary:

Example
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}

print(len(thisdict))

REMOVE ADS

Class Polymorphism
Polymorphism is often used in Class methods, where we can have multiple
classes with the same method name.

For example, say we have three classes: Car, Boat, and Plane, and they all
have a method called move():

Example
Different classes with the same method:

class Car:
def __init__(self, brand, model):
[Link] = brand
[Link] = model

def move(self):
print("Drive!")

class Boat:
def __init__(self, brand, model):
[Link] = brand
[Link] = model

def move(self):
print("Sail!")

class Plane:
def __init__(self, brand, model):
[Link] = brand
[Link] = model

def move(self):
print("Fly!")

car1 = Car("Ford", "Mustang") #Create a Car object


boat1 = Boat("Ibiza", "Touring 20") #Create a Boat object
plane1 = Plane("Boeing", "747") #Create a Plane object

for x in (car1, boat1, plane1):


[Link]()

Look at the for loop at the end. Because of polymorphism we can execute
the same method for all three classes.

Inheritance Class Polymorphism


What about classes with child classes with the same name? Can we use
polymorphism there?

Yes. If we use the example above and make a parent class called Vehicle,
and make Car, Boat, Plane child classes of Vehicle, the child classes
inherits the Vehicle methods, but can override them:

Example
Create a class called Vehicle and make Car, Boat, Plane child classes
of Vehicle:
class Vehicle:
def __init__(self, brand, model):
[Link] = brand
[Link] = model

def move(self):
print("Move!")

class Car(Vehicle):
pass

class Boat(Vehicle):
def move(self):
print("Sail!")

class Plane(Vehicle):
def move(self):
print("Fly!")

car1 = Car("Ford", "Mustang") #Create a Car object


boat1 = Boat("Ibiza", "Touring 20") #Create a Boat object
plane1 = Plane("Boeing", "747") #Create a Plane object

for x in (car1, boat1, plane1):


print([Link])
print([Link])
[Link]()

Child classes inherits the properties and methods from the parent class.

In the example above you can see that the Car class is empty, but it
inherits brand, model, and move() from Vehicle.

The Boat and Plane classes also inherit brand, model,


and move() from Vehicle, but they both override the move() method.

Because of polymorphism we can execute the same method for all classes.

Python Encapsulation

Python Encapsulation
Encapsulation is about protecting data inside a class.

It means keeping data (properties) and methods together in a class, while


controlling how the data can be accessed from outside the class.

This prevents accidental changes to your data and hides the internal details
of how your class works.

Private Properties
In Python, you can make properties private by using a double
underscore __ prefix:

ExampleGet your own Python Server


Create a private class property named __age:

class Person:
def __init__(self, name, age):
[Link] = name
self.__age = age # Private property

p1 = Person("Emil", 25)
print([Link])
print(p1.__age) # This will cause an error
Note: Private properties cannot be accessed directly from outside the class.

Get Private Property Value


To access a private property, you can create a getter method:

Example
Use a getter method to access a private property:

class Person:
def __init__(self, name, age):
[Link] = name
self.__age = age
def get_age(self):
return self.__age

p1 = Person("Tobias", 25)
print(p1.get_age())

Set Private Property Value


To modify a private property, you can create a setter method.

The setter method can also validate the value before setting it:

Example
Use a setter method to change a private property:

class Person:
def __init__(self, name, age):
[Link] = name
self.__age = age

def get_age(self):
return self.__age

def set_age(self, age):


if age > 0:
self.__age = age
else:
print("Age must be positive")

p1 = Person("Tobias", 25)
print(p1.get_age())

p1.set_age(26)
print(p1.get_age())

REMOVE ADS
Why Use Encapsulation?
Encapsulation provides several benefits:

 Data Protection: Prevents accidental modification of data


 Validation: You can validate data before setting it
 Flexibility: Internal implementation can change without affecting
external code
 Control: You have full control over how data is accessed and modified
Example
Use encapsulation to protect and validate data:

class Student:
def __init__(self, name):
[Link] = name
self.__grade = 0

def set_grade(self, grade):


if 0 <= grade <= 100:
self.__grade = grade
else:
print("Grade must be between 0 and 100")

def get_grade(self):
return self.__grade

def get_status(self):
if self.__grade >= 60:
return "Passed"
else:
return "Failed"

student = Student("Emil")
student.set_grade(85)
print(student.get_grade())
print(student.get_status())

Protected Properties
Python also has a convention for protected properties using a single
underscore _ prefix:
Example
Create a protected property:

class Person:
def __init__(self, name, salary):
[Link] = name
self._salary = salary # Protected property

p1 = Person("Linus", 50000)
print([Link])
print(p1._salary) # Can access, but shouldn't
Note: A single underscore _ is just a convention. It tells other programmers
that the property is intended for internal use, but Python doesn't enforce this
restriction.

Private Methods
You can also make methods private using the double underscore prefix:

Example
Create a private method:

class Calculator:
def __init__(self):
[Link] = 0

def __validate(self, num):


if not isinstance(num, (int, float)):
return False
return True

def add(self, num):


if self.__validate(num):
[Link] += num
else:
print("Invalid number")

calc = Calculator()
[Link](10)
[Link](5)
print([Link])
# calc.__validate(5) # This would cause an error
Note: Just like private properties with double underscores, private methods
cannot be called directly from outside the class. The __validate method can
only be used by other methods inside the class.

Name Mangling
Name mangling is how Python implements private properties and methods.

When you use double underscores __, Python automatically renames it


internally by adding _ClassName in front.

For example, __age becomes _Person__age.

Example
See how Python mangles the name:

class Person:
def __init__(self, name, age):
[Link] = name
self.__age = age

p1 = Person("Emil", 30)

# This is how Python mangles the name:


print(p1._Person__age) # Not recommended!
Note: While you can access private properties using the mangled name, it's
not recommended. It defeats the purpose of encapsulation.

Python Inner Classes


Python Inner Classes
An inner class is a class defined inside another class. The inner class can
access the properties and methods of the outer class.

Inner classes are useful for grouping classes that are only used in one place,
making your code more organized.
ExampleGet your own Python Server
Create an inner class:

class Outer:
def __init__(self):
[Link] = "Outer Class"

class Inner:
def __init__(self):
[Link] = "Inner Class"

def display(self):
print("This is the inner class")

outer = Outer()
print([Link])

Accessing Inner Class from the


Outside
To access the inner class, create an object of the outer class, and then
create an object of the inner class:

Example
Access the inner class and create an object:

class Outer:
def __init__(self):
[Link] = "Outer"

class Inner:
def __init__(self):
[Link] = "Inner"

def display(self):
print("Hello from inner class")

outer = Outer()
inner = [Link]()
[Link]()
REMOVE ADS

Accessing Outer Class from Inner


Class
Inner classes in Python do not automatically have access to the outer class
instance.

If you want the inner class to access the outer class, you need to pass the
outer class instance as a parameter:

Example
Pass the outer class instance to the inner class:

class Outer:
def __init__(self):
[Link] = "Emil"

class Inner:
def __init__(self, outer):
[Link] = outer

def display(self):
print(f"Outer class name: {[Link]}")

outer = Outer()
inner = [Link](outer)
[Link]()

Practical Example
Inner classes are useful for creating helper classes that are only used within
the context of the outer class:

Example
Use an inner class to represent a car's engine:

class Car:
def __init__(self, brand, model):
[Link] = brand
[Link] = model
[Link] = [Link]()

class Engine:
def __init__(self):
[Link] = "Off"

def start(self):
[Link] = "Running"
print("Engine started")

def stop(self):
[Link] = "Off"
print("Engine stopped")

def drive(self):
if [Link] == "Running":
print(f"Driving the {[Link]} {[Link]}")
else:
print("Start the engine first!")

car = Car("Toyota", "Corolla")


[Link]()
[Link]()
[Link]()
REMOVE ADS

Multiple Inner Classes


A class can have multiple inner classes:

Example
Create multiple inner classes:

class Computer:
def __init__(self):
[Link] = [Link]()
[Link] = [Link]()

class CPU:
def process(self):
print("Processing data...")

class RAM:
def store(self):
print("Storing data...")

computer = Computer()
[Link]()
[Link]()

Python File Open

File handling is an important part of any web application.

Python has several functions for creating, reading, updating, and


deleting files.

File Handling
The key function for working with files in Python is the open() function.

The open() function takes two parameters; filename, and mode.

There are four different methods (modes) for opening a file:

"r" - Read - Default value. Opens a file for reading, error if the file does not
exist

"a" - Append - Opens a file for appending, creates the file if it does not exist

"w" - Write - Opens a file for writing, creates the file if it does not exist

"x" - Create - Creates the specified file, returns an error if the file exists

In addition you can specify if the file should be handled as binary or text
mode
"t" - Text - Default value. Text mode

"b" - Binary - Binary mode (e.g. images)

Syntax
To open a file for reading it is enough to specify the name of the file:

f = open("[Link]")

The code above is the same as:

f = open("[Link]", "rt")

Because "r" for read, and "t" for text are the default values, you do not need
to specify them.

Note: Make sure the file exists, or else you will get an error.

Python File Open

Open a File on the Server


Assume we have the following file, located in the same folder as Python:

[Link]

Hello! Welcome to [Link]


This file is for testing purposes.
Good Luck!

To open the file, use the built-in open() function.

The open() function returns a file object, which has a read() method for
reading the content of the file:
ExampleGet your own Python Server
f = open("[Link]")
print([Link]())

If the file is located in a different location, you will have to specify the file
path, like this:

Example
Open a file on a different location:

f = open("D:\\myfiles\[Link]")
print([Link]())

Using the with statement


You can also use the with statement when opening a file:

Example
Using the with keyword:

with open("[Link]") as f:
print([Link]())

Then you do not have to worry about closing your files, the with statement
takes care of that.

Close Files
It is a good practice to always close the file when you are done with it.

If you are not using the with statement, you must write a close statement in
order to close the file:

Example
Close the file when you are finished with it:
f = open("[Link]")
print([Link]())
[Link]()
Note: You should always close your files. In some cases, due to buffering,
changes made to a file may not show until you close the file.

Read Only Parts of the File


By default the read() method returns the whole text, but you can also
specify how many characters you want to return:

Example
Return the 5 first characters of the file:

with open("[Link]") as f:
print([Link](5))

REMOVE ADS

Read Lines
You can return one line by using the readline() method:

Example
Read one line of the file:

with open("[Link]") as f:
print([Link]())

By calling readline() two times, you can read the two first lines:

Example
Read two lines of the file:
with open("[Link]") as f:
print([Link]())
print([Link]())

By looping through the lines of the file, you can read the whole file, line by
line:

Example
Loop through the file line by line:

with open("[Link]") as f:
for x in f:
print(x)

Python File Write

Write to an Existing File


To write to an existing file, you must add a parameter to
the open() function:

"a" - Append - will append to the end of the file

"w" - Write - will overwrite any existing content

ExampleGet your own Python Server


Open the file "[Link]" and append content to the file:

with open("[Link]", "a") as f:


[Link]("Now the file has more content!")

#open and read the file after the appending:


with open("[Link]") as f:
print([Link]())

Overwrite Existing Content


To overwrite the existing content to the file, use the w parameter:

Example
Open the file "[Link]" and overwrite the content:

with open("[Link]", "w") as f:


[Link]("Woops! I have deleted the content!")

#open and read the file after the overwriting:


with open("[Link]") as f:
print([Link]())
Note: the "w" method will overwrite the entire file.

Create a New File


To create a new file in Python, use the open() method, with one of the
following parameters:

"x" - Create - will create a file, returns an error if the file exists

"a" - Append - will create a file if the specified file does not exists

"w" - Write - will create a file if the specified file does not exists

Example
Create a new file called "[Link]":

f = open("[Link]", "x")

Result: a new empty file is created.

Note: If the file already exists, an error will be raised.

Python Delete File


Delete a File
To delete a file, you must import the OS module, and run
its [Link]() function:

ExampleGet your own Python Server


Remove the file "[Link]":

import os
[Link]("[Link]")

Check if File exist:


To avoid getting an error, you might want to check if the file exists before
you try to delete it:

Example
Check if file exists, then delete it:

import os
if [Link]("[Link]"):
[Link]("[Link]")
else:
print("The file does not exist")

Delete Folder
To delete an entire folder, use the [Link]() method:

Example
Remove the folder "myfolder":

import os
[Link]("myfolder")
Note: You can only remove empty folders.

Matplotlib Tutorial
What is Matplotlib?
Matplotlib is a low level graph plotting library in python that serves as a
visualization utility.

Matplotlib was created by John D. Hunter.

Matplotlib is open source and we can use it freely.

Matplotlib is mostly written in python, a few segments are written in C,


Objective-C and Javascript for Platform compatibility.

Where is the Matplotlib Codebase?


The source code for Matplotlib is located at this github
repository [Link]

Matplotlib Getting Started


Installation of Matplotlib
If you have Python and PIP already installed on a system, then installation of
Matplotlib is very easy.

Install it using this command:

C:\Users\Your Name>pip install matplotlib

If this command fails, then use a python distribution that already has
Matplotlib installed, like Anaconda, Spyder etc.
Import Matplotlib
Once Matplotlib is installed, import it in your applications by adding
the import module statement:

import matplotlib

Now Matplotlib is imported and ready to use:

Checking Matplotlib Version


The version string is stored under __version__ attribute.

ExampleGet your own Python Server


import matplotlib

print(matplotlib.__version__)
Note: two underscore characters are used in __version__.

Matplotlib Pyplot
Pyplot
Most of the Matplotlib utilities lies under the pyplot submodule, and are
usually imported under the plt alias:

import [Link] as plt

Now the Pyplot package can be referred to as plt.

ExampleGet your own Python Server


Draw a line in a diagram from position (0,0) to position (6,250):

import [Link] as plt


import numpy as np
xpoints = [Link]([0, 6])
ypoints = [Link]([0, 250])

[Link](xpoints, ypoints)
[Link]()

Result:

Matplotlib Plotting
Plotting x and y points
The plot() function is used to draw points (markers) in a diagram.

By default, the plot() function draws a line from point to point.

The function takes parameters for specifying points in the diagram.


Parameter 1 is an array containing the points on the x-axis.

Parameter 2 is an array containing the points on the y-axis.

If we need to plot a line from (1, 3) to (8, 10), we have to pass two arrays [1,
8] and [3, 10] to the plot function.

ExampleGet your own Python Server


Draw a line in a diagram from position (1, 3) to position (8, 10):

import [Link] as plt


import numpy as np

xpoints = [Link]([1, 8])


ypoints = [Link]([3, 10])

[Link](xpoints, ypoints)
[Link]()

Result:
The x-axis is the horizontal axis.

The y-axis is the vertical axis

Plotting Without Line


To plot only the markers, you can use shortcut string notation parameter 'o',
which means 'rings'.

Example
Draw two points in the diagram, one at position (1, 3) and one in position (8,
10):

import [Link] as plt


import numpy as np

xpoints = [Link]([1, 8])


ypoints = [Link]([3, 10])
[Link](xpoints, ypoints, 'o')
[Link]()

Result:

You will learn more about markers in the next chapter.

Multiple Points
You can plot as many points as you like, just make sure you have the same
number of points in both axis.

Example
Draw a line in a diagram from position (1, 3) to (2, 8) then to (6, 1) and
finally to position (8, 10):

import [Link] as plt


import numpy as np

xpoints = [Link]([1, 2, 6, 8])


ypoints = [Link]([3, 8, 1, 10])

[Link](xpoints, ypoints)
[Link]()

Result:

Default X-Points
If we do not specify the points on the x-axis, they will get the default values
0, 1, 2, 3 etc., depending on the length of the y-points.

So, if we take the same example as above, and leave out the x-points, the
diagram will look like this:

Example
Plotting without x-points:

import [Link] as plt


import numpy as np

ypoints = [Link]([3, 8, 1, 10, 5, 7])

[Link](ypoints)
[Link]()

Result:
The x-points in the example above are [0, 1, 2, 3, 4, 5].

Matplotlib Markers

Markers
You can use the keyword argument marker to emphasize each point with a
specified marker:

ExampleGet your own Python Server


Mark each point with a circle:

import [Link] as plt


import numpy as np

ypoints = [Link]([3, 8, 1, 10])

[Link](ypoints, marker = 'o')


[Link]()

Result:
Example
Mark each point with a star:

...
[Link](ypoints, marker = '*')
...

Result:
Marker Reference
You can choose any of these markers:

Marker Description

'o' Circle

'*' Star
'.' Point

',' Pixel

'x' X

'X' X (filled)

'+' Plus

'P' Plus (filled)

's' Square

'D' Diamond

'd' Diamond (thin)

'p' Pentagon

'H' Hexagon
'h' Hexagon

'v' Triangle Down

'^' Triangle Up

'<' Triangle Left

'>' Triangle Right

'1' Tri Down

'2' Tri Up

'3' Tri Left

'4' Tri Right

'|' Vline

'_' Hline
Format Strings fmt
You can also use the shortcut string notation parameter to specify the
marker.

This parameter is also called fmt, and is written with this syntax:

marker|line|color
Example
Mark each point with a circle:

import [Link] as plt


import numpy as np

ypoints = [Link]([3, 8, 1, 10])

[Link](ypoints, 'o:r')
[Link]()

Result:
The marker value can be anything from the Marker Reference above.

The line value can be one of the following:

Line Reference
Line Syntax Description

'-' Solid line

':' Dotted line


'--' Dashed line

'-.' Dashed/dotted line

Note: If you leave out the line value in the fmt parameter, no line will be
plotted.

The short color value can be one of the following:

Color Reference
Color Syntax Description

'r' Red

'g' Green

'b' Blue

'c' Cyan

'm' Magenta

'y' Yellow
'k' Black

'w' White

Marker Size
You can use the keyword argument markersize or the shorter version, ms to
set the size of the markers:

Example
Set the size of the markers to 20:

import [Link] as plt


import numpy as np

ypoints = [Link]([3, 8, 1, 10])

[Link](ypoints, marker = 'o', ms = 20)


[Link]()

Result:
Marker Color
You can use the keyword argument markeredgecolor or the shorter mec to
set the color of the edge of the markers:

Example
Set the EDGE color to red:

import [Link] as plt


import numpy as np

ypoints = [Link]([3, 8, 1, 10])

[Link](ypoints, marker = 'o', ms = 20, mec = 'r')


[Link]()
Result:

You can use the keyword argument markerfacecolor or the shorter mfc to
set the color inside the edge of the markers:

Example
Set the FACE color to red:

import [Link] as plt


import numpy as np

ypoints = [Link]([3, 8, 1, 10])

[Link](ypoints, marker = 'o', ms = 20, mfc = 'r')


[Link]()

Result:
Use both the mec and mfc arguments to color the entire marker:

Example
Set the color of both the edge and the face to red:

import [Link] as plt


import numpy as np

ypoints = [Link]([3, 8, 1, 10])

[Link](ypoints, marker = 'o', ms = 20, mec = 'r', mfc = 'r')


[Link]()

Result:
You can also use Hexadecimal color values:

Example
Mark each point with a beautiful green color:

...
[Link](ypoints, marker = 'o', ms = 20, mec = '#4CAF50', mfc
= '#4CAF50')
...

Result:
Or any of the 140 supported color names.

Example
Mark each point with the color named "hotpink":

...
[Link](ypoints, marker = 'o', ms = 20, mec = 'hotpink', mfc
= 'hotpink')
...

Result:
Matplotlib Line

Linestyle
You can use the keyword argument linestyle, or shorter ls, to change the
style of the plotted line:

ExampleGet your own Python Server


Use a dotted line:

import [Link] as plt


import numpy as np

ypoints = [Link]([3, 8, 1, 10])


[Link](ypoints, linestyle = 'dotted')
[Link]()

Result:

Example
Use a dashed line:

[Link](ypoints, linestyle = 'dashed')

Result:
REMOVE ADS

Shorter Syntax
The line style can be written in a shorter syntax:

linestyle can be written as ls.

dotted can be written as :.

dashed can be written as --.

Example
Shorter syntax:

[Link](ypoints, ls = ':')

Result:

Line Styles
You can choose any of these styles:

Style Or
'solid' (default) '-'

'dotted' ':'

'dashed' '--'

'dashdot' '-.'

'None' '' or ' '

Line Color
You can use the keyword argument color or the shorter c to set the color of
the line:

Example
Set the line color to red:

import [Link] as plt


import numpy as np

ypoints = [Link]([3, 8, 1, 10])

[Link](ypoints, color = 'r')


[Link]()

Result:
You can also use Hexadecimal color values:

Example
Plot with a beautiful green line:

...
[Link](ypoints, c = '#4CAF50')
...

Result:
Or any of the 140 supported color names.

Example
Plot with the color named "hotpink":

...
[Link](ypoints, c = 'hotpink')
...

Result:
REMOVE ADS

Line Width
You can use the keyword argument linewidth or the shorter lw to change
the width of the line.

The value is a floating number, in points:

Example
Plot with a 20.5pt wide line:

import [Link] as plt


import numpy as np
ypoints = [Link]([3, 8, 1, 10])

[Link](ypoints, linewidth = '20.5')


[Link]()

Result:

Multiple Lines
You can plot as many lines as you like by simply adding
more [Link]() functions:

Example
Draw two lines by specifying a [Link]() function for each line:
import [Link] as plt
import numpy as np

y1 = [Link]([3, 8, 1, 10])
y2 = [Link]([6, 2, 7, 11])

[Link](y1)
[Link](y2)

[Link]()

Result:

You can also plot many lines by adding the points for the x- and y-axis for
each line in the same [Link]() function.

(In the examples above we only specified the points on the y-axis, meaning
that the points on the x-axis got the the default values (0, 1, 2, 3).)

The x- and y- values come in pairs:


Example
Draw two lines by specifiyng the x- and y-point values for both lines:

import [Link] as plt


import numpy as np

x1 = [Link]([0, 1, 2, 3])
y1 = [Link]([3, 8, 1, 10])
x2 = [Link]([0, 1, 2, 3])
y2 = [Link]([6, 2, 7, 11])

[Link](x1, y1, x2, y2)


[Link]()

Result:

Matplotlib Labels and Title


Create Labels for a Plot
With Pyplot, you can use the xlabel() and ylabel() functions to set a label
for the x- and y-axis.

ExampleGet your own Python Server


Add labels to the x- and y-axis:

import numpy as np
import [Link] as plt

x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])

[Link](x, y)

[Link]("Average Pulse")
[Link]("Calorie Burnage")

[Link]()

Result:
Create a Title for a Plot
With Pyplot, you can use the title() function to set a title for the plot.

Example
Add a plot title and labels for the x- and y-axis:

import numpy as np
import [Link] as plt

x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])

[Link](x, y)
[Link]("Sports Watch Data")
[Link]("Average Pulse")
[Link]("Calorie Burnage")

[Link]()

Result:

REMOVE ADS

Set Font Properties for Title and


Labels
You can use the fontdict parameter in xlabel(), ylabel(), and title() to
set font properties for the title and labels.

Example
Set font properties for the title and labels:

import numpy as np
import [Link] as plt

x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])

font1 = {'family':'serif','color':'blue','size':20}
font2 = {'family':'serif','color':'darkred','size':15}

[Link]("Sports Watch Data", fontdict = font1)


[Link]("Average Pulse", fontdict = font2)
[Link]("Calorie Burnage", fontdict = font2)

[Link](x, y)
[Link]()

Result:
Position the Title
You can use the loc parameter in title() to position the title.

Legal values are: 'left', 'right', and 'center'. Default value is 'center'.

Example
Position the title to the left:

import numpy as np
import [Link] as plt

x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])
[Link]("Sports Watch Data", loc = 'left')
[Link]("Average Pulse")
[Link]("Calorie Burnage")

[Link](x, y)
[Link]()

Result:

Matplotlib Adding Grid Lines

Add Grid Lines to a Plot


With Pyplot, you can use the grid() function to add grid lines to the plot.
ExampleGet your own Python Server
Add grid lines to the plot:

import numpy as np
import [Link] as plt

x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])

[Link]("Sports Watch Data")


[Link]("Average Pulse")
[Link]("Calorie Burnage")

[Link](x, y)

[Link]()

[Link]()

Result:
REMOVE ADS

Specify Which Grid Lines to Display


You can use the axis parameter in the grid() function to specify which grid
lines to display.

Legal values are: 'x', 'y', and 'both'. Default value is 'both'.

Example
Display only grid lines for the x-axis:
import numpy as np
import [Link] as plt

x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])

[Link]("Sports Watch Data")


[Link]("Average Pulse")
[Link]("Calorie Burnage")

[Link](x, y)

[Link](axis = 'x')

[Link]()

Result:

Example
Display only grid lines for the y-axis:

import numpy as np
import [Link] as plt

x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])

[Link]("Sports Watch Data")


[Link]("Average Pulse")
[Link]("Calorie Burnage")

[Link](x, y)

[Link](axis = 'y')

[Link]()

Result:
Set Line Properties for the Grid
You can also set the line properties of the grid, like this: grid(color = 'color',
linestyle = 'linestyle', linewidth = number).

Example
Set the line properties of the grid:

import numpy as np
import [Link] as plt

x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])

[Link]("Sports Watch Data")


[Link]("Average Pulse")
[Link]("Calorie Burnage")

[Link](x, y)

[Link](color = 'green', linestyle = '--', linewidth = 0.5)

[Link]()

Result:
Matplotlib Subplot

Display Multiple Plots


With the subplot() function you can draw multiple plots in one figure:

ExampleGet your own Python Server


Draw 2 plots:

import [Link] as plt


import numpy as np

#plot 1:
x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])

[Link](1, 2, 1)
[Link](x,y)

#plot 2:
x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])

[Link](1, 2, 2)
[Link](x,y)

[Link]()

Result:

The subplot() Function


The subplot() function takes three arguments that describes the layout of
the figure.

The layout is organized in rows and columns, which are represented by


the first and second argument.

The third argument represents the index of the current plot.

[Link](1, 2, 1)
#the figure has 1 row, 2 columns, and this plot is
the first plot.

[Link](1, 2, 2)
#the figure has 1 row, 2 columns, and this plot is
the second plot.

So, if we want a figure with 2 rows an 1 column (meaning that the two plots
will be displayed on top of each other instead of side-by-side), we can write
the syntax like this:

Example
Draw 2 plots on top of each other:

import [Link] as plt


import numpy as np

#plot 1:
x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])

[Link](2, 1, 1)
[Link](x,y)

#plot 2:
x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])

[Link](2, 1, 2)
[Link](x,y)

[Link]()

Result:
You can draw as many plots you like on one figure, just descibe the number
of rows, columns, and the index of the plot.

Example
Draw 6 plots:

import [Link] as plt


import numpy as np

x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])

[Link](2, 3, 1)
[Link](x,y)

x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])
[Link](2, 3, 2)
[Link](x,y)

x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])

[Link](2, 3, 3)
[Link](x,y)

x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])

[Link](2, 3, 4)
[Link](x,y)

x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])

[Link](2, 3, 5)
[Link](x,y)

x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])

[Link](2, 3, 6)
[Link](x,y)

[Link]()

Result:
REMOVE ADS

Title
You can add a title to each plot with the title() function:

Example
2 plots, with titles:

import [Link] as plt


import numpy as np

#plot 1:
x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])

[Link](1, 2, 1)
[Link](x,y)
[Link]("SALES")

#plot 2:
x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])

[Link](1, 2, 2)
[Link](x,y)
[Link]("INCOME")

[Link]()

Result:
Super Title
You can add a title to the entire figure with the suptitle() function:

Example
Add a title for the entire figure:

import [Link] as plt


import numpy as np

#plot 1:
x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])

[Link](1, 2, 1)
[Link](x,y)
[Link]("SALES")

#plot 2:
x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])

[Link](1, 2, 2)
[Link](x,y)
[Link]("INCOME")

[Link]("MY SHOP")
[Link]()

Result:
Matplotlib Scatter

Creating Scatter Plots


With Pyplot, you can use the scatter() function to draw a scatter plot.

The scatter() function plots one dot for each observation. It needs two arrays
of the same length, one for the values of the x-axis, and one for values on
the y-axis:

ExampleGet your own Python Server


A simple scatter plot:
import [Link] as plt
import numpy as np

x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])

[Link](x, y)
[Link]()

Result:

The observation in the example above is the result of 13 cars passing by.

The X-axis shows how old the car is.

The Y-axis shows the speed of the car when it passes.

Are there any relationships between the observations?


It seems that the newer the car, the faster it drives, but that could be a
coincidence, after all we only registered 13 cars.

Compare Plots
In the example above, there seems to be a relationship between speed and
age, but what if we plot the observations from another day as well? Will the
scatter plot tell us something else?

Example
Draw two plots on the same figure:

import [Link] as plt


import numpy as np

#day one, the age and speed of 13 cars:


x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
[Link](x, y)

#day two, the age and speed of 15 cars:


x = [Link]([2,2,8,1,15,8,12,9,7,3,11,4,7,14,12])
y = [Link]([100,105,84,105,90,99,90,95,94,100,79,112,91,80,85])
[Link](x, y)

[Link]()

Result:
Note: The two plots are plotted with two different colors, by default blue and
orange, you will learn how to change colors later in this chapter.

By comparing the two plots, I think it is safe to say that they both gives us
the same conclusion: the newer the car, the faster it drives.

REMOVE ADS

Colors
You can set your own color for each scatter plot with the color or
the c argument:

Example
Set your own color of the markers:

import [Link] as plt


import numpy as np

x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
[Link](x, y, color = 'hotpink')

x = [Link]([2,2,8,1,15,8,12,9,7,3,11,4,7,14,12])
y = [Link]([100,105,84,105,90,99,90,95,94,100,79,112,91,80,85])
[Link](x, y, color = '#88c999')

[Link]()

Result:
Color Each Dot
You can even set a specific color for each dot by using an array of colors as
value for the c argument:

Note: You cannot use the color argument for this, only the c argument.

Example
Set your own color of the markers:

import [Link] as plt


import numpy as np

x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
colors =
[Link](["red","green","blue","yellow","pink","black","orange","purp
le","beige","brown","gray","cyan","magenta"])

[Link](x, y, c=colors)

[Link]()

Result:
REMOVE ADS

ColorMap
The Matplotlib module has a number of available colormaps.

A colormap is like a list of colors, where each color has a value that ranges
from 0 to 100.

Here is an example of a colormap:


This colormap is called 'viridis' and as you can see it ranges from 0, which is
a purple color, up to 100, which is a yellow color.

How to Use the ColorMap


You can specify the colormap with the keyword argument cmap with the value
of the colormap, in this case 'viridis' which is one of the built-in colormaps
available in Matplotlib.

In addition you have to create an array with values (from 0 to 100), one
value for each point in the scatter plot:

Example
Create a color array, and specify a colormap in the scatter plot:

import [Link] as plt


import numpy as np

x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
colors =
[Link]([0, 10, 20, 30, 40, 45, 50, 55, 60, 70, 80, 90, 100])
[Link](x, y, c=colors, cmap='viridis')

[Link]()

Result:

You can include the colormap in the drawing by including


the [Link]() statement:

Example
Include the actual colormap:

import [Link] as plt


import numpy as np

x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
colors =
[Link]([0, 10, 20, 30, 40, 45, 50, 55, 60, 70, 80, 90, 100])

[Link](x, y, c=colors, cmap='viridis')

[Link]()

[Link]()

Result:

Available ColorMaps
You can choose any of the built-in colormaps:

Name Reverse
Accent Accent_r

Blues Blues_r

BrBG BrBG_r

BuGn BuGn_r

BuPu BuPu_r

CMRmap CMRmap_r

Dark2 Dark2_r

GnBu GnBu_r

Greens Greens_r

Greys Greys_r

OrRd OrRd_r
Oranges Oranges_r

PRGn PRGn_r

Paired Paired_r

Pastel1 Pastel1_r

Pastel2 Pastel2_r

PiYG PiYG_r

PuBu PuBu_r

PuBuGn PuBuGn_r

PuOr PuOr_r

PuRd PuRd_r

Purples Purples_r
RdBu RdBu_r

RdGy RdGy_r

RdPu RdPu_r

RdYlBu RdYlBu_r

RdYlGn RdYlGn_r

Reds Reds_r

Set1 Set1_r

Set2 Set2_r

Set3 Set3_r

Spectral Spectral_r

Wistia Wistia_r
YlGn YlGn_r

YlGnBu YlGnBu_r

YlOrBr YlOrBr_r

YlOrRd YlOrRd_r

afmhot afmhot_r

autumn autumn_r

binary binary_r

bone bone_r

brg brg_r

bwr bwr_r

cividis cividis_r
cool cool_r

coolwarm coolwarm_r

copper copper_r

cubehelix cubehelix_r

flag flag_r

gist_earth gist_earth_r

gist_gray gist_gray_r

gist_heat gist_heat_r

gist_ncar gist_ncar_r

gist_rainbow gist_rainbow_r

gist_stern gist_stern_r
gist_yarg gist_yarg_r

gnuplot gnuplot_r

gnuplot2 gnuplot2_r

gray gray_r

hot hot_r

hsv hsv_r

inferno inferno_r

jet jet_r

magma magma_r

nipy_spectral nipy_spectral_r

ocean ocean_r
pink pink_r

plasma plasma_r

prism prism_r

rainbow rainbow_r

seismic seismic_r

spring spring_r

summer summer_r

tab10 tab10_r

tab20 tab20_r

tab20b tab20b_r

tab20c tab20c_r
terrain terrain_r

twilight twilight_r

twilight_shifted twilight_shifted_r

viridis viridis_r

winter winter_r

Size
You can change the size of the dots with the s argument.

Just like colors, make sure the array for sizes has the same length as the
arrays for the x- and y-axis:

Example
Set your own size for the markers:

import [Link] as plt


import numpy as np

x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
sizes = [Link]([20,50,100,200,500,1000,60,90,10,300,600,800,75])

[Link](x, y, s=sizes)

[Link]()
Result:

Alpha
You can adjust the transparency of the dots with the alpha argument.

Just like colors, make sure the array for sizes has the same length as the
arrays for the x- and y-axis:

Example
Set your own size for the markers:

import [Link] as plt


import numpy as np
x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
sizes = [Link]([20,50,100,200,500,1000,60,90,10,300,600,800,75])

[Link](x, y, s=sizes, alpha=0.5)

[Link]()

Result:

Combine Color Size and Alpha


You can combine a colormap with different sizes of the dots. This is best
visualized if the dots are transparent:
Example
Create random arrays with 100 values for x-points, y-points, colors and sizes:

import [Link] as plt


import numpy as np

x = [Link](100, size=(100))
y = [Link](100, size=(100))
colors = [Link](100, size=(100))
sizes = 10 * [Link](100, size=(100))

[Link](x, y, c=colors, s=sizes, alpha=0.5, cmap='nipy_spectral')

[Link]()

[Link]()

Result:
Matplotlib Bars

Creating Bars
With Pyplot, you can use the bar() function to draw bar graphs:

ExampleGet your own Python Server


Draw 4 bars:

import [Link] as plt


import numpy as np

x = [Link](["A", "B", "C", "D"])


y = [Link]([3, 8, 1, 10])

[Link](x,y)
[Link]()

Result:
The bar() function takes arguments that describes the layout of the bars.

The categories and their values represented by


the first and second argument as arrays.

Example
x = ["APPLES", "BANANAS"]
y = [400, 350]
[Link](x, y)

REMOVE ADS

Horizontal Bars
If you want the bars to be displayed horizontally instead of vertically, use
the barh() function:

Example
Draw 4 horizontal bars:

import [Link] as plt


import numpy as np

x = [Link](["A", "B", "C", "D"])


y = [Link]([3, 8, 1, 10])

[Link](x, y)
[Link]()

Result:
Bar Color
The bar() and barh() take the keyword argument color to set the color of
the bars:

Example
Draw 4 red bars:

import [Link] as plt


import numpy as np

x = [Link](["A", "B", "C", "D"])


y = [Link]([3, 8, 1, 10])

[Link](x, y, color = "red")


[Link]()

Result:
Color Names
You can use any of the 140 supported color names.

Example
Draw 4 "hot pink" bars:

import [Link] as plt


import numpy as np

x = [Link](["A", "B", "C", "D"])


y = [Link]([3, 8, 1, 10])

[Link](x, y, color = "hotpink")


[Link]()

Result:
Color Hex
Or you can use Hexadecimal color values:

Example
Draw 4 bars with a beautiful green color:

import [Link] as plt


import numpy as np

x = [Link](["A", "B", "C", "D"])


y = [Link]([3, 8, 1, 10])

[Link](x, y, color = "#4CAF50")


[Link]()

Result:
Bar Width
The bar() takes the keyword argument width to set the width of the bars:

Example
Draw 4 very thin bars:

import [Link] as plt


import numpy as np

x = [Link](["A", "B", "C", "D"])


y = [Link]([3, 8, 1, 10])

[Link](x, y, width = 0.1)


[Link]()
Result:

The default width value is 0.8

Note: For horizontal bars, use height instead of width.

Bar Height
The barh() takes the keyword argument height to set the height of the
bars:

Example
Draw 4 very thin bars:
import [Link] as plt
import numpy as np

x = [Link](["A", "B", "C", "D"])


y = [Link]([3, 8, 1, 10])

[Link](x, y, height = 0.1)


[Link]()

Result:

The default height value is 0.8

Matplotlib Histograms
Histogram
A histogram is a graph showing frequency distributions.

It is a graph showing the number of observations within each given interval.

Example: Say you ask for the height of 250 people, you might end up with a
histogram like this:

You can read from the histogram that there are approximately:

2 people from 140 to 145cm


5 people from 145 to 150cm
15 people from 151 to 156cm
31 people from 157 to 162cm
46 people from 163 to 168cm
53 people from 168 to 173cm
45 people from 173 to 178cm
28 people from 179 to 184cm
21 people from 185 to 190cm
4 people from 190 to 195cm

Create Histogram
In Matplotlib, we use the hist() function to create histograms.

The hist() function will use an array of numbers to create a histogram, the
array is sent into the function as an argument.

For simplicity we use NumPy to randomly generate an array with 250 values,
where the values will concentrate around 170, and the standard deviation is
10. Learn more about Normal Data Distribution in our Machine Learning
Tutorial.

ExampleGet your own Python Server


A Normal Data Distribution by NumPy:

import numpy as np

x = [Link](170, 10, 250)

print(x)

Result:
This will generate a random result, and could look like this:

[167.62255766 175.32495609 152.84661337 165.50264047


163.17457988
162.29867872 172.83638413 168.67303667 164.57361342
180.81120541
170.57782187 167.53075749 176.15356275 176.95378312
158.4125473
187.8842668 159.03730075 166.69284332 160.73882029
152.22378865
164.01255164 163.95288674 176.58146832 173.19849526
169.40206527
166.88861903 149.90348576 148.39039643 177.90349066
166.72462233
177.44776004 170.93335636 173.26312881 174.76534435
162.28791953
166.77301551 160.53785202 170.67972019 159.11594186
165.36992993
178.38979253 171.52158489 173.32636678 159.63894401
151.95735707
175.71274153 165.00458544 164.80607211 177.50988211
149.28106703
179.43586267 181.98365273 170.98196794 179.1093176
176.91855744
168.32092784 162.33939782 165.18364866 160.52300507
174.14316386
163.01947601 172.01767945 173.33491959 169.75842718
198.04834503
192.82490521 164.54557943 206.36247244 165.47748898
195.26377975
164.37569092 156.15175531 162.15564208 179.34100362
167.22138242
147.23667125 162.86940215 167.84986671 172.99302505
166.77279814
196.6137667 159.79012341 166.5840824 170.68645637
165.62204521
174.5559345 165.0079216 187.92545129 166.86186393
179.78383824
161.0973573 167.44890343 157.38075812 151.35412246
171.3107829
162.57149341 182.49985133 163.24700057 168.72639903
169.05309467
167.19232875 161.06405208 176.87667712 165.48750185
179.68799986
158.7913483 170.22465411 182.66432721 173.5675715
176.85646836
157.31299754 174.88959677 183.78323508 174.36814558
182.55474697
180.03359793 180.53094948 161.09560099 172.29179934
161.22665588
171.88382477 159.04626132 169.43886536 163.75793589
157.73710983
174.68921523 176.19843414 167.39315397 181.17128255
174.2674597
186.05053154 177.06516302 171.78523683 166.14875436
163.31607668
174.01429569 194.98819875 169.75129209 164.25748789
180.25773528
170.44784934 157.81966006 171.33315907 174.71390637
160.55423274
163.92896899 177.29159542 168.30674234 165.42853878
176.46256226
162.61719142 166.60810831 165.83648812 184.83238352
188.99833856
161.3054697 175.30396693 175.28109026 171.54765201
162.08762813
164.53011089 189.86213299 170.83784593 163.25869004
198.68079225
166.95154328 152.03381334 152.25444225 149.75522816
161.79200594
162.13535052 183.37298831 165.40405341 155.59224806
172.68678385
179.35359654 174.19668349 163.46176882 168.26621173
162.97527574
192.80170974 151.29673582 178.65251432 163.17266558
165.11172588
183.11107905 169.69556831 166.35149789 178.74419135
166.28562032
169.96465166 178.24368042 175.3035525 170.16496554
158.80682882
187.10006553 178.90542991 171.65790645 183.19289193
168.17446717
155.84544031 177.96091745 186.28887898 187.89867406
163.26716924
169.71242393 152.9410412 158.68101969 171.12655559
178.1482624
187.45272185 173.02872935 163.8047623 169.95676819
179.36887054
157.01955088 185.58143864 170.19037101 157.221245
168.90639755
178.7045601 168.64074373 172.37416382 165.61890535
163.40873027
168.98683006 149.48186389 172.20815568 172.82947206
173.71584064
189.42642762 172.79575803 177.00005573 169.24498561
171.55576698
161.36400372 176.47928342 163.02642822 165.09656415
186.70951892
153.27990317 165.59289527 180.34566865 189.19506385
183.10723435
173.48070474 170.28701875 157.24642079 157.9096498
176.4248199 ]

The hist() function will read the array and produce a histogram:
Example
A simple histogram:

import [Link] as plt


import numpy as np

x = [Link](170, 10, 250)

[Link](x)
[Link]()

Result:

Matplotlib Pie Charts


Creating Pie Charts
With Pyplot, you can use the pie() function to draw pie charts:

ExampleGet your own Python Server


A simple pie chart:

import [Link] as plt


import numpy as np

y = [Link]([35, 25, 25, 15])

[Link](y)
[Link]()

Result:
As you can see the pie chart draws one piece (called a wedge) for each value
in the array (in this case [35, 25, 25, 15]).

By default the plotting of the first wedge starts from the x-axis and
moves counterclockwise:

Note: The size of each wedge is determined by comparing the value with all
the other values, by using this formula:

The value divided by the sum of all values: x/sum(x)

REMOVE ADS

Labels
Add labels to the pie chart with the labels parameter.

The labels parameter must be an array with one label for each wedge:

Example
A simple pie chart:

import [Link] as plt


import numpy as np

y = [Link]([35, 25, 25, 15])


mylabels = ["Apples", "Bananas", "Cherries", "Dates"]

[Link](y, labels = mylabels)


[Link]()

Result:
Start Angle
As mentioned the default start angle is at the x-axis, but you can change the
start angle by specifying a startangle parameter.

The startangle parameter is defined with an angle in degrees, default angle


is 0:

Example
Start the first wedge at 90 degrees:

import [Link] as plt


import numpy as np

y = [Link]([35, 25, 25, 15])


mylabels = ["Apples", "Bananas", "Cherries", "Dates"]

[Link](y, labels = mylabels, startangle = 90)


[Link]()

Result:

Explode
Maybe you want one of the wedges to stand out? The explode parameter
allows you to do that.

The explode parameter, if specified, and not None, must be an array with
one value for each wedge.

Each value represents how far from the center each wedge is displayed:
Example
Pull the "Apples" wedge 0.2 from the center of the pie:

import [Link] as plt


import numpy as np

y = [Link]([35, 25, 25, 15])


mylabels = ["Apples", "Bananas", "Cherries", "Dates"]
myexplode = [0.2, 0, 0, 0]

[Link](y, labels = mylabels, explode = myexplode)


[Link]()

Result:
Shadow
Add a shadow to the pie chart by setting the shadows parameter to True:

Example
Add a shadow:

import [Link] as plt


import numpy as np

y = [Link]([35, 25, 25, 15])


mylabels = ["Apples", "Bananas", "Cherries", "Dates"]
myexplode = [0.2, 0, 0, 0]

[Link](y, labels = mylabels, explode = myexplode, shadow = True)


[Link]()

Result:
Colors
You can set the color of each wedge with the colors parameter.

The colors parameter, if specified, must be an array with one value for each
wedge:

Example
Specify a new color for each wedge:

import [Link] as plt


import numpy as np

y = [Link]([35, 25, 25, 15])


mylabels = ["Apples", "Bananas", "Cherries", "Dates"]
mycolors = ["black", "hotpink", "b", "#4CAF50"]

[Link](y, labels = mylabels, colors = mycolors)


[Link]()

Result:

You can use Hexadecimal color values, any of the 140 supported color
names, or one of these shortcuts:

'r' - Red
'g' - Green
'b' - Blue
'c' - Cyan
'm' - Magenta
'y' - Yellow
'k' - Black
'w' - White
Legend
To add a list of explanation for each wedge, use the legend() function:

Example
Add a legend:

import [Link] as plt


import numpy as np

y = [Link]([35, 25, 25, 15])


mylabels = ["Apples", "Bananas", "Cherries", "Dates"]

[Link](y, labels = mylabels)


[Link]()
[Link]()

Result:
Legend With Header
To add a header to the legend, add the title parameter to
the legend function.

Example
Add a legend with a header:

import [Link] as plt


import numpy as np

y = [Link]([35, 25, 25, 15])


mylabels = ["Apples", "Bananas", "Cherries", "Dates"]

[Link](y, labels = mylabels)


[Link](title = "Four Fruits:")
[Link]()
Result:

Machine Learning
Where To Start?
In this tutorial we will go back to mathematics and study statistics, and how
to calculate important numbers based on data sets.

We will also learn how to use various Python modules to get the answers we
need.

And we will learn how to make functions that are able to predict the outcome
based on what we have learned.
Data Set
In the mind of a computer, a data set is any collection of data. It can be
anything from an array to a complete database.

Example of an array:

[99,86,87,88,111,86,103,87,94,78,77,85,86]

Example of a database:

Carname Color Age Sp

BMW red 5

Volvo black 7

VW gray 8

VW white 7

Ford white 2

VW white 17

Tesla red 2
BMW black 9

Volvo gray 4

Ford white 11

Toyota gray 12

VW white 9

Toyota blue 6

By looking at the array, we can guess that the average value is probably
around 80 or 90, and we are also able to determine the highest value and
the lowest value, but what else can we do?

And by looking at the database we can see that the most popular color is
white, and the oldest car is 17 years, but what if we could predict if a car had
an AutoPass, just by looking at the other values?

That is what Machine Learning is for! Analyzing data and predicting the
outcome!

In Machine Learning it is common to work with very large data sets. In this
tutorial we will try to make it as easy as possible to understand the different
concepts of machine learning, and we will work with small easy-to-
understand data sets.

REMOVE ADS
Data Types
To analyze data, it is important to know what type of data we are dealing
with.

We can split the data types into three main categories:

 Numerical
 Categorical
 Ordinal

Numerical data are numbers, and can be split into two numerical
categories:

 Discrete Data
- counted data that are limited to integers. Example: The number of
cars passing by.
 Continuous Data
- measured data that can be any number. Example: The price of an
item, or the size of an item

Categorical data are values that cannot be measured up against each


other. Example: a color value, or any yes/no values.

Ordinal data are like categorical data, but can be measured up against each
other. Example: school grades where A is better than B and so on.

By knowing the data type of your data source, you will be able to know what
technique to use when analyzing them.

You will learn more about statistics and analyzing data in the next chapters.

Machine Learning - Mean


Median Mode

Mean, Median, and Mode


What can we learn from looking at a group of numbers?

In Machine Learning (and in mathematics) there are often three values that
interests us:

 Mean - The average value


 Median - The mid point value
 Mode - The most common value

Example: We have registered the speed of 13 cars:

speed = [99,86,87,88,111,86,103,87,94,78,77,85,86]

What is the average, the middle, or the most common speed value?

Mean
The mean value is the average value.

To calculate the mean, find the sum of all values, and divide the sum by the
number of values:

(99+86+87+88+111+86+103+87+94+78+77+85+86) / 13 = 89.77

The NumPy module has a method for this. Learn about the NumPy module in
our NumPy Tutorial.

ExampleGet your own Python Server


Use the NumPy mean() method to find the average speed:

import numpy

speed = [99,86,87,88,111,86,103,87,94,78,77,85,86]

x = [Link](speed)

print(x)

REMOVE ADS
Median
The median value is the value in the middle, after you have sorted all the
values:

77, 78, 85, 86, 86, 86, 87, 87, 88, 94, 99, 103, 111

It is important that the numbers are sorted before you can find the median.

The NumPy module has a method for this:

Example
Use the NumPy median() method to find the middle value:

import numpy

speed = [99,86,87,88,111,86,103,87,94,78,77,85,86]

x = [Link](speed)

print(x)
If there are two numbers in the middle, divide the sum of those numbers by
two.
77, 78, 85, 86, 86, 86, 87, 87, 94, 98, 99, 103

(86 + 87) / 2 = 86.5


Example
Using the NumPy module:

import numpy

speed = [99,86,87,88,86,103,87,94,78,77,85,86]

x = [Link](speed)

print(x)
Mode
The Mode value is the value that appears the most number of times:

99, 86, 87, 88, 111, 86, 103, 87, 94, 78, 77, 85, 86 = 86

The SciPy module has a method for this. Learn about the SciPy module in
our SciPy Tutorial.

Example
Use the SciPy mode() method to find the number that appears the most:

from scipy import stats

speed = [99,86,87,88,111,86,103,87,94,78,77,85,86]

x = [Link](speed)

print(x)
REMOVE ADS

Chapter Summary
The Mean, Median, and Mode are techniques that are often used in Machine
Learning, so it is important to understand the concept behind them.

Machine Learning - Standard


Deviation

What is Standard Deviation?


Standard deviation is a number that describes how spread out the values
are.
A low standard deviation means that most of the numbers are close to the
mean (average) value.

A high standard deviation means that the values are spread out over a wider
range.

Example: This time we have registered the speed of 7 cars:

speed = [86,87,88,86,87,85,86]

The standard deviation is:

0.9

Meaning that most of the values are within the range of 0.9 from the mean
value, which is 86.4.

Let us do the same with a selection of numbers with a wider range:

speed = [32,111,138,28,59,77,97]

The standard deviation is:

37.85

Meaning that most of the values are within the range of 37.85 from the mean
value, which is 77.4.

As you can see, a higher standard deviation indicates that the values are
spread out over a wider range.

The NumPy module has a method to calculate the standard deviation:

ExampleGet your own Python Server


Use the NumPy std() method to find the standard deviation:

import numpy

speed = [86,87,88,86,87,85,86]

x = [Link](speed)

print(x)

Example
import numpy

speed = [32,111,138,28,59,77,97]

x = [Link](speed)

print(x)

Variance
Variance is another number that indicates how spread out the values are.

In fact, if you take the square root of the variance, you get the standard
deviation!

Or the other way around, if you multiply the standard deviation by itself, you
get the variance!

To calculate the variance you have to do as follows:

1. Find the mean:

(32+111+138+28+59+77+97) / 7 = 77.4

2. For each value: find the difference from the mean:

32 - 77.4 = -45.4
111 - 77.4 = 33.6
138 - 77.4 = 60.6
28 - 77.4 = -49.4
59 - 77.4 = -18.4
77 - 77.4 = - 0.4
97 - 77.4 = 19.6

3. For each difference: find the square value:

(-45.4)2 = 2061.16
(33.6)2 = 1128.96
(60.6)2 = 3672.36
(-49.4)2 = 2440.36
(-18.4)2 = 338.56
(- 0.4)2 = 0.16
(19.6)2 = 384.16

4. The variance is the average number of these squared differences:


(2061.16+1128.96+3672.36+2440.36+338.56+0.16+384.16) / 7 = 1432.2

Luckily, NumPy has a method to calculate the variance:

Example
Use the NumPy var() method to find the variance:

import numpy

speed = [32,111,138,28,59,77,97]

x = [Link](speed)

print(x)

Standard Deviation
As we have learned, the formula to find the standard deviation is the square
root of the variance:

√1432.25 = 37.85

Or, as in the example from before, use the NumPy to calculate the standard
deviation:

Example
Use the NumPy std() method to find the standard deviation:

import numpy

speed = [32,111,138,28,59,77,97]

x = [Link](speed)

print(x)

Symbols
Standard Deviation is often represented by the symbol Sigma: σ
Variance is often represented by the symbol Sigma Squared: σ 2

REMOVE ADS

Chapter Summary
The Standard Deviation and Variance are terms that are often used in
Machine Learning, so it is important to understand how to get them, and the
concept behind them.

Machine Learning -
Percentiles

What are Percentiles?


Percentiles are used in statistics to give you a number that describes the
value that a given percent of the values are lower than.

Example: Let's say we have an array that contains the ages of every person
living on a street.

ages =
[5,31,43,48,50,41,7,11,15,39,80,82,32,2,8,6,25,36,27,61,31]

What is the 75. percentile? The answer is 43, meaning that 75% of the
people are 43 or younger.

The NumPy module has a method for finding the specified percentile:

ExampleGet your own Python Server


Use the NumPy percentile() method to find the percentiles:
import numpy

ages = [5,31,43,48,50,41,7,11,15,39,80,82,32,2,8,6,25,36,27,61,31]

x = [Link](ages, 75)

print(x)

Example
What is the age that 90% of the people are younger than?

import numpy

ages = [5,31,43,48,50,41,7,11,15,39,80,82,32,2,8,6,25,36,27,61,31]

x = [Link](ages, 90)

print(x)

Machine Learning - Data


Distribution

Data Distribution
Earlier in this tutorial we have worked with very small amounts of data in our
examples, just to understand the different concepts.

In the real world, the data sets are much bigger, but it can be difficult to
gather real world data, at least at an early stage of a project.

How Can we Get Big Data Sets?


To create big data sets for testing, we use the Python module NumPy, which
comes with a number of methods to create random data sets, of any size.

ExampleGet your own Python Server


Create an array containing 250 random floats between 0 and 5:
import numpy

x = [Link](0.0, 5.0, 250)

print(x)

Histogram
To visualize the data set we can draw a histogram with the data we
collected.

We will use the Python module Matplotlib to draw a histogram.

Learn about the Matplotlib module in our Matplotlib Tutorial.

Example
Draw a histogram:

import numpy
import [Link] as plt

x = [Link](0.0, 5.0, 250)

[Link](x, 5)
[Link]()

Result:
Histogram Explained
We use the array from the example above to draw a histogram with 5 bars.

The first bar represents how many values in the array are between 0 and 1.

The second bar represents how many values are between 1 and 2.

Etc.

Which gives us this result:

 52 values are between 0 and 1


 48 values are between 1 and 2
 49 values are between 2 and 3
 51 values are between 3 and 4
 50 values are between 4 and 5
Note: The array values are random numbers and will not show the exact
same result on your computer.
Big Data Distributions
An array containing 250 values is not considered very big, but now you know
how to create a random set of values, and by changing the parameters, you
can create the data set as big as you want.

Example
Create an array with 100000 random numbers, and display them using a
histogram with 100 bars:

import numpy
import [Link] as plt

x = [Link](0.0, 5.0, 100000)

[Link](x, 100)
[Link]()

Machine Learning - Normal


Data Distribution

Normal Data Distribution


In the previous chapter we learned how to create a completely random
array, of a given size, and between two given values.

In this chapter we will learn how to create an array where the values are
concentrated around a given value.

In probability theory this kind of data distribution is known as the normal


data distribution, or the Gaussian data distribution, after the mathematician
Carl Friedrich Gauss who came up with the formula of this data distribution.

ExampleGet your own Python Server


A typical normal data distribution:
import numpy
import [Link] as plt

x = [Link](5.0, 1.0, 100000)

[Link](x, 100)
[Link]()

Result:

Note: A normal distribution graph is also known as the bell curve because of
it's characteristic shape of a bell.

Histogram Explained
We use the array from the [Link]() method, with 100000
values, to draw a histogram with 100 bars.

We specify that the mean value is 5.0, and the standard deviation is 1.0.
Meaning that the values should be concentrated around 5.0, and rarely
further away than 1.0 from the mean.

And as you can see from the histogram, most values are between 4.0 and
6.0, with a top at approximately 5.0.

Machine Learning - Scatter


Plot

Scatter Plot
A scatter plot is a diagram where each value in the data set is represented
by a dot.
The Matplotlib module has a method for drawing scatter plots, it needs two
arrays of the same length, one for the values of the x-axis, and one for the
values of the y-axis:

x = [5,7,8,7,2,17,2,9,4,11,12,9,6]

y = [99,86,87,88,111,86,103,87,94,78,77,85,86]

The x array represents the age of each car.

The y array represents the speed of each car.

ExampleGet your own Python Server


Use the scatter() method to draw a scatter plot diagram:

import [Link] as plt

x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]

[Link](x, y)
[Link]()

Result:
Scatter Plot Explained
The x-axis represents ages, and the y-axis represents speeds.

What we can read from the diagram is that the two fastest cars were both 2
years old, and the slowest car was 12 years old.

Note: It seems that the newer the car, the faster it drives, but that could be
a coincidence, after all we only registered 13 cars.

REMOVE ADS

Random Data Distributions


In Machine Learning the data sets can contain thousands-, or even millions,
of values.

You might not have real world data when you are testing an algorithm, you
might have to use randomly generated values.

As we have learned in the previous chapter, the NumPy module can help us
with that!

Let us create two arrays that are both filled with 1000 random numbers from
a normal data distribution.

The first array will have the mean set to 5.0 with a standard deviation of 1.0.

The second array will have the mean set to 10.0 with a standard deviation of
2.0:

Example
A scatter plot with 1000 dots:

import numpy
import [Link] as plt

x = [Link](5.0, 1.0, 1000)


y = [Link](10.0, 2.0, 1000)

[Link](x, y)
[Link]()

Result:
Scatter Plot Explained
We can see that the dots are concentrated around the value 5 on the x-axis,
and 10 on the y-axis.

We can also see that the spread is wider on the y-axis than on the x-axis.

Machine Learning - Linear


Regression

Regression
The term regression is used when you try to find the relationship between
variables.

In Machine Learning, and in statistical modeling, that relationship is used to


predict the outcome of future events.

Linear Regression
Linear regression uses the relationship between the data-points to draw a
straight line through all them.

This line can be used to predict future values.

In Machine Learning, predicting the future is very important.


How Does it Work?
Python has methods for finding a relationship between data-points and to
draw a line of linear regression. We will show you how to use these methods
instead of going through the mathematic formula.

In the example below, the x-axis represents age, and the y-axis represents
speed. We have registered the age and speed of 13 cars as they were
passing a tollbooth. Let us see if the data we collected could be used in a
linear regression:

ExampleGet your own Python Server


Start by drawing a scatter plot:

import [Link] as plt

x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]

[Link](x, y)
[Link]()

Result:
Example
Import scipy and draw the line of Linear Regression:

import [Link] as plt


from scipy import stats

x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]

slope, intercept, r, p, std_err = [Link](x, y)

def myfunc(x):
return slope * x + intercept

mymodel = list(map(myfunc, x))

[Link](x, y)
[Link](x, mymodel)
[Link]()
Result:

Example Explained
Import the modules you need.

You can learn about the Matplotlib module in our Matplotlib Tutorial.

You can learn about the SciPy module in our SciPy Tutorial.

import [Link] as plt


from scipy import stats

Create the arrays that represent the values of the x and y axis:

x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]
Execute a method that returns some important key values of Linear
Regression:

slope, intercept, r, p, std_err = [Link](x, y)

Create a function that uses the slope and intercept values to return a new
value. This new value represents where on the y-axis the corresponding x
value will be placed:

def myfunc(x):
return slope * x + intercept

Run each value of the x array through the function. This will result in a new
array with new values for the y-axis:

mymodel = list(map(myfunc, x))

Draw the original scatter plot:

[Link](x, y)

Draw the line of linear regression:

[Link](x, mymodel)

Display the diagram:

[Link]()

REMOVE ADS

R for Relationship
It is important to know how the relationship between the values of the x-axis
and the values of the y-axis is, if there are no relationship the linear
regression can not be used to predict anything.

This relationship - the coefficient of correlation - is called r.


The r value ranges from -1 to 1, where 0 means no relationship, and 1 (and -
1) means 100% related.

Python and the Scipy module will compute this value for you, all you have to
do is feed it with the x and y values.

Example
How well does my data fit in a linear regression?

from scipy import stats

x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]

slope, intercept, r, p, std_err = [Link](x, y)

print(r)

Note: The result -0.76 shows that there is a relationship, not perfect, but it
indicates that we could use linear regression in future predictions.
REMOVE ADS

Predict Future Values


Now we can use the information we have gathered to predict future values.

Example: Let us try to predict the speed of a 10 years old car.

To do so, we need the same myfunc() function from the example above:

def myfunc(x):
return slope * x + intercept

Example
Predict the speed of a 10 years old car:

from scipy import stats

x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]
slope, intercept, r, p, std_err = [Link](x, y)

def myfunc(x):
return slope * x + intercept

speed = myfunc(10)

print(speed)

The example predicted a speed at 85.6, which we also could read from the
diagram:

Bad Fit?
Let us create an example where linear regression would not be the best
method to predict future values.

Example
These values for the x- and y-axis should result in a very bad fit for linear
regression:

import [Link] as plt


from scipy import stats

x = [89,43,36,36,95,10,66,34,38,20,26,29,48,64,6,5,36,66,72,40]
y = [21,46,3,35,67,95,53,72,58,10,26,34,90,33,38,20,56,2,47,15]

slope, intercept, r, p, std_err = [Link](x, y)

def myfunc(x):
return slope * x + intercept

mymodel = list(map(myfunc, x))

[Link](x, y)
[Link](x, mymodel)
[Link]()

Result:
And the r for relationship?

Example
You should get a very low r value.

import numpy
from scipy import stats

x = [89,43,36,36,95,10,66,34,38,20,26,29,48,64,6,5,36,66,72,40]
y = [21,46,3,35,67,95,53,72,58,10,26,34,90,33,38,20,56,2,47,15]

slope, intercept, r, p, std_err = [Link](x, y)

print(r)

The result: 0.013 indicates a very bad relationship, and tells us that this data
set is not suitable for linear regression.
Machine Learning -
Polynomial Regression

Polynomial Regression
If your data points clearly will not fit a linear regression (a straight line
through all data points), it might be ideal for polynomial regression.

Polynomial regression, like linear regression, uses the relationship between


the variables x and y to find the best way to draw a line through the data
points.
How Does it Work?
Python has methods for finding a relationship between data-points and to
draw a line of polynomial regression. We will show you how to use these
methods instead of going through the mathematic formula.

In the example below, we have registered 18 cars as they were passing a


certain tollbooth.

We have registered the car's speed, and the time of day (hour) the passing
occurred.

The x-axis represents the hours of the day and the y-axis represents the
speed:
ExampleGet your own Python Server
Start by drawing a scatter plot:

import [Link] as plt

x = [1,2,3,5,6,7,8,9,10,12,13,14,15,16,18,19,21,22]
y = [100,90,80,60,60,55,60,65,70,70,75,76,78,79,90,99,99,100]

[Link](x, y)
[Link]()

Result:

Example
Import numpy and matplotlib then draw the line of Polynomial Regression:

import numpy
import [Link] as plt
x = [1,2,3,5,6,7,8,9,10,12,13,14,15,16,18,19,21,22]
y = [100,90,80,60,60,55,60,65,70,70,75,76,78,79,90,99,99,100]

mymodel = numpy.poly1d([Link](x, y, 3))

myline = [Link](1, 22, 100)

[Link](x, y)
[Link](myline, mymodel(myline))
[Link]()

Result:

Example Explained
Import the modules you need.

You can learn about the NumPy module in our NumPy Tutorial.
You can learn about the SciPy module in our SciPy Tutorial.

import numpy
import [Link] as plt

Create the arrays that represent the values of the x and y axis:

x = [1,2,3,5,6,7,8,9,10,12,13,14,15,16,18,19,21,22]
y = [100,90,80,60,60,55,60,65,70,70,75,76,78,79,90,99,99,100]

NumPy has a method that lets us make a polynomial model:

mymodel = numpy.poly1d([Link](x, y, 3))

Then specify how the line will display, we start at position 1, and end at
position 22:

myline = [Link](1, 22, 100)

Draw the original scatter plot:

[Link](x, y)

Draw the line of polynomial regression:

[Link](myline, mymodel(myline))

Display the diagram:

[Link]()

REMOVE ADS

R-Squared
It is important to know how well the relationship between the values of the x-
and y-axis is, if there are no relationship the polynomial regression can not
be used to predict anything.

The relationship is measured with a value called the r-squared.


The r-squared value ranges from 0 to 1, where 0 means no relationship, and
1 means 100% related.

Python and the Sklearn module will compute this value for you, all you have
to do is feed it with the x and y arrays:

Example
How well does my data fit in a polynomial regression?

import numpy
from [Link] import r2_score

x = [1,2,3,5,6,7,8,9,10,12,13,14,15,16,18,19,21,22]
y = [100,90,80,60,60,55,60,65,70,70,75,76,78,79,90,99,99,100]

mymodel = numpy.poly1d([Link](x, y, 3))

print(r2_score(y, mymodel(x)))
Note: The result 0.94 shows that there is a very good relationship, and we
can use polynomial regression in future predictions.

Predict Future Values


Now we can use the information we have gathered to predict future values.

Example: Let us try to predict the speed of a car that passes the tollbooth at
around the time 17:00:

To do so, we need the same mymodel array from the example above:

mymodel = numpy.poly1d([Link](x, y, 3))

Example
Predict the speed of a car passing at 17:00:

import numpy
from [Link] import r2_score

x = [1,2,3,5,6,7,8,9,10,12,13,14,15,16,18,19,21,22]
y = [100,90,80,60,60,55,60,65,70,70,75,76,78,79,90,99,99,100]
mymodel = numpy.poly1d([Link](x, y, 3))

speed = mymodel(17)
print(speed)

The example predicted a speed to be 88.87, which we also could read from
the diagram:

REMOVE ADS

Bad Fit?
Let us create an example where polynomial regression would not be the best
method to predict future values.

Example
These values for the x- and y-axis should result in a very bad fit for
polynomial regression:

import numpy
import [Link] as plt

x = [89,43,36,36,95,10,66,34,38,20,26,29,48,64,6,5,36,66,72,40]
y = [21,46,3,35,67,95,53,72,58,10,26,34,90,33,38,20,56,2,47,15]

mymodel = numpy.poly1d([Link](x, y, 3))

myline = [Link](2, 95, 100)

[Link](x, y)
[Link](myline, mymodel(myline))
[Link]()

Result:

And the r-squared value?


Example
You should get a very low r-squared value.

import numpy
from [Link] import r2_score

x = [89,43,36,36,95,10,66,34,38,20,26,29,48,64,6,5,36,66,72,40]
y = [21,46,3,35,67,95,53,72,58,10,26,34,90,33,38,20,56,2,47,15]

mymodel = numpy.poly1d([Link](x, y, 3))

print(r2_score(y, mymodel(x)))

The result: 0.00995 indicates a very bad relationship, and tells us that this
data set is not suitable for polynomial regression.

Machine Learning - Multiple Regression

Multiple Regression

Multiple regression is like linear regression, but with more than one
independent value, meaning that we try to predict a value based on two or
more variables.

Take a look at the data set below, it contains some information about cars.

Car Model Volume Weight

Toyota Aygo 1000 790

Mitsubishi Space Star 1200 1160

Skoda Citigo 1000 929


Fiat 500 900 865

Mini Cooper 1500 1140

VW Up! 1000 929

Skoda Fabia 1400 1109

Mercedes A-Class 1500 1365

Ford Fiesta 1500 1112

Audi A1 1600 1150

Hyundai I20 1100 980

Suzuki Swift 1300 990

Ford Fiesta 1000 1112

Honda Civic 1600 1252

Hundai I30 1600 1326


Opel Astra 1600 1330

BMW 1 1600 1365

Mazda 3 2200 1280

Skoda Rapid 1600 1119

Ford Focus 2000 1328

Ford Mondeo 1600 1584

Opel Insignia 2000 1428

Mercedes C-Class 2100 1365

Skoda Octavia 1600 1415

Volvo S60 2000 1415

Mercedes CLA 1500 1465


Audi A4 2000 1490

Audi A6 2000 1725

Volvo V70 1600 1523

BMW 5 2000 1705

Mercedes E-Class 2100 1605

Volvo XC70 2000 1746

Ford B-Max 1600 1235

BMW 2 1600 1390

Opel Zafira 1600 1405

Mercedes SLK 2500 1395


We can predict the CO2 emission of a car based on the size of the engine,
but with multiple regression we can throw in more variables, like the weight
of the car, to make the prediction more accurate.

How Does it Work?

In Python we have modules that will do the work for us. Start by importing
the Pandas module.

import pandas

Learn about the Pandas module in our Pandas Tutorial.

The Pandas module allows us to read csv files and return a DataFrame
object.

The file is meant for testing purposes only, you can download it
here: [Link]

df = pandas.read_csv("[Link]")

Then make a list of the independent values and call this variable X.

Put the dependent values in a variable called y.

X = df[['Weight', 'Volume']]
y = df['CO2']

Tip: It is common to name the list of independent values with a upper case
X, and the list of dependent values with a lower case y.

We will use some methods from the sklearn module, so we will have to
import that module as well:

from sklearn import linear_model

From the sklearn module we will use the LinearRegression() method to


create a linear regression object.

This object has a method called fit() that takes the independent and
dependent values as parameters and fills the regression object with data
that describes the relationship:
regr = linear_model.LinearRegression()
[Link](X, y)

Now we have a regression object that are ready to predict CO2 values based
on a car's weight and volume:

#predict the CO2 emission of a car where the weight is 2300kg, and the
volume is 1300cm3:
predictedCO2 = [Link]([[2300, 1300]])

ExampleGet your own Python Server

See the whole example in action:

import pandas
from sklearn import linear_model

df = pandas.read_csv("[Link]")

X = df[['Weight', 'Volume']]
y = df['CO2']

regr = linear_model.LinearRegression()
[Link](X, y)

#predict the CO2 emission of a car where the weight is 2300kg, and the
volume is 1300cm3:
predictedCO2 = [Link]([[2300, 1300]])

print(predictedCO2)

Result:

[107.2087328]

We have predicted that a car with 1.3 liter engine, and a weight of 2300 kg,
will release approximately 107 grams of CO2 for every kilometer it drives.

REMOVE ADS

Coefficient
The coefficient is a factor that describes the relationship with an unknown
variable.

Example: if x is a variable, then 2x is x two times. x is the unknown variable,


and the number 2 is the coefficient.

In this case, we can ask for the coefficient value of weight against CO2, and
for volume against CO2. The answer(s) we get tells us what would happen if
we increase, or decrease, one of the independent values.

Example

Print the coefficient values of the regression object:

import pandas
from sklearn import linear_model

df = pandas.read_csv("[Link]")

X = df[['Weight', 'Volume']]
y = df['CO2']

regr = linear_model.LinearRegression()
[Link](X, y)

print(regr.coef_)

Result:

[0.00755095 0.00780526]

Result Explained

The result array represents the coefficient values of weight and volume.

Weight: 0.00755095
Volume: 0.00780526

These values tell us that if the weight increase by 1kg, the CO2 emission
increases by 0.00755095g.

And if the engine size (Volume) increases by 1cm 3, the CO2 emission
increases by 0.00780526g.

I think that is a fair guess, but let test it!


We have already predicted that if a car with a 1300cm 3 engine weighs
2300kg, the CO2 emission will be approximately 107g.

What if we increase the weight with 1000kg?

Example

Copy the example from before, but change the weight from 2300 to 3300:

import pandas
from sklearn import linear_model

df = pandas.read_csv("[Link]")

X = df[['Weight', 'Volume']]
y = df['CO2']

regr = linear_model.LinearRegression()
[Link](X, y)

predictedCO2 = [Link]([[3300, 1300]])

print(predictedCO2)

Result:

[114.75968007]

We have predicted that a car with 1.3 liter engine, and a weight of 3300 kg,
will release approximately 115 grams of CO2 for every kilometer it drives.

Which shows that the coefficient of 0.00755095 is correct:

107.2087328 + (1000 * 0.00755095) = 114.75968

Machine Learning - Scale

Scale Features
When your data has different values, and even different measurement units,
it can be difficult to compare them. What is kilograms compared to meters?
Or altitude compared to time?

The answer to this problem is scaling. We can scale data into new values
that are easier to compare.

Take a look at the table below, it is the same data set that we used in
the multiple regression chapter, but this time the volume column contains
values in liters instead of cm3 (1.0 instead of 1000).

Car Model Volume Weight

Toyota Aygo 1.0 790

Mitsubishi Space Star 1.2 1160

Skoda Citigo 1.0 929

Fiat 500 0.9 865

Mini Cooper 1.5 1140

VW Up! 1.0 929

Skoda Fabia 1.4 1109

Mercedes A-Class 1.5 1365

Ford Fiesta 1.5 1112

Audi A1 1.6 1150

Hyundai I20 1.1 980

Suzuki Swift 1.3 990

Ford Fiesta 1.0 1112

Honda Civic 1.6 1252

Hundai I30 1.6 1326

Opel Astra 1.6 1330

BMW 1 1.6 1365

Mazda 3 2.2 1280


Skoda Rapid 1.6 1119

Ford Focus 2.0 1328

Ford Mondeo 1.6 1584

Opel Insignia 2.0 1428

Mercedes C-Class 2.1 1365

Skoda Octavia 1.6 1415

Volvo S60 2.0 1415

Mercedes CLA 1.5 1465

Audi A4 2.0 1490

Audi A6 2.0 1725

Volvo V70 1.6 1523

BMW 5 2.0 1705

Mercedes E-Class 2.1 1605

Volvo XC70 2.0 1746

Ford B-Max 1.6 1235

BMW 2 1.6 1390

Opel Zafira 1.6 1405

Mercedes SLK 2.5 1395

It can be difficult to compare the volume 1.0 with the weight 790, but if we
scale them both into comparable values, we can easily see how much one
value is compared to the other.

There are different methods for scaling data, in this tutorial we will use a
method called standardization.

The standardization method uses this formula:

z = (x - u) / s
Where z is the new value, x is the original value, u is the mean and s is the
standard deviation.

If you take the weight column from the data set above, the first value is
790, and the scaled value will be:

(790 - 1292.23) / 238.74 = -2.1

If you take the volume column from the data set above, the first value is
1.0, and the scaled value will be:

(1.0 - 1.61) / 0.38 = -1.59

Now you can compare -2.1 with -1.59 instead of comparing 790 with 1.0.

You do not have to do this manually, the Python sklearn module has a
method called StandardScaler() which returns a Scaler object with methods
for transforming data sets.

ExampleGet your own Python Server


Scale all values in the Weight and Volume columns:

import pandas
from sklearn import linear_model
from [Link] import StandardScaler
scale = StandardScaler()

df = pandas.read_csv("[Link]")

X = df[['Weight', 'Volume']]

scaledX = scale.fit_transform(X)

print(scaledX)

Result:
Note that the first two values are -2.1 and -1.59, which corresponds to our
calculations:

[[-2.10389253 -1.59336644]
[-0.55407235 -1.07190106]
[-1.52166278 -1.59336644]
[-1.78973979 -1.85409913]
[-0.63784641 -0.28970299]
[-1.52166278 -1.59336644]
[-0.76769621 -0.55043568]
[ 0.3046118 -0.28970299]
[-0.7551301 -0.28970299]
[-0.59595938 -0.0289703 ]
[-1.30803892 -1.33263375]
[-1.26615189 -0.81116837]
[-0.7551301 -1.59336644]
[-0.16871166 -0.0289703 ]
[ 0.14125238 -0.0289703 ]
[ 0.15800719 -0.0289703 ]
[ 0.3046118 -0.0289703 ]
[-0.05142797 1.53542584]
[-0.72580918 -0.0289703 ]
[ 0.14962979 1.01396046]
[ 1.2219378 -0.0289703 ]
[ 0.5685001 1.01396046]
[ 0.3046118 1.27469315]
[ 0.51404696 -0.0289703 ]
[ 0.51404696 1.01396046]
[ 0.72348212 -0.28970299]
[ 0.8281997 1.01396046]
[ 1.81254495 1.01396046]
[ 0.96642691 -0.0289703 ]
[ 1.72877089 1.01396046]
[ 1.30990057 1.27469315]
[ 1.90050772 1.01396046]
[-0.23991961 -0.0289703 ]
[ 0.40932938 -0.0289703 ]
[ 0.47215993 -0.0289703 ]
[ 0.4302729 2.31762392]]

REMOVE ADS

Predict CO2 Values


The task in the Multiple Regression chapter was to predict the CO2 emission
from a car when you only knew its weight and volume.

When the data set is scaled, you will have to use the scale when you predict
values:
Example
Predict the CO2 emission from a 1.3 liter car that weighs 2300 kilograms:

import pandas
from sklearn import linear_model
from [Link] import StandardScaler
scale = StandardScaler()

df = pandas.read_csv("[Link]")

X = df[['Weight', 'Volume']]
y = df['CO2']

scaledX = scale.fit_transform(X)

regr = linear_model.LinearRegression()
[Link](scaledX, y)

scaled = [Link]([[2300, 1.3]])

predictedCO2 = [Link]([scaled[0]])
print(predictedCO2)

Result:
[107.2087328]

Machine Learning - Train/Test

Evaluate Your Model


In Machine Learning we create models to predict the outcome of certain
events, like in the previous chapter where we predicted the CO2 emission of
a car when we knew the weight and engine size.

To measure if the model is good enough, we can use a method called


Train/Test.
What is Train/Test
Train/Test is a method to measure the accuracy of your model.

It is called Train/Test because you split the data set into two sets: a training
set and a testing set.

80% for training, and 20% for testing.

You train the model using the training set.

You test the model using the testing set.

Train the model means create the model.

Test the model means test the accuracy of the model.

Start With a Data Set


Start with a data set you want to test.

Our data set illustrates 100 customers in a shop, and their shopping habits.

ExampleGet your own Python Server


import numpy
import [Link] as plt
[Link](2)

x = [Link](3, 1, 100)
y = [Link](150, 40, 100) / x

[Link](x, y)
[Link]()

Result:
The x axis represents the number of minutes before making a purchase.

The y axis represents the amount of money spent on the purchase.


REMOVE ADS

Split Into Train/Test


The training set should be a random selection of 80% of the original data.

The testing set should be the remaining 20%.

train_x = x[:80]
train_y = y[:80]

test_x = x[80:]
test_y = y[80:]
Display the Training Set
Display the same scatter plot with the training set:

Example
[Link](train_x, train_y)
[Link]()

Result:
It looks like the original data set, so it seems to be a fair selection:
Display the Testing Set
To make sure the testing set is not completely different, we will take a look
at the testing set as well.

Example
[Link](test_x, test_y)
[Link]()

Result:
The testing set also looks like the original data set:

Fit the Data Set


What does the data set look like? In my opinion I think the best fit would be
a polynomial regression, so let us draw a line of polynomial regression.

To draw a line through the data points, we use the plot() method of the
matplotlib module:

Example
Draw a polynomial regression line through the data points:

import numpy
import [Link] as plt
[Link](2)

x = [Link](3, 1, 100)
y = [Link](150, 40, 100) / x

train_x = x[:80]
train_y = y[:80]

test_x = x[80:]
test_y = y[80:]

mymodel = numpy.poly1d([Link](train_x, train_y, 4))

myline = [Link](0, 6, 100)

[Link](train_x, train_y)
[Link](myline, mymodel(myline))
[Link]()

Result:
The result can back my suggestion of the data set fitting a polynomial
regression, even though it would give us some weird results if we try to
predict values outside of the data set. Example: the line indicates that a
customer spending 6 minutes in the shop would make a purchase worth 200.
That is probably a sign of overfitting.

But what about the R-squared score? The R-squared score is a good indicator
of how well my data set is fitting the model.

R2
Remember R2, also known as R-squared?

It measures the relationship between the x axis and the y axis, and the value
ranges from 0 to 1, where 0 means no relationship, and 1 means totally
related.
The sklearn module has a method called r2_score() that will help us find this
relationship.

In this case we would like to measure the relationship between the minutes a
customer stays in the shop and how much money they spend.

Example
How well does my training data fit in a polynomial regression?

import numpy
from [Link] import r2_score
[Link](2)

x = [Link](3, 1, 100)
y = [Link](150, 40, 100) / x

train_x = x[:80]
train_y = y[:80]

test_x = x[80:]
test_y = y[80:]

mymodel = numpy.poly1d([Link](train_x, train_y, 4))

r2 = r2_score(train_y, mymodel(train_x))

print(r2)
Note: The result 0.799 shows that there is a OK relationship.

Bring in the Testing Set


Now we have made a model that is OK, at least when it comes to training
data.

Now we want to test the model with the testing data as well, to see if gives
us the same result.

Example
Let us find the R2 score when using testing data:

import numpy
from [Link] import r2_score
[Link](2)
x = [Link](3, 1, 100)
y = [Link](150, 40, 100) / x

train_x = x[:80]
train_y = y[:80]

test_x = x[80:]
test_y = y[80:]

mymodel = numpy.poly1d([Link](train_x, train_y, 4))

r2 = r2_score(test_y, mymodel(test_x))

print(r2)
Note: The result 0.809 shows that the model fits the testing set as well, and
we are confident that we can use the model to predict future values.

Predict Values
Now that we have established that our model is OK, we can start predicting
new values.

Example
How much money will a buying customer spend, if she or he stays in the
shop for 5 minutes?

print(mymodel(5))

The example predicted the customer to spend 22.88 dollars, as seems to


correspond to the diagram:

Machine Learning - Decision


Tree
Decision Tree
In this chapter we will show you how to make a "Decision Tree". A Decision
Tree is a Flow Chart, and can help you make decisions based on previous
experience.
In the example, a person will try to decide if he/she should go to a comedy
show or not.

Luckily our example person has registered every time there was a comedy
show in town, and registered some information about the comedian, and also
registered if he/she went or not.

Age Experience Rank Nationality

36 10 9 UK

42 12 4 USA

23 4 6 N

52 4 4 USA

43 21 8 USA

44 14 5 UK

66 3 7 N

35 14 9 UK

52 13 7 N
35 5 9 N

24 3 5 USA

18 3 7 UK

45 9 9 UK

Now, based on this data set, Python can create a decision tree that can be
used to decide if any new shows are worth attending to.

REMOVE ADS

How Does it Work?


First, read the dataset with pandas:

ExampleGet your own Python Server


Read and print the data set:

import pandas

df = pandas.read_csv("[Link]")

print(df)
To make a decision tree, all data has to be numerical.

We have to convert the non numerical columns 'Nationality' and 'Go' into
numerical values.

Pandas has a map() method that takes a dictionary with information on how
to convert the values.

{'UK': 0, 'USA': 1, 'N': 2}

Means convert the values 'UK' to 0, 'USA' to 1, and 'N' to 2.

Example
Change string values into numerical values:

d = {'UK': 0, 'USA': 1, 'N': 2}


df['Nationality'] = df['Nationality'].map(d)
d = {'YES': 1, 'NO': 0}
df['Go'] = df['Go'].map(d)

print(df)

Then we have to separate the feature columns from the target column.

The feature columns are the columns that we try to predict from, and the
target column is the column with the values we try to predict.

Example
X is the feature columns, y is the target column:

features = ['Age', 'Experience', 'Rank', 'Nationality']

X = df[features]
y = df['Go']

print(X)
print(y)

Now we can create the actual decision tree, fit it with our details. Start by
importing the modules we need:

Example
Create and display a Decision Tree:

import pandas
from sklearn import tree
from [Link] import DecisionTreeClassifier
import [Link] as plt

df = pandas.read_csv("[Link]")

d = {'UK': 0, 'USA': 1, 'N': 2}


df['Nationality'] = df['Nationality'].map(d)
d = {'YES': 1, 'NO': 0}
df['Go'] = df['Go'].map(d)

features = ['Age', 'Experience', 'Rank', 'Nationality']

X = df[features]
y = df['Go']

dtree = DecisionTreeClassifier()
dtree = [Link](X, y)

tree.plot_tree(dtree, feature_names=features)

Result Explained
The decision tree uses your earlier decisions to calculate the odds for you to
wanting to go see a comedian or not.

Let us read the different aspects of the decision tree:

Rank
Rank <= 6.5 means that every comedian with a rank of 6.5 or lower will follow
the True arrow (to the left), and the rest will follow the False arrow (to the
right).

gini = 0.497 refers to the quality of the split, and is always a number
between 0.0 and 0.5, where 0.0 would mean all of the samples got the same
result, and 0.5 would mean that the split is done exactly in the middle.

samples = 13means that there are 13 comedians left at this point in the
decision, which is all of them since this is the first step.

value = [6, 7] means that of these 13 comedians, 6 will get a "NO", and 7
will get a "GO".

Gini
There are many ways to split the samples, we use the GINI method in this
tutorial.

The Gini method uses this formula:

Gini = 1 - (x/n)2 - (y/n)2

Where x is the number of positive answers("GO"), n is the number of


samples, and y is the number of negative answers ("NO"), which gives us this
calculation:

1 - (7 / 13)2 - (6 / 13)2 = 0.497


The next step contains two boxes, one box for the comedians with a 'Rank' of
6.5 or lower, and one box with the rest.

True - 5 Comedians End Here:


gini = 0.0 means all of the samples got the same result.

samples = 5means that there are 5 comedians left in this branch (5 comedian
with a Rank of 6.5 or lower).

value = [5, 0] means that 5 will get a "NO" and 0 will get a "GO".

False - 8 Comedians Continue:


Nationality
Nationality <= 0.5 means that the comedians with a nationality value of less
than 0.5 will follow the arrow to the left (which means everyone from the UK,
), and the rest will follow the arrow to the right.

gini = 0.219 means that about 22% of the samples would go in one direction.

samples = 8means that there are 8 comedians left in this branch (8 comedian
with a Rank higher than 6.5).

value = [1, 7] means that of these 8 comedians, 1 will get a "NO" and 7 will
get a "GO".
True - 4 Comedians Continue:
Age
Age <= 35.5means that comedians at the age of 35.5 or younger will follow
the arrow to the left, and the rest will follow the arrow to the right.

gini = 0.375 means that about 37,5% of the samples would go in one
direction.

means that there are 4 comedians left in this branch (4


samples = 4
comedians from the UK).

value = [1, 3] means that of these 4 comedians, 1 will get a "NO" and 3 will
get a "GO".

False - 4 Comedians End Here:


gini = 0.0 means all of the samples got the same result.

means that there are 4 comedians left in this branch (4


samples = 4
comedians not from the UK).

value = [0, 4] means that of these 4 comedians, 0 will get a "NO" and 4 will
get a "GO".
True - 2 Comedians End Here:
gini = 0.0 means all of the samples got the same result.

means that there are 2 comedians left in this branch (2


samples = 2
comedians at the age 35.5 or younger).

value = [0, 2] means that of these 2 comedians, 0 will get a "NO" and 2 will
get a "GO".

False - 2 Comedians Continue:


Experience
Experience <= 9.5 means that comedians with 9.5 years of experience, or
less, will follow the arrow to the left, and the rest will follow the arrow to the
right.

gini = 0.5 means that 50% of the samples would go in one direction.

means that there are 2 comedians left in this branch (2


samples = 2
comedians older than 35.5).
value = [1, 1] means that of these 2 comedians, 1 will get a "NO" and 1 will
get a "GO".

True - 1 Comedian Ends Here:


gini = 0.0 means all of the samples got the same result.

samples = 1means that there is 1 comedian left in this branch (1 comedian


with 9.5 years of experience or less).

value = [0, 1] means that 0 will get a "NO" and 1 will get a "GO".

False - 1 Comedian Ends Here:


gini = 0.0 means all of the samples got the same result.

samples = 1means that there is 1 comedians left in this branch (1 comedian


with more than 9.5 years of experience).

value = [1, 0] means that 1 will get a "NO" and 0 will get a "GO".

Predict Values
We can use the Decision Tree to predict new values.

Example: Should I go see a show starring a 40 years old American comedian,


with 10 years of experience, and a comedy ranking of 7?

Example
Use predict() method to predict new values:

print([Link]([[40, 10, 7, 1]]))

Example
What would the answer be if the comedy rank was 6?

print([Link]([[40, 10, 6, 1]]))

Different Results
You will see that the Decision Tree gives you different results if you run it
enough times, even if you feed it with the same data.

That is because the Decision Tree does not give us a 100% certain answer. It
is based on the probability of an outcome, and the answer will vary.

Machine Learning - Confusion


Matrix

What is a confusion matrix?


It is a table that is used in classification problems to assess where errors in
the model were made.

The rows represent the actual classes the outcomes should have been. While
the columns represent the predictions we have made. Using this table it is
easy to see which predictions are wrong.
Creating a Confusion Matrix
Confusion matrixes can be created by predictions made from a logistic
regression.

For now we will generate actual and predicted values by utilizing NumPy:

import numpy

Next we will need to generate the numbers for "actual" and "predicted"
values.

actual = [Link](1, 0.9, size = 1000)


predicted = [Link](1, 0.9, size = 1000)

In order to create the confusion matrix we need to import metrics from the
sklearn module.

from sklearn import metrics

Once metrics is imported we can use the confusion matrix function on our
actual and predicted values.

confusion_matrix = metrics.confusion_matrix(actual, predicted)

To create a more interpretable visual display we need to convert the table


into a confusion matrix display.

cm_display = [Link](confusion_matrix =
confusion_matrix, display_labels = [0, 1])

Vizualizing the display requires that we import pyplot from matplotlib.

import [Link] as plt

Finally to display the plot we can use the functions plot() and show() from
pyplot.

cm_display.plot()
[Link]()

See the whole example in action:


ExampleGet your own Python Server
import [Link] as plt
import numpy
from sklearn import metrics

actual = [Link](1,.9,size = 1000)


predicted = [Link](1,.9,size = 1000)

confusion_matrix = metrics.confusion_matrix(actual, predicted)

cm_display = [Link](confusion_matrix =
confusion_matrix, display_labels = [0, 1])

cm_display.plot()
[Link]()

Result

Results Explained
The Confusion Matrix created has four different quadrants:

True Negative (Top-Left Quadrant)


False Positive (Top-Right Quadrant)
False Negative (Bottom-Left Quadrant)
True Positive (Bottom-Right Quadrant)

True means that the values were accurately predicted, False means that
there was an error or wrong prediction.

Now that we have made a Confusion Matrix, we can calculate different


measures to quantify the quality of the model. First, lets look at Accuracy.

REMOVE ADS

Created Metrics
The matrix provides us with many useful metrics that help us to evaluate our
classification model.

The different measures include: Accuracy, Precision, Sensitivity (Recall),


Specificity, and the F-score, explained below.

Accuracy
Accuracy measures how often the model is correct.

How to Calculate
(True Positive + True Negative) / Total Predictions

Example
Accuracy = metrics.accuracy_score(actual, predicted)
Precision
Of the positives predicted, what percentage is truly positive?

How to Calculate
True Positive / (True Positive + False Positive)

Precision does not evaluate the correctly predicted negative cases:

Example
Precision = metrics.precision_score(actual, predicted)

Sensitivity (Recall)
Of all the positive cases, what percentage are predicted positive?

Sensitivity (sometimes called Recall) measures how good the model is at


predicting positives.

This means it looks at true positives and false negatives (which are positives
that have been incorrectly predicted as negative).

How to Calculate
True Positive / (True Positive + False Negative)

Sensitivity is good at understanding how well the model predicts something


is positive:

Example
Sensitivity_recall = metrics.recall_score(actual, predicted)

Specificity
How well the model is at prediciting negative results?

Specificity is similar to sensitivity, but looks at it from the persepctive of


negative results.

How to Calculate
True Negative / (True Negative + False Positive)

Since it is just the opposite of Recall, we use the recall_score function, taking
the opposite position label:

Example
Specificity = metrics.recall_score(actual, predicted, pos_label=0)

F-score
F-score is the "harmonic mean" of precision and sensitivity.

It considers both false positive and false negative cases and is good for
imbalanced datasets.

How to Calculate
2 * ((Precision * Sensitivity) / (Precision + Sensitivity))

This score does not take into consideration the True Negative values:

Example
F1_score = metrics.f1_score(actual, predicted)

All calulations in one:

Example
#metrics
print({"Accuracy":Accuracy,"Precision":Precision,"Sensitivity_recall"
:Sensitivity_recall,"Specificity":Specificity,"F1_score":F1_score})
Machine Learning -
Hierarchical Clustering

Hierarchical Clustering
Hierarchical clustering is an unsupervised learning method for clustering
data points. The algorithm builds clusters by measuring the dissimilarities
between data. Unsupervised learning means that a model does not have to
be trained, and we do not need a "target" variable. This method can be used
on any data to visualize and interpret the relationship between individual
data points.

Here we will use hierarchical clustering to group data points and visualize the
clusters using both a dendrogram and scatter plot.

How does it work?


We will use Agglomerative Clustering, a type of hierarchical clustering that
follows a bottom up approach. We begin by treating each data point as its
own cluster. Then, we join clusters together that have the shortest distance
between them to create larger clusters. This step is repeated until one large
cluster is formed containing all of the data points.

Hierarchical clustering requires us to decide on both a distance and linkage


method. We will use euclidean distance and the Ward linkage method, which
attempts to minimize the variance between clusters.

ExampleGet your own Python Server


Start by visualizing some data points:

import numpy as np
import [Link] as plt

x = [4, 5, 10, 4, 3, 11, 14 , 6, 10, 12]


y = [21, 19, 24, 17, 16, 25, 24, 22, 21, 21]
[Link](x, y)
[Link]()

Result

Now we compute the ward linkage using euclidean distance, and visualize it
using a dendrogram:

Example
import numpy as np
import [Link] as plt
from [Link] import dendrogram, linkage

x = [4, 5, 10, 4, 3, 11, 14 , 6, 10, 12]


y = [21, 19, 24, 17, 16, 25, 24, 22, 21, 21]
data = list(zip(x, y))

linkage_data = linkage(data, method='ward', metric='euclidean')


dendrogram(linkage_data)

[Link]()

Result

Here, we do the same thing with Python's scikit-learn library. Then, visualize
on a 2-dimensional plot:
Example
import numpy as np
import [Link] as plt
from [Link] import AgglomerativeClustering

x = [4, 5, 10, 4, 3, 11, 14 , 6, 10, 12]


y = [21, 19, 24, 17, 16, 25, 24, 22, 21, 21]
data = list(zip(x, y))

hierarchical_cluster =
AgglomerativeClustering(n_clusters=2, linkage='ward')
labels = hierarchical_cluster.fit_predict(data)

[Link](x, y, c=labels)
[Link]()

Result

REMOVE ADS
Example Explained
Import the modules you need.

import numpy as np
import [Link] as plt
from [Link] import dendrogram, linkage
from [Link] import AgglomerativeClustering

You can learn about the Matplotlib module in our "Matplotlib Tutorial.

You can learn about the SciPy module in our SciPy Tutorial.

NumPy is a library for working with arrays and matricies in Python, you can
learn about the NumPy module in our NumPy Tutorial.

scikit-learn is a popular library for machine learning.

Create arrays that resemble two variables in a dataset. Note that while we
only use two variables here, this method will work with any number of
variables:

x = [4, 5, 10, 4, 3, 11, 14 , 6, 10, 12]


y = [21, 19, 24, 17, 16, 25, 24, 22, 21, 21]

Turn the data into a set of points:

data = list(zip(x, y))


print(data)

Result:

[(4, 21), (5, 19), (10, 24), (4, 17), (3, 16), (11, 25),
(14, 24), (6, 22), (10, 21), (12, 21)]

Compute the linkage between all of the different points. Here we use a
simple euclidean distance measure and Ward's linkage, which seeks to
minimize the variance between clusters.

linkage_data = linkage(data, method='ward', metric='euclidean')

Finally, plot the results in a dendrogram. This plot will show us the hierarchy
of clusters from the bottom (individual points) to the top (a single cluster
consisting of all data points).
[Link]() lets us visualize the dendrogram instead of just the raw linkage
data.

dendrogram(linkage_data)
[Link]()

Result:

The scikit-learn library allows us to use hierarchichal clustering in a different


manner. First, we initialize the AgglomerativeClustering class with 2 clusters
and the Ward linkage.

hierarchical_cluster = AgglomerativeClustering(n_clusters=2,
linkage='ward')

The .fit_predict method can be called on our data to compute the clusters
using the defined parameters across our chosen number of clusters.

labels = hierarchical_cluster.fit_predict(data) print(labels)


Result:

[0 0 1 0 0 1 1 0 1 1]

Finally, if we plot the same data and color the points using the labels
assigned to each index by the hierarchical clustering method, we can see the
cluster each point was assigned to:

[Link](x, y, c=labels)
[Link]()

Result:

Machine Learning - Logistic


Regression

Logistic Regression
Logistic regression aims to solve classification problems. It does this by
predicting categorical outcomes, unlike linear regression that predicts a
continuous outcome.

In the simplest case there are two outcomes, which is called binomial, an
example of which is predicting if a tumor is malignant or benign. Other cases
have more than two outcomes to classify, in this case it is called multinomial.
A common example for multinomial logistic regression would be predicting
the class of an iris flower between 3 different species.

Here we will be using basic logistic regression to predict a binomial variable.


This means it has only two possible outcomes.

How does it work?


In Python we have modules that will do the work for us. Start by importing
the NumPy module.
import numpy

Store the independent variables in X.

Store the dependent variable in y.

Below is a sample dataset:

#X represents the size of a tumor in centimeters.


X =
[Link]([3.78, 2.44, 2.09, 0.14, 1.72, 1.65, 4.92, 4.37, 4.96
, 4.52, 3.69, 5.88]).reshape(-1,1)

#Note: X has to be reshaped into a column from a row for the


LogisticRegression() function to work.
#y represents whether or not the tumor is cancerous (0 for "No",
1 for "Yes").
y = [Link]([0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1])

We will use a method from the sklearn module, so we will have to import that
module as well:

from sklearn import linear_model

From the sklearn module we will use the LogisticRegression() method to


create a logistic regression object.

This object has a method called fit() that takes the independent and
dependent values as parameters and fills the regression object with data
that describes the relationship:

logr = linear_model.LogisticRegression()
[Link](X,y)

Now we have a logistic regression object that is ready to whether a tumor is


cancerous based on the tumor size:

#predict if tumor is cancerous where the size is 3.46mm:


predicted = [Link]([Link]([3.46]).reshape(-1,1))

ExampleGet your own Python Server


See the whole example in action:

import numpy
from sklearn import linear_model
#Reshaped for Logistic function.
X =
[Link]([3.78, 2.44, 2.09, 0.14, 1.72, 1.65, 4.92, 4.37, 4.96, 4.
52, 3.69, 5.88]).reshape(-1,1)
y = [Link]([0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1])

logr = linear_model.LogisticRegression()
[Link](X,y)

#predict if tumor is cancerous where the size is 3.46mm:


predicted = [Link]([Link]([3.46]).reshape(-1,1))
print(predicted)

Result
[0]

We have predicted that a tumor with a size of 3.46mm will not be cancerous.

REMOVE ADS

Coefficient
In logistic regression the coefficient is the expected change in log-odds of
having the outcome per unit change in X.

This does not have the most intuitive understanding so let's use it to create
something that makes more sense, odds.

Example
See the whole example in action:

import numpy
from sklearn import linear_model

#Reshaped for Logistic function.


X =
[Link]([3.78, 2.44, 2.09, 0.14, 1.72, 1.65, 4.92, 4.37, 4.96, 4.
52, 3.69, 5.88]).reshape(-1,1)
y = [Link]([0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1])

logr = linear_model.LogisticRegression()
[Link](X,y)

log_odds = logr.coef_
odds = [Link](log_odds)

print(odds)

Result
[4.03541657]

This tells us that as the size of a tumor increases by 1mm the odds of it
being a cancerous tumor increases by 4x.

REMOVE ADS

Probability
The coefficient and intercept values can be used to find the probability that
each tumor is cancerous.

Create a function that uses the model's coefficient and intercept values to
return a new value. This new value represents probability that the given
observation is a tumor:

def logit2prob(logr,x):
log_odds = logr.coef_ * x + logr.intercept_
odds = [Link](log_odds)
probability = odds / (1 + odds)
return(probability)

Function Explained
To find the log-odds for each observation, we must first create a formula that
looks similar to the one from linear regression, extracting the coefficient and
the intercept.
log_odds = logr.coef_ * x + logr.intercept_

To then convert the log-odds to odds we must exponentiate the log-odds.

odds = [Link](log_odds)

Now that we have the odds, we can convert it to probability by dividing it by


1 plus the odds.

probability = odds / (1 + odds)

Let us now use the function with what we have learned to find out the
probability that each tumor is cancerous.

Example
See the whole example in action:

import numpy
from sklearn import linear_model

X =
[Link]([3.78, 2.44, 2.09, 0.14, 1.72, 1.65, 4.92, 4.37, 4.96, 4.
52, 3.69, 5.88]).reshape(-1,1)
y = [Link]([0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1])

logr = linear_model.LogisticRegression()
[Link](X,y)

def logit2prob(logr, X):


log_odds = logr.coef_ * X + logr.intercept_
odds = [Link](log_odds)
probability = odds / (1 + odds)
return(probability)

print(logit2prob(logr, X))

Result
[[0.60749955]
[0.19268876]
[0.12775886]
[0.00955221]
[0.08038616]
[0.07345637]
[0.88362743]
[0.77901378]
[0.88924409]
[0.81293497]
[0.57719129]
[0.96664243]]

Results Explained
3.78 0.61 The probability that a tumor with the size 3.78cm is cancerous is
61%.

2.44 0.19 The probability that a tumor with the size 2.44cm is cancerous is
19%.

2.09 0.13 The probability that a tumor with the size 2.09cm is cancerous is
13%

Machine Learning - Grid


Search

Grid Search
The majority of machine learning models contain parameters that can be
adjusted to vary how the model learns. For example, the logistic regression
model, from sklearn, has a parameter C that controls regularization,which
affects the complexity of the model.

How do we pick the best value for C? The best value is dependent on the
data used to train the model.

How does it work?


One method is to try out different values and then pick the value that gives
the best score. This technique is known as a grid search. If we had to select
the values for two or more parameters, we would evaluate all combinations
of the sets of values thus forming a grid of values.

Before we get into the example it is good to know what the parameter we
are changing does. Higher values of C tell the model, the training data
resembles real world information, place a greater weight on the training
data. While lower values of C do the opposite.

Using Default Parameters


First let's see what kind of results we can generate without a grid search
using only the base parameters.

To get started we must first load in the dataset we will be working with.

from sklearn import datasets


iris = datasets.load_iris()

Next in order to create the model we must have a set of independent


variables X and a dependant variable y.

X = iris['data']
y = iris['target']

Now we will load the logistic model for classifying the iris flowers.

from sklearn.linear_model import LogisticRegression

Creating the model, setting max_iter to a higher value to ensure that the
model finds a result.

Keep in mind the default value for C in a logistic regression model is 1, we


will compare this later.

In the example below, we look at the iris data set and try to train a model
with varying values for C in logistic regression.

logit = LogisticRegression(max_iter = 10000)

After we create the model, we must fit the model to the data.

print([Link](X,y))
To evaluate the model we run the score method.

print([Link](X,y))

ExampleGet your own Python Server


from sklearn import datasets
from sklearn.linear_model import LogisticRegression

iris = datasets.load_iris()

X = iris['data']
y = iris['target']

logit = LogisticRegression(max_iter = 10000)

print([Link](X,y))

print([Link](X,y))

With the default setting of C = 1, we achieved a score of 0.973.

Let's see if we can do any better by implementing a grid search with


difference values of 0.973.

REMOVE ADS

Implementing Grid Search


We will follow the same steps of before except this time we will set a range
of values for C.

Knowing which values to set for the searched parameters will take a
combination of domain knowledge and practice.

Since the default value for C is 1, we will set a range of values surrounding it.

C = [0.25, 0.5, 0.75, 1, 1.25, 1.5, 1.75, 2]

Next we will create a for loop to change out the values of C and evaluate the
model with each change.
First we will create an empty list to store the score within.

scores = []

To change the values of C we must loop over the range of values and update
the parameter each time.

for choice in C:
logit.set_params(C=choice)
[Link](X, y)
[Link]([Link](X, y))

With the scores stored in a list, we can evaluate what the best choice of C is.

print(scores)

Example
from sklearn import datasets
from sklearn.linear_model import LogisticRegression

iris = datasets.load_iris()

X = iris['data']
y = iris['target']

logit = LogisticRegression(max_iter = 10000)

C = [0.25, 0.5, 0.75, 1, 1.25, 1.5, 1.75, 2]

scores = []

for choice in C:
logit.set_params(C=choice)
[Link](X, y)
[Link]([Link](X, y))

print(scores)

Results Explained
We can see that the lower values of C performed worse than the base
parameter of 1. However, as we increased the value of C to 1.75 the model
experienced increased accuracy.
It seems that increasing C beyond this amount does not help increase model
accuracy.

Note on Best Practices


We scored our logistic regression model by using the same data that was
used to train it. If the model corresponds too closely to that data, it may not
be great at predicting unseen data. This statistical error is known as over
fitting.

To avoid being misled by the scores on the training data, we can put aside a
portion of our data and use it specifically for the purpose of testing the
model. Refer to the lecture on train/test splitting to avoid being misled and
overfitting.

Preprocessing - Categorical
Data

Categorical Data
When your data has categories represented by strings, it will be difficult to
use them to train machine learning models which often only accepts numeric
data.

Instead of ignoring the categorical data and excluding the information from
our model, you can tranform the data so it can be used in your models.

Take a look at the table below, it is the same data set that we used in
the multiple regression chapter.

ExampleGet your own Python Server


import pandas as pd

cars = pd.read_csv('[Link]')
print(cars.to_string())
Result
Car Model Volume Weight CO2
0 Toyoty Aygo 1000 790 99
1 Mitsubishi Space Star 1200 1160 95
2 Skoda Citigo 1000 929 95
3 Fiat 500 900 865 90
4 Mini Cooper 1500 1140 105
5 VW Up! 1000 929 105
6 Skoda Fabia 1400 1109 90
7 Mercedes A-Class 1500 1365 92
8 Ford Fiesta 1500 1112 98
9 Audi A1 1600 1150 99
10 Hyundai I20 1100 980 99
11 Suzuki Swift 1300 990 101
12 Ford Fiesta 1000 1112 99
13 Honda Civic 1600 1252 94
14 Hundai I30 1600 1326 97
15 Opel Astra 1600 1330 97
16 BMW 1 1600 1365 99
17 Mazda 3 2200 1280 104
18 Skoda Rapid 1600 1119 104
19 Ford Focus 2000 1328 105
20 Ford Mondeo 1600 1584 94
21 Opel Insignia 2000 1428 99
22 Mercedes C-Class 2100 1365 99
23 Skoda Octavia 1600 1415 99
24 Volvo S60 2000 1415 99
25 Mercedes CLA 1500 1465 102
26 Audi A4 2000 1490 104
27 Audi A6 2000 1725 114
28 Volvo V70 1600 1523 109
29 BMW 5 2000 1705 114
30 Mercedes E-Class 2100 1605 115
31 Volvo XC70 2000 1746 117
32 Ford B-Max 1600 1235 104
33 BMW 216 1600 1390 108
34 Opel Zafira 1600 1405 109
35 Mercedes SLK 2500 1395 120

In the multiple regression chapter, we tried to predict the CO2 emitted based
on the volume of the engine and the weight of the car but we excluded
information about the car brand and model.
The information about the car brand or the car model might help us make a
better prediction of the CO2 emitted.

One Hot Encoding


We cannot make use of the Car or Model column in our data since they are
not numeric. A linear relationship between a categorical variable, Car or
Model, and a numeric variable, CO2, cannot be determined.

To fix this issue, we must have a numeric representation of the categorical


variable. One way to do this is to have a column representing each group in
the category.

For each column, the values will be 1 or 0 where 1 represents the inclusion of
the group and 0 represents the exclusion. This transformation is called one
hot encoding.

You do not have to do this manually, the Python Pandas module has a
function that called get_dummies() which does one hot encoding.

Learn about the Pandas module in our Pandas Tutorial.

Example
One Hot Encode the Car column:

import pandas as pd

cars = pd.read_csv('[Link]')
ohe_cars = pd.get_dummies(cars[['Car']])

print(ohe_cars.to_string())

Result
Car_Audi Car_BMW Car_Fiat Car_Ford Car_Honda
Car_Hundai Car_Hyundai Car_Mazda Car_Mercedes Car_Mini
Car_Mitsubishi Car_Opel Car_Skoda Car_Suzuki Car_Toyoty
Car_VW Car_Volvo
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 1 0 0
1 0 0 0 0 0
0 0 0 0 0 1
0 0 0 0 0 0
2 0 0 0 0 0
0 0 0 0 0 0
0 1 0 0 0 0
3 0 0 1 0 0
0 0 0 0 0 0
0 0 0 0 0 0
4 0 0 0 0 0
0 0 0 0 1 0
0 0 0 0 0 0
5 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 1 0
6 0 0 0 0 0
0 0 0 0 0 0
0 1 0 0 0 0
7 0 0 0 0 0
0 0 0 1 0 0
0 0 0 0 0 0
8 0 0 0 1 0
0 0 0 0 0 0
0 0 0 0 0 0
9 1 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
10 0 0 0 0 0
0 1 0 0 0 0
0 0 0 0 0 0
11 0 0 0 0 0
0 0 0 0 0 0
0 0 1 0 0 0
12 0 0 0 1 0
0 0 0 0 0 0
0 0 0 0 0 0
13 0 0 0 0 1
0 0 0 0 0 0
0 0 0 0 0 0
14 0 0 0 0 0
1 0 0 0 0 0
0 0 0 0 0 0
15 0 0 0 0 0
0 0 0 0 0 0
1 0 0 0 0 0
16 0 1 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
17 0 0 0 0 0
0 0 1 0 0 0
0 0 0 0 0 0
18 0 0 0 0 0
0 0 0 0 0 0
0 1 0 0 0 0
19 0 0 0 1 0
0 0 0 0 0 0
0 0 0 0 0 0
20 0 0 0 1 0
0 0 0 0 0 0
0 0 0 0 0 0
21 0 0 0 0 0
0 0 0 0 0 0
1 0 0 0 0 0
22 0 0 0 0 0
0 0 0 1 0 0
0 0 0 0 0 0
23 0 0 0 0 0
0 0 0 0 0 0
0 1 0 0 0 0
24 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 1
25 0 0 0 0 0
0 0 0 1 0 0
0 0 0 0 0 0
26 1 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
27 1 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
28 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 1
29 0 1 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
30 0 0 0 0 0
0 0 0 1 0 0
0 0 0 0 0 0
31 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 1
32 0 0 0 1 0
0 0 0 0 0 0
0 0 0 0 0 0
33 0 1 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
34 0 0 0 0 0
0 0 0 0 0 0
1 0 0 0 0 0
35 0 0 0 0 0
0 0 0 1 0 0
0 0 0 0 0 0

Results
A column was created for every car brand in the Car column.

REMOVE ADS

Predict CO2
We can use this additional information alongside the volume and weight to
predict CO2

To combine the information, we can use the concat() function from pandas.

First we will need to import a couple modules.

We will start with importing the Pandas.

import pandas

The pandas module allows us to read csv files and manipulate DataFrame
objects:

cars = pandas.read_csv("[Link]")
It also allows us to create the dummy variables:

ohe_cars = pandas.get_dummies(cars[['Car']])

Then we must select the independent variables (X) and add the dummy
variables columnwise.

Also store the dependent variable in y.

X = [Link]([cars[['Volume', 'Weight']], ohe_cars], axis=1)


y = cars['CO2']

We also need to import a method from sklearn to create a linear model

Learn about linear regression.

from sklearn import linear_model

Now we can fit the data to a linear regression:

regr = linear_model.LinearRegression()
[Link](X,y)

Finally we can predict the CO2 emissions based on the car's weight, volume,
and manufacturer.

##predict the CO2 emission of a VW where the weight is 2300kg,


and the volume is 1300cm3:
predictedCO2 =
[Link]([[2300, 1300,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,0]])

Example
import pandas
from sklearn import linear_model

cars = pandas.read_csv("[Link]")
ohe_cars = pandas.get_dummies(cars[['Car']])

X = [Link]([cars[['Volume', 'Weight']], ohe_cars], axis=1)


y = cars['CO2']

regr = linear_model.LinearRegression()
[Link](X,y)

##predict the CO2 emission of a VW where the weight is 2300kg, and


the volume is 1300cm3:
predictedCO2 =
[Link]([[2300, 1300,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,1,0]])

print(predictedCO2)

Result
[122.45153299]

We now have a coefficient for the volume, the weight, and each car brand in
the data set

REMOVE ADS

Dummifying
It is not necessary to create one column for each group in your category. The
information can be retained using 1 column less than the number of groups
you have.

For example, you have a column representing colors and in that column, you
have two colors, red and blue.

Example
import pandas as pd

colors = [Link]({'color': ['blue', 'red']})

print(colors)

Result
color
0 blue
1 red

You can create 1 column called red where 1 represents red and 0 represents
not red, which means it is blue.
To do this, we can use the same function that we used for one hot encoding,
get_dummies, and then drop one of the columns. There is an argument,
drop_first, which allows us to exclude the first column from the resulting
table.

Example
import pandas as pd

colors = [Link]({'color': ['blue', 'red']})


dummies = pd.get_dummies(colors, drop_first=True)

print(dummies)

Result
color_red
0 0
1 1

What if you have more than 2 groups? How can the multiple groups be
represented by 1 less column?

Let's say we have three colors this time, red, blue and green. When we
get_dummies while dropping the first column, we get the following table.

Example
import pandas as pd

colors = [Link]({'color': ['blue', 'red', 'green']})


dummies = pd.get_dummies(colors, drop_first=True)
dummies['color'] = colors['color']

print(dummies)

Result
color_green color_red color
0 0 0 blue
1 0 1 red
2 1 0 green
Machine Learning - K-means

K-means
K-means is an unsupervised learning method for clustering data points. The
algorithm iteratively divides data points into K clusters by minimizing the
variance in each cluster.

Here, we will show you how to estimate the best value for K using the elbow
method, then use K-means clustering to group the data points into clusters.

How does it work?


First, each data point is randomly assigned to one of the K clusters. Then, we
compute the centroid (functionally the center) of each cluster, and reassign
each data point to the cluster with the closest centroid. We repeat this
process until the cluster assignments for each data point are no longer
changing.

K-means clustering requires us to select K, the number of clusters we want


to group the data into. The elbow method lets us graph the inertia (a
distance-based metric) and visualize the point at which it starts decreasing
linearly. This point is referred to as the "elbow" and is a good estimate for
the best value for K based on our data.

ExampleGet your own Python Server


Start by visualizing some data points:

import [Link] as plt

x = [4, 5, 10, 4, 3, 11, 14 , 6, 10, 12]


y = [21, 19, 24, 17, 16, 25, 24, 22, 21, 21]

[Link](x, y)
[Link]()

Result
Now we utilize the elbow method to visualize the intertia for different values
of K:

Example
from [Link] import KMeans

data = list(zip(x, y))


inertias = []

for i in range(1,11):
kmeans = KMeans(n_clusters=i)
[Link](data)
[Link](kmeans.inertia_)

[Link](range(1,11), inertias, marker='o')


[Link]('Elbow method')
[Link]('Number of clusters')
[Link]('Inertia')
[Link]()

Result

The elbow method shows that 2 is a good value for K, so we retrain and
visualize the result:

Example
kmeans = KMeans(n_clusters=2)
[Link](data)

[Link](x, y, c=kmeans.labels_)
[Link]()

Result
REMOVE ADS

Example Explained
Import the modules you need.

import [Link] as plt


from [Link] import KMeans

You can learn about the Matplotlib module in our "Matplotlib Tutorial.

scikit-learn is a popular library for machine learning.


Create arrays that resemble two variables in a dataset. Note that while we
only use two variables here, this method will work with any number of
variables:

x = [4, 5, 10, 4, 3, 11, 14 , 6, 10, 12]


y = [21, 19, 24, 17, 16, 25, 24, 22, 21, 21]

Turn the data into a set of points:

data = list(zip(x, y))


print(data)

Result:

[(4, 21), (5, 19), (10, 24), (4, 17), (3, 16), (11, 25),
(14, 24), (6, 22), (10, 21), (12, 21)]

In order to find the best value for K, we need to run K-means across our data
for a range of possible values. We only have 10 data points, so the maximum
number of clusters is 10. So for each value K in range(1,11), we train a K-
means model and plot the intertia at that number of clusters:

inertias = []

for i in range(1,11):
kmeans = KMeans(n_clusters=i)
[Link](data)
[Link](kmeans.inertia_)

[Link](range(1,11), inertias, marker='o')


[Link]('Elbow method')
[Link]('Number of clusters')
[Link]('Inertia')
[Link]()

Result:
We can see that the "elbow" on the graph above (where the interia becomes
more linear) is at K=2. We can then fit our K-means algorithm one more time
and plot the different clusters assigned to the data:

kmeans = KMeans(n_clusters=2)
[Link](data)

[Link](x, y, c=kmeans.labels_)
[Link]()

Result:
Machine Learning - Bootstrap
Aggregation (Bagging)

Bagging
Methods such as Decision Trees, can be prone to overfitting on the training
set which can lead to wrong predictions on new data.

Bootstrap Aggregation (bagging) is a ensembling method that attempts to


resolve overfitting for classification or regression problems. Bagging aims to
improve the accuracy and performance of machine learning algorithms. It
does this by taking random subsets of an original dataset, with replacement,
and fits either a classifier (for classification) or regressor (for regression) to
each subset. The predictions for each subset are then aggregated through
majority vote for classification or averaging for regression, increasing
prediction accuracy.

Evaluating a Base Classifier


To see how bagging can improve model performance, we must start by
evaluating how the base classifier performs on the dataset. If you do not
know what decision trees are review the lesson on decision trees before
moving forward, as bagging is a continuation of the concept.

We will be looking to identify different classes of wines found in Sklearn's


wine dataset.

Let's start by importing the necessary modules.


from sklearn import datasets
from sklearn.model_selection import train_test_split
from [Link] import accuracy_score
from [Link] import DecisionTreeClassifier

Next we need to load in the data and store it into X (input features) and y
(target). The parameter as_frame is set equal to True so we do not lose the
feature names when loading the data. (sklearn version older than 0.23 must
skip the as_frame argument as it is not supported)

data = datasets.load_wine(as_frame = True)

X = [Link]
y = [Link]

In order to properly evaluate our model on unseen data, we need to split X


and y into train and test sets. For information on splitting data, see the
Train/Test lesson.

X_train, X_test, y_train, y_test = train_test_split(X, y,


test_size = 0.25, random_state = 22)

With our data prepared, we can now instantiate a base classifier and fit it to
the training data.

dtree = DecisionTreeClassifier(random_state = 22)


[Link](X_train,y_train)
Result:

DecisionTreeClassifier(random_state=22)

We can now predict the class of wine the unseen test set and evaluate the
model performance.

y_pred = [Link](X_test)

print("Train data accuracy:",accuracy_score(y_true = y_train,


y_pred = [Link](X_train)))
print("Test data accuracy:",accuracy_score(y_true = y_test,
y_pred = y_pred))

Result:

Train data accuracy: 1.0


Test data accuracy: 0.8222222222222222

ExampleGet your own Python Server


Import the necessary data and evaluate base classifier performance.

from sklearn import datasets


from sklearn.model_selection import train_test_split
from [Link] import accuracy_score
from [Link] import DecisionTreeClassifier

data = datasets.load_wine(as_frame = True)

X = [Link]
y = [Link]

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size


= 0.25, random_state = 22)

dtree = DecisionTreeClassifier(random_state = 22)


[Link](X_train,y_train)

y_pred = [Link](X_test)

print("Train data accuracy:",accuracy_score(y_true = y_train, y_pred


= [Link](X_train)))
print("Test data accuracy:",accuracy_score(y_true = y_test, y_pred =
y_pred))
The base classifier performs reasonably well on the dataset achieving 82%
accuracy on the test dataset with the current parameters (Different results
may occur if you do not have the random_state parameter set).

Now that we have a baseline accuracy for the test dataset, we can see how
the Bagging Classifier out performs a single Decision Tree Classifier.

REMOVE ADS

Creating a Bagging Classifier


For bagging we need to set the parameter n_estimators, this is the number
of base classifiers that our model is going to aggregate together.

For this sample dataset the number of estimators is relatively low, it is often
the case that much larger ranges are explored. Hyperparameter tuning is
usually done with a grid search, but for now we will use a select set of values
for the number of estimators.

We start by importing the necessary model.

from [Link] import BaggingClassifier

Now lets create a range of values that represent the number of estimators
we want to use in each ensemble.

estimator_range = [2,4,6,8,10,12,14,16]

To see how the Bagging Classifier performs with differing values of


n_estimators we need a way to iterate over the range of values and store the
results from each ensemble. To do this we will create a for loop, storing the
models and scores in separate lists for later visualizations.

Note: The default parameter for the base classifier in BaggingClassifier is


the DecisionTreeClassifier therefore we do not need to set it when
instantiating the bagging model.

models = []
scores = []
for n_estimators in estimator_range:

# Create bagging classifier


clf = BaggingClassifier(n_estimators = n_estimators,
random_state = 22)

# Fit the model


[Link](X_train, y_train)

# Append the model and score to their respective list


[Link](clf)
[Link](accuracy_score(y_true = y_test, y_pred =
[Link](X_test)))

With the models and scores stored, we can now visualize the improvement in
model performance.

import [Link] as plt

# Generate the plot of scores against number of estimators


[Link](figsize=(9,6))
[Link](estimator_range, scores)

# Adjust labels and font (to make visable)


[Link]("n_estimators", fontsize = 18)
[Link]("score", fontsize = 18)
plt.tick_params(labelsize = 16)

# Visualize plot
[Link]()
Example
Import the necessary data and evaluate
the BaggingClassifier performance.

import [Link] as plt


from sklearn import datasets
from sklearn.model_selection import train_test_split
from [Link] import accuracy_score
from [Link] import BaggingClassifier

data = datasets.load_wine(as_frame = True)


X = [Link]
y = [Link]

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size


= 0.25, random_state = 22)

estimator_range = [2,4,6,8,10,12,14,16]

models = []
scores = []

for n_estimators in estimator_range:

# Create bagging classifier


clf = BaggingClassifier(n_estimators = n_estimators, random_state
= 22)

# Fit the model


[Link](X_train, y_train)

# Append the model and score to their respective list


[Link](clf)
[Link](accuracy_score(y_true = y_test, y_pred =
[Link](X_test)))

# Generate the plot of scores against number of estimators


[Link](figsize=(9,6))
[Link](estimator_range, scores)

# Adjust labels and font (to make visable)


[Link]("n_estimators", fontsize = 18)
[Link]("score", fontsize = 18)
plt.tick_params(labelsize = 16)

# Visualize plot
[Link]()

Result
Results Explained
By iterating through different values for the number of estimators we can
see an increase in model performance from 82.2% to 95.5%. After 14
estimators the accuracy begins to drop, again if you set a
different random_state the values you see will vary. That is why it is best
practice to use cross validation to ensure stable results.
In this case, we see a 13.3% increase in accuracy when it comes to
identifying the type of the wine.

REMOVE ADS

Another Form of Evaluation


As bootstrapping chooses random subsets of observations to create
classifiers, there are observations that are left out in the selection process.
These "out-of-bag" observations can then be used to evaluate the model,
similarly to that of a test set. Keep in mind, that out-of-bag estimation can
overestimate error in binary classification problems and should only be used
as a compliment to other metrics.

We saw in the last exercise that 12 estimators yielded the highest accuracy,
so we will use that to create our model. This time setting the
parameter oob_score to true to evaluate the model with out-of-bag score.

Example
Create a model with out-of-bag metric.

from sklearn import datasets


from sklearn.model_selection import train_test_split
from [Link] import BaggingClassifier

data = datasets.load_wine(as_frame = True)

X = [Link]
y = [Link]

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size


= 0.25, random_state = 22)

oob_model = BaggingClassifier(n_estimators = 12, oob_score


= True,random_state = 22)

oob_model.fit(X_train, y_train)

print(oob_model.oob_score_)

Since the samples used in OOB and the test set are different, and the
dataset is relatively small, there is a difference in the accuracy. It is rare that
they would be exactly the same, again OOB should be used quick means for
estimating error, but is not the only evaluation metric.

Generating Decision Trees from


Bagging Classifier
As was seen in the Decision Tree lesson, it is possible to graph the decision
tree the model created. It is also possible to see the individual decision trees
that went into the aggregated classifier. This helps us to gain a more
intuitive understanding on how the bagging model arrives at its predictions.

Note: This is only functional with smaller datasets, where the trees are
relatively shallow and narrow making them easy to visualize.

We will need to import plot_tree function from [Link]. The different


trees can be graphed by changing the estimator you wish to visualize.

Example
Generate Decision Trees from Bagging Classifier

from sklearn import datasets


from sklearn.model_selection import train_test_split
from [Link] import BaggingClassifier
from [Link] import plot_tree

X = [Link]
y = [Link]

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size


= 0.25, random_state = 22)

clf = BaggingClassifier(n_estimators = 12, oob_score


= True,random_state = 22)

[Link](X_train, y_train)

[Link](figsize=(30, 20))

plot_tree(clf.estimators_[0], feature_names = [Link])

Result
Here we can see just the first decision tree that was used to vote on the final
prediction. Again, by changing the index of the classifier you can see each of
the trees that have been aggregated.

Machine Learning - Cross


Validation

Cross Validation
When adjusting models we are aiming to increase overall model performance
on unseen data. Hyperparameter tuning can lead to much better
performance on test sets. However, optimizing parameters to the test set
can lead information leakage causing the model to preform worse on unseen
data. To correct for this we can perform cross validation.

To better understand CV, we will be performing different methods on the iris


dataset. Let us first load in and separate the data.

from sklearn import datasets

X, y = datasets.load_iris(return_X_y=True)

There are many methods to cross validation, we will start by looking at k-fold
cross validation.

K-Fold
The training data used in the model is split, into k number of smaller sets, to
be used to validate the model. The model is then trained on k-1 folds of
training set. The remaining fold is then used as a validation set to evaluate
the model.

As we will be trying to classify different species of iris flowers we will need to


import a classifier model, for this exercise we will be using
a DecisionTreeClassifier. We will also need to import CV modules
from sklearn.
from [Link] import DecisionTreeClassifier
from sklearn.model_selection import KFold, cross_val_score

With the data loaded we can now create and fit a model for evaluation.

clf = DecisionTreeClassifier(random_state=42)

Now let's evaluate our model and see how it performs on each k-fold.

k_folds = KFold(n_splits = 5)

scores = cross_val_score(clf, X, y, cv = k_folds)

It is also good pratice to see how CV performed overall by averaging the


scores for all folds.

ExampleGet your own Python Server


Run k-fold CV:

from sklearn import datasets


from [Link] import DecisionTreeClassifier
from sklearn.model_selection import KFold, cross_val_score

X, y = datasets.load_iris(return_X_y=True)

clf = DecisionTreeClassifier(random_state=42)

k_folds = KFold(n_splits = 5)

scores = cross_val_score(clf, X, y, cv = k_folds)

print("Cross Validation Scores: ", scores)


print("Average CV Score: ", [Link]())
print("Number of CV Scores used in Average: ", len(scores))

REMOVE ADS

Stratified K-Fold
In cases where classes are imbalanced we need a way to account for the
imbalance in both the train and validation sets. To do so we can stratify the
target classes, meaning that both sets will have an equal proportion of all
classes.

Example
from sklearn import datasets
from [Link] import DecisionTreeClassifier
from sklearn.model_selection import StratifiedKFold, cross_val_score

X, y = datasets.load_iris(return_X_y=True)

clf = DecisionTreeClassifier(random_state=42)

sk_folds = StratifiedKFold(n_splits = 5)

scores = cross_val_score(clf, X, y, cv = sk_folds)

print("Cross Validation Scores: ", scores)


print("Average CV Score: ", [Link]())
print("Number of CV Scores used in Average: ", len(scores))

While the number of folds is the same, the average CV increases from the
basic k-fold when making sure there is stratified classes.

Leave-One-Out (LOO)
Instead of selecting the number of splits in the training data set like k-fold
LeaveOneOut, utilize 1 observation to validate and n-1 observations to train.
This method is an exaustive technique.

Example
Run LOO CV:

from sklearn import datasets


from [Link] import DecisionTreeClassifier
from sklearn.model_selection import LeaveOneOut, cross_val_score

X, y = datasets.load_iris(return_X_y=True)

clf = DecisionTreeClassifier(random_state=42)
loo = LeaveOneOut()

scores = cross_val_score(clf, X, y, cv = loo)

print("Cross Validation Scores: ", scores)


print("Average CV Score: ", [Link]())
print("Number of CV Scores used in Average: ", len(scores))

We can observe that the number of cross validation scores performed is


equal to the number of observations in the dataset. In this case there are
150 observations in the iris dataset.

The average CV score is 94%.

REMOVE ADS

Leave-P-Out (LPO)
Leave-P-Out is simply a nuanced diffence to the Leave-One-Out idea, in that
we can select the number of p to use in our validation set.

Example
Run LPO CV:

from sklearn import datasets


from [Link] import DecisionTreeClassifier
from sklearn.model_selection import LeavePOut, cross_val_score

X, y = datasets.load_iris(return_X_y=True)

clf = DecisionTreeClassifier(random_state=42)

lpo = LeavePOut(p=2)

scores = cross_val_score(clf, X, y, cv = lpo)

print("Cross Validation Scores: ", scores)


print("Average CV Score: ", [Link]())
print("Number of CV Scores used in Average: ", len(scores))
As we can see this is an exhaustive method we many more scores being
calculated than Leave-One-Out, even with a p = 2, yet it achieves roughly
the same average CV score.

Shuffle Split
Unlike KFold, ShuffleSplit leaves out a percentage of the data, not to be
used in the train or validation sets. To do so we must decide what the train
and test sizes are, as well as the number of splits.

Example
Run Shuffle Split CV:

from sklearn import datasets


from [Link] import DecisionTreeClassifier
from sklearn.model_selection import ShuffleSplit, cross_val_score

X, y = datasets.load_iris(return_X_y=True)

clf = DecisionTreeClassifier(random_state=42)

ss = ShuffleSplit(train_size=0.6, test_size=0.3, n_splits = 5)

scores = cross_val_score(clf, X, y, cv = ss)

print("Cross Validation Scores: ", scores)


print("Average CV Score: ", [Link]())
print("Number of CV Scores used in Average: ", len(scores))

Ending Notes
These are just a few of the CV methods that can be applied to models. There
are many more cross validation classes, with most models having their own
class. Check out sklearns cross validation for more CV options.
Machine Learning - AUC -
ROC Curve

AUC - ROC Curve


In classification, there are many different evaluation metrics. The most
popular is accuracy, which measures how often the model is correct. This is
a great metric because it is easy to understand and getting the most correct
guesses is often desired. There are some cases where you might consider
using another evaluation metric.

Another common metric is AUC, area under the receiver operating


characteristic (ROC) curve. The Reciever operating characteristic curve plots
the true positive (TP) rate versus the false positive (FP) rate at different
classification thresholds. The thresholds are different probability cutoffs that
separate the two classes in binary classification. It uses probability to tell us
how well a model separates the classes.

Imbalanced Data
Suppose we have an imbalanced data set where the majority of our data is of
one value. We can obtain high accuracy for the model by predicting the
majority class.

ExampleGet your own Python Server


import numpy as np
from [Link] import accuracy_score, confusion_matrix,
roc_auc_score, roc_curve

n = 10000
ratio = .95
n_0 = int((1-ratio) * n)
n_1 = int(ratio * n)

y = [Link]([0] * n_0 + [1] * n_1)


# below are the probabilities obtained from a hypothetical model that
always predicts the majority class
# probability of predicting class 1 is going to be 100%
y_proba = [Link]([1]*n)
y_pred = y_proba > .5

print(f'accuracy score: {accuracy_score(y, y_pred)}')


cf_mat = confusion_matrix(y, y_pred)
print('Confusion matrix')
print(cf_mat)
print(f'class 0 accuracy: {cf_mat[0][0]/n_0}')
print(f'class 1 accuracy: {cf_mat[1][1]/n_1}')

Although we obtain a very high accuracy, the model provided no information


about the data so it's not useful. We accurately predict class 1 100% of the
time while inaccurately predict class 0 0% of the time. At the expense of
accuracy, it might be better to have a model that can somewhat separate
the two classes.

Example
# below are the probabilities obtained from a hypothetical model that
doesn't always predict the mode
y_proba_2 = [Link](
[Link](0, .7, n_0).tolist() +
[Link](.3, 1, n_1).tolist()
)
y_pred_2 = y_proba_2 > .5

print(f'accuracy score: {accuracy_score(y, y_pred_2)}')


cf_mat = confusion_matrix(y, y_pred_2)
print('Confusion matrix')
print(cf_mat)
print(f'class 0 accuracy: {cf_mat[0][0]/n_0}')
print(f'class 1 accuracy: {cf_mat[1][1]/n_1}')

For the second set of predictions, we do not have as high of an accuracy


score as the first but the accuracy for each class is more balanced. Using
accuracy as an evaluation metric we would rate the first model higher than
the second even though it doesn't tell us anything about the data.

In cases like this, using another evaluation metric like AUC would be
preferred.

import [Link] as plt


def plot_roc_curve(true_y, y_prob):
"""
plots the roc curve based of the probabilities
"""

fpr, tpr, thresholds = roc_curve(true_y, y_prob)


[Link](fpr, tpr)
[Link]('False Positive Rate')
[Link]('True Positive Rate')

Example
Model 1:

plot_roc_curve(y, y_proba)
print(f'model 1 AUC score: {roc_auc_score(y, y_proba)}')

Result

model 1 AUC score: 0.5


Example
Model 2:

plot_roc_curve(y, y_proba_2)
print(f'model 2 AUC score: {roc_auc_score(y, y_proba_2)}')

Result

model 2 AUC score: 0.8270551578947367

An AUC score of around .5 would mean that the model is unable to make a
distinction between the two classes and the curve would look like a line with
a slope of 1. An AUC score closer to 1 means that the model has the ability to
separate the two classes and the curve would come closer to the top left
corner of the graph.

REMOVE ADS
Probabilities
Because AUC is a metric that utilizes probabilities of the class predictions, we
can be more confident in a model that has a higher AUC score than one with
a lower score even if they have similar accuracies.

In the data below, we have two sets of probabilites from hypothetical


models. The first has probabilities that are not as "confident" when
predicting the two classes (the probabilities are close to .5). The second has
probabilities that are more "confident" when predicting the two classes (the
probabilities are close to the extremes of 0 or 1).

Example
import numpy as np

n = 10000
y = [Link]([0] * n + [1] * n)
#
y_prob_1 = [Link](
[Link](.25, .5, n//2).tolist() +
[Link](.3, .7, n).tolist() +
[Link](.5, .75, n//2).tolist()
)
y_prob_2 = [Link](
[Link](0, .4, n//2).tolist() +
[Link](.3, .7, n).tolist() +
[Link](.6, 1, n//2).tolist()
)

print(f'model 1 accuracy score: {accuracy_score(y, y_prob_1>.5)}')


print(f'model 2 accuracy score: {accuracy_score(y, y_prob_2>.5)}')

print(f'model 1 AUC score: {roc_auc_score(y, y_prob_1)}')


print(f'model 2 AUC score: {roc_auc_score(y, y_prob_2)}')

Example
Plot model 1:

plot_roc_curve(y, y_prob_1)

Result
Example
Plot model 2:

fpr, tpr, thresholds = roc_curve(y, y_prob_2)


[Link](fpr, tpr)

Result
Even though the accuracies for the two models are similar, the model with
the higher AUC score will be more reliable because it takes into account the
predicted probability. It is more likely to give you higher accuracy when
predicting future data.

Machine Learning - K-nearest


neighbors (KNN)

KNN
KNN is a simple, supervised machine learning (ML) algorithm that can be
used for classification or regression tasks - and is also frequently used in
missing value imputation. It is based on the idea that the observations
closest to a given data point are the most "similar" observations in a data
set, and we can therefore classify unforeseen points based on the values of
the closest existing points. By choosing K, the user can select the number of
nearby observations to use in the algorithm.

Here, we will show you how to implement the KNN algorithm for
classification, and show how different values of K affect the results.

How does it work?


K is the number of nearest neighbors to use. For classification, a majority
vote is used to determined which class a new observation should fall into.
Larger values of K are often more robust to outliers and produce more stable
decision boundaries than very small values (K=3 would be better than K=1,
which might produce undesirable results.

ExampleGet your own Python Server


Start by visualizing some data points:

import [Link] as plt

x = [4, 5, 10, 4, 3, 11, 14 , 8, 10, 12]


y = [21, 19, 24, 17, 16, 25, 24, 22, 21, 21]
classes = [0, 0, 1, 0, 0, 1, 1, 0, 1, 1]

[Link](x, y, c=classes)
[Link]()

Result
Now we fit the KNN algorithm with K=1:

from [Link] import KNeighborsClassifier

data = list(zip(x, y))


knn = KNeighborsClassifier(n_neighbors=1)

[Link](data, classes)

And use it to classify a new data point:

Example
new_x = 8
new_y = 21
new_point = [(new_x, new_y)]
prediction = [Link](new_point)

[Link](x + [new_x], y + [new_y], c=classes + [prediction[0]])


[Link](x=new_x-1.7, y=new_y-0.7, s=f"new point, class:
{prediction[0]}")
[Link]()

Result

Now we do the same thing, but with a higher K value which changes the
prediction:

Example
knn = KNeighborsClassifier(n_neighbors=5)

[Link](data, classes)

prediction = [Link](new_point)
[Link](x + [new_x], y + [new_y], c=classes + [prediction[0]])
[Link](x=new_x-1.7, y=new_y-0.7, s=f"new point, class:
{prediction[0]}")
[Link]()

Result

REMOVE ADS

Example Explained
Import the modules you need.

You can learn about the Matplotlib module in our "Matplotlib Tutorial.
scikit-learn is a popular library for machine learning in Python.

import [Link] as plt


from [Link] import KNeighborsClassifier

Create arrays that resemble variables in a dataset. We have two input


features (x and y) and then a target class (class). The input features that are
pre-labeled with our target class will be used to predict the class of new
data. Note that while we only use two input features here, this method will
work with any number of variables:

x = [4, 5, 10, 4, 3, 11, 14 , 8, 10, 12]


y = [21, 19, 24, 17, 16, 25, 24, 22, 21, 21]
classes = [0, 0, 1, 0, 0, 1, 1, 0, 1, 1]

Turn the input features into a set of points:

data = list(zip(x, y))


print(data)

Result:
[(4, 21), (5, 19), (10, 24), (4, 17), (3, 16), (11, 25),
(14, 24), (8, 22), (10, 21), (12, 21)]

Using the input features and target class, we fit a KNN model on the model
using 1 nearest neighbor:

knn = KNeighborsClassifier(n_neighbors=1)
[Link](data, classes)

Then, we can use the same KNN object to predict the class of new,
unforeseen data points. First we create new x and y features, and then
call [Link]() on the new data point to get a class of 0 or 1:

new_x = 8
new_y = 21
new_point = [(new_x, new_y)]
prediction = [Link](new_point)
print(prediction)

Result:
[0]
When we plot all the data along with the new point and class, we can see it's
been labeled blue with the 1 class. The text annotation is just to highlight the
location of the new point:

[Link](x + [new_x], y + [new_y], c=classes +


[prediction[0]])
[Link](x=new_x-1.7, y=new_y-0.7, s=f"new point, class:
{prediction[0]}")
[Link]()

Result:

However, when we changes the number of neighbors to 5, the number of


points used to classify our new point changes. As a result, so does the
classification of the new point:

knn = KNeighborsClassifier(n_neighbors=5)
[Link](data, classes)
prediction = [Link](new_point)
print(prediction)

Result:
[1]

When we plot the class of the new point along with the older points, we note
that the color has changed based on the associated class label:

[Link](x + [new_x], y + [new_y], c=classes +


[prediction[0]])
[Link](x=new_x-1.7, y=new_y-0.7, s=f"new point, class:
{prediction[0]}")
[Link]()

Result:
DSA with Python
Data Structures is about how data can be stored in different structures.

Algorithms is about how to solve different problems, often by searching


through and manipulating data structures.

Understanding DSA helps you to find the best combination of Data


Structures and Algorithms to create more efficient code.

Data Structures
Data Structures are a way of storing and organizing data in a computer.

Python has built-in support for several data structures, such as lists,
dictionaries, and sets.

Other data structures can be implemented using Python classes and objects,
such as linked lists, stacks, queues, trees, and graphs.

In this tutorial we will concentrate on these Data Structures:

 Lists and Arrays


 Stacks
 Queues
 Linked Lists
 Hash Tables
 Trees
o Binary Trees
o Binary Search Trees
o AVL Trees
 Graphs

Algorithms
Algorithms are a way of working with data in a computer and solving
problems like sorting, searching, etc.
In this tutorial we will concentrate on these search and sort Algorithms:

 Linear Search
 Binary Search
 Bubble Sort
 Selection Sort
 Insertion Sort
 Quick Sort
 Counting Sort
 Radix Sort
 Merge Sort

Why Learn DSA with Python


 Python has a clean readable syntax
 DSA allows you to improve problem-solving skills
 DSA and Python helps you write more efficient code
 DSA gives you a better understanding of memory storage
 DSA helps you handle complex programming challenges
 Python is widely used in Data Science and Machine Learning

Python Lists and Arrays

In Python, lists are the built-in data structure that serves as a dynamic
array.

Lists are ordered, mutable, and can contain elements of different


types.

Lists
A list is a built-in data structure in Python, used to store multiple elements.
Lists are used by many algorithms.

Creating Lists
Lists are created using square brackets []:

ExampleGet your own Python Server


# Empty list
x = []

# List with initial values


y = [1, 2, 3, 4, 5]

# List with mixed types


z = [1, "hello", 3.14, True]

List Methods
Python lists come with several built-in algorithms (called methods), to
perform common operations like appending, sorting, and more.

Example
Append one element to the list, and sort the list ascending:

x = [9, 12, 7, 4, 11]

# Add element:
[Link](8)

# Sort list ascending:


[Link]()

Create Algorithms
Sometimes we want to perform actions that are not built into Python.
Then we can create our own algorithms.

For example, an algorithm can be used to find the lowest value in a list, like
in the example below:

Example
Create an algorithm to find the lowest value in a list:

my_array = [7, 12, 9, 4, 11, 8]


minVal = my_array[0]

for i in my_array:
if i < minVal:
minVal = i

print('Lowest value:', minVal)

The algorithm above is very simple, and fast enough for small data sets, but
if the data is big enough, any algorithm will take time to run.

This is where optimization comes in.

Optimization is an important part of algorithm development, and of course,


an important part of DSA programming.

REMOVE ADS

Time Complexity
When exploring algorithms, we often look at how much time an algorithm
takes to run relative to the size of the data set.

In the example above, the time the algorithm needs to run is proportional, or
linear, to the size of the data set. This is because the algorithm must visit
every array element one time to find the lowest value. The loop must run 5
times since there are 5 values in the array. And if the array had 1000 values,
the loop would have to run 1000 times.

Try the simulation below to see this relationship between the number of
compare operations needed to find the lowest value, and the size of the
array.

See this page for a more thorough explanation of what time complexity is.

Each algorithm in this tutorial will be presented together with its time
complexity.

Stacks with Python

A stack is a linear data structure that follows the Last-In-First-Out


(LIFO) principle.

Think of it like a stack of pancakes - you can only add or remove


pancakes from the top.
Stacks
A stack is a data structure that can hold many elements, and the last
element added is the first one to be removed.

Like a pile of pancakes, the pancakes are both added and removed from the
top. So when removing a pancake, it will always be the last pancake you
added. This way of organizing elements is called LIFO: Last In First Out.

Basic operations we can do on a stack are:

 Push: Adds a new element on the stack.


 Pop: Removes and returns the top element from the stack.
 Peek: Returns the top (last) element on the stack.
 isEmpty: Checks if the stack is empty.
 Size: Finds the number of elements in the stack.

Stacks can be implemented by using arrays or linked lists.

Stacks can be used to implement undo mechanisms, to revert to previous


states, to create algorithms for depth-first search in graphs, or for
backtracking.

Stacks are often mentioned together with Queues, which is a similar data
structure described on the next page.

Stack Implementation using Python


Lists
For Python lists (and arrays), a stack can look and behave like this:

x = [5, 6, 2, 9, 3, 8, 4, 2]
Add: Remove:

Since Python lists has good support for functionality needed to implement
stacks, we start with creating a stack and do stack operations with just a few
lines like this:

ExampleGet your own Python Server


Using a Python list as a stack:

stack = []

# Push
[Link]('A')
[Link]('B')
[Link]('C')
print("Stack: ", stack)

# Peek
topElement = stack[-1]
print("Peek: ", topElement)

# Pop
poppedElement = [Link]()
print("Pop: ", poppedElement)

# Stack after Pop


print("Stack after Pop: ", stack)

# isEmpty
isEmpty = not bool(stack)
print("isEmpty: ", isEmpty)

# Size
print("Size: ",len(stack))

While Python lists can be used as stacks, creating a dedicated Stack


class provides better encapsulation and additional functionality:

Example
Creating a stack using class:

class Stack:
def __init__(self):
[Link] = []

def push(self, element):


[Link](element)

def pop(self):
if [Link]():
return "Stack is empty"
return [Link]()
def peek(self):
if [Link]():
return "Stack is empty"
return [Link][-1]

def isEmpty(self):
return len([Link]) == 0

def size(self):
return len([Link])

# Create a stack
myStack = Stack()

[Link]('A')
[Link]('B')
[Link]('C')

print("Stack: ", [Link])


print("Pop: ", [Link]())
print("Stack after Pop: ", [Link])
print("Peek: ", [Link]())
print("isEmpty: ", [Link]())
print("Size: ", [Link]())

Reasons to implement stacks using lists/arrays:

 Memory Efficient: Array elements do not hold the next elements


address like linked list nodes do.
 Easier to implement and understand: Using arrays to implement
stacks require less code than using linked lists, and for this reason it is
typically easier to understand as well.

A reason for not using arrays to implement stacks:

 Fixed size: An array occupies a fixed part of the memory. This means
that it could take up more memory than needed, or if the array fills up,
it cannot hold more elements.

REMOVE ADS
Stack Implementation using Linked
Lists
A linked list consists of nodes with some sort of data, and a pointer to the
next node.

A big benefit with using linked lists is that nodes are stored wherever there is
free space in memory, the nodes do not have to be stored contiguously right
after each other like elements are stored in arrays. Another nice thing with
linked lists is that when adding or removing nodes, the rest of the nodes in
the list do not have to be shifted.

To better understand the benefits with using arrays or linked lists to


implement stacks, you should check out this page that explains how arrays
and linked lists are stored in memory.

This is how a stack can be implemented using a linked list.

Example
Creating a Stack using a Linked List:

class Node:
def __init__(self, value):
[Link] = value
[Link] = None

class Stack:
def __init__(self):
[Link] = None
[Link] = 0

def push(self, value):


new_node = Node(value)
if [Link]:
new_node.next = [Link]
[Link] = new_node
[Link] += 1

def pop(self):
if [Link]():
return "Stack is empty"
popped_node = [Link]
[Link] = [Link]
[Link] -= 1
return popped_node.value

def peek(self):
if [Link]():
return "Stack is empty"
return [Link]

def isEmpty(self):
return [Link] == 0

def stackSize(self):
return [Link]

def traverseAndPrint(self):
currentNode = [Link]
while currentNode:
print([Link], end=" -> ")
currentNode = [Link]
print()

myStack = Stack()
[Link]('A')
[Link]('B')
[Link]('C')

print("LinkedList: ", end="")


[Link]()
print("Peek: ", [Link]())
print("Pop: ", [Link]())
print("LinkedList after Pop: ", end="")
[Link]()
print("isEmpty: ", [Link]())
print("Size: ", [Link]())

A reason for using linked lists to implement stacks:

 Dynamic size: The stack can grow and shrink dynamically, unlike with
arrays.

Reasons for not using linked lists to implement stacks:

 Extra memory: Each stack element must contain the address to the
next element (the next linked list node).
 Readability: The code might be harder to read and write for some
because it is longer and more complex.

Common Stack Applications


Stacks are used in many real-world scenarios:

 Undo/Redo operations in text editors


 Browser history (back/forward)
 Function call stack in programming
 Expression evaluation

Queues with Python

A queue is a linear data structure that follows the First-In-First-Out


(FIFO) principle.

Queues
Think of a queue as people standing in line in a supermarket.

The first person to stand in line is also the first who can pay and leave the
supermarket.

Basic operations we can do on a queue are:

 Enqueue: Adds a new element to the queue.


 Dequeue: Removes and returns the first (front) element from the
queue.
 Peek: Returns the first element in the queue.
 isEmpty: Checks if the queue is empty.
 Size: Finds the number of elements in the queue.

Queues can be implemented by using arrays or linked lists.


Queues can be used to implement job scheduling for an office printer, order
processing for e-tickets, or to create algorithms for breadth-first search in
graphs.

Queues are often mentioned together with Stacks, which is a similar data
structure described on the previous page.

Queue Implementation using Python


Lists
For Python lists (and arrays), a Queue can look and behave like this:

x = [5, 6, 2, 9, 3, 8, 4, 2]
Add: Remove:

Since Python lists has good support for functionality needed to implement
queues, we start with creating a queue and do queue operations with just a
few lines:

ExampleGet your own Python Server


Using a Python list as a queue:

queue = []

# Enqueue
[Link]('A')
[Link]('B')
[Link]('C')
print("Queue: ", queue)

# Peek
frontElement = queue[0]
print("Peek: ", frontElement)

# Dequeue
poppedElement = [Link](0)
print("Dequeue: ", poppedElement)

print("Queue after Dequeue: ", queue)


# isEmpty
isEmpty = not bool(queue)
print("isEmpty: ", isEmpty)

# Size
print("Size: ", len(queue))

Note: While using a list is simple, removing elements from the beginning
(dequeue operation) requires shifting all remaining elements, making it less
efficient for large queues.

REMOVE ADS

Implementing a Queue Class


Here's a complete implementation of a Queue class:

Example
Using a Python class as a queue:

class Queue:
def __init__(self):
[Link] = []

def enqueue(self, element):


[Link](element)

def dequeue(self):
if [Link]():
return "Queue is empty"
return [Link](0)

def peek(self):
if [Link]():
return "Queue is empty"
return [Link][0]

def isEmpty(self):
return len([Link]) == 0
def size(self):
return len([Link])

# Create a queue
myQueue = Queue()

[Link]('A')
[Link]('B')
[Link]('C')

print("Queue: ", [Link])


print("Peek: ", [Link]())
print("Dequeue: ", [Link]())
print("Queue after Dequeue: ", [Link])
print("isEmpty: ", [Link]())
print("Size: ", [Link]())

Queue Implementation using Linked


Lists
A linked list consists of nodes with some sort of data, and a pointer to the
next node.

A big benefit with using linked lists is that nodes are stored wherever there is
free space in memory, the nodes do not have to be stored contiguously right
after each other like elements are stored in arrays. Another nice thing with
linked lists is that when adding or removing nodes, the rest of the nodes in
the list do not have to be shifted.

To better understand the benefits with using arrays or linked lists to


implement queues, you should check out this page that explains how arrays
and linked lists are stored in memory.

This is how a queue can be implemented using a linked list.

Example
Creating a Queue using a Linked List:
class Node:
def __init__(self, data):
[Link] = data
[Link] = None

class Queue:
def __init__(self):
[Link] = None
[Link] = None
[Link] = 0

def enqueue(self, element):


new_node = Node(element)
if [Link] is None:
[Link] = [Link] = new_node
[Link] += 1
return
[Link] = new_node
[Link] = new_node
[Link] += 1

def dequeue(self):
if [Link]():
return "Queue is empty"
temp = [Link]
[Link] = [Link]
[Link] -= 1
if [Link] is None:
[Link] = None
return [Link]

def peek(self):
if [Link]():
return "Queue is empty"
return [Link]

def isEmpty(self):
return [Link] == 0

def size(self):
return [Link]

def printQueue(self):
temp = [Link]
while temp:
print([Link], end=" -> ")
temp = [Link]
print()
# Create a queue
myQueue = Queue()

[Link]('A')
[Link]('B')
[Link]('C')

print("Queue: ", end="")


[Link]()
print("Peek: ", [Link]())
print("Dequeue: ", [Link]())
print("Queue after Dequeue: ", end="")
[Link]()
print("isEmpty: ", [Link]())
print("Size: ", [Link]())

Reasons for using linked lists to implement queues:

 Dynamic size: The queue can grow and shrink dynamically, unlike
with arrays.
 No shifting: The front element of the queue can be removed
(enqueue) without having to shift other elements in the memory.

Reasons for not using linked lists to implement queues:

 Extra memory: Each queue element must contain the address to the
next element (the next linked list node).
 Readability: The code might be harder to read and write for some
because it is longer and more complex.
REMOVE ADS

Common Queue Applications


Queues are used in many real-world scenarios:

 Task scheduling in operating systems


 Breadth-first search in graphs
 Message queues in distributed systems

Linked Lists with Python


A Linked List is, as the word implies, a list where the nodes are linked
together. Each node contains data and a pointer. The way they are linked
together is that each node points to where in the memory the next node is
placed.

Linked Lists
A linked list consists of nodes with some sort of data, and a pointer, or link,
to the next node.

Linked Lists vs Arrays


The easiest way to understand linked lists is perhaps by comparing linked
lists with arrays.

Linked lists consist of nodes, and is a linear data structure we make


ourselves, unlike arrays which is an existing data structure in the
programming language that we can use.

Nodes in a linked list store links to other nodes, but array elements do not
need to store links to other elements.

Note: How linked lists and arrays are stored in memory is explained in detail
on the page Linked Lists in Memory.

The table below compares linked lists with arrays to give a better
understanding of what linked lists are.

Arrays

An existing data structure in the programming language Yes


Fixed size in memory Yes

Elements, or nodes, are stored right after each other in memory Yes
(contiguously)

Memory usage is low Yes


(each node only contains data, no links to other nodes)

Elements, or nodes, can be accessed directly (random access) Yes

Elements, or nodes, can be inserted or deleted in constant time, no No


shifting operations in memory needed.

These are some key linked list properties, compared to arrays:

 Linked lists are not allocated to a fixed size in memory like arrays are,
so linked lists do not require to move the whole list into a larger
memory space when the fixed memory space fills up, like arrays must.
 Linked list nodes are not laid out one right after the other in memory
(contiguously), so linked list nodes do not have to be shifted up or
down in memory when nodes are inserted or deleted.
 Linked list nodes require more memory to store one or more links to
other nodes. Array elements do not require that much memory,
because array elements do not contain links to other elements.
 Linked list operations are usually harder to program and require more
lines than similar array operations, because programming languages
have better built in support for arrays.
 We must traverse a linked list to find a node at a specific position, but
with arrays we can access an element directly by writing myArray[5].

REMOVE ADS
Types of Linked Lists
There are three basic forms of linked lists:

1. Singly linked lists


2. Doubly linked lists
3. Circular linked lists

A singly linked list is the simplest kind of linked lists. It takes up less space
in memory because each node has only one address to the next node, like in
the image below.

A doubly linked list has nodes with addresses to both the previous and the
next node, like in the image below, and therefore takes up more memory.
But doubly linked lists are good if you want to be able to move both up and
down in the list.

A circular linked list is like a singly or doubly linked list with the first node,
the "head", and the last node, the "tail", connected.

In singly or doubly linked lists, we can find the start and end of a list by just
checking if the links are null. But for circular linked lists, more complex code
is needed to explicitly check for start and end nodes in certain applications.

Circular linked lists are good for lists you need to cycle through continuously.

The image below is an example of a singly circular linked list:


The image below is an example of a doubly circular linked list:

Note: What kind of linked list you need depends on the problem you are
trying to solve.

Linked List Operations


Basic things we can do with linked lists are:

1. Traversal
2. Remove a node
3. Insert a node
4. Sort

For simplicity, singly linked lists will be used to explain these operations
below.

REMOVE ADS

Traversal of a Linked List


Traversing a linked list means to go through the linked list by following the
links from one node to the next.

Traversal of linked lists is typically done to search for a specific node, and
read or modify the node's content, remove the node, or insert a node right
before or after that node.

To traverse a singly linked list, we start with the first node in the list, the
head node, and follow that node's next link, and the next node's next link
and so on, until the next address is null.

The code below prints out the node values as it traverses along the linked
list, in the same way as the animation above.

ExampleGet your own Python Server


Traversal of a singly linked list in Python:

class Node:
def __init__(self, data):
[Link] = data
[Link] = None

def traverseAndPrint(head):
currentNode = head
while currentNode:
print([Link], end=" -> ")
currentNode = [Link]
print("null")

node1 = Node(7)
node2 = Node(11)
node3 = Node(3)
node4 = Node(2)
node5 = Node(9)

[Link] = node2
[Link] = node3
[Link] = node4
[Link] = node5

traverseAndPrint(node1)

Find The Lowest Value in a Linked List


Let's find the lowest value in a singly linked list by traversing it and checking
each value.

Finding the lowest value in a linked list is very similar to how we found the
lowest value in an array, except that we need to follow the next link to get to
the next node.

To find the lowest value we need to traverse the list like in the previous
code. But in addition to traversing the list, we must also update the current
lowest value when we find a node with a lower value.

In the code below, the algorithm to find the lowest value is moved into a
function called findLowestValue.
Example
Finding the lowest value in a singly linked list in Python:

class Node:
def __init__(self, data):
[Link] = data
[Link] = None

def findLowestValue(head):
minValue = [Link]
currentNode = [Link]
while currentNode:
if [Link] < minValue:
minValue = [Link]
currentNode = [Link]
return minValue

node1 = Node(7)
node2 = Node(11)
node3 = Node(3)
node4 = Node(2)
node5 = Node(9)

[Link] = node2
[Link] = node3
[Link] = node4
[Link] = node5

print("The lowest value in the linked list is:",


findLowestValue(node1))

Delete a Node in a Linked List


If you want to delete a node in a linked list, it is important to connect the
nodes on each side of the node before deleting it, so that the linked list is not
broken.

So before deleting the node, we need to get the next pointer from the
previous node, and connect the previous node to the new next node before
deleting the node in between.
Also, it is a good idea to first connect next pointer to the node after the node
we want to delete, before we delete it. This is to avoid a 'dangling' pointer, a
pointer that points to nothing, even if it is just for a brief moment.

The simulation below shows the node we want to delete, and how the list
must be traversed first to connect the list properly before deleting the node
without breaking the linked list.

Head7next11next3next2next9nextnull
Delete

In the code below, the algorithm to delete a node is moved into a function
called deleteSpecificNode.

Example
Deleting a specific node in a singly linked list in Python:

class Node:
def __init__(self, data):
[Link] = data
[Link] = None

def traverseAndPrint(head):
currentNode = head
while currentNode:
print([Link], end=" -> ")
currentNode = [Link]
print("null")

def deleteSpecificNode(head, nodeToDelete):


if head == nodeToDelete:
return [Link]

currentNode = head
while [Link] and [Link] != nodeToDelete:
currentNode = [Link]

if [Link] is None:
return head

[Link] = [Link]

return head

node1 = Node(7)
node2 = Node(11)
node3 = Node(3)
node4 = Node(2)
node5 = Node(9)

[Link] = node2
[Link] = node3
[Link] = node4
[Link] = node5

print("Before deletion:")
traverseAndPrint(node1)

# Delete node4
node1 = deleteSpecificNode(node1, node4)

print("\nAfter deletion:")
traverseAndPrint(node1)

In the deleteSpecificNode function above, the return value is the new head
of the linked list. So for example, if the node to be deleted is the first node,
the new head returned will be the next node.

Insert a Node in a Linked List


Inserting a node into a linked list is very similar to deleting a node, because
in both cases we need to take care of the next pointers to make sure we do
not break the linked list.

To insert a node in a linked list we first need to create the node, and then at
the position where we insert it, we need to adjust the pointers so that the
previous node points to the new node, and the new node points to the
correct next node.

The simulation below shows how the links are adjusted when inserting a new
node.

Head7next97next3next2next9nextnullInsert
1. New node is created
2. Node 1 is linked to new node
3. New node is linked to next node
Example
Inserting a node in a singly linked list in Python:

class Node:
def __init__(self, data):
[Link] = data
[Link] = None

def traverseAndPrint(head):
currentNode = head
while currentNode:
print([Link], end=" -> ")
currentNode = [Link]
print("null")

def insertNodeAtPosition(head, newNode, position):


if position == 1:
[Link] = head
return newNode

currentNode = head
for _ in range(position - 2):
if [Link] is None:
break
currentNode = [Link]

[Link] = [Link]
[Link] = newNode
return head

node1 = Node(7)
node2 = Node(3)
node3 = Node(2)
node4 = Node(9)

[Link] = node2
[Link] = node3
[Link] = node4

print("Original list:")
traverseAndPrint(node1)

# Insert a new node with value 97 at position 2


newNode = Node(97)
node1 = insertNodeAtPosition(node1, newNode, 2)

print("\nAfter insertion:")
traverseAndPrint(node1)
In the insertNodeAtPosition function above, the return value is the new
head of the linked list. So for example, if the node is inserted at the start of
the linked list, the new head returned will be the new node.

Time Complexity of Linked Lists


Operations
Here we discuss time complexity of linked list operations, and compare these
with the time complexity of the array algorithms that we have discussed
previously in this tutorial.

Remember that time complexity just says something about the approximate
number of operations needed by the algorithm based on a large set of
data (n), and does not tell us the exact time a specific implementation of an
algorithm takes.

This means that even though linear search is said to have the same time
complexity for arrays as for linked list: O(n), it does not mean they take the
same amount of time. The exact time it takes for an algorithm to run
depends on programming language, computer hardware, differences in time
needed for operations on arrays vs linked lists, and many other things as
well.

Linear search for linked lists works the same as for arrays. A list of unsorted
values are traversed from the head node until the node with the specific
value is found. Time complexity is O(n).

Binary search is not possible for linked lists because the algorithm is based
on jumping directly to different array elements, and that is not possible with
linked lists.

Sorting algorithms have the same time complexities as for arrays, and these
are explained earlier in this tutorial. But remember, sorting algorithms that
are based on directly accessing an array element based on an index, do not
work on linked lists.

Hash Tables with Python


Hash Table
A Hash Table is a data structure designed to be fast to work with.

The reason Hash Tables are sometimes preferred instead of arrays or linked
lists is because searching for, adding, and deleting data can be done really
quickly, even for large amounts of data.

In a Linked List, finding a person "Bob" takes time because we would have to
go from one node to the next, checking each node, until the node with "Bob"
is found.

And finding "Bob" in an list/array could be fast if we knew the index, but
when we only know the name "Bob", we need to compare each element and
that takes time.

With a Hash Table however, finding "Bob" is done really fast because there is
a way to go directly to where "Bob" is stored, using something called a hash
function.

Building A Hash Table from Scratch


To get the idea of what a Hash Table is, let's try to build one from scratch, to
store unique first names inside it.

We will build the Hash Table in 5 steps:

1. Create an empty list (it can also be a dictionary or a set).


2. Create a hash function.
3. Inserting an element using a hash function.
4. Looking up an element using a hash function.
5. Handling collisions.

Step 1: Create an Empty List


To keep it simple, let's create a list with 10 empty elements.

my_list = [None, None, None, None, None, None, None, None, None,
None]
Each of these elements is called a bucket in a Hash Table.

Step 2: Create a Hash Function


Now comes the special way we interact with Hash Tables.

We want to store a name directly into its right place in the array, and this is
where the hash function comes in.

A hash function can be made in many ways, it is up to the creator of the


Hash Table. A common way is to find a way to convert the value into a
number that equals one of the Hash Table's index numbers, in this case a
number from 0 to 9.

In our example we will use the Unicode number of each character,


summarize them and do a modulo 10 operation to get index numbers 0-9.

ExampleGet your own Python Server


Create a Hash Function that sums the Unicode numbers of each character
and return a number between 0 and 9:

def hash_function(value):
sum_of_chars = 0
for char in value:
sum_of_chars += ord(char)

return sum_of_chars % 10

print("'Bob' has hash code:", hash_function('Bob'))

The character B has Unicode number 66, o has 111, and b has 98. Adding those
together we get 275. Modulo 10 of 275 is 5, so "Bob" should be stored at
index 5.

The number returned by the hash function is called the hash code.

Unicode number: Everything in our computers are stored as numbers, and


the Unicode code number is a unique number that exist for every character.
For example, the character A has Unicode number 65.

See this page for more information about how characters are represented as
numbers.
Modulo: A modulo operation divides a number with another number, and
gives us the resulting remainder. So for example, 7 % 3 will give us the
remainder 1. (Dividing 7 apples between 3 people, means that each person
gets 2 apples, with 1 apple to spare.)

In Python and most programming languages, the modolo operator is written


as %.

REMOVE ADS

Step 3: Inserting an Element


According to our hash function, "Bob" should be stored at index 5.

Lets create a function that add items to our hash table:

Example
def add(name):
index = hash_function(name)
my_list[index] = name

add('Bob')
print(my_list)

After storing "Bob" at index 5, our array now looks like this:

my_list = [None, None, None, None, None, 'Bob', None, None, None,
None]

We can use the same functions to store "Pete", "Jones", "Lisa", and "Siri" as
well.

Example
add('Pete')
add('Jones')
add('Lisa')
add('Siri')
print(my_list)
After using the hash function to store those names in the correct position,
our array looks like this:

Example
my_list = [None, 'Jones', None, 'Lisa', None, 'Bob',
None, 'Siri', 'Pete', None]

Step 4: Looking up a name


Now that we have a super basic Hash Table, let's see how we can look up a
name from it.

To find "Pete" in the Hash Table, we give the name "Pete" to our hash
function. The hash function returns 8, meaning that "Pete" is stored at index
8.

Example
def contains(name):
index = hash_function(name)
return my_list[index] == name

print("'Pete' is in the Hash Table:", contains('Pete'))

Because we do not have to check element by element to find out if "Pete" is


in there, we can just use the hash function to go straight to the right
element!

Step 5: Handling collisions


Let's also add "Stuart" to our Hash Table.

We give "Stuart" to our hash function, which returns 3, meaning "Stuart"


should be stored at index 3.

Trying to store "Stuart" in index 3, creates what is called a collision,


because "Lisa" is already stored at index 3.
To fix the collision, we can make room for more elements in the same
bucket. Solving the collision problem in this way is called chaining, and
means giving room for more elements in the same bucket.

Start by creating a new list with the same size as the original list, but with
empty buckets:

my_list = [
[],
[],
[],
[],
[],
[],
[],
[],
[],
[]
]

Rewrite the add() function, and add the same names as before:

Example
def add(name):
index = hash_function(name)
my_list[index].append(name)

add('Bob')
add('Pete')
add('Jones')
add('Lisa')
add('Siri')
add('Stuart')
print(my_list)

After implementing each bucket as a list, "Stuart" can also be stored at index
3, and our Hash Set now looks like this:

Result
my_list = [
[None],
['Jones'],
[None],
['Lisa', 'Stuart'],
[None],
['Bob'],
[None],
['Siri'],
['Pete'],
[None]
]
Searching for "Stuart" now takes a little bit longer time, because we also find
"Lisa" in the same bucket, but still much faster than searching the entire
Hash Table.

Uses of Hash Tables


Hash Tables are great for:

 Checking if something is in a collection (like finding a book in a library).


 Storing unique items and quickly finding them (like storing phone
numbers).
 Connecting values to keys (like linking names to phone numbers).

The most important reason why Hash Tables are great for these things is
that Hash Tables are very fast compared Arrays and Linked Lists, especially
for large sets. Arrays and Linked Lists have time complexity O(n) for search
and delete, while Hash Tables have just O(1) on average.

Hash Tables Summarized


Hash Table elements are stored in storage containers called buckets.

A hash function takes the key of an element to generate a hash code.

The hash code says what bucket the element belongs to, so now we can go
directly to that Hash Table element: to modify it, or to delete it, or just to
check if it exists.

A collision happens when two Hash Table elements have the same hash
code, because that means they belong to the same bucket.

Collision can be solved by Chaining by using lists to allow more than one
element in the same bucket.
Python Trees

A tree is a hierarchical data structure consisting of nodes connected by


edges.

Each node contains a value and references to its child nodes.

Trees
The Tree data structure is similar to Linked Lists in that each node contains
data and can be linked to other nodes.

We have previously covered data structures like Arrays, Linked Lists, Stacks,
and Queues. These are all linear structures, which means that each element
follows directly after another in a sequence. Trees however, are different. In
a Tree, a single element can have multiple 'next' elements, allowing the data
structure to branch out in various directions.

The data structure is called a "tree" because it looks like a tree's structure.

RABCDEFGHI

The Tree data structure can be useful in many cases:

 Hierarchical Data: File systems, organizational models, etc.


 Databases: Used for quick data retrieval.
 Routing Tables: Used for routing data in network algorithms.
 Sorting/Searching: Used for sorting data and searching for data.
 Priority Queues: Priority queue data structures are commonly
implemented using trees, such as binary heaps.

REMOVE ADS
Types of Trees
Trees are a fundamental data structure in computer science, used to
represent hierarchical relationships. This tutorial covers several key types of
trees.

Binary Trees: Each node has up to two children, the left child node and the
right child node. This structure is the foundation for more complex tree types
like Binay Search Trees and AVL Trees.

Binary Search Trees (BSTs): A type of Binary Tree where for each node,
the left child node has a lower value, and the right child node has a higher
value.

AVL Trees: A type of Binary Search Tree that self-balances so that for every
node, the difference in height between the left and right subtrees is at most
one. This balance is maintained through rotations when nodes are inserted
or deleted.

Each of these data structures are described in detail on the next pages,
including animations and how to implement them.

Trees vs Arrays and Linked Lists


Benefits of Trees over Arrays and Linked Lists:

 Arrays are fast when you want to access an element directly, like
element number 700 in an array of 1000 elements for example. But
inserting and deleting elements require other elements to shift in
memory to make place for the new element, or to take the deleted
elements place, and that is time consuming.
 Linked Lists are fast when inserting or deleting nodes, no memory
shifting needed, but to access an element inside the list, the list must
be traversed, and that takes time.
 Trees, such as Binary Trees, Binary Search Trees and AVL Trees, are
great compared to Arrays and Linked Lists because they are BOTH fast
at accessing a node, AND fast when it comes to deleting or inserting a
node, with no shifts in memory needed.

Python Binary Trees


A tree is a hierarchical data structure consisting of nodes connected by
edges.

Each node contains a value and references to its child nodes.

Binary Trees
A Binary Tree is a type of tree data structure where each node can have a
maximum of two child nodes, a left child node and a right child node.

This restriction, that a node can have a maximum of two child nodes, gives
us many benefits:

 Algorithms like traversing, searching, insertion and deletion become


easier to understand, to implement, and run faster.
 Keeping data sorted in a Binary Search Tree (BST) makes searching
very efficient.
 Balancing trees is easier to do with a limited number of child nodes,
using an AVL Binary Tree for example.
 Binary Trees can be represented as arrays, making the tree more
memory efficient.

Binary Tree Implementation


RABCDEFG

The Binary Tree above can be implemented much like a Linked List, except
that instead of linking each node to one next node, we create a structure
where each node can be linked to both its left and right child nodes.

ExampleGet your own Python Server


Create a Binary Tree in Python:

class TreeNode:
def __init__(self, data):
[Link] = data
[Link] = None
[Link] = None

root = TreeNode('R')
nodeA = TreeNode('A')
nodeB = TreeNode('B')
nodeC = TreeNode('C')
nodeD = TreeNode('D')
nodeE = TreeNode('E')
nodeF = TreeNode('F')
nodeG = TreeNode('G')

[Link] = nodeA
[Link] = nodeB

[Link] = nodeC
[Link] = nodeD

[Link] = nodeE
[Link] = nodeF

[Link] = nodeG

# Test
print("[Link]:", [Link])

REMOVE ADS

Types of Binary Trees


There are different variants, or types, of Binary Trees worth discussing to get
a better understanding of how Binary Trees can be structured.

The different kinds of Binary Trees are also worth mentioning now as these
words and concepts will be used later in the tutorial.

Below are short explanations of different types of Binary Tree structures, and
below the explanations are drawings of these kinds of structures to make it
as easy to understand as possible.
A balanced Binary Tree has at most 1 in difference between its left and right
subtree heights, for each node in the tree.

A complete Binary Tree has all levels full of nodes, except the last level,
which is can also be full, or filled from left to right. The properties of a
complete Binary Tree means it is also balanced.

A full Binary Tree is a kind of tree where each node has either 0 or 2 child
nodes.

A perfect Binary Tree has all leaf nodes on the same level, which means
that all levels are full of nodes, and all internal nodes have two child
[Link] properties of a perfect Binary Tree means it is also full, balanced,
and complete.

1171539131918Balanced11715391319248Complete and
balanced1171513191214Full11715313199Perfect, full, balanced and
complete

Binary Tree Traversal


Going through a Tree by visiting every node, one node at a time, is called
traversal.

Since Arrays and Linked Lists are linear data structures, there is only one
obvious way to traverse these: start at the first element, or node, and
continue to visit the next until you have visited them all.

But since a Tree can branch out in different directions (non-linear), there are
different ways of traversing Trees.

There are two main categories of Tree traversal methods:

Breadth First Search (BFS) is when the nodes on the same level are
visited before going to the next level in the tree. This means that the tree is
explored in a more sideways direction.

Depth First Search (DFS) is when the traversal moves down the tree all
the way to the leaf nodes, exploring the tree branch by branch in a
downwards direction.

There are three different types of DFS traversals:


 pre-order
 in-order
 post-order
REMOVE ADS

Pre-order Traversal of Binary Trees


Pre-order Traversal is a type of Depth First Search, where each node is
visited in a certain order..

Pre-order Traversal is done by visiting the root node first, then recursively do
a pre-order traversal of the left subtree, followed by a recursive pre-order
traversal of the right subtree. It's used for creating a copy of the tree, prefix
notation of an expression tree, etc.

This traversal is "pre" order because the node is visited "before" the
recursive pre-order traversal of the left and right subtrees.

This is how the code for pre-order traversal looks like:

Example
A pre-order traversal:

def preOrderTraversal(node):
if node is None:
return
print([Link], end=", ")
preOrderTraversal([Link])
preOrderTraversal([Link])

The first node to be printed is node R, as the Pre-order Traversal works by


first visiting, or printing, the current node (line 4), before calling the left and
right child nodes recursively (line 5 and 6).

The preOrderTraversal() function keeps traversing the left subtree recursively


(line 5), before going on to traversing the right subtree (line 6). So the next
nodes that are printed are 'A' and then 'C'.

The first time the argument node is None is when the left child of node C is
given as an argument (C has no left child).
After None is returned the first time when calling C's left child, C's right child
also returns None, and then the recursive calls continue to propagate back so
that A's right child D is the next to be printed.

The code continues to propagate back so that the rest of the nodes in R's
right subtree gets printed.

In-order Traversal of Binary Trees


In-order Traversal is a type of Depth First Search, where each node is visited
in a certain order.

In-order Traversal does a recursive In-order Traversal of the left subtree,


visits the root node, and finally, does a recursive In-order Traversal of the
right subtree. This traversal is mainly used for Binary Search Trees where it
returns values in ascending order.

What makes this traversal "in" order, is that the node is visited in between
the recursive function calls. The node is visited after the In-order Traversal of
the left subtree, and before the In-order Traversal of the right subtree.

This is how the code for In-order Traversal looks like:

Example
Create an In-order Traversal:

def inOrderTraversal(node):
if node is None:
return
inOrderTraversal([Link])
print([Link], end=", ")
inOrderTraversal([Link])

The inOrderTraversal() function keeps calling itself with the current left child
node as an argument (line 4) until that argument is None and the function
returns (line 2-3).

The first time the argument node is None is when the left child of node C is
given as an argument (C has no left child).
After that, the data part of node C is printed (line 5), which means that 'C' is
the first thing that gets printed.

Then, node C's right child is given as an argument (line 6), which is None, so
the function call returns without doing anything else.

After 'C' is printed, the previous inOrderTraversal() function calls continue to


run, so that 'A' gets printed, then 'D', then 'R', and so on.

Post-order Traversal of Binary Trees


Post-order Traversal is a type of Depth First Search, where each node is
visited in a certain order..

Post-order Traversal works by recursively doing a Post-order Traversal of the


left subtree and the right subtree, followed by a visit to the root node. It is
used for deleting a tree, post-fix notation of an expression tree, etc.

What makes this traversal "post" is that visiting a node is done "after" the
left and right child nodes are called recursively.

This is how the code for Post-order Traversal looks like:

Example
Post-order Traversal:

def postOrderTraversal(node):
if node is None:
return
postOrderTraversal([Link])
postOrderTraversal([Link])
print([Link], end=", ")

The postOrderTraversal() function keeps traversing the left subtree


recursively (line 4), until None is returned when C's left child node is called as
the node argument.

After C's left child node returns None, line 5 runs and C's right child node
returns None, and then the letter 'C' is printed (line 6).
This means that C is visited, or printed, "after" its left and right child nodes
are traversed, that is why it is called "post" order traversal.

The postOrderTraversal() function continues to propagate back to previous


recursive function calls, so the next node to be printed is 'D', then 'A'.

The function continues to propagate back and printing nodes until all nodes
are printed, or visited.

Python Binary Search Trees

A Binary Search Tree is a Binary Tree where every node's left child has a
lower value, and every node's right child has a higher value.

A clear advantage with Binary Search Trees is that operations like search,
delete, and insert are fast and done without having to shift values in
memory.

Binary Search Trees


A Binary Search Tree (BST) is a type of Binary Tree data structure, where the
following properties must be true for any node "X" in the tree:

 The X node's left child and all of its descendants (children, children's
children, and so on) have lower values than X's value.
 The right child, and all its descendants have higher values than X's
value.
 Left and right subtrees must also be Binary Search Trees.

These properties makes it faster to search, add and delete values than a
regular binary tree.

To make this as easy to understand and implement as possible, let's also


assume that all values in a Binary Search Tree are unique.

The size of a tree is the number of nodes in it (n).

A subtree starts with one of the nodes in the tree as a local root, and
consists of that node and all its descendants.
The descendants of a node are all the child nodes of that node, and all their
child nodes, and so on. Just start with a node, and the descendants will be all
nodes that are connected below that node.

The node's height is the maximum number of edges between that node
and a leaf node.

A node's in-order successor is the node that comes after it if we were to


do in-order traversal. In-order traversal of the BST above would result in
node 13 coming before node 14, and so the successor of node 13 is node 14.

Traversal of a Binary Search Tree


Just to confirm that we actually have a Binary Search Tree data structure in
front of us, we can check if the properties at the top of this page are true. So
for every node in the figure above, check if all the values to the left of the
node are lower, and that all values to the right are higher.

Another way to check if a Binary Tree is BST, is to do an in-order traversal


(like we did on the previous page) and check if the resulting list of values are
in an increasing order.

The code below is an implementation of the Binary Search Tree in the figure
above, with traversal.

ExampleGet your own Python Server


Traversal of a Binary Search Tree in Python

class TreeNode:
def __init__(self, data):
[Link] = data
[Link] = None
[Link] = None

def inOrderTraversal(node):
if node is None:
return
inOrderTraversal([Link])
print([Link], end=", ")
inOrderTraversal([Link])

root = TreeNode(13)
node7 = TreeNode(7)
node15 = TreeNode(15)
node3 = TreeNode(3)
node8 = TreeNode(8)
node14 = TreeNode(14)
node19 = TreeNode(19)
node18 = TreeNode(18)

[Link] = node7
[Link] = node15

[Link] = node3
[Link] = node8

[Link] = node14
[Link] = node19

[Link] = node18

# Traverse
inOrderTraversal(root)

As we can see by running the code example above, the in-order traversal
produces a list of numbers in an increasing (ascending) order, which means
that this Binary Tree is a Binary Search Tree.

REMOVE ADS

Search for a Value in a BST


Searching for a value in a BST is very similar to how we found a value
using Binary Search on an array.

For Binary Search to work, the array must be sorted already, and searching
for a value in an array can then be done really fast.

Similarly, searching for a value in a BST can also be done really fast because
of how the nodes are placed.

How it works:
1. Start at the root node.
2. If this is the value we are looking for, return.
3. If the value we are looking for is higher, continue searching in the right
subtree.
4. If the value we are looking for is lower, continue searching in the left
subtree.
5. If the subtree we want to search does not exist, depending on the
programming language, return None, or NULL, or something similar, to
indicate that the value is not inside the BST.

The algorithm can be implemented like this:

Example
Search the Tree for the value "13"

def search(node, target):


if node is None:
return None
elif [Link] == target:
return node
elif target < [Link]:
return search([Link], target)
else:
return search([Link], target)

# Search for a value


result = search(root, 13)
if result:
print(f"Found the node with value: {[Link]}")
else:
print("Value not found in the BST.")

The time complexity for searching a BST for a value is O(h), where h is the
height of the tree.

For a BST with most nodes on the right side for example, the height of the
tree becomes larger than it needs to be, and the worst case search will take
longer. Such trees are called unbalanced.

1371538141918Balanced BST7133158191418Unbalanced BST

Both Binary Search Trees above have the same nodes, and in-order traversal
of both trees gives us the same result but the height is very different. It
takes longer time to search the unbalanced tree above because it is higher.
We will use the next page to describe a type of Binary Tree called AVL Trees.
AVL trees are self-balancing, which means that the height of the tree is kept
to a minimum so that operations like search, insertion and deletion take less
time.

Insert a Node in a BST


Inserting a node in a BST is similar to searching for a value.

How it works:

1. Start at the root node.


2. Compare each node:
o Is the value lower? Go left.
o Is the value higher? Go right.
3. Continue to compare nodes with the new value until there is no right or left
to compare with. That is where the new node is inserted.

Inserting nodes as described above means that an inserted node will always
become a new leaf node.

All nodes in the BST are unique, so in case we find the same value as the one
we want to insert, we do nothing.

This is how node insertion in BST can be implemented:

Example
Inserting a node in a BST:

def insert(node, data):


if node is None:
return TreeNode(data)
else:
if data < [Link]:
[Link] = insert([Link], data)
elif data > [Link]:
[Link] = insert([Link], data)
return node

# Inserting new value into the BST


insert(root, 10)
REMOVE ADS

Find The Lowest Value in a BST


Subtree
The next section will explain how we can delete a node in a BST, but to do
that we need a function that finds the lowest value in a node's subtree.

How it works:

1. Start at the root node of the subtree.


2. Go left as far as possible.
3. The node you end up in is the node with the lowest value in that BST
subtree.

This is how a function for finding the lowest value in the subtree of a BST
node looks like:

Example
Find the lowest value in a BST subtree

def minValueNode(node):
current = node
while [Link] is not None:
current = [Link]
return current

# Find Lowest
print("\nLowest value:",minValueNode(root).data)

We will use this minValueNode() function in the section below, to find a node's
in-order successor, and use that to delete a node.

Delete a Node in a BST


To delete a node, our function must first search the BST to find it.
After the node is found there are three different cases where deleting a node
must be done differently.

How it works:

1. If the node is a leaf node, remove it by removing the link to it.


2. If the node only has one child node, connect the parent node of the node
you want to remove to that child node.
3. If the node has both right and left child nodes: Find the node's in-order
successor, change values with that node, then delete it.

In step 3 above, the successor we find will always be a leaf node, and
because it is the node that comes right after the node we want to delete, we
can swap values with it and delete it.

This is how a BST can be implemented with functionality for deleting a node:

Example
Delete a Node in a BST

def delete(node, data):


if not node:
return None

if data < [Link]:


[Link] = delete([Link], data)
elif data > [Link]:
[Link] = delete([Link], data)
else:
# Node with only one child or no child
if not [Link]:
temp = [Link]
node = None
return temp
elif not [Link]:
temp = [Link]
node = None
return temp

# Node with two children, get the in-order successor


[Link] = minValueNode([Link]).data
[Link] = delete([Link], [Link])

return node

# Delete node 15
delete(root,15)

Line 1: The node argument here makes it possible for the function to call
itself recursively on smaller and smaller subtrees in the search for the node
with the data we want to delete.

Line 2-8: This is searching for the node with correct data that we want to
delete.

Line 9-22: The node we want to delete has been found. There are three
such cases:

1. Case 1: Node with no child nodes (leaf node). None is returned, and
that becomes the parent node's new left or right value by recursion
(line 6 or 8).
2. Case 2: Node with either left or right child node. That left or right child
node becomes the parent's new left or right child through recursion
(line 7 or 9).
3. Case 3: Node has both left and right child nodes. The in-order
successor is found using the minValueNode() function. We keep the
successor's value by setting it as the value of the node we want to
delete, and then we can delete the successor node.

Line 24: node is returned to maintain the recursive functionality.

BST Compared to Other Data


Structures
Binary Search Trees take the best from two other data structures: Arrays and
Linked Lists.

Data Structure Searching for a value Delete / Insert leads to

Sorted Array O(\log n) Yes


Linked List O(n) No

Binary Search Tree O(\log n) No

Searching a BST is just as fast as Binary Search on an array, with the same
time complexity O(log n).

And deleting and inserting new values can be done without shifting elements
in memory, just like with Linked Lists.

BST Balance and Time Complexity


On a Binary Search Tree, operations like inserting a new node, deleting a
node, or searching for a node are actually O(h). That means that the higher
the tree is (h), the longer the operation will take.

The reason why we wrote that searching for a value is O(log n) in the table
above is because that is true if the tree is "balanced", like in the image
below.

1371538141918Balanced BST

We call this tree balanced because there are approximately the same
number of nodes on the left and right side of the tree.

The exact way to tell that a Binary Tree is balanced is that the height of the
left and right subtrees of any node only differs by one. In the image above,
the left subtree of the root node has height h=2, and the right subtree has
height h=3.

For a balanced BST, with a large number of nodes (big n), we get height h ≈ \
log_2 n, and therefore the time complexity for searching, deleting, or
inserting a node can be written as O(h) = O(\log n).

But, in case the BST is completely unbalanced, like in the image below, the
height of the tree is approximately the same as the number of nodes, h ≈ n,
and we get time complexity O(h) = O(n) for searching, deleting, or inserting a
node.

7133158191418Unbalanced BST

So, to optimize operations on a BST, the height must be minimized, and to


do that the tree must be balanced.

And keeping a Binary Search Tree balanced is exactly what AVL Trees do,
which is the data structure explained on the next page.

Python AVL Trees


The AVL Tree is a type of Binary Search Tree named after two Soviet
inventors Georgy Adelson-Velsky and Evgenii Landis who invented the AVL
Tree in 1962.

AVL trees are self-balancing, which means that the tree height is kept to a
minimum so that a very fast runtime is guaranteed for searching, inserting
and deleting nodes, with time complexity O(logn).

AVL Trees
The only difference between a regular Binary Search Tree and an AVL Tree is
that AVL Trees do rotation operations in addition, to keep the tree balance.

A Binary Search Tree is in balance when the difference in height between left
and right subtrees is less than 2.

By keeping balance, the AVL Tree ensures a minimum tree height, which
means that search, insert, and delete operations can be done really fast.

BGEKFPIMBinary Search Tree


(unbalanced)
Height: 6GEKBFIPMAVL Tree
(self-balancing)
Height: 3

The two trees above are both Binary Search Trees, they have the same
nodes, and the same in-order traversal (alphabetical), but the height is very
different because the AVL Tree has balanced itself.
Step through the building of an AVL Tree in the animation below to see how
the balance factors are updated, and how rotation operations are done when
required to restore the balance.

0C0F0G0D0B0AInsert C

Continue reading to learn more about how the balance factor is calculated,
how rotation operations are done, and how AVL Trees can be implemented.

REMOVE ADS

Left and Right Rotations


To restore balance in an AVL Tree, left or right rotations are done, or a
combination of left and right rotations.

The previous animation shows one specific left rotation, and one specific
right rotation.

But in general, left and right rotations are done like in the animation below.

XYRotate Right

Notice how the subtree changes its parent. Subtrees change parent in this
way during rotation to maintain the correct in-order traversal, and to
maintain the BST property that the left child is less than the right child, for all
nodes in the tree.

Also keep in mind that it is not always the root node that become
unbalanced and need rotation.

The Balance Factor


A node's balance factor is the difference in subtree heights.
The subtree heights are stored at each node for all nodes in an AVL Tree, and
the balance factor is calculated based on its subtree heights to check if the
tree has become out of balance.

The height of a subtree is the number of edges between the root node of the
subtree and the leaf node farthest down in that subtree.

The Balance Factor (BF) for a node (X) is the difference in height between
its right and left subtrees.
BF(X)=height(rightSubtree(X))−height(leftSubtree(X))
Balance factor values

 0: The node is in balance.


 more than 0: The node is "right heavy".
 less than 0: The node is "left heavy".

If the balance factor is less than -1, or more than 1, for one or more nodes in
the tree, the tree is considered not in balance, and a rotation operation is
needed to restore balance.

Let's take a closer look at the different rotation operations that an AVL Tree
can do to regain balance.

The Four "out-of-balance" Cases


When the balance factor of just one node is less than -1, or more than 1, the
tree is regarded as out of balance, and a rotation is needed to restore
balance.

There are four different ways an AVL Tree can be out of balance, and each of
these cases require a different rotation operation.

Case Description Rotation to Restore Ba

Left-Left The unbalanced node and its left child node A single right rotation.
(LL) are both left-heavy.
Right-Right The unbalanced node and its right child node A single left rotation.
(RR) are both right-heavy.

Left-Right The unbalanced node is left heavy, and its left First do a left rotation on
(LR) child node is right heavy. rotation on the unbalance

Right-Left The unbalanced node is right heavy, and its First do a right rotation on
(RL) right child node is left heavy. left rotation on the unbala

See animations and explanations of these cases below.

The Left-Left (LL) Case


The node where the unbalance is discovered is left heavy, and the node's left
child node is also left heavy.

When this LL case happens, a single right rotation on the unbalanced node is
enough to restore balance.

Step through the animation below to see the LL case, and how the balance is
restored by a single right rotation.

-1Q0P0D0L0C0B0K0AInsert D

As you step through the animation above, two LL cases happen:

1. When D is added, the balance factor of Q becomes -2, which means


the tree is unbalanced. This is an LL case because both the unbalance
node Q and its left child node P are left heavy (negative balance
factors). A single right rotation at node Q restores the tree balance.
2. After nodes L, C, and B are added, P's balance factor is -2, which
means the tree is out of balance. This is also an LL case because both
the unbalanced node P and its left child node D are left heavy. A single
right rotation restores the balance.
Note: The second time the LL case happens in the animation above, a right
rotation is done, and L goes from being the right child of D to being the left
child of P. Rotations are done like that to keep the correct in-order traversal
('B, C, D, L, P, Q' in the animation above). Another reason for changing
parent when a rotation is done is to keep the BST property, that the left child
is always lower than the node, and that the right child always higher.

The Right-Right (RR) Case


A Right-Right case happens when a node is unbalanced and right heavy, and
the right child node is also right heavy.

A single left rotation at the unbalanced node is enough to restore balance in


the RR case.

+1A0B0D0C0E0FInsert D

The RR case happens two times in the animation above:

1. When node D is inserted, A becomes unbalanced, and bot A and B are


right heavy. A left rotation at node A restores the tree balance.
2. After nodes E, C and F are inserted, node B becomes unbalanced. This
is an RR case because both node B and its right child node D are right
heavy. A left rotation restores the tree balance.

The Left-Right (LR) Case


The Left-Right case is when the unbalanced node is left heavy, but its left
child node is right heavy.

In this LR case, a left rotation is first done on the left child node, and then a
right rotation is done on the original unbalanced node.

Step through the animation below to see how the Left-Right case can
happen, and how the rotation operations are done to restore balance.

-1Q0E0K0C0F0GInsert K
As you are building the AVL Tree in the animation above, the Left-Right case
happens 2 times, and rotation operations are required and done to restore
balance:

1. When K is inserted, node Q gets unbalanced with a balance factor of -


2, so it is left heavy, and its left child E is right heavy, so this is a Left-
Right case.
2. After nodes C, F, and G are inserted, node K becomes unbalanced and
left heavy, with its left child node E right heavy, so it is a Left-Right
case.

The Right-Left (RL) Case


The Right-Left case is when the unbalanced node is right heavy, and its right
child node is left heavy.

In this case we first do a right rotation on the unbalanced node's right child,
and then we do a left rotation on the unbalanced node itself.

Step through the animation below to see how the Right-Left case can occur,
and how rotations are done to restore the balance.

+1A0F0B0G0E0DInsert B

After inserting node B, we get a Right-Left case because node A becomes


unbalanced and right heavy, and its right child is left heavy. To restore
balance, a right rotation is first done on node F, and then a left rotation is
done on node A.

The next Right-Left case occurs after nodes G, E, and D are added. This is a
Right-Left case because B is unbalanced and right heavy, and its right child F
is left heavy. To restore balance, a right rotation is first done on node F, and
then a left rotation is done on node B.

Retracing in AVL Trees


After inserting or deleting a node in an AVL tree, the tree may become
unbalanced. To find out if the tree is unbalanced, we need to update the
heights and recalculate the balance factors of all ancestor nodes.
This process, known as retracing, is handled through recursion. As the
recursive calls propagate back to the root after an insertion or deletion, each
ancestor node's height is updated and the balance factor is recalculated. If
any ancestor node is found to have a balance factor outside the range of -1
to 1, a rotation is performed at that node to restore the tree's balance.

In the simulation below, after inserting node F, the nodes C, E and H are all
unbalanced, but since retracing works through recursion, the unbalance at
node H is discovered and fixed first, which in this case also fixes the
unbalance in nodes E and C.

0A-1B+1C0D+1E0G-1H0FInsert F

After node F is inserted, the code will retrace, calculating balancing factors
as it propagates back up towards the root node. When node H is reached and
the balancing factor -2 is calculated, a right rotation is done. Only after this
rotation is done, the code will continue to retrace, calculating balancing
factors further up on ancestor nodes E and C.

Because of the rotation, balancing factors for nodes E and C stay the same
as before node F was inserted.

AVL Tree Implementation in Python


This code is based on the BST implementation on the previous page, for
inserting nodes.

There is only one new attribute for each node in the AVL tree compared to
the BST, and that is the height, but there are many new functions and extra
code lines needed for the AVL Tree implementation because of how the AVL
Tree rebalances itself.

The implementation below builds an AVL tree based on a list of characters,


to create the AVL Tree in the simulation above. The last node to be inserted
'F', also triggers a right rotation, just like in the simulation above.

ExampleGet your own Python Server


Implement AVL Tree in Python:

class TreeNode:
def __init__(self, data):
[Link] = data
[Link] = None
[Link] = None
[Link] = 1

def getHeight(node):
if not node:
return 0
return [Link]

def getBalance(node):
if not node:
return 0
return getHeight([Link]) - getHeight([Link])

def rightRotate(y):
print('Rotate right on node',[Link])
x = [Link]
T2 = [Link]
[Link] = y
[Link] = T2
[Link] = 1 + max(getHeight([Link]), getHeight([Link]))
[Link] = 1 + max(getHeight([Link]), getHeight([Link]))
return x

def leftRotate(x):
print('Rotate left on node',[Link])
y = [Link]
T2 = [Link]
[Link] = x
[Link] = T2
[Link] = 1 + max(getHeight([Link]), getHeight([Link]))
[Link] = 1 + max(getHeight([Link]), getHeight([Link]))
return y

def insert(node, data):


if not node:
return TreeNode(data)

if data < [Link]:


[Link] = insert([Link], data)
elif data > [Link]:
[Link] = insert([Link], data)

# Update the balance factor and balance the tree


[Link] = 1 + max(getHeight([Link]), getHeight([Link]))
balance = getBalance(node)
# Balancing the tree
# Left Left
if balance > 1 and getBalance([Link]) >= 0:
return rightRotate(node)

# Left Right
if balance > 1 and getBalance([Link]) < 0:
[Link] = leftRotate([Link])
return rightRotate(node)

# Right Right
if balance < -1 and getBalance([Link]) <= 0:
return leftRotate(node)

# Right Left
if balance < -1 and getBalance([Link]) > 0:
[Link] = rightRotate([Link])
return leftRotate(node)

return node

def inOrderTraversal(node):
if node is None:
return
inOrderTraversal([Link])
print([Link], end=", ")
inOrderTraversal([Link])

# Inserting nodes
root = None
letters = ['C', 'B', 'E', 'A', 'D', 'H', 'G', 'F']
for letter in letters:
root = insert(root, letter)

inOrderTraversal(root)

AVL Delete Node Implementation


When deleting a node that is not a leaf node, the AVL Tree requires
the minValueNode() function to find a node's next node in the in-order
traversal. This is the same as when deleting a node in a Binary Search Tree,
as explained on the previous page.
To delete a node in an AVL Tree, the same code to restore balance is needed
as for the code to insert a node.

Example
Delete Node:

def minValueNode(node):
current = node
while [Link] is not None:
current = [Link]
return current

def delete(node, data):


if not node:
return node

if data < [Link]:


[Link] = delete([Link], data)
elif data > [Link]:
[Link] = delete([Link], data)
else:
if [Link] is None:
temp = [Link]
node = None
return temp
elif [Link] is None:
temp = [Link]
node = None
return temp

temp = minValueNode([Link])
[Link] = [Link]
[Link] = delete([Link], [Link])

return node

def inOrderTraversal(node):
if node is None:
return
inOrderTraversal([Link])
print([Link], end=", ")
inOrderTraversal([Link])

# Inserting nodes
root = None
letters = ['C', 'B', 'E', 'A', 'D', 'H', 'G', 'F']
for letter in letters:
root = insert(root, letter)

inOrderTraversal(root)

Time Complexity for AVL Trees


Take a look at the unbalanced Binary Search Tree below. Searching for "M"
means that all nodes except 1 must be compared. But searching for "M" in
the AVL Tree below only requires us to visit 4 nodes.

So in worst case, algorithms like search, insert, and delete must run through
the whole height of the tree. This means that keeping the height (h) of the
tree low, like we do using AVL Trees, gives us a lower runtime.

BGEKFPIMBinary Search Tree


(unbalanced)GEKBFIPMAVL Tree
(self-balancing)
See the comparison of the time complexities between Binary Search Trees
and AVL Trees below, and how the time complexities relate to the height ( h)
of the tree, and the number of nodes (n) in the tree.
 The BST is not self-balancing. This means that a BST can be very
unbalanced, almost like a long chain, where the height is nearly the
same as the number of nodes. This makes operations like searching,
deleting and inserting nodes slow, with time complexity O(h)=O(n).
 The AVL Tree however is self-balancing. That means that the height of
the tree is kept to a minimum so that operations like searching,
deleting and inserting nodes are much faster, with time
complexity O(h)=O(logn).

O(logn) Explained
The fact that the time complexity is O(h)=O(logn) for search, insert, and
delete on an AVL Tree with height h and nodes n can be explained like this:

Imagine a perfect Binary Tree where all nodes have two child nodes except
on the lowest level, like the AVL Tree below.

HDBFEGACLJNMOIK
The number of nodes on each level in such an AVL Tree are:

1,2,4,8,16,32,..

Which is the same as:

20,21,22,23,24,25,..
To get the number of nodes n in a perfect Binary Tree with height h=3, we
can add the number of nodes on each level together:
n3=20+21+22+23=15

Which is actually the same as:

n3=24−1=15
And this is actually the case for larger trees as well! If we want to get the
number of nodes n in a tree with height h=5 for example, we find the
number of nodes like this:
n5=26−1=63
So in general, the relationship between the height h of a perfect Binary Tree
and the number of nodes in it n, can be expressed like this:
nh=2h+1−1
Note: The formula above can also be found by calculating the sum of the
geometric series 20+21+22+23+...+2n
We know that the time complexity for searching, deleting, or inserting a
node in an AVL tree is O(h), but we want to argue that the time complexity
is actually O(log(n)), so we need to find the height h described by the
number of nodes n:
n=2h+1−1n+1=2h+1log2(n+1)=log2(2h+1)h=log2(n+1)−1O(h)=O(log
n)
How the last line above is derived might not be obvious, but for a Binary Tree
with a lot of nodes (big n), the "+1" and "-1" terms are not important when
we consider time complexity. For more details on how to calculate the time
complexity using Big O notation, see this page.
The math above shows that the time complexity for search, delete, and
insert operations on an AVL Tree O(h), can actually be expressed
as O(logn), which is fast, a lot faster than the time complexity for BSTs
which is O(n).

Python Graphs
Graphs
A Graph is a non-linear data structure that consists of vertices (nodes) and
edges.

F24BCAEDG

A vertex, also called a node, is a point or an object in the Graph, and an edge
is used to connect two vertices with each other.

Graphs are non-linear because the data structure allows us to have different
paths to get from one vertex to another, unlike with linear data structures
like Arrays or Linked Lists.

Graphs are used to represent and solve problems where the data consists of
objects and relationships between them, such as:

 Social Networks: Each person is a vertex, and relationships (like


friendships) are the edges. Algorithms can suggest potential friends.
 Maps and Navigation: Locations, like a town or bus stops, are stored as
vertices, and roads are stored as edges. Algorithms can find the
shortest route between two locations when stored as a Graph.
 Internet: Can be represented as a Graph, with web pages as vertices
and hyperlinks as edges.
 Biology: Graphs can model systems like neural networks or the spread
of diseases.

REMOVE ADS

Graph Representations
A Graph representation tells us how a Graph is stored in memory.

Different Graph representations can:

 take up more or less space.


 be faster or slower to search or manipulate.
 be better suited depending on what type of Graph we have (weighted,
directed, etc.), and what we want to do with the Graph.
 be easier to understand and implement than others.

Below are short introductions of the different Graph representations, but


Adjacency Matrix is the representation we will use for Graphs moving
forward in this tutorial, as it is easy to understand and implement, and works
in all cases relevant for this tutorial.

Graph representations store information about which vertices are adjacent,


and how the edges between the vertices are. Graph representations are
slightly different if the edges are directed or weighted.

Two vertices are adjacent, or neighbors, if there is an edge between them.

Adjacency Matrix Graph


Representation
Adjacency Matrix is the Graph representation (structure) we will use for this
tutorial.

How to implement an Adjacency Matrix is shown on the next page.

The Adjacency Matrix is a 2D array (matrix) where each cell on


index (i,j) stores information about the edge from vertex i to vertex j.

Below is a Graph with the Adjacency Matrix representation next to it.

ABCDABCDABCD11111111An undirected Graph


and the adjacency matrix

The adjacency matrix above represents an undirected Graph, so the values


'1' only tells us where the edges are. Also, the values in the adjacency matrix
is symmetrical because the edges go both ways (undirected Graph).

To create a directed Graph with an adjacency matrix, we must decide which


vertices the edges go from and to, by inserting the value at the correct
indexes (i,j). To represent a weighted Graph we can put other values than
'1' inside the adjacency matrix.

Below is a directed and weighted Graph with the Adjacency Matrix


representation next to it.
AB13C42DABCDABCD3214A directed and weighted Graph,
and its adjacency matrix.

In the adjacency matrix above, the value 3 on index (0,1) tells us there is an
edge from vertex A to vertex B, and the weight for that edge is 3.

As you can see, the weights are placed directly into the adjacency matrix for
the correct edge, and for a directed Graph, the adjacency matrix does not
have to be symmetric.

Adjacency List Graph Representation


In case we have a 'sparse' Graph with many vertices, we can save space by
using an Adjacency List compared to using an Adjacency Matrix, because an
Adjacency Matrix would reserve a lot of memory on empty Array elements
for edges that don't exist.

A 'sparse' Graph is a Graph where each vertex only has edges to a small
portion of the other vertices in the Graph.

An Adjacency List has an array that contains all the vertices in the Graph,
and each vertex has a Linked List (or Array) with the vertex's edges.

ABCD0123ABCD312null02null10null0nullAn undirected Graph


and its adjacency list.

In the adjacency list above, the vertices A to D are placed in an Array, and
each vertex in the array has its index written right next to it.

Each vertex in the Array has a pointer to a Linked List that represents that
vertex's edges. More specifically, the Linked List contains the indexes to the
adjacent (neighbor) vertices.

So for example, vertex A has a link to a Linked List with values 3, 1, and 2.
These values are the indexes to A's adjacent vertices D, B, and C.

An Adjacency List can also represent a directed and weighted Graph, like
this:

AB13C42D0123ABCD1,32,2nullnull1,1null0,4nullA directed and weighted Graph


and its adjacency list.
In the Adjacency List above, vertices are stored in an Array. Each vertex has
a pointer to a Linked List with edges stored as i,w, where i is the index of the
vertex the edge goes to, and w is the weight of that edge.

Node D for example, has a pointer to a Linked List with an edge to vertex A.
The values 0,4 means that vertex D has an edge to vertex on index 0 (vertex
A), and the weight of that edge is 4.

Linear Search with Python

Linear Search
Linear search (or sequential search) is the simplest search algorithm. It
checks each element one by one.

Search for the value 4

10

11

12
13

14

15

16

17

18

19

20

21

Run the simulation above to see how the Linear Search algorithm works.

This algorithm is very simple and easy to understand and implement.

How it works:

1. Go through the array value by value from the start.


2. Compare each value to check if it is equal to the value we are looking for.
3. If the value is found, return the index of that value.
4. If the end of the array is reached and the value is not found, return -1 to
indicate that the value was not found.
If the array is already sorted, it is better to use the much faster Binary
Search algorithm that we will explore on the next page.

REMOVE ADS

Implement Linear Search in Python


In Python, the fastest way check if a value exists in a list is to use
the in operator.

ExampleGet your own Python Server


Check if a value exists in a list:
mylist = [3, 7, 2, 9, 5, 1, 8, 4, 6]

if 4 in mylist:
print("Found!")
else:
print("Not found!")

But if you need to find the index of a value, you will need to implement a
linear search:

Example
Find the index of a value in a list:

def linearSearch(arr, targetVal):


for i in range(len(arr)):
if arr[i] == targetVal:
return i
return -1

mylist = [3, 7, 2, 9, 5, 1, 8, 4, 6]
x = 4

result = linearSearch(mylist, x)

if result != -1:
print("Found at index", result)
else:
print("Not found")

To implement the Linear Search algorithm we need:

1. An array with values to search through.


2. A target value to search for.
3. A loop that goes through the array from start to end.
4. An if-statement that compares the current value with the target value,
and returns the current index if the target value is found.
5. After the loop, return -1, because at this point we know the target
value has not been found.

Linear Search Time Complexity


If Linear Search runs and finds the target value as the first array value in an
array with n values, only one compare is needed.
But if Linear Search runs through the whole array of n values, without finding
the target value, n compares are needed.
This means that time complexity for Linear Search is: O(n)
If we draw how much time Linear Search needs to find a value in an array
of n values, we get this graph:

Binary Search with Python

Binary Search
The Binary Search algorithm searches through a sorted array and returns
the index of the value it searches for.

Search for value 2

10

11
12

13

14

15

16

17

18

19

20

Run the simulation to see how the Binary Search algorithm works.

Binary Search is much faster than Linear Search, but requires a sorted array
to work.

The Binary Search algorithm works by checking the value in the center of the
array. If the target value is lower, the next value to check is in the center of
the left half of the array. This way of searching means that the search area is
always half of the previous search area, and this is why the Binary Search
algorithm is so fast.

This process of halving the search area happens until the target value is
found, or until the search area of the array is empty.

How it works:

1. Check the value in the center of the array.


2. If the target value is lower, search the left half of the array. If the target
value is higher, search the right half.
3. Continue step 1 and 2 for the new reduced part of the array until the target
value is found or until the search area is empty.
4. If the value is found, return the target value index. If the target value is not
found, return -1.

Manual Run Through


Let's try to do the searching manually, just to get an even better
understanding of how Binary Search works before actually implementing it in
a Python program. We will search for value 11.
Step 1: We start with an array.

[ 2, 3, 7, 7, 11, 15, 25]

Step 2: The value in the middle of the array at index 3, is it equal to 11?

[ 2, 3, 7, 7, 11, 15, 25]

Step 3: 7 is less than 11, so we must search for 11 to the right of index 3.
The values to the right of index 3 are [ 11, 15, 25]. The next value to check is
the middle value 15, at index 5.

[ 2, 3, 7, 7, 11, 15, 25]

Step 4: 15 is higher than 11, so we must search to the left of index 5. We


have already checked index 0-3, so index 4 is only value left to check.

[ 2, 3, 7, 7, 11, 15, 25]

We have found it!

Value 11 is found at index 4.

Returning index position 4.

Binary Search is finished.

Run the simulation below to see the steps above animated:

Search for value 11


[
2,
3,
7,
7,
11,
15,
25
]

Implementing Binary Search in


Python
To implement the Binary Search algorithm we need:

1. An array with values to search through.


2. A target value to search for.
3. A loop that runs as long as left index is less than, or equal to, the right index.
4. An if-statement that compares the middle value with the target value, and
returns the index if the target value is found.
5. An if-statement that checks if the target value is less than, or larger than, the
middle value, and updates the "left" or "right" variables to narrow down the
search area.
6. After the loop, return -1, because at this point we know the target value has
not been found.

The resulting code for Binary Search looks like this:

ExampleGet your own Python Server


Create a Binary Search algorithm in Python:

def binarySearch(arr, targetVal):


left = 0
right = len(arr) - 1

while left <= right:


mid = (left + right) // 2

if arr[mid] == targetVal:
return mid

if arr[mid] < targetVal:


left = mid + 1
else:
right = mid - 1

return -1

mylist = [1, 3, 5, 7, 9, 11, 13, 15, 17, 19]


x = 11

result = binarySearch(mylist, x)

if result != -1:
print("Found at index", result)
else:
print("Not found")

Binary Search Time Complexity


Each time Binary Search checks a new value to see if it is the target value,
the search area is halved.

This means that even in the worst case scenario where Binary Search cannot
find the target value, it still only needs log2n comparisons to look through a
sorted array of n values.
Time complexity for Binary Search is: O(log2n)
Note: When writing time complexity using Big O notation we could also just
have written O(logn), but O(log2n) reminds us that the array search area is
halved for every new comparison, which is the basic concept of Binary
Search, so we will just keep the base 2 indication in this case.
If we draw how much time Binary Search needs to find a value in an array
of n values, compared to Linear Search, we get this graph:

Bubble Sort with Python

Bubble Sort
Bubble Sort is an algorithm that sorts an array from the lowest value to the
highest value.

Sort

Run the simulation to see how it looks like when the Bubble Sort algorithm
sorts an array of values. Each value in the array is represented by a column.

The word 'Bubble' comes from how this algorithm works, it makes the
highest values 'bubble up'.

How it works:

1. Go through the array, one value at a time.


2. For each value, compare the value with the next value.
3. If the value is higher than the next one, swap the values so that the
highest value comes last.
4. Go through the array as many times as there are values in the array.

REMOVE ADS

Manual Run Through


Before we implement the Bubble Sort algorithm in a programming language,
let's manually run through a short array only one time, just to get the idea.

Step 1: We start with an unsorted array.

[7, 12, 9, 11, 3]

Step 2: We look at the two first values. Does the lowest value come first?
Yes, so we don't need to swap them.

[7, 12, 9, 11, 3]


Step 3: Take one step forward and look at values 12 and 9. Does the lowest
value come first? No.

[7, 12, 9, 11, 3]

Step 4: So we need to swap them so that 9 comes first.

[7, 9, 12, 11, 3]

Step 5: Taking one step forward, looking at 12 and 11.

[7, 9, 12, 11, 3]

Step 6: We must swap so that 11 comes before 12.

[7, 9, 11, 12, 3]

Step 7: Looking at 12 and 3, do we need to swap them? Yes.

[7, 9, 11, 12, 3]

Step 8: Swapping 12 and 3 so that 3 comes first.

[7, 9, 11, 3, 12]

Repeat until no more swaps are needed and you will get a sorted array:

Bubble Sort
[
7,
12,
9,
11,
3
]

Implement Bubble Sort in Python


To implement the Bubble Sort algorithm in Python, we need:

1. An array with values to sort.


2. An inner loop that goes through the array and swaps values if the first
value is higher than the next value. This loop must loop through one
less value each time it runs.
3. An outer loop that controls how many times the inner loop must run.
For an array with n values, this outer loop must run n-1 times.

The resulting code looks like this:

ExampleGet your own Python Server


Create a Bubble Sort algorithm in Python:

mylist = [64, 34, 25, 12, 22, 11, 90, 5]

n = len(mylist)
for i in range(n-1):
for j in range(n-i-1):
if mylist[j] > mylist[j+1]:
mylist[j], mylist[j+1] = mylist[j+1], mylist[j]

print(mylist)

Bubble Sort Improvement


The Bubble Sort algorithm can be improved a little bit more.

Imagine that the array is almost sorted already, with the lowest numbers at
the start, like this for example:

mylist = [7, 3, 9, 12, 11]


In this case, the array will be sorted after the first run, but the Bubble Sort
algorithm will continue to run, without swapping elements, and that is not
necessary.

If the algorithm goes through the array one time without swapping any
values, the array must be finished sorted, and we can stop the algorithm,
like this:

Example
Improved Bubble Sort algorithm:

mylist = [7, 3, 9, 12, 11]

n = len(mylist)
for i in range(n-1):
swapped = False
for j in range(n-i-1):
if mylist[j] > mylist[j+1]:
mylist[j], mylist[j+1] = mylist[j+1], mylist[j]
swapped = True
if not swapped:
break

print(mylist)
REMOVE ADS

Bubble Sort Time Complexity


The Bubble Sort algorithm loops through every value in the array, comparing
it to the value next to it. So for an array of n values, there must be n such
comparisons in one loop.
And after one loop, the array is looped through again and again n times.
This means there are n⋅n comparisons done in total, so the time complexity
for Bubble Sort is: O(n2)

The graph describing the Bubble Sort time complexity looks like this:
As you can see, the run time increases really fast when the size of the array
is increased.

Luckily there are sorting algorithms that are faster than this, like Quicksort,
that we will look at later.

Selection Sort with Python

Selection Sort
The Selection Sort algorithm finds the lowest value in an array and moves it
to the front of the array.

Sort

The algorithm looks through the array again and again, moving the next
lowest values to the front, until the array is sorted.

How it works:

1. Go through the array to find the lowest value.


2. Move the lowest value to the front of the unsorted part of the array.
3. Go through the array again as many times as there are values in the array.
REMOVE ADS

Manual Run Through


Before we implement the Selection Sort algorithm in Python program, let's
manually run through a short array only one time, just to get the idea.

Step 1: We start with an unsorted array.

[ 7, 12, 9, 11, 3]

Step 2: Go through the array, one value at a time. Which value is the
lowest? 3, right?

[ 7, 12, 9, 11, 3]

Step 3: Move the lowest value 3 to the front of the array.

[ 3, 7, 12, 9, 11]

Step 4: Look through the rest of the values, starting with 7. 7 is the lowest
value, and already at the front of the array, so we don't need to move it.

[ 3, 7, 12, 9, 11]

Step 5: Look through the rest of the array: 12, 9 and 11. 9 is the lowest
value.

[ 3, 7, 12, 9, 11]
Step 6: Move 9 to the front.

[ 3, 7, 9, 12, 11]

Step 7: Looking at 12 and 11, 11 is the lowest.

[ 3, 7, 9, 12, 11]

Step 8: Move it to the front.

[ 3, 7, 9, 11, 12]

Finally, the array is sorted.

Run the simulation below to see the steps above animated:

Selection Sort
[
7,
12,
9,
11,
3
]

Implement Selection Sort in Python


To implement the Selection Sort algorithm in Python, we need:

1. An array with values to sort.


2. An inner loop that goes through the array, finds the lowest value, and
moves it to the front of the array. This loop must loop through one less
value each time it runs.
3. An outer loop that controls how many times the inner loop must run.
For an array with n values, this outer loop must run n−1 times.
The resulting code looks like this:

ExampleGet your own Python Server


Using the Selection sort on a Python list:

mylist = [64, 34, 25, 5, 22, 11, 90, 12]

n = len(mylist)
for i in range(n-1):
min_index = i
for j in range(i+1, n):
if mylist[j] < mylist[min_index]:
min_index = j
min_value = [Link](min_index)
[Link](i, min_value)

print(mylist)

Selection Sort Shifting Problem


The Selection Sort algorithm can be improved a little bit more.

In the code above, the lowest value element is removed, and then inserted in
front of the array.

Each time the next lowest value array element is removed, all following
elements must be shifted one place down to make up for the removal.

These shifting operation takes a lot of time, and we are not even done yet!
After the lowest value (5) is found and removed, it is inserted at the start of
the array, causing all following values to shift one position up to make space
for the new value, like the image below shows.
Note: You will not see these shifting operations happening in the code if you
are using a high level programming language such as Python or Java, but the
shifting operations are still happening in the background. Such shifting
operations require extra time for the computer to do, which can be a
problem.
REMOVE ADS

Solution: Swap Values!


Instead of all the shifting, swap the lowest value (5) with the first value (64)
like below.

We can swap values like the image above shows because the lowest value
ends up in the correct position, and it does not matter where we put the
other value we are swapping with, because it is not sorted yet.

Here is a simulation that shows how this improved Selection Sort with
swapping works:

Sort

We will insert the improvement in the Selection Sort algorithm:

Example
The improved Selection Sort algorithm, including swapping values:
mylist = [64, 34, 25, 12, 22, 11, 90, 5]

n = len(mylist)
for i in range(n):
min_index = i
for j in range(i+1, n):
if mylist[j] < mylist[min_index]:
min_index = j
mylist[i], mylist[min_index] = mylist[min_index], mylist[i]

print(mylist)

Selection Sort Time Complexity


Selection Sort sorts an array of n values.
On average, about n2 elements are compared to find the lowest value in
each loop.
And Selection Sort must run the loop to find the lowest value

We get time complexity: O(n2⋅n)=O(n2)


approximately n times.

The time complexity for the Selection Sort algorithm can be displayed in a
graph like this:

As you can see, the run time is the same as for Bubble Sort: The run time
increases really fast when the size of the array is increased.
Insertion Sort with Python

Insertion Sort
The Insertion Sort algorithm uses one part of the array to hold the sorted
values, and the other part of the array to hold values that are not sorted yet.

Sort

The algorithm takes one value at a time from the unsorted part of the array
and puts it into the right place in the sorted part of the array, until the array
is sorted.

How it works:

1. Take the first value from the unsorted part of the array.
2. Move the value into the correct place in the sorted part of the array.
3. Go through the unsorted part of the array again as many times as there
are values.

REMOVE ADS

Manual Run Through


Before we implement the Insertion Sort algorithm in a Python program, let's
manually run through a short array, just to get the idea.

Step 1: We start with an unsorted array.

[ 7, 12, 9, 11, 3]
Step 2: We can consider the first value as the initial sorted part of the array.
If it is just one value, it must be sorted, right?

[ 7, 12, 9, 11, 3]

Step 3: The next value 12 should now be moved into the correct position in
the sorted part of the array. But 12 is higher than 7, so it is already in the
correct position.

[ 7, 12, 9, 11, 3]

Step 4: Consider the next value 9.

[ 7, 12, 9, 11, 3]

Step 5: The value 9 must now be moved into the correct position inside the
sorted part of the array, so we move 9 in between 7 and 12.

[ 7, 9, 12, 11, 3]

Step 6: The next value is 11.

[ 7, 9, 12, > 11, 3]

Step 7: We move it in between 9 and 12 in the sorted part of the array.

[ 7, 9, 11, 12, 3]

Step 8: The last value to insert into the correct position is 3.

[ 7, 9, 11, 12, 3]
Step 9: We insert 3 in front of all other values because it is the lowest value.

[ 3,7, 9, 11, 12]

Finally, the array is sorted.

Run the simulation below to see the steps above animated:

Insertion Sort
[
7,
12,
9,
11,
3
]

Implement Insertion Sort in Python


To implement the Insertion Sort algorithm in a Python program, we need:

1. An array with values to sort.


2. An outer loop that picks a value to be sorted. For an array
with n values, this outer loop skips the first value, and must
run n−1 times.
3. An inner loop that goes through the sorted part of the array, to find
where to insert the value. If the value to be sorted is at index i, the
sorted part of the array starts at index 0 and ends at index i−1.

The resulting code looks like this:

ExampleGet your own Python Server


Using the Insertion Sort on a Python list:

mylist = [64, 34, 25, 12, 22, 11, 90, 5]

n = len(mylist)
for i in range(1,n):
insert_index = i
current_value = [Link](i)
for j in range(i-1, -1, -1):
if mylist[j] > current_value:
insert_index = j
[Link](insert_index, current_value)

print(mylist)

Insertion Sort Improvement


Insertion Sort can be improved a little bit more.

The way the code above first removes a value and then inserts it somewhere
else is intuitive. It is how you would do Insertion Sort physically with a hand
of cards for example. If low value cards are sorted to the left, you pick up a
new unsorted card, and insert it in the correct place between the other
already sorted cards.

The problem with this way of programming it is that when removing a value
from the array, all elements above must be shifted one index place down:

And when inserting the removed value into the array again, there are also
many shift operations that must be done: all following elements must shift
one position up to make place for the inserted value:
These shifting operations can take a lot of time, especially for an array with
many elements.

Hidden memory shifts: You will not see these shifting operations
happening in the code if you are using a high-level programming language
such as Python or JavaScript, but the shifting operations are still happening
in the background. Such shifting operations require extra time for the
computer to do, which can be a problem.

You can read more about how arrays are stored in memory here.

REMOVE ADS

Improved Solution
We can avoid most of these shift operations by only shifting the values
necessary:

In the image above, first value 7 is copied, then values 11 and 12 are shifted
one place up in the array, and at last value 7 is put where value 11 was
before.

The number of shifting operations is reduced from 12 to 2 in this case.

This improvement is implemented in the example below:

Example
Insert the improvements in the sorting algorithm:

mylist = [64, 34, 25, 12, 22, 11, 90, 5]

n = len(mylist)
for i in range(1,n):
insert_index = i
current_value = mylist[i]
for j in range(i-1, -1, -1):
if mylist[j] > current_value:
mylist[j+1] = mylist[j]
insert_index = j
else:
break
mylist[insert_index] = current_value

print(mylist)

What is also done in the code above is to break out of the inner loop. That is
because there is no need to continue comparing values when we have
already found the correct place for the current value.

Insertion Sort Time Complexity


Insertion Sort sorts an array of n values.
On average, each value must be compared to about n2 other values to find
the correct place to insert it.
Insertion Sort must run the loop to insert a value in its correct place

We get time complexity for Insertion Sort: O(n2⋅n)=O(n2)


approximately n times.

The time complexity for Insertion Sort can be displayed like this:
For Insertion Sort, there is a big difference between best, average and worst
case scenarios.

Next up is Quicksort. Finally we will see a faster sorting algorithm!

DSA Quicksort with Python

Quicksort
As the name suggests, Quicksort is one of the fastest sorting algorithms.

The Quicksort algorithm takes an array of values, chooses one of the values
as the 'pivot' element, and moves the other values so that lower values are
on the left of the pivot element, and higher values are on the right of it.

Sort

In this tutorial the last element of the array is chosen to be the pivot
element, but we could also have chosen the first element of the array, or any
element in the array really.

Then, the Quicksort algorithm does the same operation recursively on the
sub-arrays to the left and right side of the pivot element. This continues until
the array is sorted.

Recursion is when a function calls itself.

After the Quicksort algorithm has put the pivot element in between a sub-
array with lower values on the left side, and a sub-array with higher values
on the right side, the algorithm calls itself twice, so that Quicksort runs again
for the sub-array on the left side, and for the sub-array on the right side. The
Quicksort algorithm continues to call itself until the sub-arrays are too small
to be sorted.

The algorithm can be described like this:

How it works:

1. Choose a value in the array to be the pivot element.


2. Order the rest of the array so that lower values than the pivot element are
on the left, and higher values are on the right.
3. Swap the pivot element with the first element of the higher values so that
the pivot element lands in between the lower and higher values.
4. Do the same operations (recursively) for the sub-arrays on the left and
right side of the pivot element.

REMOVE ADS

Manual Run Through


Before we implement the Quicksort algorithm in a programming language,
let's manually run through a short array, just to get the idea.

Step 1: We start with an unsorted array.

[ 11, 9, 12, 7, 3]

Step 2: We choose the last value 3 as the pivot element.

[ 11, 9, 12, 7, 3]

Step 3: The rest of the values in the array are all greater than 3, and must
be on the right side of 3. Swap 3 with 11.

[ 3, 9, 12, 7, 11]

Step 4: Value 3 is now in the correct position. We need to sort the values to
the right of 3. We choose the last value 11 as the new pivot element.

[ 3, 9, 12, 7, 11]
Step 5: The value 7 must be to the left of pivot value 11, and 12 must be to
the right of it. Move 7 and 12.

[ 3, 9, 7, 12, 11]

Step 6: Swap 11 with 12 so that lower values 9 and 7 are on the left side of
11, and 12 is on the right side.

[ 3, 9, 7, 11, 12]

Step 7: 11 and 12 are in the correct positions. We choose 7 as the pivot


element in sub-array [ 9, 7], to the left of 11.

[ 3, 9, 7, 11, 12]

Step 8: We must swap 9 with 7.

[ 3, 7, 9, 11, 12]

And now, the array is sorted.

Run the simulation below to see the steps above animated:

Sort
[
11,
9,
12,
7,
3
]

Implement Quicksort in Python


To write a 'quickSort' method that splits the array into shorter and shorter
sub-arrays we use recursion. This means that the 'quickSort' method must
call itself with the new sub-arrays to the left and right of the pivot element.
Read more about recursion here.

To implement the Quicksort algorithm in a Python program, we need:

1. An array with values to sort.


2. A quickSort method that calls itself (recursion) if the sub-array has a
size larger than 1.
3. A partition method that receives a sub-array, moves values around,
swaps the pivot element into the sub-array and returns the index
where the next split in sub-arrays happens.

The resulting code looks like this:

ExampleGet your own Python Server


Using the Quicksort algorithm in a Python program:

def partition(array, low, high):


pivot = array[high]
i = low - 1

for j in range(low, high):


if array[j] <= pivot:
i += 1
array[i], array[j] = array[j], array[i]

array[i+1], array[high] = array[high], array[i+1]


return i+1

def quicksort(array, low=0, high=None):


if high is None:
high = len(array) - 1

if low < high:


pivot_index = partition(array, low, high)
quicksort(array, low, pivot_index-1)
quicksort(array, pivot_index+1, high)

mylist = [64, 34, 25, 5, 22, 11, 90, 12]


quicksort(mylist)
print(mylist)
Quicksort Time Complexity
The worst case scenario for Quicksort is O(n2). This is when the pivot
element is either the highest or lowest value in every sub-array, which leads
to a lot of recursive calls. With our implementation above, this happens when
the array is already sorted.
But on average, the time complexity for Quicksort is actually just O(nlogn),
which is a lot better than for the previous sorting algorithms we have looked
at. That is why Quicksort is so popular.
Below you can see the significant improvement in time complexity for
Quicksort in an average scenario O(nlogn), compared to the previous
sorting algorithms Bubble, Selection and Insertion Sort with time
complexity O(n2):

The recursion part of the Quicksort algorithm is actually a reason why the
average sorting scenario is so fast, because for good picks of the pivot
element, the array will be split in half somewhat evenly each time the
algorithm calls itself. So the number of recursive calls do not double, even if
the number of values n double.

DSA Counting Sort with


Python
Counting Sort
The Counting Sort algorithm sorts an array by counting the number of times
each value occurs.

Sort
0
1
0
2
0
3
0
4
0
5

Run the simulation to see how 17 integer values from 1 till 5 are sorted using
Counting Sort.

Counting Sort does not compare values like the previous sorting algorithms
we have looked at, and only works on non negative integers.

Furthermore, Counting Sort is fast when the range of possible values k is


smaller than the number of values n.
How it works:

1. Create a new array for counting how many there are of the different
values.
2. Go through the array that needs to be sorted.
3. For each value, count it by increasing the counting array at the
corresponding index.
4. After counting the values, go through the counting array to create the
sorted array.
5. For each count in the counting array, create the correct number of
elements, with values that correspond to the counting array index.

REMOVE ADS
Conditions for Counting Sort
These are the reasons why Counting Sort is said to only work for a limited
range of non-negative integer values:

 Integer values: Counting Sort relies on counting occurrences of


distinct values, so they must be integers. With integers, each value fits
with an index (for non negative values), and there is a limited number
of different values, so that the number of possible different values k is
not too big compared to the number of values n.
 Non negative values: Counting Sort is usually implemented by
creating an array for counting. When the algorithm goes through the
values to be sorted, value x is counted by increasing the counting
array value at index x. If we tried sorting negative values, we would
get in trouble with sorting value -3, because index -3 would be outside
the counting array.
 Limited range of values: If the number of possible different values
to be sorted k is larger than the number of values to be sorted n, the
counting array we need for sorting will be larger than the original array
we have that needs sorting, and the algorithm becomes ineffective.

Manual Run Through


Before we implement the Counting Sort algorithm in a programming
language, let's manually run through a short array, just to get the idea.

Step 1: We start with an unsorted array.

myArray = [ 2, 3, 0, 2, 3, 2]

Step 2: We create another array for counting how many there are of each
value. The array has 4 elements, to hold values 0 through 3.

myArray = [ 2, 3, 0, 2, 3, 2]
countArray = [ 0, 0, 0, 0]
Step 3: Now let's start counting. The first element is 2, so we must
increment the counting array element at index 2.

myArray = [ 2, 3, 0, 2, 3, 2]
countArray = [ 0, 0, 1, 0]

Step 4: After counting a value, we can remove it, and count the next value,
which is 3.

myArray = [ 3, 0, 2, 3, 2]
countArray = [ 0, 0, 1, 1]

Step 5: The next value we count is 0, so we increment index 0 in the


counting array.

myArray = [ 0, 2, 3, 2]
countArray = [ 1, 0, 1, 1]

Step 6: We continue like this until all values are counted.

myArray = [ ]
countArray = [ 1, 0, 3, 2]

Step 7: Now we will recreate the elements from the initial array, and we will
do it so that the elements are ordered lowest to highest.

The first element in the counting array tells us that we have 1 element with
value 0. So we push 1 element with value 0 into the array, and we decrease
the element at index 0 in the counting array with 1.

myArray = [ 0]
countArray = [ 0, 0, 3, 2]

Step 8: From the counting array we see that we do not need to create any
elements with value 1.
myArray = [ 0]
countArray = [ 0, 0, 3, 2]

Step 9: We push 3 elements with value 2 into the end of the array. And as
we create these elements we also decrease the counting array at index 2.

myArray = [ 0, 2, 2, 2]
countArray = [ 0, 0, 0, 2]

Step 10: At last we must add 2 elements with value 3 at the end of the
array.

myArray = [0, 2, 2, 2, 3, 3]
countArray = [ 0, 0, 0, 0]

Finally! The array is sorted.

Run the simulation below to see the steps above animated:

Sort
myArray = [
2,
3,
0,
2,
3,
2
]

countArray = [
0,
0,
0,
0
]
Implement Counting Sort in Python
To implement the Counting Sort algorithm in a Python program, we need:

1. An array with values to sort.


2. A 'countingSort' method that receives an array of integers.
3. An array inside the method to keep count of the values.
4. A loop inside the method that counts and removes values, by
incrementing elements in the counting array.
5. A loop inside the method that recreates the array by using the
counting array, so that the elements appear in the right order.

One more thing: We need to find out what the highest value in the array is,
so that the counting array can be created with the correct size. For example,
if the highest value is 5, the counting array must be 6 elements in total, to
be able count all possible non negative integers 0, 1, 2, 3, 4 and 5.

The resulting code looks like this:

ExampleGet your own Python Server


Using the Counting Sort algorithm in a Python program:

def countingSort(arr):
max_val = max(arr)
count = [0] * (max_val + 1)

while len(arr) > 0:


num = [Link](0)
count[num] += 1

for i in range(len(count)):
while count[i] > 0:
[Link](i)
count[i] -= 1

return arr

mylist = [4, 2, 2, 6, 3, 3, 1, 6, 5, 2, 3]
mysortedlist = countingSort(mylist)
print(mysortedlist)
REMOVE ADS
Counting Sort Time Complexity
How fast the Counting Sort algorithm runs depends on both the range of
possible values k and the number of values n.
In general, time complexity for Counting Sort is O(n+k).
In a best case scenario, the range of possible different values k is very small
compared to the number of values n and Counting Sort has time
complexity O(n).
But in a worst case scenario, the range of possible different values k is very
big compared to the number of values n and Counting Sort can have time
complexity O(n2) or even worse.

The plot below shows how much the time complexity for Counting Sort can
vary.

As you can see, it is important to consider the range of values compared to


the number of values to be sorted before choosing Counting Sort as your
algorithm. Also, as mentioned at the top of the page, keep in mind that
Counting Sort only works for non negative integer values.

As mentioned previously: if the numbers to be sorted varies a lot in value


(large k), and there are few numbers to sort (small n), the Counting Sort
algorithm is not effective.
If we hold n and k fixed, the "Random", "Descending" and "Ascending"
alternatives in the simulation above results in the same number of
operations. This is because the same thing happens in all three cases: A
counting array is set up, the numbers are counted, and the new sorted array
is created.

DSA Radix Sort with Python

Radix Sort
The Radix Sort algorithm sorts an array by individual digits, starting with the
least significant digit (the rightmost one).

Click the button to do Radix Sort, one step (digit) at a time.

Step
490
369
504
185
583
385
348
204
515
198

The radix (or base) is the number of unique digits in a number system. In the
decimal system we normally use, there are 10 different digits from 0 till 9.

Radix Sort uses the radix so that decimal values are put into 10 different
buckets (or containers) corresponding to the digit that is in focus, then put
back into the array before moving on to the next digit.

Radix Sort is a non comparative algorithm that only works with non negative
integers.

The Radix Sort algorithm can be described like this:

How it works:
1. Start with the least significant digit (rightmost digit).
2. Sort the values based on the digit in focus by first putting the values in the
correct bucket based on the digit in focus, and then put them back into
array in the correct order.
3. Move to the next digit, and sort again, like in the step above, until there
are no digits left.

REMOVE ADS

Stable Sorting
Radix Sort must sort the elements in a stable way for the result to be sorted
correctly.

A stable sorting algorithm is an algorithm that keeps the order of elements


with the same value before and after the sorting. Let's say we have two
elements "K" and "L", where "K" comes before "L", and they both have value
"3". A sorting algorithm is considered stable if element "K" still comes before
"L" after the array is sorted.

It makes little sense to talk about stable sorting algorithms for the previous
algorithms we have looked at individually, because the result would be same
if they are stable or not. But it is important for Radix Sort that the the sorting
is done in a stable way because the elements are sorted by just one digit at
a time.

So after sorting the elements on the least significant digit and moving to the
next digit, it is important to not destroy the sorting work that has already
been done on the previous digit position, and that is why we need to take
care that Radix Sort does the sorting on each digit position in a stable way.

In the simulation below it is revealed how the underlying sorting into buckets
is done. And to get a better understanding of how stable sorting works, you
can also choose to sort in an unstable way, that will lead to an incorrect
result. The sorting is made unstable by simply putting elements into buckets
from the end of the array instead of from the start of the array.
Stable sort? Yes

Sort
0
1
2
3
4
5
6
7
8
9
141
115
372
453
272
323
348
301
292
473

Manual Run Through


Let's try to do the sorting manually, just to get an even better understanding
of how Radix Sort works before actually implementing it in a programming
language.

Step 1: We start with an unsorted array, and an empty array to fit values
with corresponding radices 0 till 9.

myArray = [ 33, 45, 40, 25, 17, 24]


radixArray = [ [], [], [], [], [], [], [], [], [], [] ]

Step 2: We start sorting by focusing on the least significant digit.


myArray = [ 33, 45, 40, 25, 17, 24]
radixArray = [ [], [], [], [], [], [], [], [], [], [] ]

Step 3: Now we move the elements into the correct positions in the radix
array according to the digit in focus. Elements are taken from the start of
myArray and pushed into the correct position in the radixArray.

myArray = [ ]
radixArray = [ [40], [], [], [33], [24], [45, 25], [], [17], [], []
]

Step 4: We move the elements back into the initial array, and the sorting is
now done for the least significant digit. Elements are taken from the end
radixArray, and put into the start of myArray.

myArray = [ 40, 33, 24, 45, 25, 17 ]


radixArray = [ [], [], [], [], [], [], [], [], [], [] ]

Step 5: We move focus to the next digit. Notice that values 45 and 25 are
still in the same order relative to each other as they were to start with,
because we sort in a stable way.

myArray = [ 40, 33, 24, 45, 25, 17 ]


radixArray = [ [], [], [], [], [], [], [], [], [], [] ]

Step 6: We move elements into the radix array according to the focused
digit.

myArray = [ ]
radixArray = [ [], [17], [24, 25], [33], [40, 45], [], [], [], [],
[] ]

Step 7: We move elements back into the start of myArray, from the back of
radixArray.
myArray = [ 17, 24, 25, 33, 40, 45 ]
radixArray = [ [], [], [], [], [], [], [], [], [], [] ]

The sorting is finished!

Run the simulation below to see the steps above animated:

Sort
myArray = [
33,
45,
40,
25,
17,
24
]

radixArray = [ [ ], [ ], [ ], [ ], [ ], [ ], [ ], [ ], [ ], [ ], [ ]
]

Implement Radix Sort in Python


To implement the Radix Sort algorithm we need:

1. An array with non negative integers that needs to be sorted.


2. A two dimensional array with index 0 to 9 to hold values with the
current radix in focus.
3. A loop that takes values from the unsorted array and places them in
the correct position in the two dimensional radix array.
4. A loop that puts values back into the initial array from the radix array.
5. An outer loop that runs as many times as there are digits in the highest
value.

The resulting code looks like this:

ExampleGet your own Python Server


Using the Radix Sort algorithm in a Python program:
mylist = [170, 45, 75, 90, 802, 24, 2, 66]
print("Original array:", mylist)
radixArray = [[], [], [], [], [], [], [], [], [], []]
maxVal = max(mylist)
exp = 1

while maxVal // exp > 0:

while len(mylist) > 0:


val = [Link]()
radixIndex = (val // exp) % 10
radixArray[radixIndex].append(val)

for bucket in radixArray:


while len(bucket) > 0:
val = [Link]()
[Link](val)

exp *= 10

print(mylist)

On line 7, we use floor division ("//") to divide the maximum value 802 by 1
the first time the while loop runs, the next time it is divided by 10, and the
last time it is divided by 100. When using floor division "//", any number
beyond the decimal point are disregarded, and an integer is returned.

On line 11, it is decided where to put a value in the radixArray based on its
radix, or digit in focus. For example, the second time the outer while loop
runs exp will be 10. Value 170 divided by 10 will be 17. The "%10" operation
divides by 10 and returns what is left. In this case 17 is divided by 10 one
time, and 7 is left. So value 170 is placed in index 7 in the radixArray.

REMOVE ADS

Radix Sort Using Other Sorting


Algorithms
Radix Sort can actually be implemented together with any other sorting
algorithm as long as it is stable. This means that when it comes down to
sorting on a specific digit, any stable sorting algorithm will work, such as
counting sort or bubble sort.
This is an implementation of Radix Sort that uses Bubble Sort to sort on the
individual digits:

Example
A Radix Sort algorithm that uses Bubble Sort:

def bubbleSort(arr):
n = len(arr)
for i in range(n):
for j in range(0, n - i - 1):
if arr[j] > arr[j + 1]:
arr[j], arr[j + 1] = arr[j + 1], arr[j]

def radixSortWithBubbleSort(arr):
max_val = max(arr)
exp = 1

while max_val // exp > 0:


radixList = [[],[],[],[],[],[],[],[],[],[]]

for num in arr:


radixIndex = (num // exp) % 10
radixList[radixIndex].append(num)

for bucket in radixList:


bubbleSort(bucket)

i = 0
for bucket in radixList:
for num in bucket:
arr[i] = num
i += 1

exp *= 10

mylist = [170, 45, 75, 90, 802, 24, 2, 66]

radixSortWithBubbleSort(mylist)

print(mylist)

Radix Sort Time Complexity


The time complexity for Radix Sort is: O(n⋅k)
This means that Radix Sort depends both on the values that need to be
sorted n, and the number of digits in the highest value k.
A best case scenario for Radix Sort is if there are lots of values to sort, but
the values have few digits. For example if there are more than a million
values to sort, and the highest value is 999, with just three digits. In such a
case the time complexity O(n⋅k) can be simplified to just O(n).
A worst case scenario for Radix Sort would be if there are as many digits in
the highest value as there are values to sort. This is perhaps not a common
scenario, but the time complexity would be O(n2)in this case.
The most average or common case is perhaps if the number of digits k is
something like k(n)=logn. If so, Radix Sort gets time complexity O(n⋅logn).
An example of such a case would be if there are 1000000 values to sort, and
the values have 6 digits.

See different possible time complexities for Radix Sort in the image below.

DSA Merge Sort with Python

Merge Sort
The Merge Sort algorithm is a divide-and-conquer algorithm that sorts an
array by first breaking it down into smaller arrays, and then building the
array back together the correct way so that it is sorted.

Sort

Divide: The algorithm starts with breaking up the array into smaller and
smaller pieces until one such sub-array only consists of one element.

Conquer: The algorithm merges the small pieces of the array back together
by putting the lowest values first, resulting in a sorted array.

The breaking down and building up of the array to sort the array is done
recursively.

In the animation above, each time the bars are pushed down represents a
recursive call, splitting the array into smaller pieces. When the bars are lifted
up, it means that two sub-arrays have been merged together.
The Merge Sort algorithm can be described like this:

How it works:

1. Divide the unsorted array into two sub-arrays, half the size of the original.
2. Continue to divide the sub-arrays as long as the current piece of the array
has more than one element.
3. Merge two sub-arrays together by always putting the lowest value first.
4. Keep merging until there are no sub-arrays left.

Take a look at the drawing below to see how Merge Sort works from a
different perspective. As you can see, the array is split into smaller and
smaller pieces until it is merged back together. And as the merging happens,
values from each sub-array are compared so that the lowest value comes
first.
REMOVE ADS

Manual Run Through


Let's try to do the sorting manually, just to get an even better understanding
of how Merge Sort works before actually implementing it in a Python
program.
Step 1: We start with an unsorted array, and we know that it splits in half
until the sub-arrays only consist of one element. The Merge Sort function
calls itself two times, once for each half of the array. That means that the
first sub-array will split into the smallest pieces first.

[ 12, 8, 9, 3, 11, 5, 4]
[ 12, 8, 9] [ 3, 11, 5, 4]
[ 12] [ 8, 9] [ 3, 11, 5, 4]
[ 12] [ 8] [ 9] [ 3, 11, 5, 4]

Step 2: The splitting of the first sub-array is finished, and now it is time to
merge. 8 and 9 are the first two elements to be merged. 8 is the lowest
value, so that comes before 9 in the first merged sub-array.

[ 12] [ 8, 9] [ 3, 11, 5, 4]

Step 3: The next sub-arrays to be merged is [ 12] and [ 8, 9]. Values in both
arrays are compared from the start. 8 is lower than 12, so 8 comes first, and
9 is also lower than 12.

[ 8, 9, 12] [ 3, 11, 5, 4]

Step 4: Now the second big sub-array is split recursively.

[ 8, 9, 12] [ 3, 11, 5, 4]
[ 8, 9, 12] [ 3, 11] [ 5, 4]
[ 8, 9, 12] [ 3] [ 11] [ 5, 4]

Step 5: 3 and 11 are merged back together in the same order as they are
shown because 3 is lower than 11.

[ 8, 9, 12] [ 3, 11] [ 5, 4]

Step 6: Sub-array with values 5 and 4 is split, then merged so that 4 comes
before 5.
[ 8, 9, 12] [ 3, 11] [ 5] [ 4]
[ 8, 9, 12] [ 3, 11] [ 4, 5]

Step 7: The two sub-arrays on the right are merged. Comparisons are done
to create elements in the new merged array:

1. 3 is lower than 4
2. 4 is lower than 11
3. 5 is lower than 11
4. 11 is the last remaining value
[ 8, 9, 12] [ 3, 4, 5, 11]

Step 8: The two last remaining sub-arrays are merged. Let's look at how the
comparisons are done in more detail to create the new merged and finished
sorted array:

3 is lower than 8:

Before [ 8, 9, 12] [ 3, 4, 5, 11]


After: [ 3, 8, 9, 12] [ 4, 5, 11]

Step 9: 4 is lower than 8:

Before [ 3, 8, 9, 12] [ 4, 5, 11]


After: [ 3, 4, 8, 9, 12] [ 5, 11]

Step 10: 5 is lower than 8:

Before [ 3, 4, 8, 9, 12] [ 5, 11]


After: [ 3, 4, 5, 8, 9, 12] [ 11]

Step 11: 8 and 9 are lower than 11:

Before [ 3, 4, 5, 8, 9, 12] [ 11]


After: [ 3, 4, 5, 8, 9, 12] [ 11]
Step 12: 11 is lower than 12:

Before [ 3, 4, 5, 8, 9, 12] [ 11]


After: [ 3, 4, 5, 8, 9, 11, 12]

The sorting is finished!

Run the simulation below to see the steps above animated:

Sort
[
12
,
8
,
9
,
3
,
11
,
5
,
4
]

Implement Merge Sort in Python


To implement the Merge Sort algorithm we need:

1. An array with values that needs to be sorted.


2. A function that takes an array, splits it in two, and calls itself with each
half of that array so that the arrays are split again and again
recursively, until a sub-array only consist of one value.
3. Another function that merges the sub-arrays back together in a sorted
way.

The resulting code looks like this:

ExampleGet your own Python Server


Implementing the Merge Sort algorithm in Python:

def mergeSort(arr):
if len(arr) <= 1:
return arr

mid = len(arr) // 2
leftHalf = arr[:mid]
rightHalf = arr[mid:]

sortedLeft = mergeSort(leftHalf)
sortedRight = mergeSort(rightHalf)

return merge(sortedLeft, sortedRight)

def merge(left, right):


result = []
i = j = 0

while i < len(left) and j < len(right):


if left[i] < right[j]:
[Link](left[i])
i += 1
else:
[Link](right[j])
j += 1

[Link](left[i:])
[Link](right[j:])

return result

mylist = [3, 7, 6, -10, 15, 23.5, 55, -13]


mysortedlist = mergeSort(mylist)
print("Sorted array:", mysortedlist)
On line 6, arr[:mid] takes all values from the array up until, but not
including, the value on index "mid".

On line 7, arr[mid:] takes all values from the array, starting at the value on
index "mid" and all the next values.

On lines 26-27, the first part of the merging is done. At this this point the
values of the two sub-arrays are compared, and either the left sub-array or
the right sub-array is empty, so the result array can just be filled with the
remaining values from either the left or the right sub-array. These lines can
be swapped, and the result will be the same.

Merge Sort without Recursion


Since Merge Sort is a divide and conquer algorithm, recursion is the most
intuitive code to use for implementation. The recursive implementation of
Merge Sort is also perhaps easier to understand, and uses less code lines in
general.

But Merge Sort can also be implemented without the use of recursion, so
that there is no function calling itself.

Take a look at the Merge Sort implementation below, that does not use
recursion:

Example
A Merge sort without recursion

def merge(left, right):


result = []
i = j = 0

while i < len(left) and j < len(right):


if left[i] < right[j]:
[Link](left[i])
i += 1
else:
[Link](right[j])
j += 1

[Link](left[i:])
[Link](right[j:])
return result

def mergeSort(arr):
step = 1 # Starting with sub-arrays of length 1
length = len(arr)

while step < length:


for i in range(0, length, 2 * step):
left = arr[i:i + step]
right = arr[i + step:i + 2 * step]

merged = merge(left, right)

# Place the merged array back into the original array


for j, val in enumerate(merged):
arr[i + j] = val

step *= 2 # Double the sub-array length for the next iteration

return arr

mylist = [3, 7, 6, -10, 15, 23.5, 55, -13]


mysortedlist = mergeSort(mylist)
print(mysortedlist)

You might notice that the merge functions are exactly the same in the two
Merge Sort implementations above, but in the implementation right above
here the while loop inside the mergeSort function is used to replace the
recursion. The while loop does the splitting and merging of the array in
place, and that makes the code a bit harder to understand.

To put it simply, the while loop inside the mergeSort function uses short step
lengths to sort tiny pieces (sub-arrays) of the initial array using the merge
function. Then the step length is increased to merge and sort larger pieces of
the array until the whole array is sorted.

REMOVE ADS

Merge Sort Time Complexity


The time complexity for Merge Sort is: O(n⋅logn)
And the time complexity is pretty much the same for different kinds of
arrays. The algorithm needs to split the array and merge it back together
whether it is already sorted or completely shuffled.

The image below shows the time complexity for Merge Sort.

Merge Sort performs almost the same every time because the array is split,
and merged using comparison, both if the array is already sorted or not.

Python MySQL

Python can be used in database applications.

One of the most popular databases is MySQL.

MySQL Database
To be able to experiment with the code examples in this tutorial, you should
have MySQL installed on your computer.

You can download a MySQL database at [Link]

Install MySQL Driver


Python needs a MySQL driver to access the MySQL database.

In this tutorial we will use the driver "MySQL Connector".

We recommend that you use PIP to install "MySQL Connector".

PIP is most likely already installed in your Python environment.

Navigate your command line to the location of PIP, and type the following:

Download and install "MySQL Connector":

C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>python -m pip install mysql-connector-python

Now you have downloaded and installed a MySQL driver.

Test MySQL Connector


To test if the installation was successful, or if you already have "MySQL
Connector" installed, create a Python page with the following content:

demo_mysql_test.py:

import [Link]

If the above code was executed with no errors, "MySQL Connector" is


installed and ready to be used.

Create Connection
Start by creating a connection to the database.

Use the username and password from your MySQL database:

demo_mysql_connection.py:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword"
)

print(mydb)

Now you can start querying the database using SQL statements.

Python MySQL Create


Database

Creating a Database
To create a database in MySQL, use the "CREATE DATABASE" statement:

ExampleGet your own Python Server


create a database named "mydatabase":

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword"
)

mycursor = [Link]()
[Link]("CREATE DATABASE mydatabase")

If the above code was executed with no errors, you have successfully
created a database.

Check if Database Exists


You can check if a database exist by listing all databases in your system by
using the "SHOW DATABASES" statement:

Example
Return a list of your system's databases:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword"
)

mycursor = [Link]()

[Link]("SHOW DATABASES")

for x in mycursor:
print(x)

Or you can try to access the database when making the connection:

Example
Try connecting to the database "mydatabase":

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

If the database does not exist, you will get an error.

Python MySQL Create Table

Creating a Table
To create a table in MySQL, use the "CREATE TABLE" statement.

Make sure you define the name of the database when you create the
connection

ExampleGet your own Python Server


Create a table named "customers":

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

[Link]("CREATE TABLE customers (name VARCHAR(255), address


VARCHAR(255))")

If the above code was executed with no errors, you have now successfully
created a table.

Check if Table Exists


You can check if a table exist by listing all tables in your database with the
"SHOW TABLES" statement:

Example
Return a list of your system's databases:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

[Link]("SHOW TABLES")

for x in mycursor:
print(x)

REMOVE ADS

Primary Key
When creating a table, you should also create a column with a unique key for
each record.

This can be done by defining a PRIMARY KEY.

We use the statement "INT AUTO_INCREMENT PRIMARY KEY" which will insert
a unique number for each record. Starting at 1, and increased by one for
each record.

Example
Create primary key when creating the table:
import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

[Link]("CREATE TABLE customers (id INT AUTO_INCREMENT


PRIMARY KEY, name VARCHAR(255), address VARCHAR(255))")

If the table already exists, use the ALTER TABLE keyword:

Example
Create primary key on an existing table:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

[Link]("ALTER TABLE customers ADD COLUMN id INT


AUTO_INCREMENT PRIMARY KEY")

Python MySQL Insert Into


Table

Insert Into Table


To fill a table in MySQL, use the "INSERT INTO" statement.
ExampleGet your own Python Server
Insert a record in the "customers" table:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "INSERT INTO customers (name, address) VALUES (%s, %s)"


val = ("John", "Highway 21")
[Link](sql, val)

[Link]()

print([Link], "record inserted.")

Important!: Notice the statement: [Link](). It is required to make the


changes, otherwise no changes are made to the table.

Insert Multiple Rows


To insert multiple rows into a table, use the executemany() method.

The second parameter of the executemany() method is a list of tuples,


containing the data you want to insert:

Example
Fill the "customers" table with data:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "INSERT INTO customers (name, address) VALUES (%s, %s)"


val = [
('Peter', 'Lowstreet 4'),
('Amy', 'Apple st 652'),
('Hannah', 'Mountain 21'),
('Michael', 'Valley 345'),
('Sandy', 'Ocean blvd 2'),
('Betty', 'Green Grass 1'),
('Richard', 'Sky st 331'),
('Susan', 'One way 98'),
('Vicky', 'Yellow Garden 2'),
('Ben', 'Park Lane 38'),
('William', 'Central st 954'),
('Chuck', 'Main Road 989'),
('Viola', 'Sideway 1633')
]

[Link](sql, val)

[Link]()

print([Link], "was inserted.")

REMOVE ADS

Get Inserted ID
You can get the id of the row you just inserted by asking the cursor object.

Note: If you insert more than one row, the id of the last inserted row is
returned.

Example
Insert one row, and return the ID:
import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "INSERT INTO customers (name, address) VALUES (%s, %s)"


val = ("Michelle", "Blue Village")
[Link](sql, val)

[Link]()

print("1 record inserted, ID:", [Link])

Python MySQL Select From

Select From a Table


To select from a table in MySQL, use the "SELECT" statement:

ExampleGet your own Python Server


Select all records from the "customers" table, and display the result:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

[Link]("SELECT * FROM customers")


myresult = [Link]()

for x in myresult:
print(x)

Note: We use the fetchall() method, which fetches all rows from the last
executed statement.

Selecting Columns
To select only some of the columns in a table, use the "SELECT" statement
followed by the column name(s):

Example
Select only the name and address columns:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

[Link]("SELECT name, address FROM customers")

myresult = [Link]()

for x in myresult:
print(x)

REMOVE ADS
Using the fetchone() Method
If you are only interested in one row, you can use the fetchone() method.

The fetchone() method will return the first row of the result:

Example
Fetch only one row:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

[Link]("SELECT * FROM customers")

myresult = [Link]()

print(myresult)

Python MySQL Where

Select With a Filter


When selecting records from a table, you can filter the selection by using the
"WHERE" statement:

ExampleGet your own Python Server


Select record(s) where the address is "Park Lane 38": result:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "SELECT * FROM customers WHERE address ='Park Lane 38'"

[Link](sql)

myresult = [Link]()

for x in myresult:
print(x)

Wildcard Characters
You can also select the records that starts, includes, or ends with a given
letter or phrase.

Use the % to represent wildcard characters:

Example
Select records where the address contains the word "way":

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "SELECT * FROM customers WHERE address LIKE '%way%'"

[Link](sql)
myresult = [Link]()

for x in myresult:
print(x)

REMOVE ADS

Prevent SQL Injection


When query values are provided by the user, you should escape the values.

This is to prevent SQL injections, which is a common web hacking technique


to destroy or misuse your database.

The [Link] module has methods to escape query values:

Example
Escape query values by using the placholder %s method:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "SELECT * FROM customers WHERE address = %s"


adr = ("Yellow Garden 2", )

[Link](sql, adr)

myresult = [Link]()

for x in myresult:
print(x)
Python MySQL Order By

Sort the Result


Use the ORDER BY statement to sort the result in ascending or descending
order.

The ORDER BY keyword sorts the result ascending by default. To sort the
result in descending order, use the DESC keyword.

ExampleGet your own Python Server


Sort the result alphabetically by name: result:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "SELECT * FROM customers ORDER BY name"

[Link](sql)

myresult = [Link]()

for x in myresult:
print(x)

ORDER BY DESC
Use the DESC keyword to sort the result in a descending order.
Example
Sort the result reverse alphabetically by name:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "SELECT * FROM customers ORDER BY name DESC"

[Link](sql)

myresult = [Link]()

for x in myresult:
print(x)

Python MySQL Delete From


By

Delete Record
You can delete records from an existing table by using the "DELETE FROM"
statement:

ExampleGet your own Python Server


Delete any record where the address is "Mountain 21":

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "DELETE FROM customers WHERE address = 'Mountain 21'"

[Link](sql)

[Link]()

print([Link], "record(s) deleted")

Important!: Notice the statement: [Link](). It is required to make the


changes, otherwise no changes are made to the table.
Notice the WHERE clause in the DELETE syntax: The WHERE clause
specifies which record(s) that should be deleted. If you omit the WHERE
clause, all records will be deleted!

REMOVE ADS

Prevent SQL Injection


It is considered a good practice to escape the values of any query, also in
delete statements.

This is to prevent SQL injections, which is a common web hacking technique


to destroy or misuse your database.

The [Link] module uses the placeholder %s to escape values in the


delete statement:

Example
Escape values by using the placeholder %s method:

import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "DELETE FROM customers WHERE address = %s"


adr = ("Yellow Garden 2", )

[Link](sql, adr)

[Link]()

print([Link], "record(s) deleted")

Python MySQL Drop Table

Delete a Table
You can delete an existing table by using the "DROP TABLE" statement:

ExampleGet your own Python Server


Delete the table "customers":

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "DROP TABLE customers"

[Link](sql)
Drop Only if Exist
If the table you want to delete is already deleted, or for any other reason
does not exist, you can use the IF EXISTS keyword to avoid getting an error.

Example
Delete the table "customers" if it exists:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "DROP TABLE IF EXISTS customers"

[Link](sql)

Python MySQL Update Table


Update Table
You can update existing records in a table by using the "UPDATE" statement:

ExampleGet your own Python Server


Overwrite the address column from "Valley 345" to "Canyon 123":

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "UPDATE customers SET address = 'Canyon 123' WHERE address =


'Valley 345'"

[Link](sql)

[Link]()

print([Link], "record(s) affected")


Important!: Notice the statement: [Link](). It is required to make the
changes, otherwise no changes are made to the table.
Notice the WHERE clause in the UPDATE syntax: The WHERE clause
specifies which record or records that should be updated. If you omit the
WHERE clause, all records will be updated!

Prevent SQL Injection


It is considered a good practice to escape the values of any query, also in
update statements.

This is to prevent SQL injections, which is a common web hacking technique


to destroy or misuse your database.

The [Link] module uses the placeholder %s to escape values in the


update statement:

Example
Escape values by using the placeholder %s method:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "UPDATE customers SET address = %s WHERE address = %s"


val = ("Valley 345", "Canyon 123")

[Link](sql, val)

[Link]()

print([Link], "record(s) affected")

Python MySQL Limit


Limit the Result
You can limit the number of records returned from the query, by using the
"LIMIT" statement:

ExampleGet your own Python Server


Select the 5 first records in the "customers" table:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

[Link]("SELECT * FROM customers LIMIT 5")

myresult = [Link]()

for x in myresult:
print(x)

Start From Another Position


If you want to return five records, starting from the third record, you can use
the "OFFSET" keyword:
Example
Start from position 3, and return 5 records:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

[Link]("SELECT * FROM customers LIMIT 5 OFFSET 2")

myresult = [Link]()

for x in myresult:
print(x)

Python MySQL Join


Join Two or More Tables
You can combine rows from two or more tables, based on a related column
between them, by using a JOIN statement.

Consider you have a "users" table and a "products" table:

usersGet your own Python Server


{ id: 1, name: 'John', fav: 154},
{ id: 2, name: 'Peter', fav: 154},
{ id: 3, name: 'Amy', fav: 155},
{ id: 4, name: 'Hannah', fav:},
{ id: 5, name: 'Michael', fav:}

products
{ id: 154, name: 'Chocolate Heaven' },
{ id: 155, name: 'Tasty Lemons' },
{ id: 156, name: 'Vanilla Dreams' }

These two tables can be combined by using users' fav field and
products' id field.

Example
Join users and products to see the name of the users favorite product:

import [Link]

mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)

mycursor = [Link]()

sql = "SELECT \
[Link] AS user, \
[Link] AS favorite \
FROM users \
INNER JOIN products ON [Link] = [Link]"

[Link](sql)

myresult = [Link]()

for x in myresult:
print(x)
Note: You can use JOIN instead of INNER JOIN. They will both give you the
same result.

LEFT JOIN
In the example above, Hannah, and Michael were excluded from the result,
that is because INNER JOIN only shows the records where there is a match.

If you want to show all users, even if they do not have a favorite product, use
the LEFT JOIN statement:
Example
Select all users and their favorite product:

sql = "SELECT \
[Link] AS user, \
[Link] AS favorite \
FROM users \
LEFT JOIN products ON [Link] = [Link]"

RIGHT JOIN
If you want to return all products, and the users who have them as their
favorite, even if no user have them as their favorite, use the RIGHT JOIN
statement:

Example
Select all products, and the user(s) who have them as their favorite:

sql = "SELECT \
[Link] AS user, \
[Link] AS favorite \
FROM users \
RIGHT JOIN products ON [Link] = [Link]"
Note: Hannah and Michael, who have no favorite product, are not included
in the result.

Python MongoDB
Python can be used in database applications.

One of the most popular NoSQL database is MongoDB.

MongoDB
MongoDB stores data in JSON-like documents, which makes the database
very flexible and scalable.

To be able to experiment with the code examples in this tutorial, you will
need access to a MongoDB database.

You can download a free MongoDB database at [Link]

Or get started right away with a MongoDB cloud service


at [Link]

PyMongo
Python needs a MongoDB driver to access the MongoDB database.

In this tutorial we will use the MongoDB driver "PyMongo".

We recommend that you use PIP to install "PyMongo".

PIP is most likely already installed in your Python environment.

Navigate your command line to the location of PIP, and type the following:

Download and install "PyMongo":

C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>python -m pip install pymongo

Now you have downloaded and installed a mongoDB driver.

Test PyMongo
To test if the installation was successful, or if you already have "pymongo"
installed, create a Python page with the following content:

demo_mongodb_test.py:

import pymongo
If the above code was executed with no errors, "pymongo" is installed and
ready to be used.

Python MongoDB Create


Database
Creating a Database
To create a database in MongoDB, start by creating a MongoClient object,
then specify a connection URL with the correct ip address and the name of
the database you want to create.

MongoDB will create the database if it does not exist, and make a connection
to it.

ExampleGet your own Python Server


Create a database called "mydatabase":

import pymongo

myclient = [Link]("mongodb://localhost:27017/")

mydb = myclient["mydatabase"]

Important: In MongoDB, a database is not created until it gets content!

MongoDB waits until you have created a collection (table), with at least one
document (record) before it actually creates the database (and collection).

Check if Database Exists


Remember: In MongoDB, a database is not created until it gets content, so
if this is your first time creating a database, you should complete the next
two chapters (create collection and create document) before you check if the
database exists!

You can check if a database exist by listing all databases in you system:
Example
Return a list of your system's databases:

print(myclient.list_database_names())

Or you can check a specific database by name:

Example
Check if "mydatabase" exists:

dblist = myclient.list_database_names()
if "mydatabase" in dblist:
print("The database exists.")

Python MongoDB Create


Collection
A collection in MongoDB is the same as a table in SQL databases.

Creating a Collection
To create a collection in MongoDB, use database object and specify the
name of the collection you want to create.

MongoDB will create the collection if it does not exist.

ExampleGet your own Python Server


Create a collection called "customers":

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]

mycol = mydb["customers"]

Important: In MongoDB, a collection is not created until it gets content!


MongoDB waits until you have inserted a document before it actually creates
the collection.

Check if Collection Exists


Remember: In MongoDB, a collection is not created until it gets content, so
if this is your first time creating a collection, you should complete the next
chapter (create document) before you check if the collection exists!

You can check if a collection exist in a database by listing all collections:

Example
Return a list of all collections in your database:

print(mydb.list_collection_names())

Or you can check a specific collection by name:

Example
Check if the "customers" collection exists:

collist = mydb.list_collection_names()
if "customers" in collist:
print("The collection exists.")

Python MongoDB Insert


Document
A document in MongoDB is the same as a record in SQL databases.

Insert Into Collection


To insert a record, or document as it is called in MongoDB, into a collection,
we use the insert_one() method.

The first parameter of the insert_one() method is a dictionary containing


the name(s) and value(s) of each field in the document you want to insert.
ExampleGet your own Python Server
Insert a record in the "customers" collection:

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

mydict = { "name": "John", "address": "Highway 37" }

x = mycol.insert_one(mydict)

Return the _id Field


The insert_one() method returns a InsertOneResult object, which has a
property, inserted_id, that holds the id of the inserted document.

Example
Insert another record in the "customers" collection, and return the value of
the _id field:

mydict = { "name": "Peter", "address": "Lowstreet 27" }

x = mycol.insert_one(mydict)

print(x.inserted_id)

If you do not specify an _id field, then MongoDB will add one for you and
assign a unique id for each document.

In the example above no _id field was specified, so MongoDB assigned a


unique _id for the record (document).

Insert Multiple Documents


To insert multiple documents into a collection in MongoDB, we use
the insert_many() method.
The first parameter of the insert_many() method is a list containing
dictionaries with the data you want to insert:

Example
import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

mylist = [
{ "name": "Amy", "address": "Apple st 652"},
{ "name": "Hannah", "address": "Mountain 21"},
{ "name": "Michael", "address": "Valley 345"},
{ "name": "Sandy", "address": "Ocean blvd 2"},
{ "name": "Betty", "address": "Green Grass 1"},
{ "name": "Richard", "address": "Sky st 331"},
{ "name": "Susan", "address": "One way 98"},
{ "name": "Vicky", "address": "Yellow Garden 2"},
{ "name": "Ben", "address": "Park Lane 38"},
{ "name": "William", "address": "Central st 954"},
{ "name": "Chuck", "address": "Main Road 989"},
{ "name": "Viola", "address": "Sideway 1633"}
]

x = mycol.insert_many(mylist)

#print list of the _id values of the inserted documents:


print(x.inserted_ids)

The insert_many() method returns a InsertManyResult object, which has a


property, inserted_ids, that holds the ids of the inserted documents.

Insert Multiple Documents, with


Specified IDs
If you do not want MongoDB to assign unique ids for your document, you can
specify the _id field when you insert the document(s).

Remember that the values has to be unique. Two documents cannot have
the same _id.
Example
import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

mylist = [
{ "_id": 1, "name": "John", "address": "Highway 37"},
{ "_id": 2, "name": "Peter", "address": "Lowstreet 27"},
{ "_id": 3, "name": "Amy", "address": "Apple st 652"},
{ "_id": 4, "name": "Hannah", "address": "Mountain 21"},
{ "_id": 5, "name": "Michael", "address": "Valley 345"},
{ "_id": 6, "name": "Sandy", "address": "Ocean blvd 2"},
{ "_id": 7, "name": "Betty", "address": "Green Grass 1"},
{ "_id": 8, "name": "Richard", "address": "Sky st 331"},
{ "_id": 9, "name": "Susan", "address": "One way 98"},
{ "_id": 10, "name": "Vicky", "address": "Yellow Garden 2"},
{ "_id": 11, "name": "Ben", "address": "Park Lane 38"},
{ "_id": 12, "name": "William", "address": "Central st 954"},
{ "_id": 13, "name": "Chuck", "address": "Main Road 989"},
{ "_id": 14, "name": "Viola", "address": "Sideway 1633"}
]

x = mycol.insert_many(mylist)

#print list of the _id values of the inserted documents:


print(x.inserted_ids)

Python MongoDB Find


In MongoDB we use the find() and find_one() methods to find data in a
collection.

Just like the SELECT statement is used to find data in a table in a MySQL
database.

Find One
To select data from a collection in MongoDB, we can use
the find_one() method.

The find_one() method returns the first occurrence in the selection.


ExampleGet your own Python Server
Find the first document in the customers collection:

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

x = mycol.find_one()

print(x)

Find All
To select data from a table in MongoDB, we can also use the find() method.

The find() method returns all occurrences in the selection.

The first parameter of the find() method is a query object. In this example
we use an empty query object, which selects all documents in the collection.

No parameters in the find() method gives you the same result as SELECT
* in MySQL.

Example
Return all documents in the "customers" collection, and print each
document:

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

for x in [Link]():
print(x)
Return Only Some Fields
The second parameter of the find() method is an object describing which
fields to include in the result.

This parameter is optional, and if omitted, all fields will be included in the
result.

Example
Return only the names and addresses, not the _ids:

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

for x in [Link]({},{ "_id": 0, "name": 1, "address": 1 }):


print(x)
You are not allowed to specify both 0 and 1 values in the same object
(except if one of the fields is the _id field). If you specify a field with the value
0, all other fields get the value 1, and vice versa:

Example
This example will exclude "address" from the result:

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

for x in [Link]({},{ "address": 0 }):


print(x)

Example
You get an error if you specify both 0 and 1 values in the same object
(except if one of the fields is the _id field):

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

for x in [Link]({},{ "name": 1, "address": 0 }):


print(x)

Python MongoDB Query


Filter the Result
When finding documents in a collection, you can filter the result by using a
query object.

The first argument of the find() method is a query object, and is used to
limit the search.

ExampleGet your own Python Server


Find document(s) with the address "Park Lane 38":

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

myquery = { "address": "Park Lane 38" }

mydoc = [Link](myquery)

for x in mydoc:
print(x)

Advanced Query
To make advanced queries you can use modifiers as values in the query
object.
E.g. to find the documents where the "address" field starts with the letter "S"
or higher (alphabetically), use the greater than modifier: {"$gt": "S"}:

Example
Find documents where the address starts with the letter "S" or higher:

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

myquery = { "address": { "$gt": "S" } }

mydoc = [Link](myquery)

for x in mydoc:
print(x)

Filter With Regular Expressions


You can also use regular expressions as a modifier.

Regular expressions can only be used to query strings.

To find only the documents where the "address" field starts with the letter
"S", use the regular expression {"$regex": "^S"}:

Example
Find documents where the address starts with the letter "S":

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

myquery = { "address": { "$regex": "^S" } }

mydoc = [Link](myquery)
for x in mydoc:
print(x)

Python MongoDB Sort


Sort the Result
Use the sort() method to sort the result in ascending or descending order.

The sort() method takes one parameter for "fieldname" and one parameter
for "direction" (ascending is the default direction).

ExampleGet your own Python Server


Sort the result alphabetically by name:

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

mydoc = [Link]().sort("name")

for x in mydoc:
print(x)

Sort Descending
Use the value -1 as the second parameter to sort descending.

sort("name", 1) #ascending
sort("name", -1) #descending

Example
Sort the result reverse alphabetically by name:

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

mydoc = [Link]().sort("name", -1)

for x in mydoc:
print(x)

Python MongoDB Delete


Document
Delete Document
To delete one document, we use the delete_one() method.

The first parameter of the delete_one() method is a query object defining


which document to delete.

Note: If the query finds more than one document, only the first occurrence
is deleted.

ExampleGet your own Python Server


Delete the document with the address "Mountain 21":

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

myquery = { "address": "Mountain 21" }

mycol.delete_one(myquery)

Delete Many Documents


To delete more than one document, use the delete_many() method.
The first parameter of the delete_many() method is a query object defining
which documents to delete.

Example
Delete all documents were the address starts with the letter S:

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

myquery = { "address": {"$regex": "^S"} }

x = mycol.delete_many(myquery)

print(x.deleted_count, " documents deleted.")

Delete All Documents in a Collection


To delete all documents in a collection, pass an empty query object to
the delete_many() method:

Example
Delete all documents in the "customers" collection:

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

x = mycol.delete_many({})

print(x.deleted_count, " documents deleted.")

Python MongoDB Drop


Collection
Delete Collection
You can delete a table, or collection as it is called in MongoDB, by using
the drop() method.

ExampleGet your own Python Server


Delete the "customers" collection:

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

[Link]()

The drop() method returns true if the collection was dropped successfully,
and false if the collection does not exist.

Python MongoDB Update


Update Collection
You can update a record, or document as it is called in MongoDB, by using
the update_one() method.

The first parameter of the update_one() method is a query object defining


which document to update.

Note: If the query finds more than one record, only the first occurrence is
updated.

The second parameter is an object defining the new values of the document.

ExampleGet your own Python Server


Change the address from "Valley 345" to "Canyon 123":

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

myquery = { "address": "Valley 345" }


newvalues = { "$set": { "address": "Canyon 123" } }

mycol.update_one(myquery, newvalues)

#print "customers" after the update:


for x in [Link]():
print(x)

Update Many
To update all documents that meets the criteria of the query, use
the update_many() method.

Example
Update all documents where the address starts with the letter "S":

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

myquery = { "address": { "$regex": "^S" } }


newvalues = { "$set": { "name": "Minnie" } }

x = mycol.update_many(myquery, newvalues)

print(x.modified_count, "documents updated.")

Python MongoDB Limit


Limit the Result
To limit the result in MongoDB, we use the limit() method.

The limit() method takes one parameter, a number defining how many
documents to return.
Consider you have a "customers" collection:

CustomersGet your own Python Server


{'_id': 1, 'name': 'John', 'address': 'Highway37'}
{'_id': 2, 'name': 'Peter', 'address': 'Lowstreet 27'}
{'_id': 3, 'name': 'Amy', 'address': 'Apple st 652'}
{'_id': 4, 'name': 'Hannah', 'address': 'Mountain 21'}
{'_id': 5, 'name': 'Michael', 'address': 'Valley 345'}
{'_id': 6, 'name': 'Sandy', 'address': 'Ocean blvd 2'}
{'_id': 7, 'name': 'Betty', 'address': 'Green Grass 1'}
{'_id': 8, 'name': 'Richard', 'address': 'Sky st 331'}
{'_id': 9, 'name': 'Susan', 'address': 'One way 98'}
{'_id': 10, 'name': 'Vicky', 'address': 'Yellow Garden 2'}
{'_id': 11, 'name': 'Ben', 'address': 'Park Lane 38'}
{'_id': 12, 'name': 'William', 'address': 'Central st 954'}
{'_id': 13, 'name': 'Chuck', 'address': 'Main Road 989'}
{'_id': 14, 'name': 'Viola', 'address': 'Sideway 1633'}

Example
Limit the result to only return 5 documents:

import pymongo

myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]

myresult = [Link]().limit(5)

#print the result:


for x in myresult:
print(x)

Python Built in Functions

Python has a set of built-in functions.


Function Description

abs() Returns the absolute value of a number

all() Returns True if all items in an iterable object are true

any() Returns True if any item in an iterable object is true

ascii() Returns a readable version of an object. Replaces none-ascii characte

bin() Returns the binary version of a number

bool() Returns the boolean value of the specified object

bytearray() Returns an array of bytes

bytes() Returns a bytes object

callable() Returns True if the specified object is callable, otherwise False

chr() Returns a character from the specified Unicode code.


classmethod() Converts a method into a class method

compile() Returns the specified source as an object, ready to be executed

complex() Returns a complex number

delattr() Deletes the specified attribute (property or method) from the specifie

dict() Returns a dictionary (Array)

dir() Returns a list of the specified object's properties and methods

divmod() Returns the quotient and the remainder when argument1 is divided b

enumerate() Takes a collection (e.g. a tuple) and returns it as an enumerate objec

eval() Evaluates and executes an expression

exec() Executes the specified code (or object)

filter() Use a filter function to exclude items in an iterable object


float() Returns a floating point number

format() Formats a specified value

frozenset() Returns a frozenset object

getattr() Returns the value of the specified attribute (property or method)

globals() Returns the current global symbol table as a dictionary

hasattr() Returns True if the specified object has the specified attribute (prope

hash() Returns the hash value of a specified object

help() Executes the built-in help system

hex() Converts a number into a hexadecimal value

id() Returns the id of an object

input() Allowing user input


int() Returns an integer number

isinstance() Returns True if a specified object is an instance of a specified object

issubclass() Returns True if a specified class is a subclass of a specified object

iter() Returns an iterator object

len() Returns the length of an object

list() Returns a list

locals() Returns an updated dictionary of the current local symbol table

map() Returns the specified iterator with the specified function applied to e

max() Returns the largest item in an iterable

memoryview() Returns a memory view object

min() Returns the smallest item in an iterable


next() Returns the next item in an iterable

object() Returns a new object

oct() Converts a number into an octal

open() Opens a file and returns a file object

ord() Convert an integer representing the Unicode of the specified charact

pow() Returns the value of x to the power of y

print() Prints to the standard output device

property() Gets, sets, deletes a property

range() Returns a sequence of numbers, starting from 0 and increments by 1

repr() Returns a readable version of an object

reversed() Returns a reversed iterator


round() Rounds a numbers

set() Returns a new set object

setattr() Sets an attribute (property/method) of an object

slice() Returns a slice object

sorted() Returns a sorted list

staticmethod() Converts a method into a static method

str() Returns a string object

sum() Sums the items of an iterator

super() Returns an object that represents the parent class

tuple() Returns a tuple

type() Returns the type of an object


vars() Returns the __dict__ property of an object

zip() Returns an iterator, from two or more iterators

Python String Methods

Python has a set of built-in methods that you can use on strings.

Note: All string methods returns new values. They do not change the
original string.

Method Description

capitalize() Converts the first character to upper case

casefold() Converts string into lower case

center() Returns a centered string

count() Returns the number of times a specified value occurs in a string

encode() Returns an encoded version of the string

endswith() Returns true if the string ends with the specified value
expandtabs() Sets the tab size of the string

find() Searches the string for a specified value and returns the position of w

format() Formats specified values in a string

format_map() Formats specified values from a dictionary in a string

index() Searches the string for a specified value and returns the position of w

isalnum() Returns True if all characters in the string are alphanumeric

isalpha() Returns True if all characters in the string are in the alphabet

isascii() Returns True if all characters in the string are ascii characters

isdecimal() Returns True if all characters in the string are decimals

isdigit() Returns True if all characters in the string are digits


isidentifier() Returns True if the string is an identifier

islower() Returns True if all characters in the string are lower case

isnumeric() Returns True if all characters in the string are numeric

isprintable() Returns True if all characters in the string are printable

isspace() Returns True if all characters in the string are whitespaces

istitle() Returns True if the string follows the rules of a title

isupper() Returns True if all characters in the string are upper case

join() Converts the elements of an iterable into a string

ljust() Returns a left justified version of the string

lower() Converts a string into lower case

lstrip() Returns a left trim version of the string


maketrans() Returns a translation table to be used in translations

partition() Returns a tuple where the string is parted into three parts

replace() Returns a string where a specified value is replaced with a specified v

rfind() Searches the string for a specified value and returns the last position

rindex() Searches the string for a specified value and returns the last position

rjust() Returns a right justified version of the string

rpartition() Returns a tuple where the string is parted into three parts

rsplit() Splits the string at the specified separator, and returns a list

rstrip() Returns a right trim version of the string

split() Splits the string at the specified separator, and returns a list

splitlines() Splits the string at line breaks and returns a list


startswith() Returns true if the string starts with the specified value

strip() Returns a trimmed version of the string

swapcase() Swaps cases, lower case becomes upper case and vice versa

title() Converts the first character of each word to upper case

translate() Returns a translated string

upper() Converts a string into upper case

zfill() Fills the string with a specified number of 0 values at the beginning

Note: All string methods returns new values. They do not change the
original string.

Python List/Array Methods

Python has a set of built-in methods that you can use on lists/arrays.

Method Description

append() Adds an element at the end of the list


clear() Removes all the elements from the list

copy() Returns a copy of the list

count() Returns the number of elements with the specified value

extend() Add the elements of a list (or any iterable), to the end of the current lis

index() Returns the index of the first element with the specified value

insert() Adds an element at the specified position

pop() Removes the element at the specified position

remove() Removes the first item with the specified value

reverse() Reverses the order of the list

sort() Sorts the list

Note: Python does not have built-in support for Arrays, but Python Lists can
be used instead.
Python Dictionary Methods

Python has a set of built-in methods that you can use on dictionaries.

Method Description

clear() Removes all the elements from the dictionary

copy() Returns a copy of the dictionary

fromkeys() Returns a dictionary with the specified keys and value

get() Returns the value of the specified key

items() Returns a list containing a tuple for each key value pair

keys() Returns a list containing the dictionary's keys

pop() Removes the element with the specified key

popitem() Removes the last inserted key-value pair


setdefault() Returns the value of the specified key. If the key does not exist: insert the

update() Updates the dictionary with the specified key-value pairs

values() Returns a list of all the values in the dictionary

Python Tuple Methods

Python has two built-in methods that you can use on tuples.

Method Description

count() Returns the number of times a specified value occurs in a tuple

index() Searches the tuple for a specified value and returns the position

Python Set Methods

Python has a set of built-in methods that you can use on sets.

Method Shortcu Description


t
add() Adds an element to the set

clear() Removes all the elements from the set

copy() Returns a copy of the set

difference() - Returns a set containing the difference bet

difference_update() -= Removes the items in this set that are also

discard() Remove the specified item

intersection() & Returns a set, that is the intersection of tw

intersection_update() &= Removes the items in this set that are not

isdisjoint() Returns whether two sets have a intersect

issubset() <= Returns True if all items of this set is prese

< Returns True if all items of this set is prese


issuperset() >= Returns True if all items of another set is p

> Returns True if all items of another, smalle

pop() Removes an element from the set

remove() Removes the specified element

symmetric_difference() ^ Returns a set with the symmetric differenc

symmetric_difference_update() ^= Inserts the symmetric differences from this

union() | Return a set containing the union of sets

update() |= Update the set with the union of this set an

Python File Methods

Python has a set of methods available for the file object.

Method Description
close() Closes the file

detach() Returns the separated raw stream from the buffer

fileno() Returns a number that represents the stream, from the operating syste

flush() Flushes the internal buffer

isatty() Returns whether the file stream is interactive or not

read() Returns the file content

readable() Returns whether the file stream can be read or not

readline() Returns one line from the file

readlines() Returns a list of lines from the file

seek() Change the file position

seekable() Returns whether the file allows us to change the file position
tell() Returns the current file position

truncate() Resizes the file to a specified size

writable() Returns whether the file can be written to or not

write() Writes the specified string to the file

writelines() Writes a list of strings to the file

Python Keywords

Python has a set of keywords that are reserved words that cannot be used as
variable names, function names, or any other identifiers:

Keyword Description

and A logical operator

as To create an alias

assert For debugging


async Define an asynchronous function

await Wait for and get a result from an awaitable

break To break out of a loop

case Pattern in a match statement

class To define a class

continue To continue to the next iteration of a loop

def To define a function

del To delete an object

elif Used in conditional statements, same as else if

else Used in conditional statements

except Used with exceptions, what to do when an exception occu


False Boolean value, result of comparison operations

finally Used with exceptions, a block of code that will be execute


exception or not

for To create a for loop

from To import specific parts of a module

global To declare a global variable

if To make a conditional statement

import To import a module

in To check if a value is present in a list, tuple, etc.

is To test if two variables are equal

lambda To create an anonymous function

match Start a match statement (compare a value against cases)


None Represents a null value

nonlocal To declare a non-local variable

not A logical operator

or A logical operator

pass A null statement, a statement that will do nothing

raise To raise an exception

return To exit a function and return a value

True Boolean value, result of comparison operations

try To make a try...except statement

while To create a while loop

with Used to simplify exception handling


yield To return a list of values from a generator

Python Built-in Exceptions

Built-in Exceptions
The table below shows built-in exceptions that are usually raised in Python:

Exception Description

ArithmeticError Raised when an error occurs in numeric calculations

AssertionError Raised when an assert statement fails

AttributeError Raised when attribute reference or assignment fails

Exception Base class for all exceptions

EOFError Raised when the input() method hits an "end of file" con

FloatingPointError Raised when a floating point calculation fails

GeneratorExit Raised when a generator is closed (with the close() meth

ImportError Raised when an imported module does not exist

IndentationError Raised when indentation is not correct

IndexError Raised when an index of a sequence does not exist

KeyError Raised when a key does not exist in a dictionary

KeyboardInterrupt Raised when the user presses Ctrl+c, Ctrl+z or Delete

LookupError Raised when errors raised cant be found

MemoryError Raised when a program runs out of memory

NameError Raised when a variable does not exist


NotImplementedError Raised when an abstract method requires an inherited c

OSError Raised when a system related operation causes an error

OverflowError Raised when the result of a numeric calculation is too lar

ReferenceError Raised when a weak reference object does not exist

RuntimeError Raised when an error occurs that do not belong to any s

StopIteration Raised when the next() method of an iterator has no furt

SyntaxError Raised when a syntax error occurs

TabError Raised when indentation consists of tabs or spaces

SystemError Raised when a system error occurs

SystemExit Raised when the [Link]() function is called

TypeError Raised when two different types are combined

UnboundLocalError Raised when a local variable is referenced before assign

UnicodeError Raised when a unicode problem occurs

UnicodeEncodeError Raised when a unicode encoding problem occurs

UnicodeDecodeError Raised when a unicode decoding problem occurs

UnicodeTranslateError Raised when a unicode translation problem occurs

ValueError Raised when there is a wrong value in a specified data ty

ZeroDivisionError Raised when the second operator in a division is zero

Feature Description

Indentation Indentation refers to the spaces at the beginning of a co

Comments Comments are code lines that will not be executed

Multiline Comments How to insert comments on multiple lines

Creating Variables Variables are containers for storing data values

Variable Names How to name your variables


Camel Case Camel Case Variable Names

Pascal Case Pascal Case Variable Names

Snake Case Snake Case Variable Names

Assign Values to Multiple How to assign values to multiple variables


Variables

Output Variables Use the print statement to output variables

String Concatenation How to combine strings

Global Variables Global variables are variables that belongs to the global

Built-In Data Types Python has a set of built-in data types

Getting Data Type How to get the data type of an object

Setting Data Type How to set the data type of an object

Numbers There are three numeric types in Python

Int The integer number type

Float The floating number type

Complex The complex number type

Type Conversion How to convert from one number type to another

Random Number How to create a random number

Specify a Variable Type How to specify a certain data type for a variable

String Literals How to create string literals

Assigning a String to a How to assign a string value to a variable


Variable

Multiline Strings How to create a multiline string

Strings are Arrays Strings in Python are arrays of bytes representing Unicod

Slicing a String How to slice a string

Negative Indexing on a How to use negative indexing when accessing a string


String
String Length How to get the length of a string

Check In String How to check if a string contains a specified phrase

Format String How to combine two strings

Escape Characters How to use escape characters

Boolean Values True or False

Evaluate Booleans Evaluate a value or statement and return either True or

Return Boolean Value Functions that return a Boolean value

Operators Use operator to perform operations in Python

Arithmetic Operators Arithmetic operator are used to perform common mathe

Assignment Operators Assignment operators are use to assign values to variab

Comparison Operators Comparison operators are used to compare two values

Logical Operators Logical operators are used to combine conditional statem

Identity Operators Identity operators are used to see if two objects are in fa

Membership Operators Membership operators are used to test is a sequence is p

Bitwise Operators Bitwise operators are used to compare (binary) numbers

Lists A list is an ordered, and changeable, collection

Access List Items How to access items in a list

Change List Item How to change the value of a list item

Loop Through List Items How to loop through the items in a list

List Comprehension How use a list comprehensive

Check if List Item Exists How to check if a specified item is present in a list

List Length How to determine the length of a list

Add List Items How to add items to a list

Remove List Items How to remove list items


Copy a List How to copy a list

Join Two Lists How to join two lists

Tuple A tuple is an ordered, and unchangeable, collection

Access Tuple Items How to access items in a tuple

Change Tuple Item How to change the value of a tuple item

Loop List Items How to loop through the items in a tuple

Check if Tuple Item Exists How to check if a specified item is present in a tuple

Tuple Length How to determine the length of a tuple

Tuple With One Item How to create a tuple with only one item

Remove Tuple Items How to remove tuple items

Join Two Tuples How to join two tuples

Set A set is an unordered, and unchangeable, collection

Access Set Items How to access items in a set

Add Set Items How to add items to a set

Loop Set Items How to loop through the items in a set

Check if Set Item Exists How to check if a item exists

Set Length How to determine the length of a set

Remove Set Items How to remove set items

Join Two Sets How to join two sets

Dictionary A dictionary is an unordered, and changeable, collection

Access Dictionary Items How to access items in a dictionary

Change Dictionary Item How to change the value of a dictionary item

Loop Dictionary Items How to loop through the items in a tuple

Check if Dictionary Item How to check if a specified item is present in a dictionary


Exists
Dictionary Length How to determine the length of a dictionary

Add Dictionary Item How to add an item to a dictionary

Remove Dictionary Items How to remove dictionary items

Copy Dictionary How to copy a dictionary

Nested Dictionaries A dictionary within a dictionary

If Statement How to write an if statement

If Indentation If statements in Python relies on indentation (whitespace

Elif elif is the same as "else if" in other programming langua

Else How to write an if...else statement

Shorthand If How to write an if statement in one line

Shorthand If Else How to write an if...else statement in one line

If AND Use the and keyword to combine if statements

If OR Use the or keyword to combine if statements

If NOT Use the not keyword to reverse the condition

Nested If How to write an if statement inside an if statement

The pass Keyword in If Use the pass keyword inside empty if statements

While How to write a while loop

While Break How to break a while loop

While Continue How to stop the current iteration and continue wit the ne

While Else How to use an else statement in a while loop

For How to write a for loop

Loop Through a String How to loop through a string

For Break How to break a for loop

For Continue How to stop the current iteration and continue wit the ne
Looping Through a range How to loop through a range of values

For Else How to use an else statement in a for loop

Nested Loops How to write a loop inside a loop

For pass Use the pass keyword inside empty for loops

Function How to create a function in Python

Call a Function How to call a function in Python

Function Arguments How to use arguments in a function

*args To deal with an unknown number of arguments in a func


parameter name

Keyword Arguments How to use keyword arguments in a function

**kwargs To deal with an unknown number of keyword arguments


before the parameter name

Default Parameter Value How to use a default parameter value

Passing a List as an How to pass a list as an argument


Argument

Function Return Value How to return a value from a function

The pass Statement in Use the pass statement in empty functions


Functions

Function Recursion Functions that can call itself is called recursive functions

Lambda Function How to create anonymous functions in Python

Why Use Lambda Functions Learn when to use a lambda function or not

Array Lists can be used as Arrays

What is an Array Arrays are variables that can hold more than one value

Access Arrays How to access array items

Array Length How to get the length of an array

Looping Array Elements How to loop through array elements


Add Array Element How to add elements from an array

Remove Array Element How to remove elements from an array

Array Methods Python has a set of Array/Lists methods

Class A class is like an object constructor

Create Class How to create a class

The Class __init__() Function The __init__() function is executed when the class is initia

Object Methods Methods in objects are functions that belongs to the obje

self The self parameter refers to the current instance of the c

Modify Object Properties How to modify properties of an object

Delete Object Properties How to modify properties of an object

Delete Object How to delete an object

Class pass Statement Use the pass statement in empty classes

Create Parent Class How to create a parent class

Create Child Class How to create a child class

Create the __init__() Function How to create the __init__() function

super Function The super() function make the child class inherit the par

Add Class Properties How to add a property to a class

Add Class Methods How to add a method to a class

Iterators An iterator is an object that contains a countable numbe

Iterator vs Iterable What is the difference between an iterator and an iterab

Loop Through an Iterator How to loop through the elements of an iterator

Create an Iterator How to create an iterator

StopIteration How to stop an iterator

Global Scope When does a variable belong to the global scope?


Global Keyword The global keyword makes the variable global

Create a Module How to create a module

Variables in Modules How to use variables in a module

Renaming a Module How to rename a module

Built-in Modules How to import built-in modules

Using the dir() Function List all variable names and function names in a module

Import From Module How to import only parts from a module

Datetime Module How to work with dates in Python

Date Output How to output a date

Create a Date Object How to create a date object

The strftime Method How to format a date object into a readable string

Date Format Codes The datetime module has a set of legal format codes

JSON How to work with JSON in Python

Parse JSON How to parse JSON code in Python

Convert into JSON How to convert a Python object in to JSON

Format JSON How to format JSON output with indentations and line br

Sort JSON How to sort JSON

RegEx Module How to import the regex module

RegEx Functions The re module has a set of functions

Metacharacters in RegEx Metacharacters are characters with a special meaning

RegEx Special Sequences A backslash followed by a a character has a special mea

RegEx Sets A set is a set of characters inside a pair of square bracke

RegEx Match Object The Match Object is an object containing information abo

Install PIP How to install PIP


PIP Packages How to download and install a package with PIP

PIP Remove Package How to remove a package with PIP

Error Handling How to handle errors in Python

Handle Many Exceptions How to handle more than one exception

Try Else How to use the else keyword in a try statement

Try Finally How to use the finally keyword in a try statement

raise How to raise an exception in Python

Python Python Built-in


Modules

This page lists the built-in modules that ship with the Python 3.13 Standard
Library.

These modules are available without extra installation (some are platform-
dependent).

A
Module Description

abc Tools for defining Abstract Base Classes (interfaces for Py

aifc Read and write AIFF/AIFF-C audio files.


argparse Parse command line arguments and create user-friendly C

array Efficient arrays of basic numeric types (compact alternati

ast Work with Python code as an Abstract Syntax Tree (analy


code).

asyncio Write concurrent code using the async/await syntax (even

atexit Register functions to run automatically when the program

B
Module Description

base64 Encode and decode data using Base16, Base32, Base64,

bdb Debugger framework used by pdb (implements the core d

binascii Convert between binary and ASCII (hex, base64 helpers a

bisect Maintain sorted lists; insert and search with binary search
builtins Access Python's built-in objects like len, range, and excep

bz2 Read and write bzip2-compressed files and streams.

C
Module Description

calendar Work with dates as calendars; print text calendars and co

cgi Helpers for Common Gateway Interface (legacy web serve

cgitb Pretty tracebacks for CGI scripts (HTML formatted error re

cmd Build simple line-oriented command interpreters (REPL-lik

code Run an interactive interpreter or embed one in your progr

codecs Text encodings and decoding/encoding helpers.


codeop Compile Python code objects conditionally (used by code

collections High-performance container datatypes like deque, Counte

colorsys Convert between color systems: RGB, YIQ, HLS, HSV.

concurrent Concurrency framework (futures) for running callables as

configparser Read and write INI-style configuration files.

contextlib Utilities for context managers and the with-statement.

contextvars Context-local storage for async code (like thread-local, bu

copy Shallow and deep copy operations for Python objects.

copyreg Register custom pickling functions for complex objects.

csv Read and write CSV files (comma-separated values) with

ctypes Call C libraries and work with C-compatible data types.


curses Terminal handling for character-cell UIs (Unix-like system

REMOVE ADS

D
Module Description

dataclasses Decorator and helpers for classes that store data (auto-ge
etc.).

datetime Dates, times, time zones, and timedeltas made simple an

dbm Family of simple on-disk key/value databases (backed by

decimal Fixed-point and floating decimal arithmetic for money and

difflib Compare sequences and create human-readable diffs.

dis Disassemble Python bytecode for inspection and debuggi

doctest Test examples embedded in docstrings (keeps docs and c


E
Module Description

email Parse, build, and send email messages (headers, MIME, a

encodings Implementation of text encodings used by Python's codec

ensurepip Bootstraps pip into a Python installation.

enum Define enumerations (named constants) with nice seman

errno Standard error number constants from the OS.

F
Module Description
faulthandler Dump Python tracebacks on a crash or on demand (helps

filecmp Compare files and directories to see what changed.

fileinput Loop over lines from stdin or a list of files as a single strea

fnmatch Match filenames using shell-style wildcards.

fractions Rational numbers (Fractions) for exact arithmetic.

ftplib FTP client library for transferring files.

functools Higher-order functions and utilities (lru_cache, partial, wra

G
Module Description

gc Control the garbage collector and inspect tracked objects

getopt Parse command-line options (POSIX-style).


getpass Prompt for a password without echoing it to the console.

gettext Internationalization (i18n) support for translating messag

glob Find pathnames matching a pattern like *.py.

graphlib Topological sorting utilities for dependency graphs.

gzip Read and write gzip-compressed files and streams.

H
Module Description

hashlib Secure hashes and message digests (SHA, MD5, BLAKE2,

heapq Heap queue (priority queue) algorithms on plain lists.

hmac Keyed-hash message authentication codes (HMAC).

html HTML helpers (escape/unescape text).


http HTTP modules package (client, server, cookies).

I
Module Description

idlelib Support for the IDLE interactive Python environment.

imaplib IMAP4 client library for accessing email servers.

imghdr Determine the type of an image file.

imp Access the import internals (deprecated; use importlib)

importlib Programmatic interface to Python's import mechanism.

inspect Inspect live objects such as modules, classes, and functio

io Input/Output streams and buffering (text and binary).


J
Module Description

json Encode and decode JSON data (JavaScript Object Notation

K
Module Description

keyword Test for Python keywords; list all keywords.

L
Module Description

linecache Read text lines from files with random access.

locale Internationalization (i18n) support for formatting numbers

logging Flexible event logging system with various handlers.


lzma Compression using the LZMA algorithm (xz format).

M
Module Description

mailbox Work with various mailbox formats (mbox, Maildir, etc.).

mailcap Read mailcap files (MIME handlers configuration).

marshal Read and write Python values in a binary format (for .pyc

math Fast math functions: trigonometry, logarithms, constants,

mimetypes Guess a file's type based on its filename or URL.

mmap Memory-map files for efficient random access.

modulefinder Find modules used by a script (dependency scanning).


msilib Create and read Microsoft Installer (.msi) files (Windows o

msvcrt Access to Microsoft C runtime routines (Windows only).

multiprocessing Run code in parallel using processes (bypass the GIL).

N
Module Description

netrc Parse .netrc files for machine login credentials.

nntplib Client for NNTP (Usenet) news servers.

numbers Abstract base classes for numeric types.

nis Interface to Sun's NIS (Yellow Pages) service (Unix platfor

ntpath Windows path operations (used by [Link] on Windows).


O
Module Description

operator Functional interface to operators (add, mul, itemgetter, a

optparse Deprecated parser for command line options (use argpars

os Operating system interfaces: files, environment, processe

P
Module Description

pathlib Object-oriented filesystem paths.

pdb Interactive debugger for Python programs.

pickle Serialize and deserialize Python objects (not secure again

pickletools Tools for analyzing pickled data.


pipes Pipe shell commands together (Unix).

pkgutil Utilities for packages: walk packages, find loaders, etc.

platform Access to underlying platform information.

plistlib Read and write Apple .plist files.

poplib POP3 email client.

posix POSIX APIs (Unix-only, low-level).

pprint Pretty-print Python data structures.

profile Deterministic profiling of Python programs.

pstats Work with profiling results produced by profile/cProfile.

pty Pseudo-terminal utilities (Unix).

pyclbr Read module to get class browser information without im


pydoc Generate and view Python documentation.

py_compile Compile Python source files to bytecode.

Q
Module Description

queue Thread-safe FIFO/LIFO/priority queues.

quopri Encode/decode MIME quoted-printable data.

R
Module Description

random Generate pseudo-random numbers for various distribution

re Regular expression operations for pattern matching in str


reprlib Safe string representations for large or recursive structur

resource Control and query system resource limits (Unix).

rlcompleter Tab-completion support for the interactive interpreter.

runpy Run modules as scripts (like python -m).

S
Module Description

sched Event scheduler for running functions at specific times.

secrets Generate cryptographically strong random numbers and t

select Low-level I/O multiplexing (select, poll, epoll, kqueue).

selectors High-level I/O multiplexing built on select module.

shelve Simple persistent storage for Python objects (dict-like API


shlex Parse shell-like syntaxes into tokens.

shutil High-level file operations: copy, move, archive, disk usage

signal Set handlers for asynchronous signals.

site Site-specific configuration hook that runs at startup.

smtplib Send emails using the SMTP protocol.

socket Low-level networking interface.

socketserver Framework for network servers (TCP/UDP).

sqlite3 Built-in lightweight SQL database (SQLite).

ssl TLS/SSL wrapper for secure network connections.

stat Constants and helpers for interpreting [Link]() results.

statistics Basic statistics (mean, median, stdev).


string Common string constants and helpers.

stringprep String preparation for internet protocols (IDNA).

struct Convert between Python values and C structs packed as b

subprocess Spawn new processes and connect to their input/output/e

sunau Read and write Sun AU audio files.

symtable Access the compiler's internal symbol tables.

sys Access to interpreter variables and functions.

sysconfig Access Python's configuration information.

T
Module Description

tabnanny Detect ambiguous indentation in Python source files.


tarfile Read and write tar archives, including gzip/bz2/xz.

telnetlib Telnet client implementation.

tempfile Create temporary files and directories safely.

termios POSIX terminal control (Unix).

test Regression tests package for the Python standard library.

textwrap Wrap and fill text paragraphs.

threading Higher-level threading interface (locks, events, threads).

time Time access and conversions.

timeit Measure execution time of small code snippets.

tkinter Standard GUI toolkit (Tk interface) package.

token Token constants used by the Python tokenizer.


tokenize Turn Python source into tokens (lexical scanner).

tomllib Read TOML files into Python data structures.

trace Trace program execution and produce coverage reports.

traceback Print or retrieve stack traces.

tracemalloc Track memory allocations to find leaks and hotspots.

tty Terminal control functions (Unix).

turtle Simple graphics library for teaching and fun.

types Names for built-in types and helper factories.

typing Type hints and typing helpers for static analysis and tooli

U
Module Description
unicodedata Access the Unicode Character Database (properties, norm

unittest Unit testing framework (xUnit style) for Python.

urllib Package for working with URLs (requests, parsing, robots)

uuid Generate universally unique identifiers (UUIDs).

V
Module Description

venv Create lightweight isolated Python environments.

W
Module Description

warnings Issue and control warning messages.


wave Read and write WAV audio files.

weakref Weak references to objects (avoid reference cycles).

webbrowser Open URLs in a web browser.

wsgiref WSGI utilities and simple reference server for web apps.

X
Module Description

xdrlib Pack and unpack XDR data (External Data Representation

xml XML processing package.

xmlrpc XML-RPC client and server package.

Z
Module Description

zipapp Create executable Python zip applications.

zipfile Read and write ZIP archives.

zipimport Import modules from ZIP archives.

zlib Compress and decompress data using zlib.

zoneinfo IANA time zone support for datetime.

Python Random Module

Python has a built-in module that you can use to make random numbers.

The random module has a set of methods:

Method Description

seed() Initialize the random number generator

getstate() Returns the current internal state of the random number generator
setstate() Restores the internal state of the random number generator

getrandbits() Returns a number representing the random bits

randrange() Returns a random number between the given range

randint() Returns a random number between the given range

choice() Returns a random element from the given sequence

choices() Returns a list with a random selection from the given sequence

shuffle() Takes a sequence and returns the sequence in a random order

sample() Returns a given sample of a sequence

random() Returns a random float number between 0 and 1

uniform() Returns a random float number between two given parameters


triangular() Returns a random float number between two given parameters, you ca
specify the midpoint between the two other parameters

betavariate() Returns a random float number between 0 and 1 based on the Beta dis

expovariate() Returns a random float number based on the Exponential distribution (

gammavariate Returns a random float number based on the Gamma distribution (use
()

gauss() Returns a random float number based on the Gaussian distribution (us

lognormvariate Returns a random float number based on a log-normal distribution (use


()

normalvariate( Returns a random float number based on the normal distribution (used
)

vonmisesvariat Returns a random float number based on the von Mises distribution (us
e()

paretovariate() Returns a random float number based on the Pareto distribution (used

weibullvariate( Returns a random float number based on the Weibull distribution (used
)
Python Requests Module

ExampleGet your own Python Server


Make a request to a web page, and print the response text:

import requests

x = [Link]('[Link]

print([Link])

Definition and Usage


The requests module allows you to send HTTP requests using Python.

The HTTP request returns a Response Object with all the response data
(content, encoding, status, etc).

Download and Install the Requests


Module
Navigate your command line to the location of PIP, and type the following:

C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>pip install requests

Syntax
[Link](params)
Methods
Method Description

delete(url, args) Sends a DELETE request to the specified url

get(url, params, args) Sends a GET request to the specified url

head(url, args) Sends a HEAD request to the specified url

patch(url, data, args) Sends a PATCH request to the specified url

post(url, data, json, args) Sends a POST request to the specified url

put(url, data, args) Sends a PUT request to the specified url

request(method, url, args) Sends a request of the specified method to the specifi

Python statistics Module

Python statistics Module


Python has a built-in module that you can use to calculate mathematical
statistics of numeric data.

The statistics module was new in Python 3.4.

Statistics Methods
Method Description

statistics.harmonic_mean() Calculates the harmonic mean (central location) of

[Link]() Calculates the mean (average) of the given data

[Link]() Calculates the median (middle value) of the given d

statistics.median_grouped() Calculates the median of grouped continuous data

statistics.median_high() Calculates the high median of the given data

statistics.median_low() Calculates the low median of the given data

[Link]() Calculates the mode (central tendency) of the give

[Link]() Calculates the standard deviation from an entire po

[Link]() Calculates the standard deviation from a sample of

[Link]() Calculates the variance of an entire population

[Link]() Calculates the variance from a sample of data

Python math Module

Python math Module


Python has a built-in module that you can use for mathematical tasks.
The math module has a set of methods and constants.

Math Methods
Method Description

[Link]() Returns the arc cosine of a number

[Link]() Returns the inverse hyperbolic cosine of a number

[Link]() Returns the arc sine of a number

[Link]() Returns the inverse hyperbolic sine of a number

[Link]() Returns the arc tangent of a number in radians

math.atan2() Returns the arc tangent of y/x in radians

[Link]() Returns the inverse hyperbolic tangent of a number

[Link]() Rounds a number up to the nearest integer

[Link]() Returns the number of ways to choose k items from n


order

[Link]() Returns a float consisting of the value of the first para


parameter

[Link]() Returns the cosine of a number

[Link]() Returns the hyperbolic cosine of a number

[Link]() Converts an angle from radians to degrees

[Link]() Returns the Euclidean distance between two points (p


coordinates of that point

[Link]() Returns the error function of a number

[Link]() Returns the complementary error function of a numb

[Link]() Returns E raised to the power of x

math.expm1() Returns Ex - 1
[Link]() Returns the absolute value of a number

[Link]() Returns the factorial of a number

[Link]() Rounds a number down to the nearest integer

[Link]() Returns the remainder of x/y

[Link]() Returns the mantissa and the exponent, of a specifie

[Link]() Returns the sum of all items in any iterable (tuples, a

[Link]() Returns the gamma function at x

[Link]() Returns the greatest common divisor of two integers

[Link]() Returns the Euclidean norm

[Link]() Checks whether two values are close to each other, o

[Link]() Checks whether a number is finite or not

[Link]() Checks whether a number is infinite or not

[Link]() Checks whether a value is NaN (not a number) or not

[Link]() Rounds a square root number downwards to the near

[Link]() Returns the inverse of [Link]() which is x * (2**i)

[Link]() Returns the log gamma value of x

[Link]() Returns the natural logarithm of a number, or the log

math.log10() Returns the base-10 logarithm of x

math.log1p() Returns the natural logarithm of 1+x

math.log2() Returns the base-2 logarithm of x

[Link]() Returns the number of ways to choose k items from n


repetition

[Link]() Returns the value of x to the power of y

[Link]() Returns the product of all the elements in an iterable

[Link]() Converts a degree value into radians


[Link]() Returns the closest value that can make numerator c
denominator

[Link]() Returns the sine of a number

[Link]() Returns the hyperbolic sine of a number

[Link]() Returns the square root of a number

[Link]() Returns the tangent of a number

[Link]() Returns the hyperbolic tangent of a number

[Link]() Returns the truncated integer parts of a number

Math Constants
Constant Description

math.e Returns Euler's number (2.7182...)

[Link] Returns a floating-point positive infinity

[Link] Returns a floating-point NaN (Not a Number) value

[Link] Returns PI (3.1415...)

[Link] Returns tau (6.2831...)

Python cmath Module

Python cmath Module


Python has a built-in module that you can use for mathematical tasks for
complex numbers.

The methods in this module accepts int, float, and complex numbers. It even
accepts Python objects that has a __complex__() or __float__() method.
The methods in this module almost always return a complex number. If the
return value can be expressed as a real number, the return value has an
imaginary part of 0.

The cmath module has a set of methods and constants.

cMath Methods
Method Description

[Link](x) Returns the arc cosine value of x

[Link](x) Returns the hyperbolic arc cosine of x

[Link](x) Returns the arc sine of x

[Link](x) Returns the hyperbolic arc sine of x

[Link](x) Returns the arc tangent value of x

[Link](x) Returns the hyperbolic arctangent value of x

[Link](x) Returns the cosine of x

[Link](x) Returns the hyperbolic cosine of x

[Link](x) Returns the value of Ex, where E is Euler's number (ap


the number passed to it

[Link]() Checks whether two values are close, or not

[Link](x) Checks whether x is a finite number

[Link](x) Check whether x is a positive or negative infinty

[Link](x) Checks whether x is NaN (not a number)

[Link](x[, base]) Returns the logarithm of x to the base

cmath.log10(x) Returns the base-10 logarithm of x

[Link]() Return the phase of a complex number

[Link]() Convert a complex number to polar coordinates


[Link]() Convert polar coordinates to rectangular form

[Link](x) Returns the sine of x

[Link](x) Returns the hyperbolic sine of x

[Link](x) Returns the square root of x

[Link](x) Returns the tangent of x

[Link](x) Returns the hyperbolic tangent of x

REMOVE ADS

cMath Constants
Constant Description

cmath.e Returns Euler's number (2.7182...)

[Link] Returns a floating-point positive infinity value

[Link] Returns a complex infinity value

[Link] Returns floating-point NaN (Not a Number) value

[Link] Returns coplext NaN (Not a Number) value

[Link] Returns PI (3.1415...)

[Link] Returns tau (6.2831...)

How to Remove Duplicates


From a Python List

Learn how to remove duplicates from a List in Python.


ExampleGet your own Python Server
Remove any duplicates from a List:

mylist = ["a", "b", "a", "c", "c"]


mylist = list([Link](mylist))
print(mylist)

Example Explained
First we have a List that contains duplicates:

A List with Duplicates


mylist = ["a", "b", "a", "c", "c"]
mylist = list([Link](mylist))
print(mylist)

Create a dictionary, using the List items as keys. This will automatically
remove any duplicates because dictionaries cannot have duplicate keys.

Create a Dictionary
mylist = ["a", "b", "a", "c", "c"]
mylist = list( [Link](mylist) )
print(mylist)

Then, convert the dictionary back into a list:

Convert Into a List


mylist = ["a", "b", "a", "c", "c"]
mylist = list( [Link](mylist) )
print(mylist)

Now we have a List without any duplicates, and it has the same order as the
original List.

Print the List to demonstrate the result

Print the List


mylist = ["a", "b", "a", "c", "c"]
mylist = list([Link](mylist))
print(mylist)

Create a Function
If you like to have a function where you can send your lists, and get them
back without duplicates, you can create a function and insert the code from
the example above.

Example
def my_function(x):
return list([Link](x))

mylist = my_function(["a", "b", "a", "c", "c"])

print(mylist)

Example Explained
Create a function that takes a List as an argument.

Create a Function
def my_function(x):
return list([Link](x))

mylist = my_function(["a", "b", "a", "c", "c"])

print(mylist)

Create a dictionary, using this List items as keys.

Create a Dictionary
def my_function(x):
return list( [Link](x) )

mylist = my_function(["a", "b", "a", "c", "c"])

print(mylist)

Convert the dictionary into a list.


Convert Into a List
def my_function(x):
return list( [Link](x) )

mylist = my_function(["a", "b", "a", "c", "c"])

print(mylist)

Return the list

Return List
def my_function(x):
return list([Link](x))

mylist = my_function(["a", "b", "a", "c", "c"])

print(mylist)

Call the function, with a list as a parameter:

Call the Function


def my_function(x):
return list([Link](x))

mylist = my_function(["a", "b", "a", "c", "c"])

print(mylist)

Print the result:

Print the Result


def my_function(x):
return list([Link](x))

mylist = my_function(["a", "b", "a", "c", "c"])

print(mylist)

How to Reverse a String in


Python
Learn how to reverse a String in Python.

There is no built-in function to reverse a String in Python.

The fastest (and easiest?) way is to use a slice that steps backwards, -1.

ExampleGet your own Python Server


Reverse the string "Hello World":

txt = "Hello World"[::-1]


print(txt)

Example Explained
We have a string, "Hello World", which we want to reverse:

The String to Reverse


txt = "Hello World" [::-1]
print(txt)

Create a slice that starts at the end of the string, and moves backwards.

In this particular example, the slice statement [::-1] means start at the end
of the string and end at position 0, move with the step -1, negative one,
which means one step backwards.

Slice the String


txt = "Hello World" [::-1]
print(txt)

Now we have a string txt that reads "Hello World" backwards.

Print the String to demonstrate the result

Print the List


txt = "Hello World"[::-1]
print(txt)
REMOVE ADS

Create a Function
If you like to have a function where you can send your strings, and return
them backwards, you can create a function and insert the code from the
example above.

Example
def my_function(x):
return x[::-1]

mytxt = my_function("I wonder how this text looks like backwards")

print(mytxt)

Example Explained
Create a function that takes a String as an argument.

Create a Function
def my_function(x):
return x[::-1]

mytxt = my_function("I wonder how this text looks like backwards")

print(mytxt)

Slice the string starting at the end of the string and move backwards.

Slice the String


def my_function(x):
return x [::-1]

mytxt = my_function("I wonder how this text looks like backwards")

print(mytxt)
Return the backward String

Return the String


def my_function(x):
return x[::-1]

mytxt = my_function("I wonder how this text looks like backwards")

print(mytxt )

Call the function, with a string as a parameter:

Call the Function


def my_function(x):
return x[::-1]

mytxt = my_function("I wonder how this text looks like backwards")

print(mytxt)

Print the result:

Print the Result


def my_function(x):
return x[::-1]

mytxt = my_function("I wonder how this text looks like backwards")

print(mytxt)

How to Add Two Numbers in


Python

Learn how to add two numbers in Python.

Use the + operator to add two numbers:


ExampleGet your own Python Server
x = 5
y = 10
print(x + y)

Add Two Numbers with User Input


In this example, the user must input two numbers. Then we print the sum by
calculating (adding) the two numbers:

Example
x = input("Type a number: ")
y = input("Type another number: ")

sum = int(x) + int(y)

print("The sum is: ", sum)

Python Examples

Python Syntax

Python Variables

Python Numbers

Python Casting
REMOVE ADS

Python Strings

Python Operators

Python Lists

Python Tuples

Python Sets

Python Dictionaries

Python If ... Else

Python While Loop


Python For Loop

Python Functions

Python Lambda

Python Arrays

Python Classes and Objects

Python Iterators

Python Modules

Python Dates

Python Math
Python JSON

Python RegEx

Python PIP

Python Try Except

Python File Handling

Python MySQL

Python MongoDB

You might also like