What Is Python
What Is Python
It is used for:
Why Python?
Python works on different platforms (Windows, Mac, Linux, Raspberry Pi, etc).
Python has syntax that allows developers to write programs with fewer lines than some other
programming languages.
Python runs on an interpreter system, meaning that code can be executed as soon as it is written. This
means that prototyping can be very quick.
Or by creating a python file on the server, using the .py file extension, and
running it in the Command Line:
Python Indentation
Indentation refers to the spaces at the beginning of a code line.
Python Variables
In Python, a variable is created when you assign a value to it
Comments
Python has commenting capability for the purpose of in-code documentation.
Comments start with a #, and Python will render the rest of the line as a
comment:
Statements
A computer program is a list of "instructions" to be "executed" by a
computer.
In Python, a statement usually ends when the line ends. You do not need to use a
semicolon (;) like in many other programming languages (for example, Java or C).
Many Statements
Most Python programs contain many statements.
The statements are executed one by one, in the same order as they are
written:
1. print("Hello World!")
2. print("Have a good day.")
3. print("Learning Python is fun!")
However, if you put two statements on the same line without a separator (newline
or ;), Python will give an error:
Print Text
You have already learned that you can use the print() function to display text or
output values:
You can use the print() function as many times as you want. Each call prints text on a new line by
default:
Double Quotes
Text in Python must be inside quotes. You can use either " double quotes or ' single quotes:
If you forget to put the text inside quotes, Python will give an error:
If you want to print multiple words on the same line, you can use
the end parameter:
Note that we add a space after end=" " for better readability.
Print Numbers
You can also use the print() function to display numbers:
You can combine text and numbers in one output by separating them with a comma:
Python Comments
Creating a Comment
Comments starts with a #, and Python will ignore them:
Comments can be placed at the end of a line, and Python will ignore the rest of the
line:
A comment does not have to be text that explains the code, it can also be used to prevent Python from
executing code:
Multiline Comments
Since Python will ignore string literals that are not assigned to a variable, you can add a multiline string
(triple quotes) in your code, and place your comment inside it:
As long as the string is not assigned to a variable, Python will read the code, but then ignore it, and you
have made a multiline comment.
Variables
Variables are containers for storing data values.
Creating Variables
Python has no command for declaring a variable.
Variables do not need to be declared with any particular type, and can even
change type after they have been set.
Example
x = 4 # x is of type int
x = "Sally" # x is now of type str
print(x)
Casting
If you want to specify the data type of a variable, this can be done with
casting.
Example
x = str(3) # x will be '3'
y = int(3) # y will be 3
z = float(3) # z will be 3.0
Case-Sensitive
Variable names are case-sensitive.
Example
This will create two variables:
a = 4
A = "Sally"
#A will not overwrite a
Variable Names
A variable can have a short name (like x and y) or a more descriptive name
(age, carname, total_volume).
Example
Legal variable names:
myvar = "John"
my_var = "John"
_my_var = "John"
myVar = "John"
MYVAR = "John"
myvar2 = "John"
Example
Illegal variable names:
2myvar = "John"
my-var = "John"
my var = "John"
Example
x, y, z = "Orange", "Banana", "Cherry"
print(x)
print(y)
print(z)
Example
x = y = z = "Orange"
print(x)
print(y)
print(z)
Unpack a Collection
If you have a collection of values in a list, tuple etc. Python allows you to
extract the values into variables. This is called unpacking.
Example
Unpack a list:
Output Variables
The print() function is often used to output variables.
Example
x = "Python is awesome"
print(x)
Example
x = "Python"
y = "is"
z = "awesome"
print(x, y, z)
Example
x = "Python "
y = "is "
z = "awesome"
print(x + y + z)
In the print() function, when you try to combine a string and a number with
the + operator, Python will give you an error:
Example
x = 5
y = "John"
print(x + y)
Example
x = 5
y = "John"
print(x, y)
Example
Create a variable outside of a function, and use it inside the function
x = "awesome"
def myfunc():
print("Python is " + x)
myfunc()
If you create a variable with the same name inside a function, this variable will be
local, and can only be used inside the function. The global variable with the same
name will remain as it was, global and with the original value.
Example
Create a variable inside a function, with the same name as the global
variable
x = "awesome"
def myfunc():
x = "fantastic"
print("Python is " + x)
myfunc()
print("Python is " + x)
Example
If you use the global keyword, the variable belongs to the global scope:
def myfunc():
global x
x = "fantastic"
myfunc()
print("Python is " + x)
Also, use the global keyword if you want to change a global variable inside
a function.
Example
To change the value of a global variable inside a function, refer to the
variable by using the global keyword:
x = "awesome"
def myfunc():
global x
x = "fantastic"
myfunc()
print("Python is " + x)
Variables can store data of different types, and different types can do
different things.
Python has the following data types built-in by default, in these categories:
Example
Print the data type of the variable x:
x = 5
print(type(x))
Example
x = "Hello World"
x = 20
x = 20.5
x = 1j
x = range(6)
x = True
x = b"Hello"
x = byte array(5)
x = memory view(bytes(5))
x = None
Setting the Specific Data Type
If you want to specify the data type, you can use the following constructor
functions:
Example
x = str("Hello World")
x = int(20)
x = float(20.5)
x = complex(1j)
x = range(6)
x = dict(name="John", age=36)
x = set(("apple", "banana", "cherry"))
x = bool(5)
x = bytes(5)
x = byte array(5)
x = memory view(bytes(5))
Python Numbers
There are three numeric types in Python:
int
float
complex
Variables of numeric types are created when you assign a value to them:
Example
x = 1 # int
y = 2.8 # float
z = 1j # complex
To verify the type of any object in Python, use the type() function:
Example
print(type(x))
print(type(y))
print(type(z))
Int
Int, or integer, is a whole number, positive or negative, without decimals, of
unlimited length.
Example
Integers:
x = 1
y = 35656222554887711
z = -3255522
print(type(x))
print(type(y))
print(type(z))
Float
Float, or "floating point number" is a number, positive or negative,
containing one or more decimals.
Example
Floats:
x = 1.10
y = 1.0
z = -35.59
print(type(x))
print(type(y))
print(type(z))
Float can also be scientific numbers with an "e" to indicate the power of 10.
Example
Floats:
x = 35e3
y = 12E4
z = -87.7e100
print(type(x))
print(type(y))
print(type(z))
Complex
Complex numbers are written with a "j" as the imaginary part:
Example
Complex:
x = 3+5j
y = 5j
z = -5j
print(type(x))
print(type(y))
print(type(z))
Type Conversion
You can convert from one type to another with the int(), float(),
and complex() methods:
Example
Convert from one type to another:
x = 1 # int
y = 2.8 # float
z = 1j # complex
print(a)
print(b)
print(c)
print(type(a))
print(type(b))
print(type(c))
Random Number
Python does not have a random() function to make a random number, but
Python has a built-in module called random that can be used to make random
numbers:
Example
Import the random module, and display a random number from 1 to 9:
import random
print([Link](1, 10))
Python Casting
Specify a Variable Type
There may be times when you want to specify a type on to a variable. This
can be done with casting. Python is an object-orientated language, and as
such it uses classes to define data types, including its primitive types.
Example
x = int(1) # x will be 1
y = int(2.8) # y will be 2
z = int("3") # z will be 3
Example
Floats:
x = float(1) # x will be 1.0
y = float(2.8) # y will be 2.8
z = float("3") # z will be 3.0
w = float("4.2") # w will be 4.2
Example
Strings:
Python Strings
Strings
Strings in python are surrounded by either single quotation marks, or double
quotation marks.
Example
print("Hello")
print('Hello')
Quotes Inside Quotes
You can use quotes inside a string, as long as they don't match the quotes
surrounding the string:
Example
print("It's alright")
print("He is called 'Johnny'")
print('He is called "Johnny"')
Example
a = "Hello"
print(a)
Multiline Strings
You can assign a multiline string to a variable by using three quotes:
Example
You can use three double quotes:
Example
a = '''Lorem ipsum dolor sit amet,
consectetur adipiscing elit,
sed do eiusmod tempor incididunt
ut labore et dolore magna aliqua.'''
print(a)
However, Python does not have a character data type, a single character is
simply a string with a length of 1.
Example
Get the character at position 1 (remember that the first character has the
position 0):
a = "Hello, World!"
print(a[1])
Example
Loop through the letters in the word "banana":
for x in "banana":
print(x)
String Length
To get the length of a string, use the len() function.
Example
The len() function returns the length of a string:
a = "Hello, World!"
print(len(a))
Check String
To check if a certain phrase or character is present in a string, we can use
the keyword in.
Example
Check if "free" is present in the following text:
Use it in an if statement:
Example
Print only if "free" is present:
Check if NOT
To check if a certain phrase or character is NOT present in a string, we can
use the keyword not in.
Example
Check if "expensive" is NOT present in the following text:
Use it in an if statement:
Example
print only if "expensive" is NOT present:
Specify the start index and the end index, separated by a colon, to return a
part of the string.
Example
Get the characters from position 2 to position 5 (not included):
b = "Hello, World!"
print(b[2:5])
Example
Get the characters from the start to position 5 (not included):
b = "Hello, World!"
print(b[:5])
Example
Get the characters from position 2, and all the way to the end:
b = "Hello, World!"
print(b[2:])
Negative Indexing
Use negative indexes to start the slice from the end of the string:
Example
Get the characters:
b = "Hello, World!"
print(b[-5:-2])
Upper Case
Example
The upper() method returns the string in upper case:
a = "Hello, World!"
print([Link]())
Lower Case
Example
The lower() method returns the string in lower case:
a = "Hello, World!"
print([Link]())
Remove Whitespace
Whitespace is the space before and/or after the actual text, and very often
you want to remove this space.
Example
The strip() method removes any whitespace from the beginning or the
end:
Replace String
Example
The replace() method replaces a string with another string:
a = "Hello, World!"
print([Link]("H", "J"))
Split String
The split() method returns a list where the text between the specified
separator becomes the list items.
Example
The split() method splits the string into substrings if it finds instances of
the separator:
a = "Hello, World!"
print([Link](",")) # returns ['Hello', ' World!']
Python - String
Concatenation
String Concatenation
To concatenate, or combine, two strings you can use the + operator.
Example
Merge variable a with variable b into variable c:
a = "Hello"
b = "World"
c = a + b
print(c)
Example
To add a space between them, add a " ":
a = "Hello"
b = "World"
c = a + " " + b
print(c)
Example
age = 36
#This will produce an error:
txt = "My name is John, I am " + age
print(txt)
F-Strings
F-String was introduced in Python 3.6, and is now the preferred way of
formatting strings.
Example
Create an f-string:
age = 36
txt = f"My name is John, I am {age}"
print(txt)
Example
Add a placeholder for the price variable:
price = 59
txt = f"The price is {price} dollars"
print(txt)
Example
Display the price with 2 decimals:
price = 59
txt = f"The price is {price:.2f} dollars"
print(txt)
Example
Perform a math operation in the placeholder, and return the result:
Example
You will get an error if you use double quotes inside a string that is
surrounded by double quotes:
Example
The escape character allows you to use double quotes when you normally
would not be allowed:
Escape Characters
Other escape characters used in Python:
Code Result
\\ Backslash
\n New Line
\r Carriage Return
\t Tab
\b Backspace
\f Form Feed
Method Description
endswith() Returns true if the string ends with the specified value
find() Searches the string for a specified value and returns the position of w
index() Searches the string for a specified value and returns the position of w
isalpha() Returns True if all characters in the string are in the alphabet
isascii() Returns True if all characters in the string are ascii characters
islower() Returns True if all characters in the string are lower case
isupper() Returns True if all characters in the string are upper case
partition() Returns a tuple where the string is parted into three parts
rfind() Searches the string for a specified value and returns the last position
rindex() Searches the string for a specified value and returns the last position
rpartition() Returns a tuple where the string is parted into three parts
rsplit() Splits the string at the specified separator, and returns a list
rstrip() Returns a right trim version of the string
split() Splits the string at the specified separator, and returns a list
split lines() Splits the string at line breaks and returns a list
starts with() Returns true if the string starts with the specified value
swap case() Swaps cases, lower case becomes upper case and vice versa
zfill() Fills the string with a specified number of 0 values at the beginning
Python Booleans
booleans represent one of two values: True or False.
Boolean Values
In programming you often need to know if an expression is True or False.
You can evaluate any expression in Python, and get one of two
answers, True or False.
When you compare two values, the expression is evaluated and Python
returns the Boolean answer:
Example
print(10 > 9)
print(10 == 9)
print(10 < 9
Example
Print a message based on whether the condition is True or False:
a = 200
b = 33
if b > a:
print("b is greater than a")
else:
print("b is not greater than a")
Example
Evaluate a string and a number:
print(bool("Hello"))
print(bool(15))
Example
Evaluate two variables:
x = "Hello"
y = 15
print(bool(x))
print(bool(y))
Any list, tuple, set, and dictionary are True, except empty ones.
Example
The following will return True:
bool("abc")
bool(123)
bool(["apple", "cherry", "banana"])
Example
The following will return False:
bool(False)
bool(None)
bool(0)
bool("")
bool(())
bool([])
bool({})
One more value, or object in this case, evaluates to False, and that is if you
have an object that is made from a class with a __len__ function that
returns 0 or False:
Example
class my class():
def __len__(self):
return 0
myobj = myclass()
print(bool(myobj))
Example
Print the answer of a function:
def myFunction() :
return True
print(myFunction())
Example
Print "YES!" if the function returns True, otherwise print "NO!":
def myFunction() :
return True
if myFunction():
print("YES!")
else:
print("NO!")
Python also has many built-in functions that return a boolean value, like
the isinstance() function, which can be used to determine if an object is of
a certain data type:
Example
Check if an object is an integer or not:
x = 200
print(isinstance(x, int))
Python Operators
Python Operators
Operators are used to perform operations on variables and values.
In the example below, we use the + operator to add together two values:
Example
print (10 + 5)
Although the + operator is often used to add together two values, like in the
example above, it can also be used to add together a variable and a value,
or two variables:
Example
sum1 = 100 + 50 # 150 (100 + 50)
sum2 = sum1 + 250 # 400 (150 + 250)
sum3 = sum2 + sum2 # 800 (400 + 400)
Arithmetic operators
Assignment operators
Comparison operators
Logical operators
Identity operators
Membership operators
Bitwise operators
Python Arithmetic Operators
Arithmetic Operators
Arithmetic operators are used with numeric values to perform common
mathematical operations:
+ Addition x+y
- Subtraction x-y
* Multiplication x*y
/ Division x/y
% Modulus x%y
** Exponentiation x ** y
// Floor division x // y
Examples
Here is an example using different arithmetic operators:
Example
x = 15
y = 4
print(x + y)
print(x - y)
print(x * y)
print(x / y)
print(x % y)
print(x ** y)
print(x // y)
Division in Python
Python has two division operators:
x = 12
y = 5
print(x / y)
Example
Floor division always returns an integer.
x = 12
y = 5
print(x // y)
Python Assignment
Operators
Assignment Operators
Assignment operators are used to assign values to variables:
= x=5 x=5
+= x += 3 x=x+3
-= x -= 3 x=x-3
*= x *= 3 x=x*3
/= x /= 3 x=x/3
%= x %= 3 x=x%3
//= x //= 3 x = x // 3
**= x **= 3 x = x ** 3
|= x |= 3 x=x|3
^= x ^= 3 x=x^3
:= print(x := 3) x=3
print(x)
Example
The count variable is assigned in the if statement, and given the value 5:
numbers = [1, 2, 3, 4, 5]
== Equal x == y
!= Not equal x != y
Examples
Comparison operators return True or False based on the comparison:
Example
x = 5
y = 3
print(x == y)
print(x != y)
print(x > y)
print(x < y)
print(x >= y)
print(x <= y)
Example
x = 5
Examples
Example
Test if a number is greater than 0 and less than 10:
x = 5
Example
Test if a number is less than 5 or greater than 10:
x = 5
x = 5
Examples
Example
The is operator returns True if both variables point to the same object:
x = ["apple", "banana"]
y = ["apple", "banana"]
z = x
print(x is z)
print(x is y)
print(x == y)
Example
The is not operator returns True if both variables do not point to the same
object:
x = ["apple", "banana"]
y = ["apple", "banana"]
print(x is not y)
Python Membership
Operators
Membership Operators
Membership operators are used to test if a sequence is presented in an
object:
Examples
Example
Check if "banana" is present in a list:
print("banana" in fruits)
Example
Check if "pineapple" is NOT present in a list:
fruits = ["apple", "banana", "cherry"]
Membership in Strings
The membership operators also work with strings:
Example
text = "Hello World"
print("H" in text)
print("hello" in text)
print("z" not in text)
<< Zero fill left Shift left by pushing zeros in from the right and let the leftmos
shift bits fall off
>> Signed right Shift right by pushing copies of the leftmost bit in from the lef
shift and let the rightmost bits fall off
Examples
Example
The & operator compares each bit and set it to 1 if both are 1, otherwise it is
set to 0:
print (6 & 3)
Then the & operator compares the bits and returns 0010, which is 2 in
decimal.
Example
The | operator compares each bit and set it to 1 if one or both is 1, otherwise
it is set to 0:
print (6 | 3)
Then the | operator compares the bits and returns 0111, which is 7 in
decimal.
Example
The ^ operator compares each bit and set it to 1 if only one is 1, otherwise
(if both are 1 or both are 0) it is set to 0:
print (6 ^ 3)
Then the ^ operator compares the bits and returns 0101, which is 5 in
decimal.
Example
Parentheses has the highest precedence, meaning that expressions inside
parentheses must be evaluated first:
print((6 + 3) - (6 + 3))
Example
Multiplication * has higher precedence than addition +, and therefore
multiplications are evaluated before additions:
print(100 + 5 * 3)
Precedence Order
The precedence order is described in the table below, starting with the
highest precedence at the top:
Operator Description
() Parentheses
** Exponentiation
^ Bitwise XOR
| Bitwise OR
or OR
Left-to-Right Evaluation
If two operators have the same precedence, the expression is evaluated
from left to right.
Example
Addition + and subtraction - has the same precedence, and therefore we
evaluate the expression from left to right:
print (5 + 4 - 7 + 3)
Python Lists
mylist = ["apple", "banana", "cherry"]
List
Lists are used to store multiple items in a single variable.
Lists are one of 4 built-in data types in Python used to store collections of
data, the other 3 are Tuple, Set, and Dictionary, all with different qualities
and usage.
Example
Create a List:
List items are indexed, the first item has index [0], the second item has
index [1] etc.
Ordered
When we say that lists are ordered, it means that the items have a defined
order, and that order will not change.
If you add new items to a list, the new items will be placed at the end of the
list.
Changeable
The list is changeable, meaning that we can change, add, and remove items
in a list after it has been created.
Allow Duplicates
Since lists are indexed, lists can have items with the same value:
Example
Lists allow duplicate values:
List Length
To determine how many items a list has, use the len() function:
Example
Print the number of items in the list:
Example
String, int and boolean data types:
Example
A list with strings, integers and boolean values:
type()
From Python's perspective, lists are defined as objects with the data type
'list':
<class 'list'>
Example
What is the data type of a list?
Example
Using the list() constructor to make a List:
Example
Print the second item of the list:
-1 refers to the last item, -2 refers to the second last item etc.
Example
Print the last item of the list:
Range of Indexes
You can specify a range of indexes by specifying where to start and where to
end the range.
When specifying a range, the return value will be a new list with the
specified items.
Example
Return the third, fourth, and fifth item:
thislist =
["apple", "banana", "cherry", "orange", "kiwi", "melon", "mango"]
print(thislist[2:5])
By leaving out the start value, the range will start at the first item:
Example
This example returns the items from the beginning to, but NOT including,
"kiwi":
thislist =
["apple", "banana", "cherry", "orange", "kiwi", "melon", "mango"]
print(thislist[:4])
By leaving out the end value, the range will go on to the end of the list:
Example
This example returns the items from "cherry" to the end:
thislist =
["apple", "banana", "cherry", "orange", "kiwi", "melon", "mango"]
print(thislist[2:])
Example
This example returns the items from "orange" (-4) to, but NOT including
"mango" (-1):
thislist =
["apple", "banana", "cherry", "orange", "kiwi", "melon", "mango"]
print(thislist[-4:-1])
Example
Check if "apple" is present in the list:
Example
Change the second item:
thislist = ["apple", "banana", "cherry"]
thislist[1] = "blackcurrant"
print(thislist)
Example
Change the values "banana" and "cherry" with the values "blackcurrant" and
"watermelon":
If you insert more items than you replace, the new items will be inserted
where you specified, and the remaining items will move accordingly:
Example
Change the second value by replacing it with two new values:
If you insert less items than you replace, the new items will be inserted
where you specified, and the remaining items will move accordingly:
Example
Change the second and third value by replacing it with one value:
Insert Items
To insert a new list item, without replacing any of the existing values, we can
use the insert() method.
Example
Insert "watermelon" as the third item:
Example
Using the append() method to append an item:
Insert Items
To insert a list item at a specified index, use the insert() method.
Example
Insert an item as the second position:
Example
Add the elements of tropical to thislist:
Example
Add elements of a tuple to a list:
Example
Remove "banana":
thislist = ["apple", "banana", "cherry"]
[Link]("banana")
print(thislist)
If there are more than one item with the specified value,
the remove() method removes the first occurrence:
Example
Remove the first occurrence of "banana":
Example
Remove the second item:
If you do not specify the index, the pop() method removes the last item.
Example
Remove the last item:
Example
Remove the first item:
thislist = ["apple", "banana", "cherry"]
del thislist[0]
print(thislist)
Example
Delete the entire list:
Example
Clear the list content:
Example
Print all items in the list, one by one:
Learn more about for loops in our Python For Loops Chapter.
Loop Through the Index Numbers
You can also loop through the list items by referring to their index number.
Example
Print all items by referring to their index number:
Use the len() function to determine the length of the list, then start at 0 and
loop your way through the list items by referring to their indexes.
Example
Print all items, using a while loop to go through all the index numbers
Learn more about while loops in our Python While Loops Chapter.
Example
A short hand for loop that will print all items in a list:
Example:
Based on a list of fruits, you want a new list, containing only the fruits with
the letter "a" in the name.
Without list comprehension you will have to write a for statement with a
conditional test inside:
Example
fruits = ["apple", "banana", "cherry", "kiwi", "mango"]
newlist = []
for x in fruits:
if "a" in x:
[Link](x)
print(newlist)
With list comprehension you can do all that with only one line of code:
Example
fruits = ["apple", "banana", "cherry", "kiwi", "mango"]
print(newlist)
The Syntax
newlist = [expression for item in iterable if condition == True]
The return value is a new list, leaving the old list unchanged.
Condition
The condition is like a filter that only accepts the items that evaluate to True.
Example
Only accept items that are not "apple":
The condition if x != "apple" will return True for all elements other than
"apple", making the new list contain all fruits except "apple".
Example
With no if statement:
Iterable
The iterable can be any iterable object, like a list, tuple, set etc.
Example
You can use the range() function to create an iterable:
Expression
The expression is the current item in the iteration, but it is also the outcome,
which you can manipulate before it ends up like a list item in the new list:
Example
Set the values in the new list to upper case:
Example
Set all values in the new list to 'hello':
The expression can also contain conditions, not like a filter, but as a way to
manipulate the outcome:
Example
Return "orange" instead of "banana":
Example
Sort the list alphabetically:
Example
Sort the list numerically:
Sort Descending
To sort descending, use the keyword argument reverse = True:
Example
Sort the list descending:
Example
Sort the list descending:
The function will return a number that will be used to sort the list (the lowest
number first):
Example
Sort the list based on how close the number is to 50:
def myfunc(n):
return abs(n - 50)
Example
Case sensitive sorting can give an unexpected result:
Luckily we can use built-in functions as key functions when sorting a list.
Example
Perform a case-insensitive sort of the list:
thislist = ["banana", "Orange", "Kiwi", "cherry"]
[Link](key = [Link])
print(thislist)
Reverse Order
What if you want to reverse the order of a list, regardless of the alphabet?
The reverse() method reverses the current sorting order of the elements.
Example
Reverse the order of the list items:
Example
Make a copy of a list with the copy() method:
Example
Make a copy of a list with the list() method:
Example
Make a copy of a list with the : operator:
Example
Join two list:
list1 = ["a", "b", "c"]
list2 = [1, 2, 3]
Another way to join two lists is by appending all the items from list2 into
list1, one by one:
Example
Append list2 into list1:
for x in list2:
[Link](x)
print(list1)
Or you can use the extend() method, where the purpose is to add elements
from one list to another list:
Example
Use the extend() method to add list2 at the end of list1:
[Link](list2)
print(list1)
extend() Add the elements of a list (or any iterable), to the end of the current lis
index() Returns the index of the first element with the specified value
Python Tuples
mytuple = ("apple", "banana", "cherry")
Tuple
Tuples are used to store multiple items in a single variable.
Example
Create a Tuple:
thistuple = ("apple", "banana", "cherry")
print(thistuple)
Tuple Items
Tuple items are ordered, unchangeable, and allow duplicate values.
Tuple items are indexed, the first item has index [0], the second item has
index [1] etc.
Ordered
When we say that tuples are ordered, it means that the items have a defined
order, and that order will not change.
Unchangeable
Tuples are unchangeable, meaning that we cannot change, add or remove
items after the tuple has been created.
Allow Duplicates
Since tuples are indexed, they can have items with the same value:
Example
Tuples allow duplicate values:
Tuple Length
To determine how many items a tuple has, use the len() function:
Example
Print the number of items in the tuple:
Example
One item tuple, remember the comma:
thistuple = ("apple",)
print(type(thistuple))
#NOT a tuple
thistuple = ("apple")
print(type(thistuple))
Example
String, int and boolean data types:
Example
A tuple with strings, integers and boolean values:
type()
From Python's perspective, tuples are defined as objects with the data type
'tuple':
<class 'tuple'>
Example
What is the data type of a tuple?
Example
Using the tuple() method to make a tuple:
Example
Print the second item in the tuple:
Negative Indexing
Negative indexing means start from the end.
-1 refers to the last item, -2 refers to the second last item etc.
Example
Print the last item of the tuple:
Range of Indexes
You can specify a range of indexes by specifying where to start and where to
end the range.
When specifying a range, the return value will be a new tuple with the
specified items.
Example
Return the third, fourth, and fifth item:
thistuple =
("apple", "banana", "cherry", "orange", "kiwi", "melon", "mango")
print(thistuple[2:5])
By leaving out the start value, the range will start at the first item:
Example
This example returns the items from the beginning to, but NOT included,
"kiwi":
thistuple =
("apple", "banana", "cherry", "orange", "kiwi", "melon", "mango")
print(thistuple[:4])
By leaving out the end value, the range will go on to the end of the tuple:
Example
This example returns the items from "cherry" and to the end:
thistuple =
("apple", "banana", "cherry", "orange", "kiwi", "melon", "mango")
print(thistuple[2:])
Example
This example returns the items from index -4 (included) to index -1
(excluded)
thistuple =
("apple", "banana", "cherry", "orange", "kiwi", "melon", "mango")
print(thistuple[-4:-1])
Example
Check if "apple" is present in the tuple:
But there is a workaround. You can convert the tuple into a list, change the
list, and convert the list back into a tuple.
Example
Convert the tuple into a list to be able to change it:
print(x)
Add Items
Since tuples are immutable, they do not have a built-in append() method, but
there are other ways to add items to a tuple.
1. Convert into a list: Just like the workaround for changing a tuple, you
can convert it into a list, add your item(s), and convert it back into a tuple.
Example
Convert the tuple into a list, add "orange", and convert it back into a tuple:
2. Add tuple to a tuple. You are allowed to add tuples to tuples, so if you
want to add one item, (or many), create a new tuple with the item(s), and
add it to the existing tuple:
Example
Create a new tuple with the value "orange", and add that tuple:
print(thistuple)
Remove Items
Note: You cannot remove items in a tuple.
Tuples are unchangeable, so you cannot remove items from it, but you can
use the same workaround as we used for changing and adding tuple items:
Example
Convert the tuple into a list, remove "apple", and convert it back into a tuple:
Example
The del keyword can delete the tuple completely:
Example
Packing a tuple:
But, in Python, we are also allowed to extract the values back into variables.
This is called "unpacking":
Example
Unpacking a tuple:
print(green)
print(yellow)
print(red)
Using Asterisk*
If the number of variables is less than the number of values, you can add
an * to the variable name and the values will be assigned to the variable as
a list:
Example
Assign the rest of the values as a list called "red":
print(green)
print(yellow)
print(red)
If the asterisk is added to another variable name than the last, Python will
assign values to the variable until the number of values left matches the
number of variables left.
Example
Add a list of values the "tropic" variable:
print(green)
print(tropic)
print(red)
Example
Iterate through the items and print the values:
Example
Print all items by referring to their index number:
Use the len() function to determine the length of the tuple, then start at 0
and loop your way through the tuple items by referring to their indexes.
Example
Print all items, using a while loop to go through all the index numbers:
Example
Join two tuples:
tuple1 = ("a", "b" , "c")
tuple2 = (1, 2, 3)
Multiply Tuples
If you want to multiply the content of a tuple a given number of times, you
can use the * operator:
Example
Multiply the fruits tuple by 2:
print(mytuple)
Method Description
index() Searches the tuple for a specified value and returns the position
Python Sets
myset = {"apple", "banana", "cherry"}
Set
Sets are used to store multiple items in a single variable.
Set is one of 4 built-in data types in Python used to store collections of data,
the other 3 are List, Tuple, and Dictionary, all with different qualities and
usage.
Set Items
Set items are unordered, unchangeable, and do not allow duplicate values.
Unordered
Unordered means that the items in a set do not have a defined order.
Set items can appear in a different order every time you use them, and
cannot be referred to by index or key.
Unchangeable
Set items are unchangeable, meaning that we cannot change the items after
the set has been created.
Once a set is created, you cannot change its items, but you can remove
items and add new items.
Example
Duplicate values will be ignored:
print(thisset)
Note: The values True and 1 are considered the same value in sets, and are
treated as duplicates:
Example
True and 1 is considered the same value:
print(thisset)
Note: The values False and 0 are considered the same value in sets, and
are treated as duplicates:
Example
False and 0 is considered the same value:
print(thisset)
Get the Length of a Set
To determine how many items a set has, use the len() function.
Example
Get the number of items in a set:
print(len(thisset))
Example
String, int and boolean data types:
Example
A set with strings, integers and boolean values:
type()
From Python's perspective, sets are defined as objects with the data type
'set':
<class 'set'>
Example
What is the data type of a set?
Example
Using the set() constructor to make a set:
**As of Python version 3.7, dictionaries are ordered. In Python 3.6 and
earlier, dictionaries are unordered.
But you can loop through the set items using a for loop, or ask if a specified
value is present in a set, by using the in keyword.
Example
Loop through the set, and print the values:
for x in thisset:
print(x)
Example
Check if "banana" is present in the set:
print("banana" in thisset)
Example
Check if "banana" is NOT present in the set:
Change Items
Once a set is created, you cannot change its items, but you can add new
items.
Example
Add an item to a set, using the add() method:
[Link]("orange")
print(thisset)
Add Sets
To add items from another set into the current set, use
the update() method.
Example
Add elements from tropical into thisset:
[Link](tropical)
print(thisset)
Add Any Iterable
The object in the update() method does not have to be a set, it can be any
iterable object (tuples, lists, dictionaries etc.).
Example
Add elements of a list to a set:
[Link](mylist)
print(thisset)
Example
Remove "banana" by using the remove() method:
[Link]("banana")
print(thisset)
Note: If the item to remove does not exist, remove() will raise an error.
Example
Remove "banana" by using the discard() method:
[Link]("banana")
print(thisset)
Note: If the item to remove does not exist, discard() will NOT raise an
error.
You can also use the pop() method to remove an item, but this method will
remove a random item, so you cannot be sure what item that gets removed.
Example
Remove a random item by using the pop() method:
x = [Link]()
print(x)
print(thisset)
Note: Sets are unordered, so when using the pop() method, you do not
know which item that gets removed.
Example
The clear() method empties the set:
[Link]()
print(thisset)
Example
The del keyword will delete the set completely:
del thisset
print(thisset)
Python - Loop Sets
Loop Items
You can loop through the set items by using a for loop:
Example
Loop through the set, and print the values:
for x in thisset:
print(x)
The union() and update() methods joins all items from both sets.
The difference() method keeps the items from the first set that are not in
the other set(s).
Union
The union() method returns a new set with all items from both sets.
Example
Join set1 and set2 into a new set:
set1 = {"a", "b", "c"}
set2 = {1, 2, 3}
set3 = [Link](set2)
print(set3)
You can use the | operator instead of the union() method, and you will get
the same result.
Example
Use | to join two sets:
When using a method, just add more sets in the parentheses, separated by
commas:
Example
Join multiple sets with the union() method:
When using the | operator, separate the sets with more | operators:
Example
Use | to join two sets:
Example
Join a set with a tuple:
z = [Link](y)
print(z)
Note: The | operator only allows you to join sets with sets, and not with
other data types like you can with the union() method.
Update
The update() method inserts all items from one set into another.
The update() changes the original set, and does not return a new set.
Example
The update() method inserts the items in set2 into set1:
set1 = {"a", "b" , "c"}
set2 = {1, 2, 3}
[Link](set2)
print(set1)
Note: Both union() and update() will exclude any duplicate items
Intersection
Keep ONLY the duplicates
The intersection() method will return a new set, that only contains the
items that are present in both sets.
Example
Join set1 and set2, but keep only the duplicates:
set3 = [Link](set2)
print(set3)
You can use the & operator instead of the intersection() method, and you
will get the same result.
Example
Use & to join two sets:
The intersection_update() method will also keep ONLY the duplicates, but
it will change the original set instead of returning a new set.
Example
Keep the items that exist in both set1, and set2:
set1.intersection_update(set2)
print(set1)
The values True and 1 are considered the same value. The same goes
for False and 0.
Example
Join sets that contains the values True, False, 1, and 0, and see what is
considered as duplicates:
set3 = [Link](set2)
print(set3)
Difference
The difference() method will return a new set that will contain only the
items from the first set that are not present in the other set.
Example
Keep all items from set1 that are not in set2:
set3 = [Link](set2)
print(set3)
You can use the - operator instead of the difference() method, and you
will get the same result.
Example
Use - to join two sets:
The difference_update() method will keep the items from the first set that
are not in the other set, but it will change the original set instead of returning
a new set.
Example
Use the difference_update() method to keep only the items from the first
set that are not present in the other set:
set1.difference_update(set2)
print(set1)
Symmetric Differences
The symmetric_difference() method will keep only the elements that are
NOT present in both sets.
Example
Keep the items that are not present in both sets:
set3 = set1.symmetric_difference(set2)
print(set3)
Example
Use ^ to join two sets:
Example
Use the symmetric_difference_update() method to keep the items that
are not present in both sets:
set1.symmetric_difference_update(set2)
print(set1)
Python frozenset
Python frozenset
frozenset is an immutable version of a set.
Example
Create a frozenset and check its type:
Frozenset Methods
Being immutable means you cannot add or remove elements. However,
frozensets support all non-mutating operations of sets.
intersection_update() &= Removes the items in this set that are not
Python Dictionaries
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
Dictionary
Dictionaries are used to store data values in key:value pairs.
As of Python version 3.7, dictionaries are ordered. In Python 3.6 and earlier,
dictionaries are unordered.
Dictionaries are written with curly brackets, and have keys and values:
Example
Create and print a dictionary:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
print(thisdict)
Dictionary Items
Dictionary items are ordered, changeable, and do not allow duplicates.
Example
Print the "brand" value of the dictionary:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
print(thisdict["brand"])
Ordered or Unordered?
As of Python version 3.7, dictionaries are ordered. In Python 3.6 and earlier,
dictionaries are unordered.
When we say that dictionaries are ordered, it means that the items have a
defined order, and that order will not change.
Unordered means that the items do not have a defined order, you cannot
refer to an item by using an index.
Changeable
Dictionaries are changeable, meaning that we can change, add or remove
items after the dictionary has been created.
Example
Duplicate values will overwrite existing values:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964,
"year": 2020
}
print(thisdict)
Dictionary Length
To determine how many items a dictionary has, use the len() function:
Example
Print the number of items in the dictionary:
print(len(thisdict))
Dictionary Items - Data Types
The values in dictionary items can be of any data type:
Example
String, int, boolean, and list data types:
thisdict = {
"brand": "Ford",
"electric": False,
"year": 1964,
"colors": ["red", "white", "blue"]
}
type()
From Python's perspective, dictionaries are defined as objects with the data
type 'dict':
<class 'dict'>
Example
Print the data type of a dictionary:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
print(type(thisdict))
Example
Using the dict() method to make a dictionary:
**As of Python version 3.7, dictionaries are ordered. In Python 3.6 and
earlier, dictionaries are unordered.
Example
Get the value of the "model" key:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
x = thisdict["model"]
There is also a method called get() that will give you the same result:
Example
Get the value of the "model" key:
x = [Link]("model")
Get Keys
The keys() method will return a list of all the keys in the dictionary.
Example
Get a list of the keys:
x = [Link]()
The list of the keys is a view of the dictionary, meaning that any changes
done to the dictionary will be reflected in the keys list.
Example
Add a new item to the original dictionary, and see that the keys list gets
updated as well:
car = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
x = [Link]()
car["color"] = "white"
Get Values
The values() method will return a list of all the values in the dictionary.
Example
Get a list of the values:
x = [Link]()
The list of the values is a view of the dictionary, meaning that any changes
done to the dictionary will be reflected in the values list.
Example
Make a change in the original dictionary, and see that the values list gets
updated as well:
car = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
x = [Link]()
car["year"] = 2020
Example
Add a new item to the original dictionary, and see that the values list gets
updated as well:
car = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
x = [Link]()
car["color"] = "red"
Get Items
The items() method will return each item in a dictionary, as tuples in a list.
Example
Get a list of the key:value pairs
x = [Link]()
The returned list is a view of the items of the dictionary, meaning that any
changes done to the dictionary will be reflected in the items list.
Example
Make a change in the original dictionary, and see that the items list gets
updated as well:
car = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
x = [Link]()
car["year"] = 2020
print(x) #after the change
Example
Add a new item to the original dictionary, and see that the items list gets
updated as well:
car = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
x = [Link]()
car["color"] = "red"
Example
Check if "model" is present in the dictionary:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
if "model" in thisdict:
print("Yes, 'model' is one of the keys in the thisdict dictionary")
Example
Change the "year" to 2018:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
thisdict["year"] = 2018
Update Dictionary
The update() method will update the dictionary with the items from the
given argument.
Example
Update the "year" of the car by using the update() method:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
[Link]({"year": 2020})
Example
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
thisdict["color"] = "red"
print(thisdict)
Update Dictionary
The update() method will update the dictionary with the items from a given
argument. If the item does not exist, the item will be added.
Example
Add a color item to the dictionary by using the update() method:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
[Link]({"color": "red"})
Example
The pop() method removes the item with the specified key name:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
[Link]("model")
print(thisdict)
Example
The popitem() method removes the last inserted item (in versions before
3.7, a random item is removed instead):
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
[Link]()
print(thisdict)
Example
The del keyword removes the item with the specified key name:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
del thisdict["model"]
print(thisdict)
Example
The del keyword can also delete the dictionary completely:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
del thisdict
print(thisdict) #this will cause an error because "thisdict" no
longer exists.
Example
The clear() method empties the dictionary:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
[Link]()
print(thisdict)
When looping through a dictionary, the return value are the keys of the
dictionary, but there are methods to return the values as well.
Example
Print all key names in the dictionary, one by one:
for x in thisdict:
print(x)
Example
Print all values in the dictionary, one by one:
for x in thisdict:
print(thisdict[x])
Example
You can also use the values() method to return values of a dictionary:
for x in [Link]():
print(x)
Example
You can use the keys() method to return the keys of a dictionary:
for x in [Link]():
print(x)
Example
Loop through both keys and values, by using the items() method:
for x, y in [Link]():
print(x, y)
There are ways to make a copy, one way is to use the built-in Dictionary
method copy().
Example
Make a copy of a dictionary with the copy() method:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
mydict = [Link]()
print(mydict)
Another way to make a copy is to use the built-in function dict().
Example
Make a copy of a dictionary with the dict() function:
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
mydict = dict(thisdict)
print(mydict)
Example
Create a dictionary that contain three dictionaries:
myfamily = {
"child1" : {
"name" : "Emil",
"year" : 2004
},
"child2" : {
"name" : "Tobias",
"year" : 2007
},
"child3" : {
"name" : "Linus",
"year" : 2011
}
}
child1 = {
"name" : "Emil",
"year" : 2004
}
child2 = {
"name" : "Tobias",
"year" : 2007
}
child3 = {
"name" : "Linus",
"year" : 2011
}
myfamily = {
"child1" : child1,
"child2" : child2,
"child3" : child3
}
Example
Print the name of child 2:
print(myfamily["child2"]["name"])
for y in obj:
print(y + ':', obj[y])
Method Description
items() Returns a list containing a tuple for each key value pair
setdefault() Returns the value of the specified key. If the key does not exist: insert the
Python If Statement
Python Conditions and If statements
Python supports the usual logical conditions from mathematics:
Equals: a == b
Not Equals: a != b
Less than: a < b
Less than or equal to: a <= b
Greater than: a > b
Greater than or equal to: a >= b
Example
If statement:
a = 33
b = 200
if b > a:
print("b is greater than a")
In this example we use two variables, a and b, which are used as part of the
if statement to test whether b is greater than a. As a is 33, and b is 200, we
know that 200 is greater than 33, and so we print to screen that "b is greater
than a".
Example
Checking if a number is positive:
number = 15
if number > 0:
print("The number is positive")
Indentation
Python relies on indentation (whitespace at the beginning of a line) to define
scope in the code. Other programming languages often use curly-brackets
for this purpose.
Example
If statement, without indentation (will raise an error):
a = 33
b = 200
if b > a:
print("b is greater than a") # you will get an error
Note: You can use spaces or tabs for indentation, but you must use the
same amount of indentation for all statements within the same code block.
Example
Multiple statements in an if block:
age = 20
if age >= 18:
print("You are an adult")
print("You can vote")
print("You have full legal rights")
Example
Using a boolean variable:
is_logged_in = True
if is_logged_in:
print("Welcome back!")
Zero (0), empty strings (""), None, and empty collections are treated
as False. Everything else is treated as True.
This includes positive numbers (5), negative numbers (-3), and any non-
empty string (even "False" is treated as True because it's a non-empty
string).
Python Elif Statement
The Elif Keyword
The elif keyword is Python's way of saying "if the previous conditions were
not true, then try this condition".
The elif keyword allows you to check multiple expressions for True and
execute a block of code as soon as one of the conditions evaluates to True.
Example
Testing multiple conditions:
score = 75
Important: Only the first true condition will be executed. Even if multiple
conditions are true, Python stops after executing the first matching block.
Example
Categorizing age groups:
age = 25
Example
Day of the week checker:
day = 3
if day == 1:
print("Monday")
elif day == 2:
print("Tuesday")
elif day == 3:
print("Wednesday")
elif day == 4:
print("Thursday")
elif day == 5:
print("Friday")
elif day == 6:
print("Saturday")
elif day == 7:
print("Sunday")
Example
a=200
b = 33
if b > a:
print("b is greater than a")
elif a == b:
print("a and b are equal")
else:
print("a is greater than b")
In this example a is greater than b, so the first condition is not true, also
the elif condition is not true, so we go to the else condition and print to
screen that "a is greater than b".
This creates a simple two-way choice: if the condition is true, execute one
block; otherwise, execute the else block.
Note: The else statement must come last. You cannot have an elif after
an else.
Example
Checking even or odd numbers:
number = 7
if number % 2 == 0:
print("The number is even")
else:
print("The number is odd")
Example
Temperature classifier:
temperature = 22
Else as Fallback
The else statement acts as a fallback that executes when none of the
preceding conditions are true. This makes it useful for error handling,
validation, and providing default values.
Example
Validating user input:
username = "Emil"
if len(username) > 0:
print(f"Welcome, {username}!")
else:
print("Error: Username cannot be empty")
Python Shorthand If
Short Hand If
If you have only one statement to execute, you can put it on the same line
as the if statement.
Example
One-line if statement:
a = 5
b = 2
if a > b: print("a is greater than b")
Note: You still need the colon : after the condition.
Example
One-line if/else that prints a value:
a = 2
b = 330
print("A") if a > b else print("B")
This is called a conditional expression (sometimes known as a "ternary
operator").
Example
a = 10
b = 20
bigger = a if a > b else b
print("Bigger is", bigger)
Example
One line, three outcomes:
a = 330
b = 330
print("A") if a > b else print("=") if a == b else print("B")
Practical Examples
Ternary operators are particularly useful for simple assignments and return
statements.
Example
Finding the maximum of two numbers:
x = 15
y = 20
max_value = x if x > y else y
print("Maximum value:", max_value)
Example
Setting a default value:
username = ""
display_name = username if username else "Guest"
print("Welcome,", display_name)
Example
Test if a is greater than b, AND if c is greater than a:
a = 200
b = 33
c = 500
if a > b and c > a:
print("Both conditions are True")
The or Operator
The or keyword is a logical operator, and is used to combine conditional
statements. At least one condition must be true for the entire expression to
be true.
Example
Test if a is greater than b, OR if a is greater than c:
a = 200
b = 33
c = 500
if a > b or a > c:
print("At least one of the conditions is True")
Example
Test if a is NOT greater than b:
a = 33
b = 200
if not a > b:
print("a is NOT greater than b")
Example
Combining and, or, and not:
age = 25
is_student = False
has_discount_code = True
Truth Tables
Understanding how logical operators work with different values:
and Operator Truth Table
Example
Using parentheses for complex conditions:
temperature = 25
is_raining = False
is_weekend = True
More Examples
Example
User authentication check:
username = "Tobias"
password = "secret123"
is_verified = True
Example
Range checking with logical operators:
score = 85
Python Nested If
Nested If Statements
You can have if statements inside if statements. This is
called nested if statements.
Example
x = 41
if x > 10:
print("Above ten,")
if x > 20:
print("and also above 20!")
else:
print("but not above 20.")
In this example, the inner if statement only runs if the outer condition (x >
10) is true.
Example
Checking multiple conditions with nesting:
age = 25
has_license = True
Example
Three levels of nesting:
score = 85
attendance = 90
submitted = True
Example
This nested if:
temperature = 25
is_sunny = True
if temperature > 20:
if is_sunny:
print("Perfect beach weather!")
Example
Could also be written with and:
temperature = 25
is_sunny = True
Both approaches produce the same result. Use nested if statements when
the inner logic is complex or depends on the outer condition. Use and when
both conditions are simple and equally important.
More Examples
Example
Login validation with nested checks:
username = "Emil"
password = "python123"
is_active = True
if username:
if password:
if is_active:
print("Login successful")
else:
print("Account is not active")
else:
print("Password required")
else:
print("Username required")
Example
Grade calculation with nested logic:
score = 92
extra_credit = 5
Example
a = 33
b = 200
if b > a:
pass
Example
Placeholder for future implementation:
age = 16
pass vs Comments
A comment is ignored by Python, but pass is an actual statement that gets
executed (though it does nothing). You need pass where Python expects a
statement, not just a comment.
Example
This will cause an error (empty code block):
score = 85
Example
This works correctly with pass:
score = 85
Example
Using pass in different branches:
value = 50
if value < 0:
print("Negative value")
elif value == 0:
pass # Zero case - no action needed
else:
print("Positive value")
Example
Using pass with functions:
def calculate_discount(price):
pass # TODO: Implement discount logic
Python Match
The match statement is used to perform different actions based on
different conditions.
Syntax
match expression:
case x:
code block
case y:
code block
case z:
code block
The example below uses the weekday number to print the weekday name:
Example
day = 4
match day:
case 1:
print("Monday")
case 2:
print("Tuesday")
case 3:
print("Wednesday")
case 4:
print("Thursday")
case 5:
print("Friday")
case 6:
print("Saturday")
case 7:
print("Sunday")
Default Value
Use the underscore character _ as the last case value if you want a code
block to execute when there are not other matches:
Example
day = 4
match day:
case 6:
print("Today is Saturday")
case 7:
print("Today is Sunday")
case _:
print("Looking forward to the Weekend")
Combine Values
Use the pipe character | as an or operator in the case evaluation to check
for more than one value match in one case:
Example
day = 4
match day:
case 1 | 2 | 3 | 4 | 5:
print("Today is a weekday")
case 6 | 7:
print("I love weekends!")
If Statements as Guards
You can add if statements in the case evaluation as an extra condition-
check:
Example
month = 5
day = 4
match day:
case 1 | 2 | 3 | 4 | 5 if month == 4:
print("A weekday in April")
case 1 | 2 | 3 | 4 | 5 if month == 5:
print("A weekday in May")
case _:
print("No match")
while loops
for loops
i = 1
while i < 6:
print(i)
i += 1
Example
Exit the loop when i is 3:
i = 1
while i < 6:
print(i)
if i == 3:
break
i += 1
Example
Continue to the next iteration if i is 3:
i = 0
while i < 6:
i += 1
if i == 3:
continue
print(i)
Example
Print a message once the condition is false:
i = 1
while i < 6:
print(i)
i += 1
else:
print("i is no longer less than 6")
Note: The else block will NOT be executed if the loop is stopped by
a break statement.
This is less like the for keyword in other programming languages, and works
more like an iterator method as found in other object-orientated
programming languages.
With the for loop we can execute a set of statements, once for each item in
a list, tuple, set etc.
The for loop does not require an indexing variable to set beforehand.
Example
Loop through the letters in the word "banana":
for x in "banana":
print(x)
Example
Exit the loop when x is "banana":
Example
Exit the loop when x is "banana", but this time the break comes before the
print:
Example
Do not print banana:
fruits = ["apple", "banana", "cherry"]
for x in fruits:
if x == "banana":
continue
print(x)
Example
Using the range() function:
for x in range(6):
print(x)
Note that range(6) is not the values of 0 to 6, but the values 0 to 5.
Example
Using the start parameter:
Example
Increment the sequence with 3 (default is 1):
for x in range(2, 30, 3):
print(x)
Example
Print all numbers from 0 to 5, and print a message when the loop has ended:
for x in range(6):
print(x)
else:
print("Finally finished!")
Note: The else block will NOT be executed if the loop is stopped by
a break statement.
Example
Break the loop when x is 3, and see what happens with the else block:
for x in range(6):
if x == 3: break
print(x)
else:
print("Finally finished!")
Nested Loops
A nested loop is a loop inside a loop.
The "inner loop" will be executed one time for each iteration of the "outer
loop":
Example
Print each adjective for every fruit:
adj = ["red", "big", "tasty"]
fruits = ["apple", "banana", "cherry"]
for x in adj:
for y in fruits:
print(x, y)
Example
for x in [0, 1, 2]:
pass
Python Functions
Python Functions
A function is a block of code which only runs when it is called.
Creating a Function
In Python, a function is defined using the def keyword, followed by a function
name and parentheses:
Example
def my_function():
print("Hello from a function")
This creates a function named my_function that prints "Hello from a
function" when called.
The code inside the function must be indented. Python uses indentation to
define code blocks.
Calling a Function
To call a function, write its name followed by parentheses:
Example
def my_function():
print("Hello from a function")
my_function()
Example
def my_function():
print("Hello from a function")
my_function()
my_function()
my_function()
Function Names
Function names follow the same rules as variable names in Python:
Example
Without functions - repetitive code:
temp1 = 77
celsius1 = (temp1 - 32) * 5 / 9
print(celsius1)
temp2 = 95
celsius2 = (temp2 - 32) * 5 / 9
print(celsius2)
temp3 = 50
celsius3 = (temp3 - 32) * 5 / 9
print(celsius3)
With functions, you write the code once and reuse it:
Example
With functions - reusable code:
def fahrenheit_to_celsius(fahrenheit):
return (fahrenheit - 32) * 5 / 9
print(fahrenheit_to_celsius(77))
print(fahrenheit_to_celsius(95))
print(fahrenheit_to_celsius(50))
Return Values
Functions can send data back to the code that called them using
the return statement.
Example
A function that returns a value:
def get_greeting():
return "Hello from a function"
message = get_greeting()
print(message)
Example
Using the return value directly:
def get_greeting():
return "Hello from a function"
print(get_greeting())
If a function doesn't have a return statement, it returns None by default.
Example
def my_function():
pass
The pass statement is often used when developing, allowing you to define
the structure first and implement details later.
Arguments are specified after the function name, inside the parentheses.
You can add as many arguments as you want, just separate them with a
comma.
The following example has a function with one argument (fname). When the
function is called, we pass along a first name, which is used inside the
function to print the full name:
def my_function(fname):
print(fname + " Refsnes")
my_function("Emil")
my_function("Tobias")
my_function("Linus")
Parameters vs Arguments
The terms parameter and argument can be used for the same thing:
information that are passed into a function.
Number of Arguments
By default, a function must be called with the correct number of arguments.
Example
This function expects 2 arguments, and gets 2 arguments::
my_function("Emil", "Refsnes")
If you try to call the function with the wrong number of arguments, you will
get an error:
Example
This function expects 2 arguments, but gets only 1:
my_function("Emil")
Example
def my_function(name = "friend"):
print("Hello", name)
my_function("Emil")
my_function("Tobias")
my_function()
my_function("Linus")
Example
Default value for country parameter:
my_function("Sweden")
my_function("India")
my_function()
my_function("Brazil")
Keyword Arguments
You can send arguments with the key = value syntax.
Example
def my_function(animal, name):
print("I have a", animal)
print("My", animal + "'s name is", name)
This way, with keyword arguments, the order of the arguments does not
matter.
Example
def my_function(animal, name):
print("I have a", animal)
print("My", animal + "'s name is", name)
Example
def my_function(animal, name):
print("I have a", animal)
print("My", animal + "'s name is", name)
my_function("dog", "Buddy")
Example
Switching the order changes the result:
my_function("Buddy", "dog")
Example
def my_function(animal, name, age):
print("I have a", age, "year old", animal, "named", name)
Example
Sending a list as an argument:
def my_function(fruits):
for fruit in fruits:
print(fruit)
Example
Sending a dictionary as an argument:
def my_function(person):
print("Name:", person["name"])
print("Age:", person["age"])
Return Values
Functions can return values using the return statement:
Example
def my_function(x, y):
return x + y
result = my_function(5, 3)
print(result)
Returning Different Data Types
Functions can return any data type, including lists, tuples, dictionaries, and
more.
Example
A function that returns a list:
def my_function():
return ["apple", "banana", "cherry"]
fruits = my_function()
print(fruits[0])
print(fruits[1])
print(fruits[2])
Example
A function that returns a tuple:
def my_function():
return (10, 20)
x, y = my_function()
print("x:", x)
print("y:", y)
Positional-Only Arguments
You can specify that a function can have ONLY positional arguments.
Example
def my_function(name, /):
print("Hello", name)
my_function("Emil")
Without the , / you are actually allowed to use keyword arguments even if
the function expects positional arguments:
Example
def my_function(name):
print("Hello", name)
my_function(name = "Emil")
With , /, you will get an error if you try to use keyword arguments:
Example
def my_function(name, /):
print("Hello", name)
my_function(name = "Emil")
Keyword-Only Arguments
To specify that a function can have only keyword arguments,
add *, before the arguments:
Example
def my_function(*, name):
print("Hello", name)
my_function(name = "Emil")
Without *,, you are allowed to use positional arguments even if the function
expects keyword arguments:
Example
def my_function(name):
print("Hello", name)
my_function("Emil")
With *,, you will get an error if you try to use positional arguments:
Example
def my_function(*, name):
print("Hello", name)
my_function("Emil")
Example
def my_function(a, b, /, *, c, d):
return a + b + c + d
However, sometimes you may not know how many arguments that will be
passed into your function.
This way, the function will receive a tuple of arguments and can access the
items accordingly:
Example
Using *args to accept any number of arguments:
def my_function(*kids):
print("The youngest child is " + kids[2])
What is *args?
The *args parameter allows a function to accept any number of positional
arguments.
Inside the function, args becomes a tuple containing all the passed
arguments:
Example
Accessing individual arguments from *args:
def my_function(*args):
print("Type:", type(args))
print("First argument:", args[0])
print("Second argument:", args[1])
print("All arguments:", args)
Example
def my_function(greeting, *names):
for name in names:
print(greeting, name)
In this example, "Hello" is assigned to greeting, and the rest are collected
in names.
Example
A function that calculates the sum of any number of values:
def my_function(*numbers):
total = 0
for num in numbers:
total += num
return total
print(my_function(1, 2, 3))
print(my_function(10, 20, 30, 40))
print(my_function(5))
Example
Finding the maximum value:
def my_function(*numbers):
if len(numbers) == 0:
return None
max_num = numbers[0]
for num in numbers:
if num > max_num:
max_num = num
return max_num
print(my_function(3, 7, 2, 9, 1))
This way, the function will receive a dictionary of arguments and can access
the items accordingly:
Example
Using **kwargs to accept any number of keyword arguments:
def my_function(**kid):
print("His last name is " + kid["lname"])
What is **kwargs?
The **kwargs parameter allows a function to accept any number of keyword
arguments.
Inside the function, kwargs becomes a dictionary containing all the keyword
arguments:
Example
Accessing values from **kwargs:
def my_function(**myvar):
print("Type:", type(myvar))
print("Name:", myvar["name"])
print("Age:", myvar["age"])
print("All data:", myvar)
Example
def my_function(username, **details):
print("Username:", username)
print("Additional details:")
for key, value in [Link]():
print(" ", key + ":", value)
1. regular parameters
2. *args
3. **kwargs
Example
def my_function(title, *args, **kwargs):
print("Title:", title)
print("Positional arguments:", args)
print("Keyword arguments:", kwargs)
Example
Using * to unpack a list into arguments:
numbers = [1, 2, 3]
result = my_function(*numbers) # Same as: my_function(1, 2, 3)
print(result)
Example
Using ** to unpack a dictionary into keyword arguments:
Python Scope
Scope
A variable is only available from inside the region it is created. This is
called scope.
Local Scope
A variable created inside a function belongs to the local scope of that
function, and can only be used inside that function.
Example
A variable created inside a function is available inside that function:
def myfunc():
x = 300
print(x)
myfunc()
Example
The local variable can be accessed from a function within the function:
def myfunc():
x = 300
def myinnerfunc():
print(x)
myinnerfunc()
myfunc()
Global Scope
A variable created in the main body of the Python code is a global variable
and belongs to the global scope.
Global variables are available from within any scope, global and local.
Example
A variable created outside of a function is global and can be used by anyone:
x = 300
def myfunc():
print(x)
myfunc()
print(x)
Naming Variables
If you operate with the same variable name inside and outside of a function,
Python will treat them as two separate variables, one available in the global
scope (outside the function) and one available in the local scope (inside the
function):
Example
The function will print the local x, and then the code will print the global x:
x = 300
def myfunc():
x = 200
print(x)
myfunc()
print(x)
Global Keyword
If you need to create a global variable, but are stuck in the local scope, you
can use the global keyword.
The global keyword makes the variable global.
Example
If you use the global keyword, the variable belongs to the global scope:
def myfunc():
global x
x = 300
myfunc()
print(x)
Also, use the global keyword if you want to make a change to a global
variable inside a function.
Example
To change the value of a global variable inside a function, refer to the
variable by using the global keyword:
x = 300
def myfunc():
global x
x = 200
myfunc()
print(x)
Nonlocal Keyword
The nonlocal keyword is used to work with variables inside nested
functions.
The nonlocal keyword makes the variable belong to the outer function.
Example
If you use the nonlocal keyword, the variable will belong to the outer
function:
def myfunc1():
x = "Jane"
def myfunc2():
nonlocal x
x = "hello"
myfunc2()
return x
print(myfunc1())
x = "global"
def outer():
x = "enclosing"
def inner():
x = "local"
print("Inner:", x)
inner()
print("Outer:", x)
outer()
print("Global:", x)
Python Decorators
Decorators let you add extra behavior to a function, without changing
the function's code.
A decorator is a function that takes another function as input and
returns a new function.
Basic Decorator
Define the decorator first, then apply it with @decorator_name above the
function.
Exampl
A basic decorator that uppercases the return value of the decorated function.
def changecase(func):
def myinner():
return func().upper()
return myinner
@changecase
def myfunction():
return "Hello Sally"
print(myfunction())
Example
Using the @changecase decorator on two functions:
def changecase(func):
def myinner():
return func().upper()
return myinner
@changecase
def myfunction():
return "Hello Sally"
@changecase
def otherfunction():
return "I am speed!"
print(myfunction())
print(otherfunction())
Example
Functions with arguments can also be decorated:
def changecase(func):
def myinner(x):
return func(x).upper()
return myinner
@changecase
def myfunction(nam):
return "Hello " + nam
print(myfunction("John"))
def changecase(func):
def myinner(*args, **kwargs):
return func(*args, **kwargs).upper()
return myinner
@changecase
def myfunction(nam):
return "Hello " + nam
print(myfunction("John"))
Example
A decorator factory that takes an argument and transforms the casing based
on the argument value.
def changecase(n):
def changecase(func):
def myinner():
if n == 1:
a = func().lower()
else:
a = func().upper()
return a
return myinner
return changecase
@changecase(1)
def myfunction():
return "Hello Linus"
print(myfunction())
Multiple Decorators
You can use multiple decorators on one function.
Decorators are called in the reverse order, starting with the one closest to
the function.
Example
One decorator for upper case, and one for adding a greeting:
def changecase(func):
def myinner():
return func().upper()
return myinner
def addgreeting(func):
def myinner():
return "Hello " + func() + " Have a good day!"
return myinner
@changecase
@addgreeting
def myfunction():
return "Tobias"
print(myfunction())
Example
Normally, a function's name can be returned with the __name__ attribute:
def myfunction():
return "Have a great day!"
print(myfunction.__name__)
But, when a function is decorated, the metadata of the original function is
lost.
Example
Try returning the name from a decorated function and you will not get the
same result:
def changecase(func):
def myinner():
return func().upper()
return myinner
@changecase
def myfunction():
return "Have a great day!"
print(myfunction.__name__)
To fix this, Python has a built-in function called [Link] that can be
used to preserve the original function's name and docstring.
Example
Import [Link] to preserve the original function name and docstring.
import functools
def changecase(func):
@[Link](func)
def myinner():
return func().upper()
return myinner
@changecase
def myfunction():
return "Have a great day!"
print(myfunction.__name__)
Python Lambda
Lambda Functions
A lambda function is a small anonymous function.
A lambda function can take any number of arguments, but can only have one
expression.
Syntax
lambda arguments : expression
Example
Add 10 to argument a, and return the result:
x = lambda a : a + 10
print(x(5))
Example
Multiply argument a with argument b and return the result:
x = lambda a, b : a * b
print(x(5, 6))
Example
Summarize argument a, b, and c and return the result:
x = lambda a, b, c : a + b + c
print(x(5, 6, 2))
Say you have a function definition that takes one argument, and that
argument will be multiplied with an unknown number:
def myfunc(n):
return lambda a : a * n
Use that function definition to make a function that always doubles the
number you send in:
Example
def myfunc(n):
return lambda a : a * n
mydoubler = myfunc(2)
print(mydoubler(11))
Example
def myfunc(n):
return lambda a : a * n
mytripler = myfunc(3)
print(mytripler(11))
Or, use the same function definition to make both functions, in the same
program:
Example
def myfunc(n):
return lambda a : a * n
mydoubler = myfunc(2)
mytripler = myfunc(3)
print(mydoubler(11))
print(mytripler(11))
Use lambda functions when an anonymous function is required for a short
period of time.
Example
Double all numbers in a list:
numbers = [1, 2, 3, 4, 5]
doubled = list(map(lambda x: x * 2, numbers))
print(doubled)
Example
Filter out odd numbers from a list:
numbers = [1, 2, 3, 4, 5, 6, 7, 8]
odd_numbers = list(filter(lambda x: x % 2 != 0, numbers))
print(odd_numbers)
Example
Sort a list of tuples by the second element:
Example
Sort strings by length:
words = ["apple", "pie", "banana", "cherry"]
sorted_words = sorted(words, key=lambda x: len(x))
print(sorted_words)
Python Recursion
Recursion
Recursion is when a function calls itself.
The developer should be very careful with recursion as it can be quite easy
to slip into writing a function which never terminates, or one that uses
excess amounts of memory or processor power. However, when written
correctly recursion can be a very efficient and mathematically-elegant
approach to programming.
def countdown(n):
if n <= 0:
print("Done!")
else:
print(n)
countdown(n - 1)
countdown(5)
Example
Identifying base case and recursive case:
def factorial(n):
# Base case
if n == 0 or n == 1:
return 1
# Recursive case
else:
return n * factorial(n - 1)
print(factorial(5))
The base case is crucial. Always make sure your recursive function has a
condition that will eventually be met.
Fibonacci Sequence
The Fibonacci sequence is a classic example where each number is the sum
of the two preceding ones. The sequence starts with 0 and 1:
0, 1, 1, 2, 3, 5, 8, 13, ...
The sequence continues indefinitely, with each number being the sum of the
two preceding ones.
Example
Find the 7th number in the Fibonacci sequence:
def fibonacci(n):
if n <= 1:
return n
else:
return fibonacci(n - 1) + fibonacci(n - 2)
print(fibonacci(7))
Recursion with Lists
Recursion can be used to process lists by handling one element at a time:
Example
Calculate the sum of all elements in a list:
def sum_list(numbers):
if len(numbers) == 0:
return 0
else:
return numbers[0] + sum_list(numbers[1:])
my_list = [1, 2, 3, 4, 5]
print(sum_list(my_list))
Example
Find the maximum value in a list:
def find_max(numbers):
if len(numbers) == 1:
return numbers[0]
else:
max_of_rest = find_max(numbers[1:])
return numbers[0] if numbers[0] > max_of_rest else max_of_rest
my_list = [3, 7, 2, 9, 1]
print(find_max(my_list))
Example
Check the recursion limit:
import sys
print([Link]())
If you need deeper recursion, you can increase the limit, but be careful as
this can cause crashes:
Example
import sys
[Link](2000)
print([Link]())
Increasing the recursion limit should be done with caution. For very deep
recursion, consider using iteration instead.
Python Generators
Generators
Generators are functions that can pause and resume their execution.
The code inside the function is not executed yet, it is only compiled. The
function only executes when you iterate over the generator.
def my_generator():
yield 1
yield 2
yield 3
When yield is encountered, the function's state is saved, and the value is
returned. The next time the generator is called, it continues from where it
left off.
Example
Generator that yields numbers:
def count_up_to(n):
count = 1
while count <= n:
yield count
count += 1
Example
Generator for large sequences:
def large_sequence(n):
for i in range(n):
yield i
Example
def simple_gen():
yield "Emil"
yield "Tobias"
yield "Linus"
gen = simple_gen()
print(next(gen))
print(next(gen))
print(next(gen))
Example
def simple_gen():
yield 1
yield 2
gen = simple_gen()
print(next(gen))
print(next(gen))
print(next(gen)) # This will raise StopIteration
Generator Expressions
Similar to list comprehensions, you can create generators using generator
expressions with parentheses instead of square brackets:
Example
List comprehension vs generator expression:
Example
Using a generator expression with sum:
Example
Generate 100 Fibonacci numbers:
def fibonacci():
a, b = 0, 1
while True:
yield a
a, b = b, a + b
Generator Methods
Generators have special methods for advanced control:
send() Method
The send() method allows you to send a value to the generator:
Example
def echo_generator():
while True:
received = yield
print("Received:", received)
gen = echo_generator()
next(gen) # Prime the generator
[Link]("Hello")
[Link]("World")
close() Method
The close() method stops the generator:
Example
def my_gen():
try:
yield 1
yield 2
yield 3
finally:
print("Generator closed")
gen = my_gen()
print(next(gen))
[Link]()
Python range
Python range
The built-in range() function returns an immutable sequence of numbers,
commonly used for looping a specific number of times.
This set of numbers has its own data type called range.
x = range(10)
Example
Create a range of numbers from 3 to 9:
x = range(3, 10)
The step value means the difference between each number in the sequence.
It is optional, and if not provided, it defaults to 1.
Example
Create a range of numbers from 3 to 9:
x = range(3, 10, 2)
Using ranges
Ranges are often used in for loops to iterate over a sequence of numbers.
Example
Iterate over each value in a range:
for i in range(10):
print(i)
Example
Convert different ranges to lists:
print(list(range(5)))
print(list(range(1, 6)))
print(list(range(5, 20, 3)))
Slicing Ranges
Like other sequences, ranges can be sliced to extract a subsequence.
Example
Extract a subsequence from a range:
r = range(10)
print(r[2])
print(r[:3])
Note: The first print statement returns the value at index 2, and the second
print statement returns a new range object, from index 0 to 3.
Membership Testing
Ranges support membership testing with the in operator.
Example
Test if the numbers 6 and 7 are present in a range:
r = range(0, 10, 2)
print(6 in r)
print(7 in r)
The return value is True when the number is present in the range,
and False when it is not.
Length
Ranges support the len() function to get the number of elements in the
range.
Example
Get the length of a range:
r = range(0, 10, 2)
print(len(r))
Python Arrays
Note: Python does not have built-in support for Arrays, but Python Lists can
be used instead.
Arrays
Note: This page shows you how to use LISTS as ARRAYS, however, to work
with arrays in Python you will have to import a library, like the NumPy library.
What is an Array?
An array is a special variable, which can hold more than one value at a time.
If you have a list of items (a list of car names, for example), storing the cars
in single variables could look like this:
car1 = "Ford"
car2 = "Volvo"
car3 = "BMW"
However, what if you want to loop through the cars and find a specific one?
And what if you had not 3 cars, but 300?
An array can hold many values under a single name, and you can access the
values by referring to an index number.
Example
Get the value of the first array item:
x = cars[0]
Example
Modify the value of the first array item:
cars[0] = "Toyota"
Example
Return the number of elements in the cars array:
x = len(cars)
Note: The length of an array is always one more than the highest array
index.
Example
Print each item in the cars array:
for x in cars:
print(x)
Example
Add one more element to the cars array:
[Link]("Honda")
Example
Delete the second element of the cars array:
[Link](1)
You can also use the remove() method to remove an element from the
array.
Example
Delete the element that has the value "Volvo":
[Link]("Volvo")
Note: The list's remove() method only removes the first occurrence of the
specified value.
Array Methods
Python has a set of built-in methods that you can use on lists/arrays.
Method Description
extend() Add the elements of a list (or any iterable), to the end of the current lis
index() Returns the index of the first element with the specified value
Note: Python does not have built-in support for Arrays, but Python Lists can
be used instead.
Python Iterators
Python Iterators
An iterator is an object that contains a countable number of values.
An iterator is an object that can be iterated upon, meaning that you can
traverse through all the values.
Iterator vs Iterable
Lists, tuples, dictionaries, and sets are all iterable objects. They are
iterable containers which you can get an iterator from.
All these objects have a iter() method which is used to get an iterator:
print(next(myit))
print(next(myit))
print(next(myit))
Example
Strings are also iterable objects, containing a sequence of characters:
mystr = "banana"
myit = iter(mystr)
print(next(myit))
print(next(myit))
print(next(myit))
print(next(myit))
print(next(myit))
print(next(myit))
Example
Iterate the values of a tuple:
mytuple = ("apple", "banana", "cherry")
for x in mytuple:
print(x)
Example
Iterate the characters of a string:
mystr = "banana"
for x in mystr:
print(x)
Create an Iterator
To create an object/class as an iterator you have to implement the
methods __iter__() and __next__() to your object.
As you will learn in the Python Classes/Objects chapter, all classes have a
function called __init__(), which allows you to do some initializing when
the object is being created.
The __iter__() method acts similar, you can do operations (initializing etc.),
but must always return the iterator object itself.
The __next__() method also allows you to do operations, and must return
the next item in the sequence.
Example
Create an iterator that returns numbers, starting with 1, and each sequence
will increase by one (returning 1,2,3,4,5 etc.):
class MyNumbers:
def __iter__(self):
self.a = 1
return self
def __next__(self):
x = self.a
self.a += 1
return x
myclass = MyNumbers()
myiter = iter(myclass)
print(next(myiter))
print(next(myiter))
print(next(myiter))
print(next(myiter))
print(next(myiter))
StopIteration
The example above would continue forever if you had enough next()
statements, or if it was used in a for loop.
Example
Stop after 20 iterations:
class MyNumbers:
def __iter__(self):
self.a = 1
return self
def __next__(self):
if self.a <= 20:
x = self.a
self.a += 1
return x
else:
raise StopIteration
myclass = MyNumbers()
myiter = iter(myclass)
for x in myiter:
print(x)
Python Modules
What is a Module?
Consider a module to be the same as a code library.
Create a Module
To create a module just save the code you want in a file with the file
extension .py:
def greeting(name):
print("Hello, " + name)
Use a Module
Now we can use the module we just created, by using the import statement:
Example
Import the module named mymodule, and call the greeting function:
import mymodule
[Link]("Jonathan")
Note: When using a function from a module, use the
syntax: module_name.function_name.
Variables in Module
The module can contain functions, as already described, but also variables of
all types (arrays, dictionaries, objects etc):
Example
Save this code in the file [Link]
person1 = {
"name": "John",
"age": 36,
"country": "Norway"
}
Example
Import the module named mymodule, and access the person1 dictionary:
import mymodule
a = mymodule.person1["age"]
print(a)
Naming a Module
You can name the module file whatever you like, but it must have the file
extension .py
Re-naming a Module
You can create an alias when you import a module, by using the as keyword:
Example
Create an alias for mymodule called mx:
import mymodule as mx
a = mx.person1["age"]
print(a)
Built-in Modules
There are several built-in modules in Python, which you can import whenever
you like.
Example
Import and use the platform module:
import platform
x = [Link]()
print(x)
Example
List all the defined names belonging to the platform module:
import platform
x = dir(platform)
print(x)
Note: The dir() function can be used on all modules, also the ones you
create yourself.
Example
The module named mymodule has one function and one dictionary:
def greeting(name):
print("Hello, " + name)
person1 = {
"name": "John",
"age": 36,
"country": "Norway"
}
Example
Import only the person1 dictionary from the module:
print (person1["age"])
Note: When importing using the from keyword, do not use the module name
when referring to elements in the module.
Example: person1["age"], not mymodule.person1["age"]
Python Datetime
Python Dates
A date in Python is not a data type of its own, but we can import a module
named datetime to work with dates as date objects.
import datetime
x = [Link]()
print(x)
Date Output
When we execute the code from the example above the result will be:
2026-05-24 14:22:42.674664
The date contains year, month, day, hour, minute, second, and microsecond.
The datetime module has many methods to return information about the
date object.
Here are a few examples, you will learn more about them later in this
chapter:
Example
Return the year and name of weekday:
import datetime
x = [Link]()
print([Link])
print([Link]("%A"))
Example
Create a date object:
import datetime
x = [Link](2020, 5, 17)
print(x)
The datetime() class also takes parameters for time and timezone (hour,
minute, second, microsecond, tzone), but they are optional, and has a
default value of 0, (None for timezone).
Example
Display the name of the month:
import datetime
x = [Link](2018, 6, 1)
print([Link]("%B"))
%H Hour 00-23 17
%I Hour 00-12 05
%p AM/PM PM
%M Minute 00-59 41
%S Second 00-59 08
%f Microsecond 000000-999999 548513
%Z Timezone CST
%C Century 20
%% A % character %
%G ISO 8601 year 2018
Python Math
Python has a set of built-in math functions, including an extensive
math module, that allows you to perform mathematical tasks on
numbers.
print(x)
print(y)
The abs() function returns the absolute (positive) value of the specified
number:
Example
x = abs(-7.25)
print(x)
x = pow(4, 3)
print(x)
import math
When you have imported the math module, you can start using methods and
constants of the module.
The [Link]() method for example, returns the square root of a number:
Example
import math
x = [Link](64)
print(x)
Example
import math
x = [Link](1.4)
y = [Link](1.4)
print(x) # returns 2
print(y) # returns 1
The [Link] constant, returns the value of PI (3.14...):
Example
import math
x = [Link]
print(x)
Python JSON
JSON is a syntax for storing and exchanging data.
JSON in Python
Python has a built-in package called json, which can be used to work with
JSON data.
import json
Example
Convert from JSON to Python:
import json
# some JSON:
x = '{ "name":"John", "age":30, "city":"New York"}'
# parse x:
y = [Link](x)
Example
Convert from Python to JSON:
import json
You can convert Python objects of the following types, into JSON strings:
dict
list
tuple
string
int
float
True
False
None
Example
Convert Python objects into JSON strings, and print the values:
import json
When you convert from Python to JSON, Python objects are converted into
the JSON (JavaScript) equivalent:
Python JSON
dict Object
list Array
tuple Array
str String
int Number
float Number
True true
False false
None null
Example
Convert a Python object containing all the legal data types:
import json
x = {
"name": "John",
"age": 30,
"married": True,
"divorced": False,
"children": ("Ann","Billy"),
"pets": None,
"cars": [
{"model": "BMW 230", "mpg": 27.5},
{"model": "Ford Edge", "mpg": 24.1}
]
}
print([Link](x))
The [Link]() method has parameters to make it easier to read the result:
Example
Use the indent parameter to define the numbers of indents:
[Link](x, indent=4)
You can also define the separators, default value is (", ", ": "), which means
using a comma and a space to separate each object, and a colon and a
space to separate keys from values:
Example
Use the separators parameter to change the default separator:
Example
Use the sort_keys parameter to specify if the result should be sorted or not:
Python RegEx
A RegEx, or Regular Expression, is a sequence of characters that forms
a search pattern.
RegEx Module
Python has a built-in package called re, which can be used to work with
Regular Expressions.
import re
RegEx in Python
When you have imported the re module, you can start using regular
expressions:
import re
RegEx Functions
The re module offers a set of functions that allows us to search a string for a
match:
Function Description
split Returns a list where the string has been split at each match
Metacharacters
Metacharacters are characters with a special meaning:
Character Description
[] A set of characters
$ Ends with
| Either or
Flags
You can add flags to the pattern when using regular expressions.
[Link] re.S Makes the . character match all characters (including new
Special Sequences
A special sequence is a \ followed by one of the characters in the list below,
and has a special meaning:
Character Description
\A Returns a match if the specified characters are at the beginning of the str
\B Returns a match where the specified characters are present, but NOT at th
beginning (or at the end) of a word
(the "r" in the beginning is making sure that the string is being treated as
string")
\d Returns a match where the string contains digits (numbers from 0-9)
\S Returns a match where the string DOES NOT contain a white space charac
\w Returns a match where the string contains any word characters (characte
to Z, digits from 0-9, and the underscore _ character)
\W Returns a match where the string DOES NOT contain any word characters
\Z Returns a match if the specified characters are at the end of the string
Sets
A set is a set of characters inside a pair of square brackets [] with a special
meaning:
Set Description
[arn] Returns a match where one of the specified characters (a, r, or n) is prese
[a-n] Returns a match for any lower case character, alphabetically between a a
[0123] Returns a match where any of the specified digits (0, 1, 2, or 3) are presen
[+] In sets, +, *, ., |, (), $,{} has no special meaning, so [+] means: return a
any + character in the string
Example
Print a list of all matches:
import re
The list contains the matches in the order they are found.
Example
Return an empty list if no match was found:
import re
If there is more than one match, only the first occurrence of the match will
be returned:
Example
Search for the first white-space character in the string:
import re
Example
Make a search that returns no match:
import re
Example
Split at each white-space character:
import re
Example
Split the string only at the first occurrence:
import re
Example
Replace every white-space character with the number 9:
import re
Example
Replace the first 2 occurrences:
import re
Match Object
A Match Object is an object containing information about the search and the
result.
Note: If there is no match, the value None will be returned, instead of the
Match Object.
Example
Do a search that will return a Match Object:
import re
The Match object has properties and methods used to retrieve information
about the search, and the result:
.span() returns a tuple containing the start-, and end positions of the
match.
.string returns the string passed into the function
.group() returns the part of the string where there was a match
Example
Print the position (start- and end-position) of the first match occurrence.
The regular expression looks for any words that starts with an upper case
"S":
import re
Example
Print the string passed into the function:
import re
Example
Print the part of the string where there was a match.
The regular expression looks for any words that starts with an upper case
"S":
import re
Python PIP
What is PIP?
PIP is a package manager for Python packages, or modules if you like.
Note: If you have Python version 3.4 or later, PIP is included by default.
What is a Package?
A package contains all the files you need for a module.
Modules are Python code libraries you can include in your project.
C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>pip --version
Install PIP
If you do not have PIP installed, you can download and install it from this
page: [Link]
Download a Package
Downloading a package is very easy.
Open the command line interface and tell PIP to download the package you
want.
Navigate your command line to the location of Python's script directory, and
type the following:
Example
Download a package named "camelcase":
C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>pip install camelcase
REMOVE ADS
Using a Package
Once the package is installed, it is ready to use.
Example
Import and use "camelcase":
import camelcase
c = [Link]()
print([Link](txt))
Find Packages
Find more packages at [Link]
Remove a Package
Use the uninstall command to remove a package:
Example
Uninstall the package named "camelcase":
C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>pip uninstall camelcase
The PIP Package Manager will ask you to confirm that you want to remove
the camelcase package:
Uninstalling camelcase-02.1:
Would remove:
c:\users\Your Name\appdata\local\programs\python\python36-32\lib\
site-packages\[Link]-info
c:\users\Your Name\appdata\local\programs\python\python36-32\lib\
site-packages\camelcase\*
Proceed (y/n)?
List Packages
Use the list command to list all the packages installed on your system:
Example
List installed packages:
C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>pip list
Result:
Package Version
-----------------------
camelcase 0.2
mysql-connector 2.1.6
pip 18.1
pymongo 3.6.1
setuptools 39.0.1
The try block lets you test a block of code for errors.
The else block lets you execute code when there is no error.
The finally block lets you execute code, regardless of the result of
the try- and except blocks.
Exception Handling
When an error occurs, or exception as we call it, Python will normally stop
and generate an error message.
try:
print(x)
except:
print("An exception occurred")
Since the try block raises an error, the except block will be executed.
Without the try block, the program will crash and raise an error:
Example
This statement will raise an error, because x is not defined:
print(x)
Many Exceptions
You can define as many exception blocks as you want, e.g. if you want to
execute a special block of code for a special kind of error:
Example
Print one message if the try block raises a NameError and another for other
errors:
try:
print(x)
except NameError:
print("Variable x is not defined")
except:
print("Something else went wrong")
Else
You can use the else keyword to define a block of code to be executed if no
errors were raised:
Example
In this example, the try block does not generate any error:
try:
print("Hello")
except:
print("Something went wrong")
else:
print("Nothing went wrong")
Finally
The finally block, if specified, will be executed regardless if the try block
raises an error or not.
Example
try:
print(x)
except:
print("Something went wrong")
finally:
print("The 'try except' is finished")
Example
Try to open and write to a file that is not writable:
try:
f = open("[Link]")
try:
[Link]("Lorum Ipsum")
except:
print("Something went wrong when writing to the file")
finally:
[Link]()
except:
print("Something went wrong when opening the file")
The program can continue, without leaving the file object open.
Raise an exception
As a Python developer you can choose to throw an exception if a condition
occurs.
Example
Raise an error and stop the program if x is lower than 0:
x = -1
if x < 0:
raise Exception("Sorry, no numbers below zero")
You can define what kind of error to raise, and the text to print to the user.
Example
Raise a TypeError if x is not an integer:
x = "hello"
F-String was introduced in Python 3.6, and is now the preferred way of
formatting strings.
F-Strings
F-string allows you to format selected parts of a string.
Example
Add a placeholder for the price variable:
price = 59
txt = f"The price is {price} dollars"
print(txt)
Example
Display the price with 2 decimals:
price = 59
txt = f"The price is {price:.2f} dollars"
print(txt)
Example
Display the value 95 with 2 decimals:
Example
Perform a math operation in the placeholder, and return the result:
Example
Add taxes before displaying the price:
price = 59
tax = 0.25
txt = f"The price is {price + (price * tax)} dollars"
print(txt)
Example
Return "Expensive" if the price is over 50, otherwise return "Cheap":
price = 49
txt = f"It is very {'Expensive' if price>50 else 'Cheap'}"
print(txt)
Example
Use the string method upper()to convert a value into upper case letters:
fruit = "apples"
txt = f"I love {[Link]()}"
print(txt)
The function does not have to be a built-in Python method, you can create
your own functions and use them:
Example
Create a function that converts feet into meters:
def myconverter(x):
return x * 0.3048
More Modifiers
At the beginning of this chapter we explained how to use the .2f modifier to
format a number into a fixed point number with 2 decimals.
There are several other modifiers that can be used to format values:
Example
Use a comma as a thousand separator:
price = 59000
txt = f"The price is {price:,} dollars"
print(txt)
Formatting Types
: Use a space to insert an extra space before positive numbers (and a minus sign
:b Binary format
:d Decimal format
:F Fix point number format, in uppercase format (show inf and nan as INF and NA
:g General format
:o Octal format
:n Number format
:% Percentage format
String format()
Before Python 3.6 we used the format() method to format strings.
The format() method can still be used, but f-strings are faster and the
preferred way to format strings.
The next examples in this page demonstrates how to format strings with
the format() method.
The format() method also uses curly brackets as placeholders {}, but the
syntax is slightly different:
Example
Add a placeholder where you want to display the price:
price = 49
txt = "The price is {} dollars"
print([Link](price))
You can add parameters inside the curly brackets to specify how to convert
the value:
Example
Format the price to be displayed as a number with two decimals:
Multiple Values
If you want to use more values, just add more values to
the format() method:
Example
quantity = 3
itemno = 567
price = 49
myorder = "I want {} pieces of item number {} for {:.2f} dollars."
print([Link](quantity, itemno, price))
Index Numbers
You can use index numbers (a number inside the curly brackets {0}) to be
sure the values are placed in the correct placeholders:
Example
quantity = 3
itemno = 567
price = 49
myorder = "I want {0} pieces of item number {1} for {2:.2f} dollars."
print([Link](quantity, itemno, price))
Also, if you want to refer to the same value more than once, use the index
number:
Example
age = 36
name = "John"
txt = "His name is {1}. {1} is {0} years old."
print([Link](age, name))
Named Indexes
You can also use named indexes by entering a name inside the curly
brackets {carname}, but then you must use names when you pass the
parameter values [Link](carname = "Ford"):
Example
myorder = "I have a {carname}, it is a {model}."
print([Link](carname = "Ford", model = "Mustang"))
Python None
Python None
None is a special constant in Python that represents the absence of a value.
Its data type is NoneType, and None is the only instance of a NoneType object.
NoneType
Variables can be assigned None to indicate "no value" or "not set".
x = None
print(x)
Example
Assign and print the data type of a None value:
x = None
print(type(x))
Comparing to None
To compare a value to None, use the identity operator is or is not
Example
Use the identity operator is for comparisons with None:
result = None
if result is None:
print("No result yet")
else:
print("Result is ready")
Example
Similar example, but using is not instead:
result = None
if result is not None:
print("Result is ready")
else:
print("No result yet")
True or False
None evaluates to False in a boolean context.
Example
Check truthiness:
print(bool(None))
Example
A function without a return statement returns None:
def myfunc():
x = 5
x = myfunc()
print(x)
Python User Input
User Input
Python allows for user input.
The following example asks for your name, and when you enter a name, it
gets printed on the screen:
Using prompt
In the example above, the user had to input their name on a new line. The
Python input() function has a prompt parameter, which acts as a message
you can put in front of the user input, on the same line:
Example
Add a message in front of the user input:
Example
Multiple inputs:
Input Number
The input from the user is treated as a string. Even if, in the example above,
you can input a number, the Python interpreter will still treat it as a string.
You can convert the input into a number with the float() function:
Example
To find the square root, the input has to be converted into a number:
x = input("Enter a number:")
Validate Input
It is a good practice to validate any input from the user. In the example
above, an error will occur if the user inputs something other than a number.
To avoid getting an error, we can test the input, and if it is not a number, the
user could get a message like "Wrong input, please try again", and allowed
to make a new input:
Example
Keep asking until you get a number:
y = True
while y == True:
x = input("Enter a number:")
try:
x = float(x);
y = False
except:
print("Wrong input, please try again.")
print("Thank you!")
Windows
macOS/Linux
C:\Users\Your Name> python -m venv myfirstproject
Result
The file/folder structure will look like this:
myfirstproject
Include
Lib
Scripts
.gitignore
[Link]
Example
Activate the virtual environment:
Windows
macOS/Linux
C:\Users\Your Name> myfirstproject\Scripts\activate
After activation, your prompt will change to show that you are now working
in the active environment:
Result
The command line will look like this when the virtual environment is active:
Windows
macOS/Linux
(myfirstproject) C:\Users\Your Name>
Install Packages
Once your virtual environment is activated, you can install packages in it,
using pip.
Example
Install 'cowsay' in the virtual environment:
Windows
macOS/Linux
(myfirstproject) C:\Users\Your Name> pip install cowsay
Result
'cowsay' is installed only in the virtual environment:
Collecting cowsay
Downloading [Link] (5.6 kB)
Downloading [Link] (25 kB)
Installing collected packages: cowsay
Successfully installed cowsay-6.1
[notice] A new release of pip is available: 25.0.1 -> 25.1.1
[notice] To update, run: [Link] -m pip install --upgrade pip
Using Package
Now that the 'cowsay' module is installed in your virtual environment, lets
use it to display a talking cow.
Create a file called [Link] on your computer. You can place it wherever you
want, but I will place it in the same location as the myfirstproject folder -
not in the folder, but in the same location.
Example
Insert two lines in [Link]:
[Link]
import cowsay
[Link]("Good Mooooorning!")
Then, try to execute the file while you are in the virtual environment:
Example
Execute [Link] in the virtual environment:
Windows
macOS/Linux
(myfirstproject) C:\Users\Your Name> python [Link]
Result
The purpose of the 'cowsay' module is to draw a cow that says whatever
input you give it:
_________________
| Good Mooooorning! |
=================
\
\
^__^
(oo)\_______
(__)\ )\/\
||----w |
|| ||
Example
Deactivate the virtual environment:
Windows
macOS/Linux
(myfirstproject) C:\Users\Your Name> deactivate
As a result, you are now back in the normal command line interface:
Result
Normal command line interface:
Windows
macOS/Linux
C:\Users\Your Name>
If you try to execute the [Link] file outside of the virtual environment, you
will get an error because 'cowsay' is missing. It was only installed in the
virtual environment:
Example
Execute [Link] outside of the virtual environment:
Windows
macOS/Linux
C:\Users\Your Name> python [Link]
Result
Error because 'cowsay' is missing:
To delete a virtual environment, you can simply delete its folder with all its
content. Either directly in the file system, or use the command line interface
like this:
Example
Delete myfirstproject from the command line interface:
Windows
macOS/Linux
C:\Users\Your Name> rmdir /s /q myfirstproject
Python OOP
What is OOP?
OOP stands for Object-Oriented Programming.
Tip: The DRY principle means you should avoid writing the same code more
than once. Move repeated code into functions or classes and reuse it.
A class defines what an object should look like, and an object is created
based on that class. For example:
Class Objects
When you create an object from a class, it inherits all the variables and
functions defined inside that class.
Create a Class
To create a class, use the keyword class:
class MyClass:
x = 5
Create Object
Now we can use the class named MyClass to create objects:
Example
Create an object named p1, and print the value of x:
p1 = MyClass()
print(p1.x)
Delete Objects
You can delete objects by using the del keyword:
Example
Delete the p1 object:
del p1
Multiple Objects
You can create multiple objects from the same class:
Example
Create three objects from the MyClass class:
p1 = MyClass()
p2 = MyClass()
p3 = MyClass()
print(p1.x)
print(p2.x)
print(p3.x)
Note: Each object is independent and has its own copy of the class
properties.
Example
class Person:
pass
class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age
p1 = Person("Emil", 36)
print([Link])
print([Link])
Note: The __init__() method is called automatically every time the class is
being used to create a new object.
Example
Create a class without __init__():
class Person:
pass
p1 = Person()
[Link] = "Tobias"
[Link] = 25
print([Link])
print([Link])
Using __init__() makes it easier to create objects with initial values:
Example
With __init__(), you can set initial values when creating the object:
class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age
p1 = Person("Linus", 28)
print([Link])
print([Link])
Example
Set a default value for the age parameter:
class Person:
def __init__(self, name, age=18):
[Link] = name
[Link] = age
p1 = Person("Emil")
p2 = Person("Tobias", 25)
print([Link], [Link])
print([Link], [Link])
Multiple Parameters
The __init__() method can have as many parameters as you need:
Example
Create a Person class with multiple parameters:
class Person:
def __init__(self, name, age, city, country):
[Link] = name
[Link] = age
[Link] = city
[Link] = country
print([Link])
print([Link])
print([Link])
print([Link])
class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age
def greet(self):
print("Hello, my name is " + [Link])
p1 = Person("Emil", 25)
[Link]()
Note: The self parameter must be the first parameter of any method in the
class.
Why Use self?
Without self, Python would not know which object's properties you want to
access:
Example
The self parameter links the method to the specific object:
class Person:
def __init__(self, name):
[Link] = name
def printname(self):
print([Link])
p1 = Person("Tobias")
p2 = Person("Linus")
[Link]()
[Link]()
REMOVE ADS
Example
Use the words myobject and abc instead of self:
class Person:
def __init__(myobject, name, age):
[Link] = name
[Link] = age
def greet(abc):
print("Hello, my name is " + [Link])
p1 = Person("Emil", 36)
[Link]()
Note: While you can use a different name, it is strongly recommended to
use self as it is the convention in Python and makes your code more
readable to others.
Example
Access multiple properties using self:
class Car:
def __init__(self, brand, model, year):
[Link] = brand
[Link] = model
[Link] = year
def display_info(self):
print(f"{[Link]} {[Link]} {[Link]}")
Example
Call one method from another method using self:
class Person:
def __init__(self, name):
[Link] = name
def greet(self):
return "Hello, " + [Link]
def welcome(self):
message = [Link]()
print(message + "! Welcome to our website.")
p1 = Person("Tobias")
[Link]()
class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age
p1 = Person("Emil", 36)
print([Link])
print([Link])
Access Properties
You can access object properties using dot notation:
Example
Access the properties of an object:
class Car:
def __init__(self, brand, model):
[Link] = brand
[Link] = model
print([Link])
print([Link])
REMOVE ADS
Modify Properties
You can modify the value of properties on objects:
Example
Change the age property:
class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age
p1 = Person("Tobias", 25)
print([Link])
[Link] = 26
print([Link])
Delete Properties
You can delete properties from objects using the del keyword:
Example
Delete the age property:
class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age
p1 = Person("Linus", 30)
del [Link]
Example
Class property vs instance property:
class Person:
species = "Human" # Class property
p1 = Person("Emil")
p2 = Person("Tobias")
print([Link])
print([Link])
print([Link])
print([Link])
Modifying Class Properties
When you modify a class property, it affects all objects:
Example
Change a class property:
class Person:
lastname = ""
p1 = Person("Linus")
p2 = Person("Emil")
[Link] = "Refsnes"
print([Link])
print([Link])
Example
Add a new property to an object:
class Person:
def __init__(self, name):
[Link] = name
p1 = Person("Tobias")
[Link] = 25
[Link] = "Oslo"
print([Link])
print([Link])
print([Link])
Note: Adding properties this way only adds them to that specific object, not
to all objects of the class.
Class Methods
Methods are functions that belong to a class. They define the behavior of
objects created from the class.
class Person:
def __init__(self, name):
[Link] = name
def greet(self):
print("Hello, my name is " + [Link])
p1 = Person("Emil")
[Link]()
Note: All methods must have self as the first parameter.
Example
Create a method with parameters:
class Calculator:
def add(self, a, b):
return a + b
calc = Calculator()
print([Link](5, 3))
print([Link](4, 7))
REMOVE ADS
Example
A method that accesses object properties:
class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age
def get_info(self):
return f"{[Link]} is {[Link]} years old"
p1 = Person("Tobias", 28)
print(p1.get_info())
Example
A method that changes a property value:
class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age
def celebrate_birthday(self):
[Link] += 1
print(f"Happy birthday! You are now {[Link]}")
p1 = Person("Linus", 25)
p1.celebrate_birthday()
p1.celebrate_birthday()
Example
Without the __str__() method:
class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age
p1 = Person("Emil", 36)
print(p1)
Example
With the __str__() method:
class Person:
def __init__(self, name, age):
[Link] = name
[Link] = age
def __str__(self):
return f"{[Link]} ({[Link]})"
p1 = Person("Tobias", 36)
print(p1)
Multiple Methods
A class can have multiple methods that work together:
Example
Create multiple methods in a class:
class Playlist:
def __init__(self, name):
[Link] = name
[Link] = []
def show_songs(self):
print(f"Playlist '{[Link]}':")
for song in [Link]:
print(f"- {song}")
my_playlist = Playlist("Favorites")
my_playlist.add_song("Bohemian Rhapsody")
my_playlist.add_song("Stairway to Heaven")
my_playlist.show_songs()
Delete Methods
You can delete methods from a class using the del keyword:
Example
Delete a method from a class:
class Person:
def __init__(self, name):
[Link] = name
def greet(self):
print("Hello!")
p1 = Person("Emil")
del [Link]
Python Inheritance
Python Inheritance
Inheritance allows us to define a class that inherits all the methods and
properties from another class.
Parent class is the class being inherited from, also called base class.
Child class is the class that inherits from another class, also called derived
class.
class Person:
def __init__(self, fname, lname):
[Link] = fname
[Link] = lname
def printname(self):
print([Link], [Link])
#Use the Person class to create an object, and then execute the
printname method:
x = Person("John", "Doe")
[Link]()
Example
Create a class named Student, which will inherit the properties and methods
from the Person class:
class Student(Person):
pass
Note: Use the pass keyword when you do not want to add any other
properties or methods to the class.
Now the Student class has the same properties and methods as the Person
class.
Example
Use the Student class to create an object, and then execute
the printname method:
x = Student("Mike", "Olsen")
[Link]()
REMOVE ADS
Note: The __init__() function is called automatically every time the class is
being used to create a new object.
Example
Add the __init__() function to the Student class:
class Student(Person):
def __init__(self, fname, lname):
#add properties etc.
When you add the __init__() function, the child class will no longer inherit
the parent's __init__() function.
To keep the inheritance of the parent's __init__() function, add a call to the
parent's __init__() function:
Example
class Student(Person):
def __init__(self, fname, lname):
Person.__init__(self, fname, lname)
Now we have successfully added the __init__() function, and kept the
inheritance of the parent class, and we are ready to add functionality in
the __init__() function.
REMOVE ADS
By using the super() function, you do not have to use the name of the
parent element, it will automatically inherit the methods and properties from
its parent.
Add Properties
Example
Add a property called graduationyear to the Student class:
class Student(Person):
def __init__(self, fname, lname):
super().__init__(fname, lname)
[Link] = 2019
In the example below, the year 2019 should be a variable, and passed into
the Student class when creating student objects. To do so, add another
parameter in the __init__() function:
Example
Add a year parameter, and pass the correct year when creating objects:
class Student(Person):
def __init__(self, fname, lname, year):
super().__init__(fname, lname)
[Link] = year
Add Methods
Example
Add a method called welcome to the Student class:
class Student(Person):
def __init__(self, fname, lname, year):
super().__init__(fname, lname)
[Link] = year
def welcome(self):
print("Welcome", [Link], [Link], "to the class
of", [Link])
If you add a method in the child class with the same name as a function in
the parent class, the inheritance of the parent method will be overridden.
Python Polymorphism
Function Polymorphism
An example of a Python function that can be used on different objects is
the len() function.
String
For strings len() returns the number of characters:
print(len(x))
Tuple
For tuples len() returns the number of items in the tuple:
Example
mytuple = ("apple", "banana", "cherry")
print(len(mytuple))
Dictionary
For dictionaries len() returns the number of key/value pairs in the
dictionary:
Example
thisdict = {
"brand": "Ford",
"model": "Mustang",
"year": 1964
}
print(len(thisdict))
REMOVE ADS
Class Polymorphism
Polymorphism is often used in Class methods, where we can have multiple
classes with the same method name.
For example, say we have three classes: Car, Boat, and Plane, and they all
have a method called move():
Example
Different classes with the same method:
class Car:
def __init__(self, brand, model):
[Link] = brand
[Link] = model
def move(self):
print("Drive!")
class Boat:
def __init__(self, brand, model):
[Link] = brand
[Link] = model
def move(self):
print("Sail!")
class Plane:
def __init__(self, brand, model):
[Link] = brand
[Link] = model
def move(self):
print("Fly!")
Look at the for loop at the end. Because of polymorphism we can execute
the same method for all three classes.
Yes. If we use the example above and make a parent class called Vehicle,
and make Car, Boat, Plane child classes of Vehicle, the child classes
inherits the Vehicle methods, but can override them:
Example
Create a class called Vehicle and make Car, Boat, Plane child classes
of Vehicle:
class Vehicle:
def __init__(self, brand, model):
[Link] = brand
[Link] = model
def move(self):
print("Move!")
class Car(Vehicle):
pass
class Boat(Vehicle):
def move(self):
print("Sail!")
class Plane(Vehicle):
def move(self):
print("Fly!")
Child classes inherits the properties and methods from the parent class.
In the example above you can see that the Car class is empty, but it
inherits brand, model, and move() from Vehicle.
Because of polymorphism we can execute the same method for all classes.
Python Encapsulation
Python Encapsulation
Encapsulation is about protecting data inside a class.
This prevents accidental changes to your data and hides the internal details
of how your class works.
Private Properties
In Python, you can make properties private by using a double
underscore __ prefix:
class Person:
def __init__(self, name, age):
[Link] = name
self.__age = age # Private property
p1 = Person("Emil", 25)
print([Link])
print(p1.__age) # This will cause an error
Note: Private properties cannot be accessed directly from outside the class.
Example
Use a getter method to access a private property:
class Person:
def __init__(self, name, age):
[Link] = name
self.__age = age
def get_age(self):
return self.__age
p1 = Person("Tobias", 25)
print(p1.get_age())
The setter method can also validate the value before setting it:
Example
Use a setter method to change a private property:
class Person:
def __init__(self, name, age):
[Link] = name
self.__age = age
def get_age(self):
return self.__age
p1 = Person("Tobias", 25)
print(p1.get_age())
p1.set_age(26)
print(p1.get_age())
REMOVE ADS
Why Use Encapsulation?
Encapsulation provides several benefits:
class Student:
def __init__(self, name):
[Link] = name
self.__grade = 0
def get_grade(self):
return self.__grade
def get_status(self):
if self.__grade >= 60:
return "Passed"
else:
return "Failed"
student = Student("Emil")
student.set_grade(85)
print(student.get_grade())
print(student.get_status())
Protected Properties
Python also has a convention for protected properties using a single
underscore _ prefix:
Example
Create a protected property:
class Person:
def __init__(self, name, salary):
[Link] = name
self._salary = salary # Protected property
p1 = Person("Linus", 50000)
print([Link])
print(p1._salary) # Can access, but shouldn't
Note: A single underscore _ is just a convention. It tells other programmers
that the property is intended for internal use, but Python doesn't enforce this
restriction.
Private Methods
You can also make methods private using the double underscore prefix:
Example
Create a private method:
class Calculator:
def __init__(self):
[Link] = 0
calc = Calculator()
[Link](10)
[Link](5)
print([Link])
# calc.__validate(5) # This would cause an error
Note: Just like private properties with double underscores, private methods
cannot be called directly from outside the class. The __validate method can
only be used by other methods inside the class.
Name Mangling
Name mangling is how Python implements private properties and methods.
Example
See how Python mangles the name:
class Person:
def __init__(self, name, age):
[Link] = name
self.__age = age
p1 = Person("Emil", 30)
Inner classes are useful for grouping classes that are only used in one place,
making your code more organized.
ExampleGet your own Python Server
Create an inner class:
class Outer:
def __init__(self):
[Link] = "Outer Class"
class Inner:
def __init__(self):
[Link] = "Inner Class"
def display(self):
print("This is the inner class")
outer = Outer()
print([Link])
Example
Access the inner class and create an object:
class Outer:
def __init__(self):
[Link] = "Outer"
class Inner:
def __init__(self):
[Link] = "Inner"
def display(self):
print("Hello from inner class")
outer = Outer()
inner = [Link]()
[Link]()
REMOVE ADS
If you want the inner class to access the outer class, you need to pass the
outer class instance as a parameter:
Example
Pass the outer class instance to the inner class:
class Outer:
def __init__(self):
[Link] = "Emil"
class Inner:
def __init__(self, outer):
[Link] = outer
def display(self):
print(f"Outer class name: {[Link]}")
outer = Outer()
inner = [Link](outer)
[Link]()
Practical Example
Inner classes are useful for creating helper classes that are only used within
the context of the outer class:
Example
Use an inner class to represent a car's engine:
class Car:
def __init__(self, brand, model):
[Link] = brand
[Link] = model
[Link] = [Link]()
class Engine:
def __init__(self):
[Link] = "Off"
def start(self):
[Link] = "Running"
print("Engine started")
def stop(self):
[Link] = "Off"
print("Engine stopped")
def drive(self):
if [Link] == "Running":
print(f"Driving the {[Link]} {[Link]}")
else:
print("Start the engine first!")
Example
Create multiple inner classes:
class Computer:
def __init__(self):
[Link] = [Link]()
[Link] = [Link]()
class CPU:
def process(self):
print("Processing data...")
class RAM:
def store(self):
print("Storing data...")
computer = Computer()
[Link]()
[Link]()
File Handling
The key function for working with files in Python is the open() function.
"r" - Read - Default value. Opens a file for reading, error if the file does not
exist
"a" - Append - Opens a file for appending, creates the file if it does not exist
"w" - Write - Opens a file for writing, creates the file if it does not exist
"x" - Create - Creates the specified file, returns an error if the file exists
In addition you can specify if the file should be handled as binary or text
mode
"t" - Text - Default value. Text mode
Syntax
To open a file for reading it is enough to specify the name of the file:
f = open("[Link]")
f = open("[Link]", "rt")
Because "r" for read, and "t" for text are the default values, you do not need
to specify them.
Note: Make sure the file exists, or else you will get an error.
[Link]
The open() function returns a file object, which has a read() method for
reading the content of the file:
ExampleGet your own Python Server
f = open("[Link]")
print([Link]())
If the file is located in a different location, you will have to specify the file
path, like this:
Example
Open a file on a different location:
f = open("D:\\myfiles\[Link]")
print([Link]())
Example
Using the with keyword:
with open("[Link]") as f:
print([Link]())
Then you do not have to worry about closing your files, the with statement
takes care of that.
Close Files
It is a good practice to always close the file when you are done with it.
If you are not using the with statement, you must write a close statement in
order to close the file:
Example
Close the file when you are finished with it:
f = open("[Link]")
print([Link]())
[Link]()
Note: You should always close your files. In some cases, due to buffering,
changes made to a file may not show until you close the file.
Example
Return the 5 first characters of the file:
with open("[Link]") as f:
print([Link](5))
REMOVE ADS
Read Lines
You can return one line by using the readline() method:
Example
Read one line of the file:
with open("[Link]") as f:
print([Link]())
By calling readline() two times, you can read the two first lines:
Example
Read two lines of the file:
with open("[Link]") as f:
print([Link]())
print([Link]())
By looping through the lines of the file, you can read the whole file, line by
line:
Example
Loop through the file line by line:
with open("[Link]") as f:
for x in f:
print(x)
Example
Open the file "[Link]" and overwrite the content:
"x" - Create - will create a file, returns an error if the file exists
"a" - Append - will create a file if the specified file does not exists
"w" - Write - will create a file if the specified file does not exists
Example
Create a new file called "[Link]":
f = open("[Link]", "x")
import os
[Link]("[Link]")
Example
Check if file exists, then delete it:
import os
if [Link]("[Link]"):
[Link]("[Link]")
else:
print("The file does not exist")
Delete Folder
To delete an entire folder, use the [Link]() method:
Example
Remove the folder "myfolder":
import os
[Link]("myfolder")
Note: You can only remove empty folders.
Matplotlib Tutorial
What is Matplotlib?
Matplotlib is a low level graph plotting library in python that serves as a
visualization utility.
If this command fails, then use a python distribution that already has
Matplotlib installed, like Anaconda, Spyder etc.
Import Matplotlib
Once Matplotlib is installed, import it in your applications by adding
the import module statement:
import matplotlib
print(matplotlib.__version__)
Note: two underscore characters are used in __version__.
Matplotlib Pyplot
Pyplot
Most of the Matplotlib utilities lies under the pyplot submodule, and are
usually imported under the plt alias:
[Link](xpoints, ypoints)
[Link]()
Result:
Matplotlib Plotting
Plotting x and y points
The plot() function is used to draw points (markers) in a diagram.
If we need to plot a line from (1, 3) to (8, 10), we have to pass two arrays [1,
8] and [3, 10] to the plot function.
[Link](xpoints, ypoints)
[Link]()
Result:
The x-axis is the horizontal axis.
Example
Draw two points in the diagram, one at position (1, 3) and one in position (8,
10):
Result:
Multiple Points
You can plot as many points as you like, just make sure you have the same
number of points in both axis.
Example
Draw a line in a diagram from position (1, 3) to (2, 8) then to (6, 1) and
finally to position (8, 10):
[Link](xpoints, ypoints)
[Link]()
Result:
Default X-Points
If we do not specify the points on the x-axis, they will get the default values
0, 1, 2, 3 etc., depending on the length of the y-points.
So, if we take the same example as above, and leave out the x-points, the
diagram will look like this:
Example
Plotting without x-points:
[Link](ypoints)
[Link]()
Result:
The x-points in the example above are [0, 1, 2, 3, 4, 5].
Matplotlib Markers
Markers
You can use the keyword argument marker to emphasize each point with a
specified marker:
Result:
Example
Mark each point with a star:
...
[Link](ypoints, marker = '*')
...
Result:
Marker Reference
You can choose any of these markers:
Marker Description
'o' Circle
'*' Star
'.' Point
',' Pixel
'x' X
'X' X (filled)
'+' Plus
's' Square
'D' Diamond
'p' Pentagon
'H' Hexagon
'h' Hexagon
'^' Triangle Up
'2' Tri Up
'|' Vline
'_' Hline
Format Strings fmt
You can also use the shortcut string notation parameter to specify the
marker.
This parameter is also called fmt, and is written with this syntax:
marker|line|color
Example
Mark each point with a circle:
[Link](ypoints, 'o:r')
[Link]()
Result:
The marker value can be anything from the Marker Reference above.
Line Reference
Line Syntax Description
Note: If you leave out the line value in the fmt parameter, no line will be
plotted.
Color Reference
Color Syntax Description
'r' Red
'g' Green
'b' Blue
'c' Cyan
'm' Magenta
'y' Yellow
'k' Black
'w' White
Marker Size
You can use the keyword argument markersize or the shorter version, ms to
set the size of the markers:
Example
Set the size of the markers to 20:
Result:
Marker Color
You can use the keyword argument markeredgecolor or the shorter mec to
set the color of the edge of the markers:
Example
Set the EDGE color to red:
You can use the keyword argument markerfacecolor or the shorter mfc to
set the color inside the edge of the markers:
Example
Set the FACE color to red:
Result:
Use both the mec and mfc arguments to color the entire marker:
Example
Set the color of both the edge and the face to red:
Result:
You can also use Hexadecimal color values:
Example
Mark each point with a beautiful green color:
...
[Link](ypoints, marker = 'o', ms = 20, mec = '#4CAF50', mfc
= '#4CAF50')
...
Result:
Or any of the 140 supported color names.
Example
Mark each point with the color named "hotpink":
...
[Link](ypoints, marker = 'o', ms = 20, mec = 'hotpink', mfc
= 'hotpink')
...
Result:
Matplotlib Line
Linestyle
You can use the keyword argument linestyle, or shorter ls, to change the
style of the plotted line:
Result:
Example
Use a dashed line:
Result:
REMOVE ADS
Shorter Syntax
The line style can be written in a shorter syntax:
Example
Shorter syntax:
[Link](ypoints, ls = ':')
Result:
Line Styles
You can choose any of these styles:
Style Or
'solid' (default) '-'
'dotted' ':'
'dashed' '--'
'dashdot' '-.'
Line Color
You can use the keyword argument color or the shorter c to set the color of
the line:
Example
Set the line color to red:
Result:
You can also use Hexadecimal color values:
Example
Plot with a beautiful green line:
...
[Link](ypoints, c = '#4CAF50')
...
Result:
Or any of the 140 supported color names.
Example
Plot with the color named "hotpink":
...
[Link](ypoints, c = 'hotpink')
...
Result:
REMOVE ADS
Line Width
You can use the keyword argument linewidth or the shorter lw to change
the width of the line.
Example
Plot with a 20.5pt wide line:
Result:
Multiple Lines
You can plot as many lines as you like by simply adding
more [Link]() functions:
Example
Draw two lines by specifying a [Link]() function for each line:
import [Link] as plt
import numpy as np
y1 = [Link]([3, 8, 1, 10])
y2 = [Link]([6, 2, 7, 11])
[Link](y1)
[Link](y2)
[Link]()
Result:
You can also plot many lines by adding the points for the x- and y-axis for
each line in the same [Link]() function.
(In the examples above we only specified the points on the y-axis, meaning
that the points on the x-axis got the the default values (0, 1, 2, 3).)
x1 = [Link]([0, 1, 2, 3])
y1 = [Link]([3, 8, 1, 10])
x2 = [Link]([0, 1, 2, 3])
y2 = [Link]([6, 2, 7, 11])
Result:
import numpy as np
import [Link] as plt
x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])
[Link](x, y)
[Link]("Average Pulse")
[Link]("Calorie Burnage")
[Link]()
Result:
Create a Title for a Plot
With Pyplot, you can use the title() function to set a title for the plot.
Example
Add a plot title and labels for the x- and y-axis:
import numpy as np
import [Link] as plt
x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])
[Link](x, y)
[Link]("Sports Watch Data")
[Link]("Average Pulse")
[Link]("Calorie Burnage")
[Link]()
Result:
REMOVE ADS
Example
Set font properties for the title and labels:
import numpy as np
import [Link] as plt
x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])
font1 = {'family':'serif','color':'blue','size':20}
font2 = {'family':'serif','color':'darkred','size':15}
[Link](x, y)
[Link]()
Result:
Position the Title
You can use the loc parameter in title() to position the title.
Legal values are: 'left', 'right', and 'center'. Default value is 'center'.
Example
Position the title to the left:
import numpy as np
import [Link] as plt
x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])
[Link]("Sports Watch Data", loc = 'left')
[Link]("Average Pulse")
[Link]("Calorie Burnage")
[Link](x, y)
[Link]()
Result:
import numpy as np
import [Link] as plt
x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])
[Link](x, y)
[Link]()
[Link]()
Result:
REMOVE ADS
Legal values are: 'x', 'y', and 'both'. Default value is 'both'.
Example
Display only grid lines for the x-axis:
import numpy as np
import [Link] as plt
x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])
[Link](x, y)
[Link](axis = 'x')
[Link]()
Result:
Example
Display only grid lines for the y-axis:
import numpy as np
import [Link] as plt
x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])
[Link](x, y)
[Link](axis = 'y')
[Link]()
Result:
Set Line Properties for the Grid
You can also set the line properties of the grid, like this: grid(color = 'color',
linestyle = 'linestyle', linewidth = number).
Example
Set the line properties of the grid:
import numpy as np
import [Link] as plt
x = [Link]([80, 85, 90, 95, 100, 105, 110, 115, 120, 125])
y = [Link]([240, 250, 260, 270, 280, 290, 300, 310, 320, 330])
[Link](x, y)
[Link]()
Result:
Matplotlib Subplot
#plot 1:
x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])
[Link](1, 2, 1)
[Link](x,y)
#plot 2:
x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])
[Link](1, 2, 2)
[Link](x,y)
[Link]()
Result:
[Link](1, 2, 1)
#the figure has 1 row, 2 columns, and this plot is
the first plot.
[Link](1, 2, 2)
#the figure has 1 row, 2 columns, and this plot is
the second plot.
So, if we want a figure with 2 rows an 1 column (meaning that the two plots
will be displayed on top of each other instead of side-by-side), we can write
the syntax like this:
Example
Draw 2 plots on top of each other:
#plot 1:
x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])
[Link](2, 1, 1)
[Link](x,y)
#plot 2:
x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])
[Link](2, 1, 2)
[Link](x,y)
[Link]()
Result:
You can draw as many plots you like on one figure, just descibe the number
of rows, columns, and the index of the plot.
Example
Draw 6 plots:
x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])
[Link](2, 3, 1)
[Link](x,y)
x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])
[Link](2, 3, 2)
[Link](x,y)
x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])
[Link](2, 3, 3)
[Link](x,y)
x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])
[Link](2, 3, 4)
[Link](x,y)
x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])
[Link](2, 3, 5)
[Link](x,y)
x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])
[Link](2, 3, 6)
[Link](x,y)
[Link]()
Result:
REMOVE ADS
Title
You can add a title to each plot with the title() function:
Example
2 plots, with titles:
#plot 1:
x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])
[Link](1, 2, 1)
[Link](x,y)
[Link]("SALES")
#plot 2:
x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])
[Link](1, 2, 2)
[Link](x,y)
[Link]("INCOME")
[Link]()
Result:
Super Title
You can add a title to the entire figure with the suptitle() function:
Example
Add a title for the entire figure:
#plot 1:
x = [Link]([0, 1, 2, 3])
y = [Link]([3, 8, 1, 10])
[Link](1, 2, 1)
[Link](x,y)
[Link]("SALES")
#plot 2:
x = [Link]([0, 1, 2, 3])
y = [Link]([10, 20, 30, 40])
[Link](1, 2, 2)
[Link](x,y)
[Link]("INCOME")
[Link]("MY SHOP")
[Link]()
Result:
Matplotlib Scatter
The scatter() function plots one dot for each observation. It needs two arrays
of the same length, one for the values of the x-axis, and one for values on
the y-axis:
x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
[Link](x, y)
[Link]()
Result:
The observation in the example above is the result of 13 cars passing by.
Compare Plots
In the example above, there seems to be a relationship between speed and
age, but what if we plot the observations from another day as well? Will the
scatter plot tell us something else?
Example
Draw two plots on the same figure:
[Link]()
Result:
Note: The two plots are plotted with two different colors, by default blue and
orange, you will learn how to change colors later in this chapter.
By comparing the two plots, I think it is safe to say that they both gives us
the same conclusion: the newer the car, the faster it drives.
REMOVE ADS
Colors
You can set your own color for each scatter plot with the color or
the c argument:
Example
Set your own color of the markers:
x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
[Link](x, y, color = 'hotpink')
x = [Link]([2,2,8,1,15,8,12,9,7,3,11,4,7,14,12])
y = [Link]([100,105,84,105,90,99,90,95,94,100,79,112,91,80,85])
[Link](x, y, color = '#88c999')
[Link]()
Result:
Color Each Dot
You can even set a specific color for each dot by using an array of colors as
value for the c argument:
Note: You cannot use the color argument for this, only the c argument.
Example
Set your own color of the markers:
x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
colors =
[Link](["red","green","blue","yellow","pink","black","orange","purp
le","beige","brown","gray","cyan","magenta"])
[Link](x, y, c=colors)
[Link]()
Result:
REMOVE ADS
ColorMap
The Matplotlib module has a number of available colormaps.
A colormap is like a list of colors, where each color has a value that ranges
from 0 to 100.
In addition you have to create an array with values (from 0 to 100), one
value for each point in the scatter plot:
Example
Create a color array, and specify a colormap in the scatter plot:
x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
colors =
[Link]([0, 10, 20, 30, 40, 45, 50, 55, 60, 70, 80, 90, 100])
[Link](x, y, c=colors, cmap='viridis')
[Link]()
Result:
Example
Include the actual colormap:
x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
colors =
[Link]([0, 10, 20, 30, 40, 45, 50, 55, 60, 70, 80, 90, 100])
[Link]()
[Link]()
Result:
Available ColorMaps
You can choose any of the built-in colormaps:
Name Reverse
Accent Accent_r
Blues Blues_r
BrBG BrBG_r
BuGn BuGn_r
BuPu BuPu_r
CMRmap CMRmap_r
Dark2 Dark2_r
GnBu GnBu_r
Greens Greens_r
Greys Greys_r
OrRd OrRd_r
Oranges Oranges_r
PRGn PRGn_r
Paired Paired_r
Pastel1 Pastel1_r
Pastel2 Pastel2_r
PiYG PiYG_r
PuBu PuBu_r
PuBuGn PuBuGn_r
PuOr PuOr_r
PuRd PuRd_r
Purples Purples_r
RdBu RdBu_r
RdGy RdGy_r
RdPu RdPu_r
RdYlBu RdYlBu_r
RdYlGn RdYlGn_r
Reds Reds_r
Set1 Set1_r
Set2 Set2_r
Set3 Set3_r
Spectral Spectral_r
Wistia Wistia_r
YlGn YlGn_r
YlGnBu YlGnBu_r
YlOrBr YlOrBr_r
YlOrRd YlOrRd_r
afmhot afmhot_r
autumn autumn_r
binary binary_r
bone bone_r
brg brg_r
bwr bwr_r
cividis cividis_r
cool cool_r
coolwarm coolwarm_r
copper copper_r
cubehelix cubehelix_r
flag flag_r
gist_earth gist_earth_r
gist_gray gist_gray_r
gist_heat gist_heat_r
gist_ncar gist_ncar_r
gist_rainbow gist_rainbow_r
gist_stern gist_stern_r
gist_yarg gist_yarg_r
gnuplot gnuplot_r
gnuplot2 gnuplot2_r
gray gray_r
hot hot_r
hsv hsv_r
inferno inferno_r
jet jet_r
magma magma_r
nipy_spectral nipy_spectral_r
ocean ocean_r
pink pink_r
plasma plasma_r
prism prism_r
rainbow rainbow_r
seismic seismic_r
spring spring_r
summer summer_r
tab10 tab10_r
tab20 tab20_r
tab20b tab20b_r
tab20c tab20c_r
terrain terrain_r
twilight twilight_r
twilight_shifted twilight_shifted_r
viridis viridis_r
winter winter_r
Size
You can change the size of the dots with the s argument.
Just like colors, make sure the array for sizes has the same length as the
arrays for the x- and y-axis:
Example
Set your own size for the markers:
x = [Link]([5,7,8,7,2,17,2,9,4,11,12,9,6])
y = [Link]([99,86,87,88,111,86,103,87,94,78,77,85,86])
sizes = [Link]([20,50,100,200,500,1000,60,90,10,300,600,800,75])
[Link](x, y, s=sizes)
[Link]()
Result:
Alpha
You can adjust the transparency of the dots with the alpha argument.
Just like colors, make sure the array for sizes has the same length as the
arrays for the x- and y-axis:
Example
Set your own size for the markers:
[Link]()
Result:
x = [Link](100, size=(100))
y = [Link](100, size=(100))
colors = [Link](100, size=(100))
sizes = 10 * [Link](100, size=(100))
[Link]()
[Link]()
Result:
Matplotlib Bars
Creating Bars
With Pyplot, you can use the bar() function to draw bar graphs:
[Link](x,y)
[Link]()
Result:
The bar() function takes arguments that describes the layout of the bars.
Example
x = ["APPLES", "BANANAS"]
y = [400, 350]
[Link](x, y)
REMOVE ADS
Horizontal Bars
If you want the bars to be displayed horizontally instead of vertically, use
the barh() function:
Example
Draw 4 horizontal bars:
[Link](x, y)
[Link]()
Result:
Bar Color
The bar() and barh() take the keyword argument color to set the color of
the bars:
Example
Draw 4 red bars:
Result:
Color Names
You can use any of the 140 supported color names.
Example
Draw 4 "hot pink" bars:
Result:
Color Hex
Or you can use Hexadecimal color values:
Example
Draw 4 bars with a beautiful green color:
Result:
Bar Width
The bar() takes the keyword argument width to set the width of the bars:
Example
Draw 4 very thin bars:
Bar Height
The barh() takes the keyword argument height to set the height of the
bars:
Example
Draw 4 very thin bars:
import [Link] as plt
import numpy as np
Result:
Matplotlib Histograms
Histogram
A histogram is a graph showing frequency distributions.
Example: Say you ask for the height of 250 people, you might end up with a
histogram like this:
You can read from the histogram that there are approximately:
Create Histogram
In Matplotlib, we use the hist() function to create histograms.
The hist() function will use an array of numbers to create a histogram, the
array is sent into the function as an argument.
For simplicity we use NumPy to randomly generate an array with 250 values,
where the values will concentrate around 170, and the standard deviation is
10. Learn more about Normal Data Distribution in our Machine Learning
Tutorial.
import numpy as np
print(x)
Result:
This will generate a random result, and could look like this:
The hist() function will read the array and produce a histogram:
Example
A simple histogram:
[Link](x)
[Link]()
Result:
[Link](y)
[Link]()
Result:
As you can see the pie chart draws one piece (called a wedge) for each value
in the array (in this case [35, 25, 25, 15]).
By default the plotting of the first wedge starts from the x-axis and
moves counterclockwise:
Note: The size of each wedge is determined by comparing the value with all
the other values, by using this formula:
REMOVE ADS
Labels
Add labels to the pie chart with the labels parameter.
The labels parameter must be an array with one label for each wedge:
Example
A simple pie chart:
Result:
Start Angle
As mentioned the default start angle is at the x-axis, but you can change the
start angle by specifying a startangle parameter.
Example
Start the first wedge at 90 degrees:
Result:
Explode
Maybe you want one of the wedges to stand out? The explode parameter
allows you to do that.
The explode parameter, if specified, and not None, must be an array with
one value for each wedge.
Each value represents how far from the center each wedge is displayed:
Example
Pull the "Apples" wedge 0.2 from the center of the pie:
Result:
Shadow
Add a shadow to the pie chart by setting the shadows parameter to True:
Example
Add a shadow:
Result:
Colors
You can set the color of each wedge with the colors parameter.
The colors parameter, if specified, must be an array with one value for each
wedge:
Example
Specify a new color for each wedge:
Result:
You can use Hexadecimal color values, any of the 140 supported color
names, or one of these shortcuts:
'r' - Red
'g' - Green
'b' - Blue
'c' - Cyan
'm' - Magenta
'y' - Yellow
'k' - Black
'w' - White
Legend
To add a list of explanation for each wedge, use the legend() function:
Example
Add a legend:
Result:
Legend With Header
To add a header to the legend, add the title parameter to
the legend function.
Example
Add a legend with a header:
Machine Learning
Where To Start?
In this tutorial we will go back to mathematics and study statistics, and how
to calculate important numbers based on data sets.
We will also learn how to use various Python modules to get the answers we
need.
And we will learn how to make functions that are able to predict the outcome
based on what we have learned.
Data Set
In the mind of a computer, a data set is any collection of data. It can be
anything from an array to a complete database.
Example of an array:
[99,86,87,88,111,86,103,87,94,78,77,85,86]
Example of a database:
BMW red 5
Volvo black 7
VW gray 8
VW white 7
Ford white 2
VW white 17
Tesla red 2
BMW black 9
Volvo gray 4
Ford white 11
Toyota gray 12
VW white 9
Toyota blue 6
By looking at the array, we can guess that the average value is probably
around 80 or 90, and we are also able to determine the highest value and
the lowest value, but what else can we do?
And by looking at the database we can see that the most popular color is
white, and the oldest car is 17 years, but what if we could predict if a car had
an AutoPass, just by looking at the other values?
That is what Machine Learning is for! Analyzing data and predicting the
outcome!
In Machine Learning it is common to work with very large data sets. In this
tutorial we will try to make it as easy as possible to understand the different
concepts of machine learning, and we will work with small easy-to-
understand data sets.
REMOVE ADS
Data Types
To analyze data, it is important to know what type of data we are dealing
with.
Numerical
Categorical
Ordinal
Numerical data are numbers, and can be split into two numerical
categories:
Discrete Data
- counted data that are limited to integers. Example: The number of
cars passing by.
Continuous Data
- measured data that can be any number. Example: The price of an
item, or the size of an item
Ordinal data are like categorical data, but can be measured up against each
other. Example: school grades where A is better than B and so on.
By knowing the data type of your data source, you will be able to know what
technique to use when analyzing them.
You will learn more about statistics and analyzing data in the next chapters.
In Machine Learning (and in mathematics) there are often three values that
interests us:
speed = [99,86,87,88,111,86,103,87,94,78,77,85,86]
What is the average, the middle, or the most common speed value?
Mean
The mean value is the average value.
To calculate the mean, find the sum of all values, and divide the sum by the
number of values:
(99+86+87+88+111+86+103+87+94+78+77+85+86) / 13 = 89.77
The NumPy module has a method for this. Learn about the NumPy module in
our NumPy Tutorial.
import numpy
speed = [99,86,87,88,111,86,103,87,94,78,77,85,86]
x = [Link](speed)
print(x)
REMOVE ADS
Median
The median value is the value in the middle, after you have sorted all the
values:
77, 78, 85, 86, 86, 86, 87, 87, 88, 94, 99, 103, 111
It is important that the numbers are sorted before you can find the median.
Example
Use the NumPy median() method to find the middle value:
import numpy
speed = [99,86,87,88,111,86,103,87,94,78,77,85,86]
x = [Link](speed)
print(x)
If there are two numbers in the middle, divide the sum of those numbers by
two.
77, 78, 85, 86, 86, 86, 87, 87, 94, 98, 99, 103
import numpy
speed = [99,86,87,88,86,103,87,94,78,77,85,86]
x = [Link](speed)
print(x)
Mode
The Mode value is the value that appears the most number of times:
99, 86, 87, 88, 111, 86, 103, 87, 94, 78, 77, 85, 86 = 86
The SciPy module has a method for this. Learn about the SciPy module in
our SciPy Tutorial.
Example
Use the SciPy mode() method to find the number that appears the most:
speed = [99,86,87,88,111,86,103,87,94,78,77,85,86]
x = [Link](speed)
print(x)
REMOVE ADS
Chapter Summary
The Mean, Median, and Mode are techniques that are often used in Machine
Learning, so it is important to understand the concept behind them.
A high standard deviation means that the values are spread out over a wider
range.
speed = [86,87,88,86,87,85,86]
0.9
Meaning that most of the values are within the range of 0.9 from the mean
value, which is 86.4.
speed = [32,111,138,28,59,77,97]
37.85
Meaning that most of the values are within the range of 37.85 from the mean
value, which is 77.4.
As you can see, a higher standard deviation indicates that the values are
spread out over a wider range.
import numpy
speed = [86,87,88,86,87,85,86]
x = [Link](speed)
print(x)
Example
import numpy
speed = [32,111,138,28,59,77,97]
x = [Link](speed)
print(x)
Variance
Variance is another number that indicates how spread out the values are.
In fact, if you take the square root of the variance, you get the standard
deviation!
Or the other way around, if you multiply the standard deviation by itself, you
get the variance!
(32+111+138+28+59+77+97) / 7 = 77.4
32 - 77.4 = -45.4
111 - 77.4 = 33.6
138 - 77.4 = 60.6
28 - 77.4 = -49.4
59 - 77.4 = -18.4
77 - 77.4 = - 0.4
97 - 77.4 = 19.6
(-45.4)2 = 2061.16
(33.6)2 = 1128.96
(60.6)2 = 3672.36
(-49.4)2 = 2440.36
(-18.4)2 = 338.56
(- 0.4)2 = 0.16
(19.6)2 = 384.16
Example
Use the NumPy var() method to find the variance:
import numpy
speed = [32,111,138,28,59,77,97]
x = [Link](speed)
print(x)
Standard Deviation
As we have learned, the formula to find the standard deviation is the square
root of the variance:
√1432.25 = 37.85
Or, as in the example from before, use the NumPy to calculate the standard
deviation:
Example
Use the NumPy std() method to find the standard deviation:
import numpy
speed = [32,111,138,28,59,77,97]
x = [Link](speed)
print(x)
Symbols
Standard Deviation is often represented by the symbol Sigma: σ
Variance is often represented by the symbol Sigma Squared: σ 2
REMOVE ADS
Chapter Summary
The Standard Deviation and Variance are terms that are often used in
Machine Learning, so it is important to understand how to get them, and the
concept behind them.
Machine Learning -
Percentiles
Example: Let's say we have an array that contains the ages of every person
living on a street.
ages =
[5,31,43,48,50,41,7,11,15,39,80,82,32,2,8,6,25,36,27,61,31]
What is the 75. percentile? The answer is 43, meaning that 75% of the
people are 43 or younger.
The NumPy module has a method for finding the specified percentile:
ages = [5,31,43,48,50,41,7,11,15,39,80,82,32,2,8,6,25,36,27,61,31]
x = [Link](ages, 75)
print(x)
Example
What is the age that 90% of the people are younger than?
import numpy
ages = [5,31,43,48,50,41,7,11,15,39,80,82,32,2,8,6,25,36,27,61,31]
x = [Link](ages, 90)
print(x)
Data Distribution
Earlier in this tutorial we have worked with very small amounts of data in our
examples, just to understand the different concepts.
In the real world, the data sets are much bigger, but it can be difficult to
gather real world data, at least at an early stage of a project.
print(x)
Histogram
To visualize the data set we can draw a histogram with the data we
collected.
Example
Draw a histogram:
import numpy
import [Link] as plt
[Link](x, 5)
[Link]()
Result:
Histogram Explained
We use the array from the example above to draw a histogram with 5 bars.
The first bar represents how many values in the array are between 0 and 1.
The second bar represents how many values are between 1 and 2.
Etc.
Example
Create an array with 100000 random numbers, and display them using a
histogram with 100 bars:
import numpy
import [Link] as plt
[Link](x, 100)
[Link]()
In this chapter we will learn how to create an array where the values are
concentrated around a given value.
[Link](x, 100)
[Link]()
Result:
Note: A normal distribution graph is also known as the bell curve because of
it's characteristic shape of a bell.
Histogram Explained
We use the array from the [Link]() method, with 100000
values, to draw a histogram with 100 bars.
We specify that the mean value is 5.0, and the standard deviation is 1.0.
Meaning that the values should be concentrated around 5.0, and rarely
further away than 1.0 from the mean.
And as you can see from the histogram, most values are between 4.0 and
6.0, with a top at approximately 5.0.
Scatter Plot
A scatter plot is a diagram where each value in the data set is represented
by a dot.
The Matplotlib module has a method for drawing scatter plots, it needs two
arrays of the same length, one for the values of the x-axis, and one for the
values of the y-axis:
x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]
x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]
[Link](x, y)
[Link]()
Result:
Scatter Plot Explained
The x-axis represents ages, and the y-axis represents speeds.
What we can read from the diagram is that the two fastest cars were both 2
years old, and the slowest car was 12 years old.
Note: It seems that the newer the car, the faster it drives, but that could be
a coincidence, after all we only registered 13 cars.
REMOVE ADS
You might not have real world data when you are testing an algorithm, you
might have to use randomly generated values.
As we have learned in the previous chapter, the NumPy module can help us
with that!
Let us create two arrays that are both filled with 1000 random numbers from
a normal data distribution.
The first array will have the mean set to 5.0 with a standard deviation of 1.0.
The second array will have the mean set to 10.0 with a standard deviation of
2.0:
Example
A scatter plot with 1000 dots:
import numpy
import [Link] as plt
[Link](x, y)
[Link]()
Result:
Scatter Plot Explained
We can see that the dots are concentrated around the value 5 on the x-axis,
and 10 on the y-axis.
We can also see that the spread is wider on the y-axis than on the x-axis.
Regression
The term regression is used when you try to find the relationship between
variables.
Linear Regression
Linear regression uses the relationship between the data-points to draw a
straight line through all them.
In the example below, the x-axis represents age, and the y-axis represents
speed. We have registered the age and speed of 13 cars as they were
passing a tollbooth. Let us see if the data we collected could be used in a
linear regression:
x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]
[Link](x, y)
[Link]()
Result:
Example
Import scipy and draw the line of Linear Regression:
x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]
def myfunc(x):
return slope * x + intercept
[Link](x, y)
[Link](x, mymodel)
[Link]()
Result:
Example Explained
Import the modules you need.
You can learn about the Matplotlib module in our Matplotlib Tutorial.
You can learn about the SciPy module in our SciPy Tutorial.
Create the arrays that represent the values of the x and y axis:
x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]
Execute a method that returns some important key values of Linear
Regression:
Create a function that uses the slope and intercept values to return a new
value. This new value represents where on the y-axis the corresponding x
value will be placed:
def myfunc(x):
return slope * x + intercept
Run each value of the x array through the function. This will result in a new
array with new values for the y-axis:
[Link](x, y)
[Link](x, mymodel)
[Link]()
REMOVE ADS
R for Relationship
It is important to know how the relationship between the values of the x-axis
and the values of the y-axis is, if there are no relationship the linear
regression can not be used to predict anything.
Python and the Scipy module will compute this value for you, all you have to
do is feed it with the x and y values.
Example
How well does my data fit in a linear regression?
x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]
print(r)
Note: The result -0.76 shows that there is a relationship, not perfect, but it
indicates that we could use linear regression in future predictions.
REMOVE ADS
To do so, we need the same myfunc() function from the example above:
def myfunc(x):
return slope * x + intercept
Example
Predict the speed of a 10 years old car:
x = [5,7,8,7,2,17,2,9,4,11,12,9,6]
y = [99,86,87,88,111,86,103,87,94,78,77,85,86]
slope, intercept, r, p, std_err = [Link](x, y)
def myfunc(x):
return slope * x + intercept
speed = myfunc(10)
print(speed)
The example predicted a speed at 85.6, which we also could read from the
diagram:
Bad Fit?
Let us create an example where linear regression would not be the best
method to predict future values.
Example
These values for the x- and y-axis should result in a very bad fit for linear
regression:
x = [89,43,36,36,95,10,66,34,38,20,26,29,48,64,6,5,36,66,72,40]
y = [21,46,3,35,67,95,53,72,58,10,26,34,90,33,38,20,56,2,47,15]
def myfunc(x):
return slope * x + intercept
[Link](x, y)
[Link](x, mymodel)
[Link]()
Result:
And the r for relationship?
Example
You should get a very low r value.
import numpy
from scipy import stats
x = [89,43,36,36,95,10,66,34,38,20,26,29,48,64,6,5,36,66,72,40]
y = [21,46,3,35,67,95,53,72,58,10,26,34,90,33,38,20,56,2,47,15]
print(r)
The result: 0.013 indicates a very bad relationship, and tells us that this data
set is not suitable for linear regression.
Machine Learning -
Polynomial Regression
Polynomial Regression
If your data points clearly will not fit a linear regression (a straight line
through all data points), it might be ideal for polynomial regression.
We have registered the car's speed, and the time of day (hour) the passing
occurred.
The x-axis represents the hours of the day and the y-axis represents the
speed:
ExampleGet your own Python Server
Start by drawing a scatter plot:
x = [1,2,3,5,6,7,8,9,10,12,13,14,15,16,18,19,21,22]
y = [100,90,80,60,60,55,60,65,70,70,75,76,78,79,90,99,99,100]
[Link](x, y)
[Link]()
Result:
Example
Import numpy and matplotlib then draw the line of Polynomial Regression:
import numpy
import [Link] as plt
x = [1,2,3,5,6,7,8,9,10,12,13,14,15,16,18,19,21,22]
y = [100,90,80,60,60,55,60,65,70,70,75,76,78,79,90,99,99,100]
[Link](x, y)
[Link](myline, mymodel(myline))
[Link]()
Result:
Example Explained
Import the modules you need.
You can learn about the NumPy module in our NumPy Tutorial.
You can learn about the SciPy module in our SciPy Tutorial.
import numpy
import [Link] as plt
Create the arrays that represent the values of the x and y axis:
x = [1,2,3,5,6,7,8,9,10,12,13,14,15,16,18,19,21,22]
y = [100,90,80,60,60,55,60,65,70,70,75,76,78,79,90,99,99,100]
Then specify how the line will display, we start at position 1, and end at
position 22:
[Link](x, y)
[Link](myline, mymodel(myline))
[Link]()
REMOVE ADS
R-Squared
It is important to know how well the relationship between the values of the x-
and y-axis is, if there are no relationship the polynomial regression can not
be used to predict anything.
Python and the Sklearn module will compute this value for you, all you have
to do is feed it with the x and y arrays:
Example
How well does my data fit in a polynomial regression?
import numpy
from [Link] import r2_score
x = [1,2,3,5,6,7,8,9,10,12,13,14,15,16,18,19,21,22]
y = [100,90,80,60,60,55,60,65,70,70,75,76,78,79,90,99,99,100]
print(r2_score(y, mymodel(x)))
Note: The result 0.94 shows that there is a very good relationship, and we
can use polynomial regression in future predictions.
Example: Let us try to predict the speed of a car that passes the tollbooth at
around the time 17:00:
To do so, we need the same mymodel array from the example above:
Example
Predict the speed of a car passing at 17:00:
import numpy
from [Link] import r2_score
x = [1,2,3,5,6,7,8,9,10,12,13,14,15,16,18,19,21,22]
y = [100,90,80,60,60,55,60,65,70,70,75,76,78,79,90,99,99,100]
mymodel = numpy.poly1d([Link](x, y, 3))
speed = mymodel(17)
print(speed)
The example predicted a speed to be 88.87, which we also could read from
the diagram:
REMOVE ADS
Bad Fit?
Let us create an example where polynomial regression would not be the best
method to predict future values.
Example
These values for the x- and y-axis should result in a very bad fit for
polynomial regression:
import numpy
import [Link] as plt
x = [89,43,36,36,95,10,66,34,38,20,26,29,48,64,6,5,36,66,72,40]
y = [21,46,3,35,67,95,53,72,58,10,26,34,90,33,38,20,56,2,47,15]
[Link](x, y)
[Link](myline, mymodel(myline))
[Link]()
Result:
import numpy
from [Link] import r2_score
x = [89,43,36,36,95,10,66,34,38,20,26,29,48,64,6,5,36,66,72,40]
y = [21,46,3,35,67,95,53,72,58,10,26,34,90,33,38,20,56,2,47,15]
print(r2_score(y, mymodel(x)))
The result: 0.00995 indicates a very bad relationship, and tells us that this
data set is not suitable for polynomial regression.
Multiple Regression
Multiple regression is like linear regression, but with more than one
independent value, meaning that we try to predict a value based on two or
more variables.
Take a look at the data set below, it contains some information about cars.
In Python we have modules that will do the work for us. Start by importing
the Pandas module.
import pandas
The Pandas module allows us to read csv files and return a DataFrame
object.
The file is meant for testing purposes only, you can download it
here: [Link]
df = pandas.read_csv("[Link]")
Then make a list of the independent values and call this variable X.
X = df[['Weight', 'Volume']]
y = df['CO2']
Tip: It is common to name the list of independent values with a upper case
X, and the list of dependent values with a lower case y.
We will use some methods from the sklearn module, so we will have to
import that module as well:
This object has a method called fit() that takes the independent and
dependent values as parameters and fills the regression object with data
that describes the relationship:
regr = linear_model.LinearRegression()
[Link](X, y)
Now we have a regression object that are ready to predict CO2 values based
on a car's weight and volume:
#predict the CO2 emission of a car where the weight is 2300kg, and the
volume is 1300cm3:
predictedCO2 = [Link]([[2300, 1300]])
import pandas
from sklearn import linear_model
df = pandas.read_csv("[Link]")
X = df[['Weight', 'Volume']]
y = df['CO2']
regr = linear_model.LinearRegression()
[Link](X, y)
#predict the CO2 emission of a car where the weight is 2300kg, and the
volume is 1300cm3:
predictedCO2 = [Link]([[2300, 1300]])
print(predictedCO2)
Result:
[107.2087328]
We have predicted that a car with 1.3 liter engine, and a weight of 2300 kg,
will release approximately 107 grams of CO2 for every kilometer it drives.
REMOVE ADS
Coefficient
The coefficient is a factor that describes the relationship with an unknown
variable.
In this case, we can ask for the coefficient value of weight against CO2, and
for volume against CO2. The answer(s) we get tells us what would happen if
we increase, or decrease, one of the independent values.
Example
import pandas
from sklearn import linear_model
df = pandas.read_csv("[Link]")
X = df[['Weight', 'Volume']]
y = df['CO2']
regr = linear_model.LinearRegression()
[Link](X, y)
print(regr.coef_)
Result:
[0.00755095 0.00780526]
Result Explained
The result array represents the coefficient values of weight and volume.
Weight: 0.00755095
Volume: 0.00780526
These values tell us that if the weight increase by 1kg, the CO2 emission
increases by 0.00755095g.
And if the engine size (Volume) increases by 1cm 3, the CO2 emission
increases by 0.00780526g.
Example
Copy the example from before, but change the weight from 2300 to 3300:
import pandas
from sklearn import linear_model
df = pandas.read_csv("[Link]")
X = df[['Weight', 'Volume']]
y = df['CO2']
regr = linear_model.LinearRegression()
[Link](X, y)
print(predictedCO2)
Result:
[114.75968007]
We have predicted that a car with 1.3 liter engine, and a weight of 3300 kg,
will release approximately 115 grams of CO2 for every kilometer it drives.
Scale Features
When your data has different values, and even different measurement units,
it can be difficult to compare them. What is kilograms compared to meters?
Or altitude compared to time?
The answer to this problem is scaling. We can scale data into new values
that are easier to compare.
Take a look at the table below, it is the same data set that we used in
the multiple regression chapter, but this time the volume column contains
values in liters instead of cm3 (1.0 instead of 1000).
It can be difficult to compare the volume 1.0 with the weight 790, but if we
scale them both into comparable values, we can easily see how much one
value is compared to the other.
There are different methods for scaling data, in this tutorial we will use a
method called standardization.
z = (x - u) / s
Where z is the new value, x is the original value, u is the mean and s is the
standard deviation.
If you take the weight column from the data set above, the first value is
790, and the scaled value will be:
If you take the volume column from the data set above, the first value is
1.0, and the scaled value will be:
Now you can compare -2.1 with -1.59 instead of comparing 790 with 1.0.
You do not have to do this manually, the Python sklearn module has a
method called StandardScaler() which returns a Scaler object with methods
for transforming data sets.
import pandas
from sklearn import linear_model
from [Link] import StandardScaler
scale = StandardScaler()
df = pandas.read_csv("[Link]")
X = df[['Weight', 'Volume']]
scaledX = scale.fit_transform(X)
print(scaledX)
Result:
Note that the first two values are -2.1 and -1.59, which corresponds to our
calculations:
[[-2.10389253 -1.59336644]
[-0.55407235 -1.07190106]
[-1.52166278 -1.59336644]
[-1.78973979 -1.85409913]
[-0.63784641 -0.28970299]
[-1.52166278 -1.59336644]
[-0.76769621 -0.55043568]
[ 0.3046118 -0.28970299]
[-0.7551301 -0.28970299]
[-0.59595938 -0.0289703 ]
[-1.30803892 -1.33263375]
[-1.26615189 -0.81116837]
[-0.7551301 -1.59336644]
[-0.16871166 -0.0289703 ]
[ 0.14125238 -0.0289703 ]
[ 0.15800719 -0.0289703 ]
[ 0.3046118 -0.0289703 ]
[-0.05142797 1.53542584]
[-0.72580918 -0.0289703 ]
[ 0.14962979 1.01396046]
[ 1.2219378 -0.0289703 ]
[ 0.5685001 1.01396046]
[ 0.3046118 1.27469315]
[ 0.51404696 -0.0289703 ]
[ 0.51404696 1.01396046]
[ 0.72348212 -0.28970299]
[ 0.8281997 1.01396046]
[ 1.81254495 1.01396046]
[ 0.96642691 -0.0289703 ]
[ 1.72877089 1.01396046]
[ 1.30990057 1.27469315]
[ 1.90050772 1.01396046]
[-0.23991961 -0.0289703 ]
[ 0.40932938 -0.0289703 ]
[ 0.47215993 -0.0289703 ]
[ 0.4302729 2.31762392]]
REMOVE ADS
When the data set is scaled, you will have to use the scale when you predict
values:
Example
Predict the CO2 emission from a 1.3 liter car that weighs 2300 kilograms:
import pandas
from sklearn import linear_model
from [Link] import StandardScaler
scale = StandardScaler()
df = pandas.read_csv("[Link]")
X = df[['Weight', 'Volume']]
y = df['CO2']
scaledX = scale.fit_transform(X)
regr = linear_model.LinearRegression()
[Link](scaledX, y)
predictedCO2 = [Link]([scaled[0]])
print(predictedCO2)
Result:
[107.2087328]
It is called Train/Test because you split the data set into two sets: a training
set and a testing set.
Our data set illustrates 100 customers in a shop, and their shopping habits.
x = [Link](3, 1, 100)
y = [Link](150, 40, 100) / x
[Link](x, y)
[Link]()
Result:
The x axis represents the number of minutes before making a purchase.
train_x = x[:80]
train_y = y[:80]
test_x = x[80:]
test_y = y[80:]
Display the Training Set
Display the same scatter plot with the training set:
Example
[Link](train_x, train_y)
[Link]()
Result:
It looks like the original data set, so it seems to be a fair selection:
Display the Testing Set
To make sure the testing set is not completely different, we will take a look
at the testing set as well.
Example
[Link](test_x, test_y)
[Link]()
Result:
The testing set also looks like the original data set:
To draw a line through the data points, we use the plot() method of the
matplotlib module:
Example
Draw a polynomial regression line through the data points:
import numpy
import [Link] as plt
[Link](2)
x = [Link](3, 1, 100)
y = [Link](150, 40, 100) / x
train_x = x[:80]
train_y = y[:80]
test_x = x[80:]
test_y = y[80:]
[Link](train_x, train_y)
[Link](myline, mymodel(myline))
[Link]()
Result:
The result can back my suggestion of the data set fitting a polynomial
regression, even though it would give us some weird results if we try to
predict values outside of the data set. Example: the line indicates that a
customer spending 6 minutes in the shop would make a purchase worth 200.
That is probably a sign of overfitting.
But what about the R-squared score? The R-squared score is a good indicator
of how well my data set is fitting the model.
R2
Remember R2, also known as R-squared?
It measures the relationship between the x axis and the y axis, and the value
ranges from 0 to 1, where 0 means no relationship, and 1 means totally
related.
The sklearn module has a method called r2_score() that will help us find this
relationship.
In this case we would like to measure the relationship between the minutes a
customer stays in the shop and how much money they spend.
Example
How well does my training data fit in a polynomial regression?
import numpy
from [Link] import r2_score
[Link](2)
x = [Link](3, 1, 100)
y = [Link](150, 40, 100) / x
train_x = x[:80]
train_y = y[:80]
test_x = x[80:]
test_y = y[80:]
r2 = r2_score(train_y, mymodel(train_x))
print(r2)
Note: The result 0.799 shows that there is a OK relationship.
Now we want to test the model with the testing data as well, to see if gives
us the same result.
Example
Let us find the R2 score when using testing data:
import numpy
from [Link] import r2_score
[Link](2)
x = [Link](3, 1, 100)
y = [Link](150, 40, 100) / x
train_x = x[:80]
train_y = y[:80]
test_x = x[80:]
test_y = y[80:]
r2 = r2_score(test_y, mymodel(test_x))
print(r2)
Note: The result 0.809 shows that the model fits the testing set as well, and
we are confident that we can use the model to predict future values.
Predict Values
Now that we have established that our model is OK, we can start predicting
new values.
Example
How much money will a buying customer spend, if she or he stays in the
shop for 5 minutes?
print(mymodel(5))
Luckily our example person has registered every time there was a comedy
show in town, and registered some information about the comedian, and also
registered if he/she went or not.
36 10 9 UK
42 12 4 USA
23 4 6 N
52 4 4 USA
43 21 8 USA
44 14 5 UK
66 3 7 N
35 14 9 UK
52 13 7 N
35 5 9 N
24 3 5 USA
18 3 7 UK
45 9 9 UK
Now, based on this data set, Python can create a decision tree that can be
used to decide if any new shows are worth attending to.
REMOVE ADS
import pandas
df = pandas.read_csv("[Link]")
print(df)
To make a decision tree, all data has to be numerical.
We have to convert the non numerical columns 'Nationality' and 'Go' into
numerical values.
Pandas has a map() method that takes a dictionary with information on how
to convert the values.
Example
Change string values into numerical values:
print(df)
Then we have to separate the feature columns from the target column.
The feature columns are the columns that we try to predict from, and the
target column is the column with the values we try to predict.
Example
X is the feature columns, y is the target column:
X = df[features]
y = df['Go']
print(X)
print(y)
Now we can create the actual decision tree, fit it with our details. Start by
importing the modules we need:
Example
Create and display a Decision Tree:
import pandas
from sklearn import tree
from [Link] import DecisionTreeClassifier
import [Link] as plt
df = pandas.read_csv("[Link]")
X = df[features]
y = df['Go']
dtree = DecisionTreeClassifier()
dtree = [Link](X, y)
tree.plot_tree(dtree, feature_names=features)
Result Explained
The decision tree uses your earlier decisions to calculate the odds for you to
wanting to go see a comedian or not.
Rank
Rank <= 6.5 means that every comedian with a rank of 6.5 or lower will follow
the True arrow (to the left), and the rest will follow the False arrow (to the
right).
gini = 0.497 refers to the quality of the split, and is always a number
between 0.0 and 0.5, where 0.0 would mean all of the samples got the same
result, and 0.5 would mean that the split is done exactly in the middle.
samples = 13means that there are 13 comedians left at this point in the
decision, which is all of them since this is the first step.
value = [6, 7] means that of these 13 comedians, 6 will get a "NO", and 7
will get a "GO".
Gini
There are many ways to split the samples, we use the GINI method in this
tutorial.
samples = 5means that there are 5 comedians left in this branch (5 comedian
with a Rank of 6.5 or lower).
value = [5, 0] means that 5 will get a "NO" and 0 will get a "GO".
gini = 0.219 means that about 22% of the samples would go in one direction.
samples = 8means that there are 8 comedians left in this branch (8 comedian
with a Rank higher than 6.5).
value = [1, 7] means that of these 8 comedians, 1 will get a "NO" and 7 will
get a "GO".
True - 4 Comedians Continue:
Age
Age <= 35.5means that comedians at the age of 35.5 or younger will follow
the arrow to the left, and the rest will follow the arrow to the right.
gini = 0.375 means that about 37,5% of the samples would go in one
direction.
value = [1, 3] means that of these 4 comedians, 1 will get a "NO" and 3 will
get a "GO".
value = [0, 4] means that of these 4 comedians, 0 will get a "NO" and 4 will
get a "GO".
True - 2 Comedians End Here:
gini = 0.0 means all of the samples got the same result.
value = [0, 2] means that of these 2 comedians, 0 will get a "NO" and 2 will
get a "GO".
gini = 0.5 means that 50% of the samples would go in one direction.
value = [0, 1] means that 0 will get a "NO" and 1 will get a "GO".
value = [1, 0] means that 1 will get a "NO" and 0 will get a "GO".
Predict Values
We can use the Decision Tree to predict new values.
Example
Use predict() method to predict new values:
Example
What would the answer be if the comedy rank was 6?
Different Results
You will see that the Decision Tree gives you different results if you run it
enough times, even if you feed it with the same data.
That is because the Decision Tree does not give us a 100% certain answer. It
is based on the probability of an outcome, and the answer will vary.
The rows represent the actual classes the outcomes should have been. While
the columns represent the predictions we have made. Using this table it is
easy to see which predictions are wrong.
Creating a Confusion Matrix
Confusion matrixes can be created by predictions made from a logistic
regression.
For now we will generate actual and predicted values by utilizing NumPy:
import numpy
Next we will need to generate the numbers for "actual" and "predicted"
values.
In order to create the confusion matrix we need to import metrics from the
sklearn module.
Once metrics is imported we can use the confusion matrix function on our
actual and predicted values.
cm_display = [Link](confusion_matrix =
confusion_matrix, display_labels = [0, 1])
Finally to display the plot we can use the functions plot() and show() from
pyplot.
cm_display.plot()
[Link]()
cm_display = [Link](confusion_matrix =
confusion_matrix, display_labels = [0, 1])
cm_display.plot()
[Link]()
Result
Results Explained
The Confusion Matrix created has four different quadrants:
True means that the values were accurately predicted, False means that
there was an error or wrong prediction.
REMOVE ADS
Created Metrics
The matrix provides us with many useful metrics that help us to evaluate our
classification model.
Accuracy
Accuracy measures how often the model is correct.
How to Calculate
(True Positive + True Negative) / Total Predictions
Example
Accuracy = metrics.accuracy_score(actual, predicted)
Precision
Of the positives predicted, what percentage is truly positive?
How to Calculate
True Positive / (True Positive + False Positive)
Example
Precision = metrics.precision_score(actual, predicted)
Sensitivity (Recall)
Of all the positive cases, what percentage are predicted positive?
This means it looks at true positives and false negatives (which are positives
that have been incorrectly predicted as negative).
How to Calculate
True Positive / (True Positive + False Negative)
Example
Sensitivity_recall = metrics.recall_score(actual, predicted)
Specificity
How well the model is at prediciting negative results?
How to Calculate
True Negative / (True Negative + False Positive)
Since it is just the opposite of Recall, we use the recall_score function, taking
the opposite position label:
Example
Specificity = metrics.recall_score(actual, predicted, pos_label=0)
F-score
F-score is the "harmonic mean" of precision and sensitivity.
It considers both false positive and false negative cases and is good for
imbalanced datasets.
How to Calculate
2 * ((Precision * Sensitivity) / (Precision + Sensitivity))
This score does not take into consideration the True Negative values:
Example
F1_score = metrics.f1_score(actual, predicted)
Example
#metrics
print({"Accuracy":Accuracy,"Precision":Precision,"Sensitivity_recall"
:Sensitivity_recall,"Specificity":Specificity,"F1_score":F1_score})
Machine Learning -
Hierarchical Clustering
Hierarchical Clustering
Hierarchical clustering is an unsupervised learning method for clustering
data points. The algorithm builds clusters by measuring the dissimilarities
between data. Unsupervised learning means that a model does not have to
be trained, and we do not need a "target" variable. This method can be used
on any data to visualize and interpret the relationship between individual
data points.
Here we will use hierarchical clustering to group data points and visualize the
clusters using both a dendrogram and scatter plot.
import numpy as np
import [Link] as plt
Result
Now we compute the ward linkage using euclidean distance, and visualize it
using a dendrogram:
Example
import numpy as np
import [Link] as plt
from [Link] import dendrogram, linkage
[Link]()
Result
Here, we do the same thing with Python's scikit-learn library. Then, visualize
on a 2-dimensional plot:
Example
import numpy as np
import [Link] as plt
from [Link] import AgglomerativeClustering
hierarchical_cluster =
AgglomerativeClustering(n_clusters=2, linkage='ward')
labels = hierarchical_cluster.fit_predict(data)
[Link](x, y, c=labels)
[Link]()
Result
REMOVE ADS
Example Explained
Import the modules you need.
import numpy as np
import [Link] as plt
from [Link] import dendrogram, linkage
from [Link] import AgglomerativeClustering
You can learn about the Matplotlib module in our "Matplotlib Tutorial.
You can learn about the SciPy module in our SciPy Tutorial.
NumPy is a library for working with arrays and matricies in Python, you can
learn about the NumPy module in our NumPy Tutorial.
Create arrays that resemble two variables in a dataset. Note that while we
only use two variables here, this method will work with any number of
variables:
Result:
[(4, 21), (5, 19), (10, 24), (4, 17), (3, 16), (11, 25),
(14, 24), (6, 22), (10, 21), (12, 21)]
Compute the linkage between all of the different points. Here we use a
simple euclidean distance measure and Ward's linkage, which seeks to
minimize the variance between clusters.
Finally, plot the results in a dendrogram. This plot will show us the hierarchy
of clusters from the bottom (individual points) to the top (a single cluster
consisting of all data points).
[Link]() lets us visualize the dendrogram instead of just the raw linkage
data.
dendrogram(linkage_data)
[Link]()
Result:
hierarchical_cluster = AgglomerativeClustering(n_clusters=2,
linkage='ward')
The .fit_predict method can be called on our data to compute the clusters
using the defined parameters across our chosen number of clusters.
[0 0 1 0 0 1 1 0 1 1]
Finally, if we plot the same data and color the points using the labels
assigned to each index by the hierarchical clustering method, we can see the
cluster each point was assigned to:
[Link](x, y, c=labels)
[Link]()
Result:
Logistic Regression
Logistic regression aims to solve classification problems. It does this by
predicting categorical outcomes, unlike linear regression that predicts a
continuous outcome.
In the simplest case there are two outcomes, which is called binomial, an
example of which is predicting if a tumor is malignant or benign. Other cases
have more than two outcomes to classify, in this case it is called multinomial.
A common example for multinomial logistic regression would be predicting
the class of an iris flower between 3 different species.
We will use a method from the sklearn module, so we will have to import that
module as well:
This object has a method called fit() that takes the independent and
dependent values as parameters and fills the regression object with data
that describes the relationship:
logr = linear_model.LogisticRegression()
[Link](X,y)
import numpy
from sklearn import linear_model
#Reshaped for Logistic function.
X =
[Link]([3.78, 2.44, 2.09, 0.14, 1.72, 1.65, 4.92, 4.37, 4.96, 4.
52, 3.69, 5.88]).reshape(-1,1)
y = [Link]([0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1])
logr = linear_model.LogisticRegression()
[Link](X,y)
Result
[0]
We have predicted that a tumor with a size of 3.46mm will not be cancerous.
REMOVE ADS
Coefficient
In logistic regression the coefficient is the expected change in log-odds of
having the outcome per unit change in X.
This does not have the most intuitive understanding so let's use it to create
something that makes more sense, odds.
Example
See the whole example in action:
import numpy
from sklearn import linear_model
logr = linear_model.LogisticRegression()
[Link](X,y)
log_odds = logr.coef_
odds = [Link](log_odds)
print(odds)
Result
[4.03541657]
This tells us that as the size of a tumor increases by 1mm the odds of it
being a cancerous tumor increases by 4x.
REMOVE ADS
Probability
The coefficient and intercept values can be used to find the probability that
each tumor is cancerous.
Create a function that uses the model's coefficient and intercept values to
return a new value. This new value represents probability that the given
observation is a tumor:
def logit2prob(logr,x):
log_odds = logr.coef_ * x + logr.intercept_
odds = [Link](log_odds)
probability = odds / (1 + odds)
return(probability)
Function Explained
To find the log-odds for each observation, we must first create a formula that
looks similar to the one from linear regression, extracting the coefficient and
the intercept.
log_odds = logr.coef_ * x + logr.intercept_
odds = [Link](log_odds)
Let us now use the function with what we have learned to find out the
probability that each tumor is cancerous.
Example
See the whole example in action:
import numpy
from sklearn import linear_model
X =
[Link]([3.78, 2.44, 2.09, 0.14, 1.72, 1.65, 4.92, 4.37, 4.96, 4.
52, 3.69, 5.88]).reshape(-1,1)
y = [Link]([0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1])
logr = linear_model.LogisticRegression()
[Link](X,y)
print(logit2prob(logr, X))
Result
[[0.60749955]
[0.19268876]
[0.12775886]
[0.00955221]
[0.08038616]
[0.07345637]
[0.88362743]
[0.77901378]
[0.88924409]
[0.81293497]
[0.57719129]
[0.96664243]]
Results Explained
3.78 0.61 The probability that a tumor with the size 3.78cm is cancerous is
61%.
2.44 0.19 The probability that a tumor with the size 2.44cm is cancerous is
19%.
2.09 0.13 The probability that a tumor with the size 2.09cm is cancerous is
13%
Grid Search
The majority of machine learning models contain parameters that can be
adjusted to vary how the model learns. For example, the logistic regression
model, from sklearn, has a parameter C that controls regularization,which
affects the complexity of the model.
How do we pick the best value for C? The best value is dependent on the
data used to train the model.
Before we get into the example it is good to know what the parameter we
are changing does. Higher values of C tell the model, the training data
resembles real world information, place a greater weight on the training
data. While lower values of C do the opposite.
To get started we must first load in the dataset we will be working with.
X = iris['data']
y = iris['target']
Now we will load the logistic model for classifying the iris flowers.
Creating the model, setting max_iter to a higher value to ensure that the
model finds a result.
In the example below, we look at the iris data set and try to train a model
with varying values for C in logistic regression.
After we create the model, we must fit the model to the data.
print([Link](X,y))
To evaluate the model we run the score method.
print([Link](X,y))
iris = datasets.load_iris()
X = iris['data']
y = iris['target']
print([Link](X,y))
print([Link](X,y))
REMOVE ADS
Knowing which values to set for the searched parameters will take a
combination of domain knowledge and practice.
Since the default value for C is 1, we will set a range of values surrounding it.
Next we will create a for loop to change out the values of C and evaluate the
model with each change.
First we will create an empty list to store the score within.
scores = []
To change the values of C we must loop over the range of values and update
the parameter each time.
for choice in C:
logit.set_params(C=choice)
[Link](X, y)
[Link]([Link](X, y))
With the scores stored in a list, we can evaluate what the best choice of C is.
print(scores)
Example
from sklearn import datasets
from sklearn.linear_model import LogisticRegression
iris = datasets.load_iris()
X = iris['data']
y = iris['target']
scores = []
for choice in C:
logit.set_params(C=choice)
[Link](X, y)
[Link]([Link](X, y))
print(scores)
Results Explained
We can see that the lower values of C performed worse than the base
parameter of 1. However, as we increased the value of C to 1.75 the model
experienced increased accuracy.
It seems that increasing C beyond this amount does not help increase model
accuracy.
To avoid being misled by the scores on the training data, we can put aside a
portion of our data and use it specifically for the purpose of testing the
model. Refer to the lecture on train/test splitting to avoid being misled and
overfitting.
Preprocessing - Categorical
Data
Categorical Data
When your data has categories represented by strings, it will be difficult to
use them to train machine learning models which often only accepts numeric
data.
Instead of ignoring the categorical data and excluding the information from
our model, you can tranform the data so it can be used in your models.
Take a look at the table below, it is the same data set that we used in
the multiple regression chapter.
cars = pd.read_csv('[Link]')
print(cars.to_string())
Result
Car Model Volume Weight CO2
0 Toyoty Aygo 1000 790 99
1 Mitsubishi Space Star 1200 1160 95
2 Skoda Citigo 1000 929 95
3 Fiat 500 900 865 90
4 Mini Cooper 1500 1140 105
5 VW Up! 1000 929 105
6 Skoda Fabia 1400 1109 90
7 Mercedes A-Class 1500 1365 92
8 Ford Fiesta 1500 1112 98
9 Audi A1 1600 1150 99
10 Hyundai I20 1100 980 99
11 Suzuki Swift 1300 990 101
12 Ford Fiesta 1000 1112 99
13 Honda Civic 1600 1252 94
14 Hundai I30 1600 1326 97
15 Opel Astra 1600 1330 97
16 BMW 1 1600 1365 99
17 Mazda 3 2200 1280 104
18 Skoda Rapid 1600 1119 104
19 Ford Focus 2000 1328 105
20 Ford Mondeo 1600 1584 94
21 Opel Insignia 2000 1428 99
22 Mercedes C-Class 2100 1365 99
23 Skoda Octavia 1600 1415 99
24 Volvo S60 2000 1415 99
25 Mercedes CLA 1500 1465 102
26 Audi A4 2000 1490 104
27 Audi A6 2000 1725 114
28 Volvo V70 1600 1523 109
29 BMW 5 2000 1705 114
30 Mercedes E-Class 2100 1605 115
31 Volvo XC70 2000 1746 117
32 Ford B-Max 1600 1235 104
33 BMW 216 1600 1390 108
34 Opel Zafira 1600 1405 109
35 Mercedes SLK 2500 1395 120
In the multiple regression chapter, we tried to predict the CO2 emitted based
on the volume of the engine and the weight of the car but we excluded
information about the car brand and model.
The information about the car brand or the car model might help us make a
better prediction of the CO2 emitted.
For each column, the values will be 1 or 0 where 1 represents the inclusion of
the group and 0 represents the exclusion. This transformation is called one
hot encoding.
You do not have to do this manually, the Python Pandas module has a
function that called get_dummies() which does one hot encoding.
Example
One Hot Encode the Car column:
import pandas as pd
cars = pd.read_csv('[Link]')
ohe_cars = pd.get_dummies(cars[['Car']])
print(ohe_cars.to_string())
Result
Car_Audi Car_BMW Car_Fiat Car_Ford Car_Honda
Car_Hundai Car_Hyundai Car_Mazda Car_Mercedes Car_Mini
Car_Mitsubishi Car_Opel Car_Skoda Car_Suzuki Car_Toyoty
Car_VW Car_Volvo
0 0 0 0 0 0
0 0 0 0 0 0
0 0 0 1 0 0
1 0 0 0 0 0
0 0 0 0 0 1
0 0 0 0 0 0
2 0 0 0 0 0
0 0 0 0 0 0
0 1 0 0 0 0
3 0 0 1 0 0
0 0 0 0 0 0
0 0 0 0 0 0
4 0 0 0 0 0
0 0 0 0 1 0
0 0 0 0 0 0
5 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 1 0
6 0 0 0 0 0
0 0 0 0 0 0
0 1 0 0 0 0
7 0 0 0 0 0
0 0 0 1 0 0
0 0 0 0 0 0
8 0 0 0 1 0
0 0 0 0 0 0
0 0 0 0 0 0
9 1 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
10 0 0 0 0 0
0 1 0 0 0 0
0 0 0 0 0 0
11 0 0 0 0 0
0 0 0 0 0 0
0 0 1 0 0 0
12 0 0 0 1 0
0 0 0 0 0 0
0 0 0 0 0 0
13 0 0 0 0 1
0 0 0 0 0 0
0 0 0 0 0 0
14 0 0 0 0 0
1 0 0 0 0 0
0 0 0 0 0 0
15 0 0 0 0 0
0 0 0 0 0 0
1 0 0 0 0 0
16 0 1 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
17 0 0 0 0 0
0 0 1 0 0 0
0 0 0 0 0 0
18 0 0 0 0 0
0 0 0 0 0 0
0 1 0 0 0 0
19 0 0 0 1 0
0 0 0 0 0 0
0 0 0 0 0 0
20 0 0 0 1 0
0 0 0 0 0 0
0 0 0 0 0 0
21 0 0 0 0 0
0 0 0 0 0 0
1 0 0 0 0 0
22 0 0 0 0 0
0 0 0 1 0 0
0 0 0 0 0 0
23 0 0 0 0 0
0 0 0 0 0 0
0 1 0 0 0 0
24 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 1
25 0 0 0 0 0
0 0 0 1 0 0
0 0 0 0 0 0
26 1 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
27 1 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
28 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 1
29 0 1 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
30 0 0 0 0 0
0 0 0 1 0 0
0 0 0 0 0 0
31 0 0 0 0 0
0 0 0 0 0 0
0 0 0 0 0 1
32 0 0 0 1 0
0 0 0 0 0 0
0 0 0 0 0 0
33 0 1 0 0 0
0 0 0 0 0 0
0 0 0 0 0 0
34 0 0 0 0 0
0 0 0 0 0 0
1 0 0 0 0 0
35 0 0 0 0 0
0 0 0 1 0 0
0 0 0 0 0 0
Results
A column was created for every car brand in the Car column.
REMOVE ADS
Predict CO2
We can use this additional information alongside the volume and weight to
predict CO2
To combine the information, we can use the concat() function from pandas.
import pandas
The pandas module allows us to read csv files and manipulate DataFrame
objects:
cars = pandas.read_csv("[Link]")
It also allows us to create the dummy variables:
ohe_cars = pandas.get_dummies(cars[['Car']])
Then we must select the independent variables (X) and add the dummy
variables columnwise.
regr = linear_model.LinearRegression()
[Link](X,y)
Finally we can predict the CO2 emissions based on the car's weight, volume,
and manufacturer.
Example
import pandas
from sklearn import linear_model
cars = pandas.read_csv("[Link]")
ohe_cars = pandas.get_dummies(cars[['Car']])
regr = linear_model.LinearRegression()
[Link](X,y)
print(predictedCO2)
Result
[122.45153299]
We now have a coefficient for the volume, the weight, and each car brand in
the data set
REMOVE ADS
Dummifying
It is not necessary to create one column for each group in your category. The
information can be retained using 1 column less than the number of groups
you have.
For example, you have a column representing colors and in that column, you
have two colors, red and blue.
Example
import pandas as pd
print(colors)
Result
color
0 blue
1 red
You can create 1 column called red where 1 represents red and 0 represents
not red, which means it is blue.
To do this, we can use the same function that we used for one hot encoding,
get_dummies, and then drop one of the columns. There is an argument,
drop_first, which allows us to exclude the first column from the resulting
table.
Example
import pandas as pd
print(dummies)
Result
color_red
0 0
1 1
What if you have more than 2 groups? How can the multiple groups be
represented by 1 less column?
Let's say we have three colors this time, red, blue and green. When we
get_dummies while dropping the first column, we get the following table.
Example
import pandas as pd
print(dummies)
Result
color_green color_red color
0 0 0 blue
1 0 1 red
2 1 0 green
Machine Learning - K-means
K-means
K-means is an unsupervised learning method for clustering data points. The
algorithm iteratively divides data points into K clusters by minimizing the
variance in each cluster.
Here, we will show you how to estimate the best value for K using the elbow
method, then use K-means clustering to group the data points into clusters.
[Link](x, y)
[Link]()
Result
Now we utilize the elbow method to visualize the intertia for different values
of K:
Example
from [Link] import KMeans
for i in range(1,11):
kmeans = KMeans(n_clusters=i)
[Link](data)
[Link](kmeans.inertia_)
Result
The elbow method shows that 2 is a good value for K, so we retrain and
visualize the result:
Example
kmeans = KMeans(n_clusters=2)
[Link](data)
[Link](x, y, c=kmeans.labels_)
[Link]()
Result
REMOVE ADS
Example Explained
Import the modules you need.
You can learn about the Matplotlib module in our "Matplotlib Tutorial.
Result:
[(4, 21), (5, 19), (10, 24), (4, 17), (3, 16), (11, 25),
(14, 24), (6, 22), (10, 21), (12, 21)]
In order to find the best value for K, we need to run K-means across our data
for a range of possible values. We only have 10 data points, so the maximum
number of clusters is 10. So for each value K in range(1,11), we train a K-
means model and plot the intertia at that number of clusters:
inertias = []
for i in range(1,11):
kmeans = KMeans(n_clusters=i)
[Link](data)
[Link](kmeans.inertia_)
Result:
We can see that the "elbow" on the graph above (where the interia becomes
more linear) is at K=2. We can then fit our K-means algorithm one more time
and plot the different clusters assigned to the data:
kmeans = KMeans(n_clusters=2)
[Link](data)
[Link](x, y, c=kmeans.labels_)
[Link]()
Result:
Machine Learning - Bootstrap
Aggregation (Bagging)
Bagging
Methods such as Decision Trees, can be prone to overfitting on the training
set which can lead to wrong predictions on new data.
Next we need to load in the data and store it into X (input features) and y
(target). The parameter as_frame is set equal to True so we do not lose the
feature names when loading the data. (sklearn version older than 0.23 must
skip the as_frame argument as it is not supported)
X = [Link]
y = [Link]
With our data prepared, we can now instantiate a base classifier and fit it to
the training data.
DecisionTreeClassifier(random_state=22)
We can now predict the class of wine the unseen test set and evaluate the
model performance.
y_pred = [Link](X_test)
Result:
X = [Link]
y = [Link]
y_pred = [Link](X_test)
Now that we have a baseline accuracy for the test dataset, we can see how
the Bagging Classifier out performs a single Decision Tree Classifier.
REMOVE ADS
For this sample dataset the number of estimators is relatively low, it is often
the case that much larger ranges are explored. Hyperparameter tuning is
usually done with a grid search, but for now we will use a select set of values
for the number of estimators.
Now lets create a range of values that represent the number of estimators
we want to use in each ensemble.
estimator_range = [2,4,6,8,10,12,14,16]
models = []
scores = []
for n_estimators in estimator_range:
With the models and scores stored, we can now visualize the improvement in
model performance.
# Visualize plot
[Link]()
Example
Import the necessary data and evaluate
the BaggingClassifier performance.
estimator_range = [2,4,6,8,10,12,14,16]
models = []
scores = []
# Visualize plot
[Link]()
Result
Results Explained
By iterating through different values for the number of estimators we can
see an increase in model performance from 82.2% to 95.5%. After 14
estimators the accuracy begins to drop, again if you set a
different random_state the values you see will vary. That is why it is best
practice to use cross validation to ensure stable results.
In this case, we see a 13.3% increase in accuracy when it comes to
identifying the type of the wine.
REMOVE ADS
We saw in the last exercise that 12 estimators yielded the highest accuracy,
so we will use that to create our model. This time setting the
parameter oob_score to true to evaluate the model with out-of-bag score.
Example
Create a model with out-of-bag metric.
X = [Link]
y = [Link]
oob_model.fit(X_train, y_train)
print(oob_model.oob_score_)
Since the samples used in OOB and the test set are different, and the
dataset is relatively small, there is a difference in the accuracy. It is rare that
they would be exactly the same, again OOB should be used quick means for
estimating error, but is not the only evaluation metric.
Note: This is only functional with smaller datasets, where the trees are
relatively shallow and narrow making them easy to visualize.
Example
Generate Decision Trees from Bagging Classifier
X = [Link]
y = [Link]
[Link](X_train, y_train)
[Link](figsize=(30, 20))
Result
Here we can see just the first decision tree that was used to vote on the final
prediction. Again, by changing the index of the classifier you can see each of
the trees that have been aggregated.
Cross Validation
When adjusting models we are aiming to increase overall model performance
on unseen data. Hyperparameter tuning can lead to much better
performance on test sets. However, optimizing parameters to the test set
can lead information leakage causing the model to preform worse on unseen
data. To correct for this we can perform cross validation.
X, y = datasets.load_iris(return_X_y=True)
There are many methods to cross validation, we will start by looking at k-fold
cross validation.
K-Fold
The training data used in the model is split, into k number of smaller sets, to
be used to validate the model. The model is then trained on k-1 folds of
training set. The remaining fold is then used as a validation set to evaluate
the model.
With the data loaded we can now create and fit a model for evaluation.
clf = DecisionTreeClassifier(random_state=42)
Now let's evaluate our model and see how it performs on each k-fold.
k_folds = KFold(n_splits = 5)
X, y = datasets.load_iris(return_X_y=True)
clf = DecisionTreeClassifier(random_state=42)
k_folds = KFold(n_splits = 5)
REMOVE ADS
Stratified K-Fold
In cases where classes are imbalanced we need a way to account for the
imbalance in both the train and validation sets. To do so we can stratify the
target classes, meaning that both sets will have an equal proportion of all
classes.
Example
from sklearn import datasets
from [Link] import DecisionTreeClassifier
from sklearn.model_selection import StratifiedKFold, cross_val_score
X, y = datasets.load_iris(return_X_y=True)
clf = DecisionTreeClassifier(random_state=42)
sk_folds = StratifiedKFold(n_splits = 5)
While the number of folds is the same, the average CV increases from the
basic k-fold when making sure there is stratified classes.
Leave-One-Out (LOO)
Instead of selecting the number of splits in the training data set like k-fold
LeaveOneOut, utilize 1 observation to validate and n-1 observations to train.
This method is an exaustive technique.
Example
Run LOO CV:
X, y = datasets.load_iris(return_X_y=True)
clf = DecisionTreeClassifier(random_state=42)
loo = LeaveOneOut()
REMOVE ADS
Leave-P-Out (LPO)
Leave-P-Out is simply a nuanced diffence to the Leave-One-Out idea, in that
we can select the number of p to use in our validation set.
Example
Run LPO CV:
X, y = datasets.load_iris(return_X_y=True)
clf = DecisionTreeClassifier(random_state=42)
lpo = LeavePOut(p=2)
Shuffle Split
Unlike KFold, ShuffleSplit leaves out a percentage of the data, not to be
used in the train or validation sets. To do so we must decide what the train
and test sizes are, as well as the number of splits.
Example
Run Shuffle Split CV:
X, y = datasets.load_iris(return_X_y=True)
clf = DecisionTreeClassifier(random_state=42)
Ending Notes
These are just a few of the CV methods that can be applied to models. There
are many more cross validation classes, with most models having their own
class. Check out sklearns cross validation for more CV options.
Machine Learning - AUC -
ROC Curve
Imbalanced Data
Suppose we have an imbalanced data set where the majority of our data is of
one value. We can obtain high accuracy for the model by predicting the
majority class.
n = 10000
ratio = .95
n_0 = int((1-ratio) * n)
n_1 = int(ratio * n)
Example
# below are the probabilities obtained from a hypothetical model that
doesn't always predict the mode
y_proba_2 = [Link](
[Link](0, .7, n_0).tolist() +
[Link](.3, 1, n_1).tolist()
)
y_pred_2 = y_proba_2 > .5
In cases like this, using another evaluation metric like AUC would be
preferred.
Example
Model 1:
plot_roc_curve(y, y_proba)
print(f'model 1 AUC score: {roc_auc_score(y, y_proba)}')
Result
plot_roc_curve(y, y_proba_2)
print(f'model 2 AUC score: {roc_auc_score(y, y_proba_2)}')
Result
An AUC score of around .5 would mean that the model is unable to make a
distinction between the two classes and the curve would look like a line with
a slope of 1. An AUC score closer to 1 means that the model has the ability to
separate the two classes and the curve would come closer to the top left
corner of the graph.
REMOVE ADS
Probabilities
Because AUC is a metric that utilizes probabilities of the class predictions, we
can be more confident in a model that has a higher AUC score than one with
a lower score even if they have similar accuracies.
Example
import numpy as np
n = 10000
y = [Link]([0] * n + [1] * n)
#
y_prob_1 = [Link](
[Link](.25, .5, n//2).tolist() +
[Link](.3, .7, n).tolist() +
[Link](.5, .75, n//2).tolist()
)
y_prob_2 = [Link](
[Link](0, .4, n//2).tolist() +
[Link](.3, .7, n).tolist() +
[Link](.6, 1, n//2).tolist()
)
Example
Plot model 1:
plot_roc_curve(y, y_prob_1)
Result
Example
Plot model 2:
Result
Even though the accuracies for the two models are similar, the model with
the higher AUC score will be more reliable because it takes into account the
predicted probability. It is more likely to give you higher accuracy when
predicting future data.
KNN
KNN is a simple, supervised machine learning (ML) algorithm that can be
used for classification or regression tasks - and is also frequently used in
missing value imputation. It is based on the idea that the observations
closest to a given data point are the most "similar" observations in a data
set, and we can therefore classify unforeseen points based on the values of
the closest existing points. By choosing K, the user can select the number of
nearby observations to use in the algorithm.
Here, we will show you how to implement the KNN algorithm for
classification, and show how different values of K affect the results.
[Link](x, y, c=classes)
[Link]()
Result
Now we fit the KNN algorithm with K=1:
[Link](data, classes)
Example
new_x = 8
new_y = 21
new_point = [(new_x, new_y)]
prediction = [Link](new_point)
Result
Now we do the same thing, but with a higher K value which changes the
prediction:
Example
knn = KNeighborsClassifier(n_neighbors=5)
[Link](data, classes)
prediction = [Link](new_point)
[Link](x + [new_x], y + [new_y], c=classes + [prediction[0]])
[Link](x=new_x-1.7, y=new_y-0.7, s=f"new point, class:
{prediction[0]}")
[Link]()
Result
REMOVE ADS
Example Explained
Import the modules you need.
You can learn about the Matplotlib module in our "Matplotlib Tutorial.
scikit-learn is a popular library for machine learning in Python.
Result:
[(4, 21), (5, 19), (10, 24), (4, 17), (3, 16), (11, 25),
(14, 24), (8, 22), (10, 21), (12, 21)]
Using the input features and target class, we fit a KNN model on the model
using 1 nearest neighbor:
knn = KNeighborsClassifier(n_neighbors=1)
[Link](data, classes)
Then, we can use the same KNN object to predict the class of new,
unforeseen data points. First we create new x and y features, and then
call [Link]() on the new data point to get a class of 0 or 1:
new_x = 8
new_y = 21
new_point = [(new_x, new_y)]
prediction = [Link](new_point)
print(prediction)
Result:
[0]
When we plot all the data along with the new point and class, we can see it's
been labeled blue with the 1 class. The text annotation is just to highlight the
location of the new point:
Result:
knn = KNeighborsClassifier(n_neighbors=5)
[Link](data, classes)
prediction = [Link](new_point)
print(prediction)
Result:
[1]
When we plot the class of the new point along with the older points, we note
that the color has changed based on the associated class label:
Result:
DSA with Python
Data Structures is about how data can be stored in different structures.
Data Structures
Data Structures are a way of storing and organizing data in a computer.
Python has built-in support for several data structures, such as lists,
dictionaries, and sets.
Other data structures can be implemented using Python classes and objects,
such as linked lists, stacks, queues, trees, and graphs.
Algorithms
Algorithms are a way of working with data in a computer and solving
problems like sorting, searching, etc.
In this tutorial we will concentrate on these search and sort Algorithms:
Linear Search
Binary Search
Bubble Sort
Selection Sort
Insertion Sort
Quick Sort
Counting Sort
Radix Sort
Merge Sort
In Python, lists are the built-in data structure that serves as a dynamic
array.
Lists
A list is a built-in data structure in Python, used to store multiple elements.
Lists are used by many algorithms.
Creating Lists
Lists are created using square brackets []:
List Methods
Python lists come with several built-in algorithms (called methods), to
perform common operations like appending, sorting, and more.
Example
Append one element to the list, and sort the list ascending:
# Add element:
[Link](8)
Create Algorithms
Sometimes we want to perform actions that are not built into Python.
Then we can create our own algorithms.
For example, an algorithm can be used to find the lowest value in a list, like
in the example below:
Example
Create an algorithm to find the lowest value in a list:
for i in my_array:
if i < minVal:
minVal = i
The algorithm above is very simple, and fast enough for small data sets, but
if the data is big enough, any algorithm will take time to run.
REMOVE ADS
Time Complexity
When exploring algorithms, we often look at how much time an algorithm
takes to run relative to the size of the data set.
In the example above, the time the algorithm needs to run is proportional, or
linear, to the size of the data set. This is because the algorithm must visit
every array element one time to find the lowest value. The loop must run 5
times since there are 5 values in the array. And if the array had 1000 values,
the loop would have to run 1000 times.
Try the simulation below to see this relationship between the number of
compare operations needed to find the lowest value, and the size of the
array.
See this page for a more thorough explanation of what time complexity is.
Each algorithm in this tutorial will be presented together with its time
complexity.
Like a pile of pancakes, the pancakes are both added and removed from the
top. So when removing a pancake, it will always be the last pancake you
added. This way of organizing elements is called LIFO: Last In First Out.
Stacks are often mentioned together with Queues, which is a similar data
structure described on the next page.
x = [5, 6, 2, 9, 3, 8, 4, 2]
Add: Remove:
Since Python lists has good support for functionality needed to implement
stacks, we start with creating a stack and do stack operations with just a few
lines like this:
stack = []
# Push
[Link]('A')
[Link]('B')
[Link]('C')
print("Stack: ", stack)
# Peek
topElement = stack[-1]
print("Peek: ", topElement)
# Pop
poppedElement = [Link]()
print("Pop: ", poppedElement)
# isEmpty
isEmpty = not bool(stack)
print("isEmpty: ", isEmpty)
# Size
print("Size: ",len(stack))
Example
Creating a stack using class:
class Stack:
def __init__(self):
[Link] = []
def pop(self):
if [Link]():
return "Stack is empty"
return [Link]()
def peek(self):
if [Link]():
return "Stack is empty"
return [Link][-1]
def isEmpty(self):
return len([Link]) == 0
def size(self):
return len([Link])
# Create a stack
myStack = Stack()
[Link]('A')
[Link]('B')
[Link]('C')
Fixed size: An array occupies a fixed part of the memory. This means
that it could take up more memory than needed, or if the array fills up,
it cannot hold more elements.
REMOVE ADS
Stack Implementation using Linked
Lists
A linked list consists of nodes with some sort of data, and a pointer to the
next node.
A big benefit with using linked lists is that nodes are stored wherever there is
free space in memory, the nodes do not have to be stored contiguously right
after each other like elements are stored in arrays. Another nice thing with
linked lists is that when adding or removing nodes, the rest of the nodes in
the list do not have to be shifted.
Example
Creating a Stack using a Linked List:
class Node:
def __init__(self, value):
[Link] = value
[Link] = None
class Stack:
def __init__(self):
[Link] = None
[Link] = 0
def pop(self):
if [Link]():
return "Stack is empty"
popped_node = [Link]
[Link] = [Link]
[Link] -= 1
return popped_node.value
def peek(self):
if [Link]():
return "Stack is empty"
return [Link]
def isEmpty(self):
return [Link] == 0
def stackSize(self):
return [Link]
def traverseAndPrint(self):
currentNode = [Link]
while currentNode:
print([Link], end=" -> ")
currentNode = [Link]
print()
myStack = Stack()
[Link]('A')
[Link]('B')
[Link]('C')
Dynamic size: The stack can grow and shrink dynamically, unlike with
arrays.
Extra memory: Each stack element must contain the address to the
next element (the next linked list node).
Readability: The code might be harder to read and write for some
because it is longer and more complex.
Queues
Think of a queue as people standing in line in a supermarket.
The first person to stand in line is also the first who can pay and leave the
supermarket.
Queues are often mentioned together with Stacks, which is a similar data
structure described on the previous page.
x = [5, 6, 2, 9, 3, 8, 4, 2]
Add: Remove:
Since Python lists has good support for functionality needed to implement
queues, we start with creating a queue and do queue operations with just a
few lines:
queue = []
# Enqueue
[Link]('A')
[Link]('B')
[Link]('C')
print("Queue: ", queue)
# Peek
frontElement = queue[0]
print("Peek: ", frontElement)
# Dequeue
poppedElement = [Link](0)
print("Dequeue: ", poppedElement)
# Size
print("Size: ", len(queue))
Note: While using a list is simple, removing elements from the beginning
(dequeue operation) requires shifting all remaining elements, making it less
efficient for large queues.
REMOVE ADS
Example
Using a Python class as a queue:
class Queue:
def __init__(self):
[Link] = []
def dequeue(self):
if [Link]():
return "Queue is empty"
return [Link](0)
def peek(self):
if [Link]():
return "Queue is empty"
return [Link][0]
def isEmpty(self):
return len([Link]) == 0
def size(self):
return len([Link])
# Create a queue
myQueue = Queue()
[Link]('A')
[Link]('B')
[Link]('C')
A big benefit with using linked lists is that nodes are stored wherever there is
free space in memory, the nodes do not have to be stored contiguously right
after each other like elements are stored in arrays. Another nice thing with
linked lists is that when adding or removing nodes, the rest of the nodes in
the list do not have to be shifted.
Example
Creating a Queue using a Linked List:
class Node:
def __init__(self, data):
[Link] = data
[Link] = None
class Queue:
def __init__(self):
[Link] = None
[Link] = None
[Link] = 0
def dequeue(self):
if [Link]():
return "Queue is empty"
temp = [Link]
[Link] = [Link]
[Link] -= 1
if [Link] is None:
[Link] = None
return [Link]
def peek(self):
if [Link]():
return "Queue is empty"
return [Link]
def isEmpty(self):
return [Link] == 0
def size(self):
return [Link]
def printQueue(self):
temp = [Link]
while temp:
print([Link], end=" -> ")
temp = [Link]
print()
# Create a queue
myQueue = Queue()
[Link]('A')
[Link]('B')
[Link]('C')
Dynamic size: The queue can grow and shrink dynamically, unlike
with arrays.
No shifting: The front element of the queue can be removed
(enqueue) without having to shift other elements in the memory.
Extra memory: Each queue element must contain the address to the
next element (the next linked list node).
Readability: The code might be harder to read and write for some
because it is longer and more complex.
REMOVE ADS
Linked Lists
A linked list consists of nodes with some sort of data, and a pointer, or link,
to the next node.
Nodes in a linked list store links to other nodes, but array elements do not
need to store links to other elements.
Note: How linked lists and arrays are stored in memory is explained in detail
on the page Linked Lists in Memory.
The table below compares linked lists with arrays to give a better
understanding of what linked lists are.
Arrays
Elements, or nodes, are stored right after each other in memory Yes
(contiguously)
Linked lists are not allocated to a fixed size in memory like arrays are,
so linked lists do not require to move the whole list into a larger
memory space when the fixed memory space fills up, like arrays must.
Linked list nodes are not laid out one right after the other in memory
(contiguously), so linked list nodes do not have to be shifted up or
down in memory when nodes are inserted or deleted.
Linked list nodes require more memory to store one or more links to
other nodes. Array elements do not require that much memory,
because array elements do not contain links to other elements.
Linked list operations are usually harder to program and require more
lines than similar array operations, because programming languages
have better built in support for arrays.
We must traverse a linked list to find a node at a specific position, but
with arrays we can access an element directly by writing myArray[5].
REMOVE ADS
Types of Linked Lists
There are three basic forms of linked lists:
A singly linked list is the simplest kind of linked lists. It takes up less space
in memory because each node has only one address to the next node, like in
the image below.
A doubly linked list has nodes with addresses to both the previous and the
next node, like in the image below, and therefore takes up more memory.
But doubly linked lists are good if you want to be able to move both up and
down in the list.
A circular linked list is like a singly or doubly linked list with the first node,
the "head", and the last node, the "tail", connected.
In singly or doubly linked lists, we can find the start and end of a list by just
checking if the links are null. But for circular linked lists, more complex code
is needed to explicitly check for start and end nodes in certain applications.
Circular linked lists are good for lists you need to cycle through continuously.
Note: What kind of linked list you need depends on the problem you are
trying to solve.
1. Traversal
2. Remove a node
3. Insert a node
4. Sort
For simplicity, singly linked lists will be used to explain these operations
below.
REMOVE ADS
Traversal of linked lists is typically done to search for a specific node, and
read or modify the node's content, remove the node, or insert a node right
before or after that node.
To traverse a singly linked list, we start with the first node in the list, the
head node, and follow that node's next link, and the next node's next link
and so on, until the next address is null.
The code below prints out the node values as it traverses along the linked
list, in the same way as the animation above.
class Node:
def __init__(self, data):
[Link] = data
[Link] = None
def traverseAndPrint(head):
currentNode = head
while currentNode:
print([Link], end=" -> ")
currentNode = [Link]
print("null")
node1 = Node(7)
node2 = Node(11)
node3 = Node(3)
node4 = Node(2)
node5 = Node(9)
[Link] = node2
[Link] = node3
[Link] = node4
[Link] = node5
traverseAndPrint(node1)
Finding the lowest value in a linked list is very similar to how we found the
lowest value in an array, except that we need to follow the next link to get to
the next node.
To find the lowest value we need to traverse the list like in the previous
code. But in addition to traversing the list, we must also update the current
lowest value when we find a node with a lower value.
In the code below, the algorithm to find the lowest value is moved into a
function called findLowestValue.
Example
Finding the lowest value in a singly linked list in Python:
class Node:
def __init__(self, data):
[Link] = data
[Link] = None
def findLowestValue(head):
minValue = [Link]
currentNode = [Link]
while currentNode:
if [Link] < minValue:
minValue = [Link]
currentNode = [Link]
return minValue
node1 = Node(7)
node2 = Node(11)
node3 = Node(3)
node4 = Node(2)
node5 = Node(9)
[Link] = node2
[Link] = node3
[Link] = node4
[Link] = node5
So before deleting the node, we need to get the next pointer from the
previous node, and connect the previous node to the new next node before
deleting the node in between.
Also, it is a good idea to first connect next pointer to the node after the node
we want to delete, before we delete it. This is to avoid a 'dangling' pointer, a
pointer that points to nothing, even if it is just for a brief moment.
The simulation below shows the node we want to delete, and how the list
must be traversed first to connect the list properly before deleting the node
without breaking the linked list.
Head7next11next3next2next9nextnull
Delete
In the code below, the algorithm to delete a node is moved into a function
called deleteSpecificNode.
Example
Deleting a specific node in a singly linked list in Python:
class Node:
def __init__(self, data):
[Link] = data
[Link] = None
def traverseAndPrint(head):
currentNode = head
while currentNode:
print([Link], end=" -> ")
currentNode = [Link]
print("null")
currentNode = head
while [Link] and [Link] != nodeToDelete:
currentNode = [Link]
if [Link] is None:
return head
[Link] = [Link]
return head
node1 = Node(7)
node2 = Node(11)
node3 = Node(3)
node4 = Node(2)
node5 = Node(9)
[Link] = node2
[Link] = node3
[Link] = node4
[Link] = node5
print("Before deletion:")
traverseAndPrint(node1)
# Delete node4
node1 = deleteSpecificNode(node1, node4)
print("\nAfter deletion:")
traverseAndPrint(node1)
In the deleteSpecificNode function above, the return value is the new head
of the linked list. So for example, if the node to be deleted is the first node,
the new head returned will be the next node.
To insert a node in a linked list we first need to create the node, and then at
the position where we insert it, we need to adjust the pointers so that the
previous node points to the new node, and the new node points to the
correct next node.
The simulation below shows how the links are adjusted when inserting a new
node.
Head7next97next3next2next9nextnullInsert
1. New node is created
2. Node 1 is linked to new node
3. New node is linked to next node
Example
Inserting a node in a singly linked list in Python:
class Node:
def __init__(self, data):
[Link] = data
[Link] = None
def traverseAndPrint(head):
currentNode = head
while currentNode:
print([Link], end=" -> ")
currentNode = [Link]
print("null")
currentNode = head
for _ in range(position - 2):
if [Link] is None:
break
currentNode = [Link]
[Link] = [Link]
[Link] = newNode
return head
node1 = Node(7)
node2 = Node(3)
node3 = Node(2)
node4 = Node(9)
[Link] = node2
[Link] = node3
[Link] = node4
print("Original list:")
traverseAndPrint(node1)
print("\nAfter insertion:")
traverseAndPrint(node1)
In the insertNodeAtPosition function above, the return value is the new
head of the linked list. So for example, if the node is inserted at the start of
the linked list, the new head returned will be the new node.
Remember that time complexity just says something about the approximate
number of operations needed by the algorithm based on a large set of
data (n), and does not tell us the exact time a specific implementation of an
algorithm takes.
This means that even though linear search is said to have the same time
complexity for arrays as for linked list: O(n), it does not mean they take the
same amount of time. The exact time it takes for an algorithm to run
depends on programming language, computer hardware, differences in time
needed for operations on arrays vs linked lists, and many other things as
well.
Linear search for linked lists works the same as for arrays. A list of unsorted
values are traversed from the head node until the node with the specific
value is found. Time complexity is O(n).
Binary search is not possible for linked lists because the algorithm is based
on jumping directly to different array elements, and that is not possible with
linked lists.
Sorting algorithms have the same time complexities as for arrays, and these
are explained earlier in this tutorial. But remember, sorting algorithms that
are based on directly accessing an array element based on an index, do not
work on linked lists.
The reason Hash Tables are sometimes preferred instead of arrays or linked
lists is because searching for, adding, and deleting data can be done really
quickly, even for large amounts of data.
In a Linked List, finding a person "Bob" takes time because we would have to
go from one node to the next, checking each node, until the node with "Bob"
is found.
And finding "Bob" in an list/array could be fast if we knew the index, but
when we only know the name "Bob", we need to compare each element and
that takes time.
With a Hash Table however, finding "Bob" is done really fast because there is
a way to go directly to where "Bob" is stored, using something called a hash
function.
my_list = [None, None, None, None, None, None, None, None, None,
None]
Each of these elements is called a bucket in a Hash Table.
We want to store a name directly into its right place in the array, and this is
where the hash function comes in.
def hash_function(value):
sum_of_chars = 0
for char in value:
sum_of_chars += ord(char)
return sum_of_chars % 10
The character B has Unicode number 66, o has 111, and b has 98. Adding those
together we get 275. Modulo 10 of 275 is 5, so "Bob" should be stored at
index 5.
The number returned by the hash function is called the hash code.
See this page for more information about how characters are represented as
numbers.
Modulo: A modulo operation divides a number with another number, and
gives us the resulting remainder. So for example, 7 % 3 will give us the
remainder 1. (Dividing 7 apples between 3 people, means that each person
gets 2 apples, with 1 apple to spare.)
REMOVE ADS
Example
def add(name):
index = hash_function(name)
my_list[index] = name
add('Bob')
print(my_list)
After storing "Bob" at index 5, our array now looks like this:
my_list = [None, None, None, None, None, 'Bob', None, None, None,
None]
We can use the same functions to store "Pete", "Jones", "Lisa", and "Siri" as
well.
Example
add('Pete')
add('Jones')
add('Lisa')
add('Siri')
print(my_list)
After using the hash function to store those names in the correct position,
our array looks like this:
Example
my_list = [None, 'Jones', None, 'Lisa', None, 'Bob',
None, 'Siri', 'Pete', None]
To find "Pete" in the Hash Table, we give the name "Pete" to our hash
function. The hash function returns 8, meaning that "Pete" is stored at index
8.
Example
def contains(name):
index = hash_function(name)
return my_list[index] == name
Start by creating a new list with the same size as the original list, but with
empty buckets:
my_list = [
[],
[],
[],
[],
[],
[],
[],
[],
[],
[]
]
Rewrite the add() function, and add the same names as before:
Example
def add(name):
index = hash_function(name)
my_list[index].append(name)
add('Bob')
add('Pete')
add('Jones')
add('Lisa')
add('Siri')
add('Stuart')
print(my_list)
After implementing each bucket as a list, "Stuart" can also be stored at index
3, and our Hash Set now looks like this:
Result
my_list = [
[None],
['Jones'],
[None],
['Lisa', 'Stuart'],
[None],
['Bob'],
[None],
['Siri'],
['Pete'],
[None]
]
Searching for "Stuart" now takes a little bit longer time, because we also find
"Lisa" in the same bucket, but still much faster than searching the entire
Hash Table.
The most important reason why Hash Tables are great for these things is
that Hash Tables are very fast compared Arrays and Linked Lists, especially
for large sets. Arrays and Linked Lists have time complexity O(n) for search
and delete, while Hash Tables have just O(1) on average.
The hash code says what bucket the element belongs to, so now we can go
directly to that Hash Table element: to modify it, or to delete it, or just to
check if it exists.
A collision happens when two Hash Table elements have the same hash
code, because that means they belong to the same bucket.
Collision can be solved by Chaining by using lists to allow more than one
element in the same bucket.
Python Trees
Trees
The Tree data structure is similar to Linked Lists in that each node contains
data and can be linked to other nodes.
We have previously covered data structures like Arrays, Linked Lists, Stacks,
and Queues. These are all linear structures, which means that each element
follows directly after another in a sequence. Trees however, are different. In
a Tree, a single element can have multiple 'next' elements, allowing the data
structure to branch out in various directions.
The data structure is called a "tree" because it looks like a tree's structure.
RABCDEFGHI
REMOVE ADS
Types of Trees
Trees are a fundamental data structure in computer science, used to
represent hierarchical relationships. This tutorial covers several key types of
trees.
Binary Trees: Each node has up to two children, the left child node and the
right child node. This structure is the foundation for more complex tree types
like Binay Search Trees and AVL Trees.
Binary Search Trees (BSTs): A type of Binary Tree where for each node,
the left child node has a lower value, and the right child node has a higher
value.
AVL Trees: A type of Binary Search Tree that self-balances so that for every
node, the difference in height between the left and right subtrees is at most
one. This balance is maintained through rotations when nodes are inserted
or deleted.
Each of these data structures are described in detail on the next pages,
including animations and how to implement them.
Arrays are fast when you want to access an element directly, like
element number 700 in an array of 1000 elements for example. But
inserting and deleting elements require other elements to shift in
memory to make place for the new element, or to take the deleted
elements place, and that is time consuming.
Linked Lists are fast when inserting or deleting nodes, no memory
shifting needed, but to access an element inside the list, the list must
be traversed, and that takes time.
Trees, such as Binary Trees, Binary Search Trees and AVL Trees, are
great compared to Arrays and Linked Lists because they are BOTH fast
at accessing a node, AND fast when it comes to deleting or inserting a
node, with no shifts in memory needed.
Binary Trees
A Binary Tree is a type of tree data structure where each node can have a
maximum of two child nodes, a left child node and a right child node.
This restriction, that a node can have a maximum of two child nodes, gives
us many benefits:
The Binary Tree above can be implemented much like a Linked List, except
that instead of linking each node to one next node, we create a structure
where each node can be linked to both its left and right child nodes.
class TreeNode:
def __init__(self, data):
[Link] = data
[Link] = None
[Link] = None
root = TreeNode('R')
nodeA = TreeNode('A')
nodeB = TreeNode('B')
nodeC = TreeNode('C')
nodeD = TreeNode('D')
nodeE = TreeNode('E')
nodeF = TreeNode('F')
nodeG = TreeNode('G')
[Link] = nodeA
[Link] = nodeB
[Link] = nodeC
[Link] = nodeD
[Link] = nodeE
[Link] = nodeF
[Link] = nodeG
# Test
print("[Link]:", [Link])
REMOVE ADS
The different kinds of Binary Trees are also worth mentioning now as these
words and concepts will be used later in the tutorial.
Below are short explanations of different types of Binary Tree structures, and
below the explanations are drawings of these kinds of structures to make it
as easy to understand as possible.
A balanced Binary Tree has at most 1 in difference between its left and right
subtree heights, for each node in the tree.
A complete Binary Tree has all levels full of nodes, except the last level,
which is can also be full, or filled from left to right. The properties of a
complete Binary Tree means it is also balanced.
A full Binary Tree is a kind of tree where each node has either 0 or 2 child
nodes.
A perfect Binary Tree has all leaf nodes on the same level, which means
that all levels are full of nodes, and all internal nodes have two child
[Link] properties of a perfect Binary Tree means it is also full, balanced,
and complete.
1171539131918Balanced11715391319248Complete and
balanced1171513191214Full11715313199Perfect, full, balanced and
complete
Since Arrays and Linked Lists are linear data structures, there is only one
obvious way to traverse these: start at the first element, or node, and
continue to visit the next until you have visited them all.
But since a Tree can branch out in different directions (non-linear), there are
different ways of traversing Trees.
Breadth First Search (BFS) is when the nodes on the same level are
visited before going to the next level in the tree. This means that the tree is
explored in a more sideways direction.
Depth First Search (DFS) is when the traversal moves down the tree all
the way to the leaf nodes, exploring the tree branch by branch in a
downwards direction.
Pre-order Traversal is done by visiting the root node first, then recursively do
a pre-order traversal of the left subtree, followed by a recursive pre-order
traversal of the right subtree. It's used for creating a copy of the tree, prefix
notation of an expression tree, etc.
This traversal is "pre" order because the node is visited "before" the
recursive pre-order traversal of the left and right subtrees.
Example
A pre-order traversal:
def preOrderTraversal(node):
if node is None:
return
print([Link], end=", ")
preOrderTraversal([Link])
preOrderTraversal([Link])
The first time the argument node is None is when the left child of node C is
given as an argument (C has no left child).
After None is returned the first time when calling C's left child, C's right child
also returns None, and then the recursive calls continue to propagate back so
that A's right child D is the next to be printed.
The code continues to propagate back so that the rest of the nodes in R's
right subtree gets printed.
What makes this traversal "in" order, is that the node is visited in between
the recursive function calls. The node is visited after the In-order Traversal of
the left subtree, and before the In-order Traversal of the right subtree.
Example
Create an In-order Traversal:
def inOrderTraversal(node):
if node is None:
return
inOrderTraversal([Link])
print([Link], end=", ")
inOrderTraversal([Link])
The inOrderTraversal() function keeps calling itself with the current left child
node as an argument (line 4) until that argument is None and the function
returns (line 2-3).
The first time the argument node is None is when the left child of node C is
given as an argument (C has no left child).
After that, the data part of node C is printed (line 5), which means that 'C' is
the first thing that gets printed.
Then, node C's right child is given as an argument (line 6), which is None, so
the function call returns without doing anything else.
What makes this traversal "post" is that visiting a node is done "after" the
left and right child nodes are called recursively.
Example
Post-order Traversal:
def postOrderTraversal(node):
if node is None:
return
postOrderTraversal([Link])
postOrderTraversal([Link])
print([Link], end=", ")
After C's left child node returns None, line 5 runs and C's right child node
returns None, and then the letter 'C' is printed (line 6).
This means that C is visited, or printed, "after" its left and right child nodes
are traversed, that is why it is called "post" order traversal.
The function continues to propagate back and printing nodes until all nodes
are printed, or visited.
A Binary Search Tree is a Binary Tree where every node's left child has a
lower value, and every node's right child has a higher value.
A clear advantage with Binary Search Trees is that operations like search,
delete, and insert are fast and done without having to shift values in
memory.
The X node's left child and all of its descendants (children, children's
children, and so on) have lower values than X's value.
The right child, and all its descendants have higher values than X's
value.
Left and right subtrees must also be Binary Search Trees.
These properties makes it faster to search, add and delete values than a
regular binary tree.
A subtree starts with one of the nodes in the tree as a local root, and
consists of that node and all its descendants.
The descendants of a node are all the child nodes of that node, and all their
child nodes, and so on. Just start with a node, and the descendants will be all
nodes that are connected below that node.
The node's height is the maximum number of edges between that node
and a leaf node.
The code below is an implementation of the Binary Search Tree in the figure
above, with traversal.
class TreeNode:
def __init__(self, data):
[Link] = data
[Link] = None
[Link] = None
def inOrderTraversal(node):
if node is None:
return
inOrderTraversal([Link])
print([Link], end=", ")
inOrderTraversal([Link])
root = TreeNode(13)
node7 = TreeNode(7)
node15 = TreeNode(15)
node3 = TreeNode(3)
node8 = TreeNode(8)
node14 = TreeNode(14)
node19 = TreeNode(19)
node18 = TreeNode(18)
[Link] = node7
[Link] = node15
[Link] = node3
[Link] = node8
[Link] = node14
[Link] = node19
[Link] = node18
# Traverse
inOrderTraversal(root)
As we can see by running the code example above, the in-order traversal
produces a list of numbers in an increasing (ascending) order, which means
that this Binary Tree is a Binary Search Tree.
REMOVE ADS
For Binary Search to work, the array must be sorted already, and searching
for a value in an array can then be done really fast.
Similarly, searching for a value in a BST can also be done really fast because
of how the nodes are placed.
How it works:
1. Start at the root node.
2. If this is the value we are looking for, return.
3. If the value we are looking for is higher, continue searching in the right
subtree.
4. If the value we are looking for is lower, continue searching in the left
subtree.
5. If the subtree we want to search does not exist, depending on the
programming language, return None, or NULL, or something similar, to
indicate that the value is not inside the BST.
Example
Search the Tree for the value "13"
The time complexity for searching a BST for a value is O(h), where h is the
height of the tree.
For a BST with most nodes on the right side for example, the height of the
tree becomes larger than it needs to be, and the worst case search will take
longer. Such trees are called unbalanced.
Both Binary Search Trees above have the same nodes, and in-order traversal
of both trees gives us the same result but the height is very different. It
takes longer time to search the unbalanced tree above because it is higher.
We will use the next page to describe a type of Binary Tree called AVL Trees.
AVL trees are self-balancing, which means that the height of the tree is kept
to a minimum so that operations like search, insertion and deletion take less
time.
How it works:
Inserting nodes as described above means that an inserted node will always
become a new leaf node.
All nodes in the BST are unique, so in case we find the same value as the one
we want to insert, we do nothing.
Example
Inserting a node in a BST:
How it works:
This is how a function for finding the lowest value in the subtree of a BST
node looks like:
Example
Find the lowest value in a BST subtree
def minValueNode(node):
current = node
while [Link] is not None:
current = [Link]
return current
# Find Lowest
print("\nLowest value:",minValueNode(root).data)
We will use this minValueNode() function in the section below, to find a node's
in-order successor, and use that to delete a node.
How it works:
In step 3 above, the successor we find will always be a leaf node, and
because it is the node that comes right after the node we want to delete, we
can swap values with it and delete it.
This is how a BST can be implemented with functionality for deleting a node:
Example
Delete a Node in a BST
return node
# Delete node 15
delete(root,15)
Line 1: The node argument here makes it possible for the function to call
itself recursively on smaller and smaller subtrees in the search for the node
with the data we want to delete.
Line 2-8: This is searching for the node with correct data that we want to
delete.
Line 9-22: The node we want to delete has been found. There are three
such cases:
1. Case 1: Node with no child nodes (leaf node). None is returned, and
that becomes the parent node's new left or right value by recursion
(line 6 or 8).
2. Case 2: Node with either left or right child node. That left or right child
node becomes the parent's new left or right child through recursion
(line 7 or 9).
3. Case 3: Node has both left and right child nodes. The in-order
successor is found using the minValueNode() function. We keep the
successor's value by setting it as the value of the node we want to
delete, and then we can delete the successor node.
Searching a BST is just as fast as Binary Search on an array, with the same
time complexity O(log n).
And deleting and inserting new values can be done without shifting elements
in memory, just like with Linked Lists.
The reason why we wrote that searching for a value is O(log n) in the table
above is because that is true if the tree is "balanced", like in the image
below.
1371538141918Balanced BST
We call this tree balanced because there are approximately the same
number of nodes on the left and right side of the tree.
The exact way to tell that a Binary Tree is balanced is that the height of the
left and right subtrees of any node only differs by one. In the image above,
the left subtree of the root node has height h=2, and the right subtree has
height h=3.
For a balanced BST, with a large number of nodes (big n), we get height h ≈ \
log_2 n, and therefore the time complexity for searching, deleting, or
inserting a node can be written as O(h) = O(\log n).
But, in case the BST is completely unbalanced, like in the image below, the
height of the tree is approximately the same as the number of nodes, h ≈ n,
and we get time complexity O(h) = O(n) for searching, deleting, or inserting a
node.
7133158191418Unbalanced BST
And keeping a Binary Search Tree balanced is exactly what AVL Trees do,
which is the data structure explained on the next page.
AVL trees are self-balancing, which means that the tree height is kept to a
minimum so that a very fast runtime is guaranteed for searching, inserting
and deleting nodes, with time complexity O(logn).
AVL Trees
The only difference between a regular Binary Search Tree and an AVL Tree is
that AVL Trees do rotation operations in addition, to keep the tree balance.
A Binary Search Tree is in balance when the difference in height between left
and right subtrees is less than 2.
By keeping balance, the AVL Tree ensures a minimum tree height, which
means that search, insert, and delete operations can be done really fast.
The two trees above are both Binary Search Trees, they have the same
nodes, and the same in-order traversal (alphabetical), but the height is very
different because the AVL Tree has balanced itself.
Step through the building of an AVL Tree in the animation below to see how
the balance factors are updated, and how rotation operations are done when
required to restore the balance.
0C0F0G0D0B0AInsert C
Continue reading to learn more about how the balance factor is calculated,
how rotation operations are done, and how AVL Trees can be implemented.
REMOVE ADS
The previous animation shows one specific left rotation, and one specific
right rotation.
But in general, left and right rotations are done like in the animation below.
XYRotate Right
Notice how the subtree changes its parent. Subtrees change parent in this
way during rotation to maintain the correct in-order traversal, and to
maintain the BST property that the left child is less than the right child, for all
nodes in the tree.
Also keep in mind that it is not always the root node that become
unbalanced and need rotation.
The height of a subtree is the number of edges between the root node of the
subtree and the leaf node farthest down in that subtree.
The Balance Factor (BF) for a node (X) is the difference in height between
its right and left subtrees.
BF(X)=height(rightSubtree(X))−height(leftSubtree(X))
Balance factor values
If the balance factor is less than -1, or more than 1, for one or more nodes in
the tree, the tree is considered not in balance, and a rotation operation is
needed to restore balance.
Let's take a closer look at the different rotation operations that an AVL Tree
can do to regain balance.
There are four different ways an AVL Tree can be out of balance, and each of
these cases require a different rotation operation.
Left-Left The unbalanced node and its left child node A single right rotation.
(LL) are both left-heavy.
Right-Right The unbalanced node and its right child node A single left rotation.
(RR) are both right-heavy.
Left-Right The unbalanced node is left heavy, and its left First do a left rotation on
(LR) child node is right heavy. rotation on the unbalance
Right-Left The unbalanced node is right heavy, and its First do a right rotation on
(RL) right child node is left heavy. left rotation on the unbala
When this LL case happens, a single right rotation on the unbalanced node is
enough to restore balance.
Step through the animation below to see the LL case, and how the balance is
restored by a single right rotation.
-1Q0P0D0L0C0B0K0AInsert D
+1A0B0D0C0E0FInsert D
In this LR case, a left rotation is first done on the left child node, and then a
right rotation is done on the original unbalanced node.
Step through the animation below to see how the Left-Right case can
happen, and how the rotation operations are done to restore balance.
-1Q0E0K0C0F0GInsert K
As you are building the AVL Tree in the animation above, the Left-Right case
happens 2 times, and rotation operations are required and done to restore
balance:
In this case we first do a right rotation on the unbalanced node's right child,
and then we do a left rotation on the unbalanced node itself.
Step through the animation below to see how the Right-Left case can occur,
and how rotations are done to restore the balance.
+1A0F0B0G0E0DInsert B
The next Right-Left case occurs after nodes G, E, and D are added. This is a
Right-Left case because B is unbalanced and right heavy, and its right child F
is left heavy. To restore balance, a right rotation is first done on node F, and
then a left rotation is done on node B.
In the simulation below, after inserting node F, the nodes C, E and H are all
unbalanced, but since retracing works through recursion, the unbalance at
node H is discovered and fixed first, which in this case also fixes the
unbalance in nodes E and C.
0A-1B+1C0D+1E0G-1H0FInsert F
After node F is inserted, the code will retrace, calculating balancing factors
as it propagates back up towards the root node. When node H is reached and
the balancing factor -2 is calculated, a right rotation is done. Only after this
rotation is done, the code will continue to retrace, calculating balancing
factors further up on ancestor nodes E and C.
Because of the rotation, balancing factors for nodes E and C stay the same
as before node F was inserted.
There is only one new attribute for each node in the AVL tree compared to
the BST, and that is the height, but there are many new functions and extra
code lines needed for the AVL Tree implementation because of how the AVL
Tree rebalances itself.
class TreeNode:
def __init__(self, data):
[Link] = data
[Link] = None
[Link] = None
[Link] = 1
def getHeight(node):
if not node:
return 0
return [Link]
def getBalance(node):
if not node:
return 0
return getHeight([Link]) - getHeight([Link])
def rightRotate(y):
print('Rotate right on node',[Link])
x = [Link]
T2 = [Link]
[Link] = y
[Link] = T2
[Link] = 1 + max(getHeight([Link]), getHeight([Link]))
[Link] = 1 + max(getHeight([Link]), getHeight([Link]))
return x
def leftRotate(x):
print('Rotate left on node',[Link])
y = [Link]
T2 = [Link]
[Link] = x
[Link] = T2
[Link] = 1 + max(getHeight([Link]), getHeight([Link]))
[Link] = 1 + max(getHeight([Link]), getHeight([Link]))
return y
# Left Right
if balance > 1 and getBalance([Link]) < 0:
[Link] = leftRotate([Link])
return rightRotate(node)
# Right Right
if balance < -1 and getBalance([Link]) <= 0:
return leftRotate(node)
# Right Left
if balance < -1 and getBalance([Link]) > 0:
[Link] = rightRotate([Link])
return leftRotate(node)
return node
def inOrderTraversal(node):
if node is None:
return
inOrderTraversal([Link])
print([Link], end=", ")
inOrderTraversal([Link])
# Inserting nodes
root = None
letters = ['C', 'B', 'E', 'A', 'D', 'H', 'G', 'F']
for letter in letters:
root = insert(root, letter)
inOrderTraversal(root)
Example
Delete Node:
def minValueNode(node):
current = node
while [Link] is not None:
current = [Link]
return current
temp = minValueNode([Link])
[Link] = [Link]
[Link] = delete([Link], [Link])
return node
def inOrderTraversal(node):
if node is None:
return
inOrderTraversal([Link])
print([Link], end=", ")
inOrderTraversal([Link])
# Inserting nodes
root = None
letters = ['C', 'B', 'E', 'A', 'D', 'H', 'G', 'F']
for letter in letters:
root = insert(root, letter)
inOrderTraversal(root)
So in worst case, algorithms like search, insert, and delete must run through
the whole height of the tree. This means that keeping the height (h) of the
tree low, like we do using AVL Trees, gives us a lower runtime.
O(logn) Explained
The fact that the time complexity is O(h)=O(logn) for search, insert, and
delete on an AVL Tree with height h and nodes n can be explained like this:
Imagine a perfect Binary Tree where all nodes have two child nodes except
on the lowest level, like the AVL Tree below.
HDBFEGACLJNMOIK
The number of nodes on each level in such an AVL Tree are:
1,2,4,8,16,32,..
20,21,22,23,24,25,..
To get the number of nodes n in a perfect Binary Tree with height h=3, we
can add the number of nodes on each level together:
n3=20+21+22+23=15
n3=24−1=15
And this is actually the case for larger trees as well! If we want to get the
number of nodes n in a tree with height h=5 for example, we find the
number of nodes like this:
n5=26−1=63
So in general, the relationship between the height h of a perfect Binary Tree
and the number of nodes in it n, can be expressed like this:
nh=2h+1−1
Note: The formula above can also be found by calculating the sum of the
geometric series 20+21+22+23+...+2n
We know that the time complexity for searching, deleting, or inserting a
node in an AVL tree is O(h), but we want to argue that the time complexity
is actually O(log(n)), so we need to find the height h described by the
number of nodes n:
n=2h+1−1n+1=2h+1log2(n+1)=log2(2h+1)h=log2(n+1)−1O(h)=O(log
n)
How the last line above is derived might not be obvious, but for a Binary Tree
with a lot of nodes (big n), the "+1" and "-1" terms are not important when
we consider time complexity. For more details on how to calculate the time
complexity using Big O notation, see this page.
The math above shows that the time complexity for search, delete, and
insert operations on an AVL Tree O(h), can actually be expressed
as O(logn), which is fast, a lot faster than the time complexity for BSTs
which is O(n).
Python Graphs
Graphs
A Graph is a non-linear data structure that consists of vertices (nodes) and
edges.
F24BCAEDG
A vertex, also called a node, is a point or an object in the Graph, and an edge
is used to connect two vertices with each other.
Graphs are non-linear because the data structure allows us to have different
paths to get from one vertex to another, unlike with linear data structures
like Arrays or Linked Lists.
Graphs are used to represent and solve problems where the data consists of
objects and relationships between them, such as:
REMOVE ADS
Graph Representations
A Graph representation tells us how a Graph is stored in memory.
In the adjacency matrix above, the value 3 on index (0,1) tells us there is an
edge from vertex A to vertex B, and the weight for that edge is 3.
As you can see, the weights are placed directly into the adjacency matrix for
the correct edge, and for a directed Graph, the adjacency matrix does not
have to be symmetric.
A 'sparse' Graph is a Graph where each vertex only has edges to a small
portion of the other vertices in the Graph.
An Adjacency List has an array that contains all the vertices in the Graph,
and each vertex has a Linked List (or Array) with the vertex's edges.
In the adjacency list above, the vertices A to D are placed in an Array, and
each vertex in the array has its index written right next to it.
Each vertex in the Array has a pointer to a Linked List that represents that
vertex's edges. More specifically, the Linked List contains the indexes to the
adjacent (neighbor) vertices.
So for example, vertex A has a link to a Linked List with values 3, 1, and 2.
These values are the indexes to A's adjacent vertices D, B, and C.
An Adjacency List can also represent a directed and weighted Graph, like
this:
Node D for example, has a pointer to a Linked List with an edge to vertex A.
The values 0,4 means that vertex D has an edge to vertex on index 0 (vertex
A), and the weight of that edge is 4.
Linear Search
Linear search (or sequential search) is the simplest search algorithm. It
checks each element one by one.
10
11
12
13
14
15
16
17
18
19
20
21
Run the simulation above to see how the Linear Search algorithm works.
How it works:
REMOVE ADS
if 4 in mylist:
print("Found!")
else:
print("Not found!")
But if you need to find the index of a value, you will need to implement a
linear search:
Example
Find the index of a value in a list:
mylist = [3, 7, 2, 9, 5, 1, 8, 4, 6]
x = 4
result = linearSearch(mylist, x)
if result != -1:
print("Found at index", result)
else:
print("Not found")
Binary Search
The Binary Search algorithm searches through a sorted array and returns
the index of the value it searches for.
10
11
12
13
14
15
16
17
18
19
20
Run the simulation to see how the Binary Search algorithm works.
Binary Search is much faster than Linear Search, but requires a sorted array
to work.
The Binary Search algorithm works by checking the value in the center of the
array. If the target value is lower, the next value to check is in the center of
the left half of the array. This way of searching means that the search area is
always half of the previous search area, and this is why the Binary Search
algorithm is so fast.
This process of halving the search area happens until the target value is
found, or until the search area of the array is empty.
How it works:
Step 2: The value in the middle of the array at index 3, is it equal to 11?
Step 3: 7 is less than 11, so we must search for 11 to the right of index 3.
The values to the right of index 3 are [ 11, 15, 25]. The next value to check is
the middle value 15, at index 5.
if arr[mid] == targetVal:
return mid
return -1
result = binarySearch(mylist, x)
if result != -1:
print("Found at index", result)
else:
print("Not found")
This means that even in the worst case scenario where Binary Search cannot
find the target value, it still only needs log2n comparisons to look through a
sorted array of n values.
Time complexity for Binary Search is: O(log2n)
Note: When writing time complexity using Big O notation we could also just
have written O(logn), but O(log2n) reminds us that the array search area is
halved for every new comparison, which is the basic concept of Binary
Search, so we will just keep the base 2 indication in this case.
If we draw how much time Binary Search needs to find a value in an array
of n values, compared to Linear Search, we get this graph:
Bubble Sort
Bubble Sort is an algorithm that sorts an array from the lowest value to the
highest value.
Sort
Run the simulation to see how it looks like when the Bubble Sort algorithm
sorts an array of values. Each value in the array is represented by a column.
The word 'Bubble' comes from how this algorithm works, it makes the
highest values 'bubble up'.
How it works:
REMOVE ADS
Step 2: We look at the two first values. Does the lowest value come first?
Yes, so we don't need to swap them.
Repeat until no more swaps are needed and you will get a sorted array:
Bubble Sort
[
7,
12,
9,
11,
3
]
n = len(mylist)
for i in range(n-1):
for j in range(n-i-1):
if mylist[j] > mylist[j+1]:
mylist[j], mylist[j+1] = mylist[j+1], mylist[j]
print(mylist)
Imagine that the array is almost sorted already, with the lowest numbers at
the start, like this for example:
If the algorithm goes through the array one time without swapping any
values, the array must be finished sorted, and we can stop the algorithm,
like this:
Example
Improved Bubble Sort algorithm:
n = len(mylist)
for i in range(n-1):
swapped = False
for j in range(n-i-1):
if mylist[j] > mylist[j+1]:
mylist[j], mylist[j+1] = mylist[j+1], mylist[j]
swapped = True
if not swapped:
break
print(mylist)
REMOVE ADS
The graph describing the Bubble Sort time complexity looks like this:
As you can see, the run time increases really fast when the size of the array
is increased.
Luckily there are sorting algorithms that are faster than this, like Quicksort,
that we will look at later.
Selection Sort
The Selection Sort algorithm finds the lowest value in an array and moves it
to the front of the array.
Sort
The algorithm looks through the array again and again, moving the next
lowest values to the front, until the array is sorted.
How it works:
[ 7, 12, 9, 11, 3]
Step 2: Go through the array, one value at a time. Which value is the
lowest? 3, right?
[ 7, 12, 9, 11, 3]
[ 3, 7, 12, 9, 11]
Step 4: Look through the rest of the values, starting with 7. 7 is the lowest
value, and already at the front of the array, so we don't need to move it.
[ 3, 7, 12, 9, 11]
Step 5: Look through the rest of the array: 12, 9 and 11. 9 is the lowest
value.
[ 3, 7, 12, 9, 11]
Step 6: Move 9 to the front.
[ 3, 7, 9, 12, 11]
[ 3, 7, 9, 12, 11]
[ 3, 7, 9, 11, 12]
Selection Sort
[
7,
12,
9,
11,
3
]
n = len(mylist)
for i in range(n-1):
min_index = i
for j in range(i+1, n):
if mylist[j] < mylist[min_index]:
min_index = j
min_value = [Link](min_index)
[Link](i, min_value)
print(mylist)
In the code above, the lowest value element is removed, and then inserted in
front of the array.
Each time the next lowest value array element is removed, all following
elements must be shifted one place down to make up for the removal.
These shifting operation takes a lot of time, and we are not even done yet!
After the lowest value (5) is found and removed, it is inserted at the start of
the array, causing all following values to shift one position up to make space
for the new value, like the image below shows.
Note: You will not see these shifting operations happening in the code if you
are using a high level programming language such as Python or Java, but the
shifting operations are still happening in the background. Such shifting
operations require extra time for the computer to do, which can be a
problem.
REMOVE ADS
We can swap values like the image above shows because the lowest value
ends up in the correct position, and it does not matter where we put the
other value we are swapping with, because it is not sorted yet.
Here is a simulation that shows how this improved Selection Sort with
swapping works:
Sort
Example
The improved Selection Sort algorithm, including swapping values:
mylist = [64, 34, 25, 12, 22, 11, 90, 5]
n = len(mylist)
for i in range(n):
min_index = i
for j in range(i+1, n):
if mylist[j] < mylist[min_index]:
min_index = j
mylist[i], mylist[min_index] = mylist[min_index], mylist[i]
print(mylist)
The time complexity for the Selection Sort algorithm can be displayed in a
graph like this:
As you can see, the run time is the same as for Bubble Sort: The run time
increases really fast when the size of the array is increased.
Insertion Sort with Python
Insertion Sort
The Insertion Sort algorithm uses one part of the array to hold the sorted
values, and the other part of the array to hold values that are not sorted yet.
Sort
The algorithm takes one value at a time from the unsorted part of the array
and puts it into the right place in the sorted part of the array, until the array
is sorted.
How it works:
1. Take the first value from the unsorted part of the array.
2. Move the value into the correct place in the sorted part of the array.
3. Go through the unsorted part of the array again as many times as there
are values.
REMOVE ADS
[ 7, 12, 9, 11, 3]
Step 2: We can consider the first value as the initial sorted part of the array.
If it is just one value, it must be sorted, right?
[ 7, 12, 9, 11, 3]
Step 3: The next value 12 should now be moved into the correct position in
the sorted part of the array. But 12 is higher than 7, so it is already in the
correct position.
[ 7, 12, 9, 11, 3]
[ 7, 12, 9, 11, 3]
Step 5: The value 9 must now be moved into the correct position inside the
sorted part of the array, so we move 9 in between 7 and 12.
[ 7, 9, 12, 11, 3]
[ 7, 9, 11, 12, 3]
[ 7, 9, 11, 12, 3]
Step 9: We insert 3 in front of all other values because it is the lowest value.
Insertion Sort
[
7,
12,
9,
11,
3
]
n = len(mylist)
for i in range(1,n):
insert_index = i
current_value = [Link](i)
for j in range(i-1, -1, -1):
if mylist[j] > current_value:
insert_index = j
[Link](insert_index, current_value)
print(mylist)
The way the code above first removes a value and then inserts it somewhere
else is intuitive. It is how you would do Insertion Sort physically with a hand
of cards for example. If low value cards are sorted to the left, you pick up a
new unsorted card, and insert it in the correct place between the other
already sorted cards.
The problem with this way of programming it is that when removing a value
from the array, all elements above must be shifted one index place down:
And when inserting the removed value into the array again, there are also
many shift operations that must be done: all following elements must shift
one position up to make place for the inserted value:
These shifting operations can take a lot of time, especially for an array with
many elements.
Hidden memory shifts: You will not see these shifting operations
happening in the code if you are using a high-level programming language
such as Python or JavaScript, but the shifting operations are still happening
in the background. Such shifting operations require extra time for the
computer to do, which can be a problem.
You can read more about how arrays are stored in memory here.
REMOVE ADS
Improved Solution
We can avoid most of these shift operations by only shifting the values
necessary:
In the image above, first value 7 is copied, then values 11 and 12 are shifted
one place up in the array, and at last value 7 is put where value 11 was
before.
Example
Insert the improvements in the sorting algorithm:
n = len(mylist)
for i in range(1,n):
insert_index = i
current_value = mylist[i]
for j in range(i-1, -1, -1):
if mylist[j] > current_value:
mylist[j+1] = mylist[j]
insert_index = j
else:
break
mylist[insert_index] = current_value
print(mylist)
What is also done in the code above is to break out of the inner loop. That is
because there is no need to continue comparing values when we have
already found the correct place for the current value.
The time complexity for Insertion Sort can be displayed like this:
For Insertion Sort, there is a big difference between best, average and worst
case scenarios.
Quicksort
As the name suggests, Quicksort is one of the fastest sorting algorithms.
The Quicksort algorithm takes an array of values, chooses one of the values
as the 'pivot' element, and moves the other values so that lower values are
on the left of the pivot element, and higher values are on the right of it.
Sort
In this tutorial the last element of the array is chosen to be the pivot
element, but we could also have chosen the first element of the array, or any
element in the array really.
Then, the Quicksort algorithm does the same operation recursively on the
sub-arrays to the left and right side of the pivot element. This continues until
the array is sorted.
After the Quicksort algorithm has put the pivot element in between a sub-
array with lower values on the left side, and a sub-array with higher values
on the right side, the algorithm calls itself twice, so that Quicksort runs again
for the sub-array on the left side, and for the sub-array on the right side. The
Quicksort algorithm continues to call itself until the sub-arrays are too small
to be sorted.
How it works:
REMOVE ADS
[ 11, 9, 12, 7, 3]
[ 11, 9, 12, 7, 3]
Step 3: The rest of the values in the array are all greater than 3, and must
be on the right side of 3. Swap 3 with 11.
[ 3, 9, 12, 7, 11]
Step 4: Value 3 is now in the correct position. We need to sort the values to
the right of 3. We choose the last value 11 as the new pivot element.
[ 3, 9, 12, 7, 11]
Step 5: The value 7 must be to the left of pivot value 11, and 12 must be to
the right of it. Move 7 and 12.
[ 3, 9, 7, 12, 11]
Step 6: Swap 11 with 12 so that lower values 9 and 7 are on the left side of
11, and 12 is on the right side.
[ 3, 9, 7, 11, 12]
[ 3, 9, 7, 11, 12]
[ 3, 7, 9, 11, 12]
Sort
[
11,
9,
12,
7,
3
]
The recursion part of the Quicksort algorithm is actually a reason why the
average sorting scenario is so fast, because for good picks of the pivot
element, the array will be split in half somewhat evenly each time the
algorithm calls itself. So the number of recursive calls do not double, even if
the number of values n double.
Sort
0
1
0
2
0
3
0
4
0
5
Run the simulation to see how 17 integer values from 1 till 5 are sorted using
Counting Sort.
Counting Sort does not compare values like the previous sorting algorithms
we have looked at, and only works on non negative integers.
1. Create a new array for counting how many there are of the different
values.
2. Go through the array that needs to be sorted.
3. For each value, count it by increasing the counting array at the
corresponding index.
4. After counting the values, go through the counting array to create the
sorted array.
5. For each count in the counting array, create the correct number of
elements, with values that correspond to the counting array index.
REMOVE ADS
Conditions for Counting Sort
These are the reasons why Counting Sort is said to only work for a limited
range of non-negative integer values:
myArray = [ 2, 3, 0, 2, 3, 2]
Step 2: We create another array for counting how many there are of each
value. The array has 4 elements, to hold values 0 through 3.
myArray = [ 2, 3, 0, 2, 3, 2]
countArray = [ 0, 0, 0, 0]
Step 3: Now let's start counting. The first element is 2, so we must
increment the counting array element at index 2.
myArray = [ 2, 3, 0, 2, 3, 2]
countArray = [ 0, 0, 1, 0]
Step 4: After counting a value, we can remove it, and count the next value,
which is 3.
myArray = [ 3, 0, 2, 3, 2]
countArray = [ 0, 0, 1, 1]
myArray = [ 0, 2, 3, 2]
countArray = [ 1, 0, 1, 1]
myArray = [ ]
countArray = [ 1, 0, 3, 2]
Step 7: Now we will recreate the elements from the initial array, and we will
do it so that the elements are ordered lowest to highest.
The first element in the counting array tells us that we have 1 element with
value 0. So we push 1 element with value 0 into the array, and we decrease
the element at index 0 in the counting array with 1.
myArray = [ 0]
countArray = [ 0, 0, 3, 2]
Step 8: From the counting array we see that we do not need to create any
elements with value 1.
myArray = [ 0]
countArray = [ 0, 0, 3, 2]
Step 9: We push 3 elements with value 2 into the end of the array. And as
we create these elements we also decrease the counting array at index 2.
myArray = [ 0, 2, 2, 2]
countArray = [ 0, 0, 0, 2]
Step 10: At last we must add 2 elements with value 3 at the end of the
array.
myArray = [0, 2, 2, 2, 3, 3]
countArray = [ 0, 0, 0, 0]
Sort
myArray = [
2,
3,
0,
2,
3,
2
]
countArray = [
0,
0,
0,
0
]
Implement Counting Sort in Python
To implement the Counting Sort algorithm in a Python program, we need:
One more thing: We need to find out what the highest value in the array is,
so that the counting array can be created with the correct size. For example,
if the highest value is 5, the counting array must be 6 elements in total, to
be able count all possible non negative integers 0, 1, 2, 3, 4 and 5.
def countingSort(arr):
max_val = max(arr)
count = [0] * (max_val + 1)
for i in range(len(count)):
while count[i] > 0:
[Link](i)
count[i] -= 1
return arr
mylist = [4, 2, 2, 6, 3, 3, 1, 6, 5, 2, 3]
mysortedlist = countingSort(mylist)
print(mysortedlist)
REMOVE ADS
Counting Sort Time Complexity
How fast the Counting Sort algorithm runs depends on both the range of
possible values k and the number of values n.
In general, time complexity for Counting Sort is O(n+k).
In a best case scenario, the range of possible different values k is very small
compared to the number of values n and Counting Sort has time
complexity O(n).
But in a worst case scenario, the range of possible different values k is very
big compared to the number of values n and Counting Sort can have time
complexity O(n2) or even worse.
The plot below shows how much the time complexity for Counting Sort can
vary.
Radix Sort
The Radix Sort algorithm sorts an array by individual digits, starting with the
least significant digit (the rightmost one).
Step
490
369
504
185
583
385
348
204
515
198
The radix (or base) is the number of unique digits in a number system. In the
decimal system we normally use, there are 10 different digits from 0 till 9.
Radix Sort uses the radix so that decimal values are put into 10 different
buckets (or containers) corresponding to the digit that is in focus, then put
back into the array before moving on to the next digit.
Radix Sort is a non comparative algorithm that only works with non negative
integers.
How it works:
1. Start with the least significant digit (rightmost digit).
2. Sort the values based on the digit in focus by first putting the values in the
correct bucket based on the digit in focus, and then put them back into
array in the correct order.
3. Move to the next digit, and sort again, like in the step above, until there
are no digits left.
REMOVE ADS
Stable Sorting
Radix Sort must sort the elements in a stable way for the result to be sorted
correctly.
It makes little sense to talk about stable sorting algorithms for the previous
algorithms we have looked at individually, because the result would be same
if they are stable or not. But it is important for Radix Sort that the the sorting
is done in a stable way because the elements are sorted by just one digit at
a time.
So after sorting the elements on the least significant digit and moving to the
next digit, it is important to not destroy the sorting work that has already
been done on the previous digit position, and that is why we need to take
care that Radix Sort does the sorting on each digit position in a stable way.
In the simulation below it is revealed how the underlying sorting into buckets
is done. And to get a better understanding of how stable sorting works, you
can also choose to sort in an unstable way, that will lead to an incorrect
result. The sorting is made unstable by simply putting elements into buckets
from the end of the array instead of from the start of the array.
Stable sort? Yes
Sort
0
1
2
3
4
5
6
7
8
9
141
115
372
453
272
323
348
301
292
473
Step 1: We start with an unsorted array, and an empty array to fit values
with corresponding radices 0 till 9.
Step 3: Now we move the elements into the correct positions in the radix
array according to the digit in focus. Elements are taken from the start of
myArray and pushed into the correct position in the radixArray.
myArray = [ ]
radixArray = [ [40], [], [], [33], [24], [45, 25], [], [17], [], []
]
Step 4: We move the elements back into the initial array, and the sorting is
now done for the least significant digit. Elements are taken from the end
radixArray, and put into the start of myArray.
Step 5: We move focus to the next digit. Notice that values 45 and 25 are
still in the same order relative to each other as they were to start with,
because we sort in a stable way.
Step 6: We move elements into the radix array according to the focused
digit.
myArray = [ ]
radixArray = [ [], [17], [24, 25], [33], [40, 45], [], [], [], [],
[] ]
Step 7: We move elements back into the start of myArray, from the back of
radixArray.
myArray = [ 17, 24, 25, 33, 40, 45 ]
radixArray = [ [], [], [], [], [], [], [], [], [], [] ]
Sort
myArray = [
33,
45,
40,
25,
17,
24
]
radixArray = [ [ ], [ ], [ ], [ ], [ ], [ ], [ ], [ ], [ ], [ ], [ ]
]
exp *= 10
print(mylist)
On line 7, we use floor division ("//") to divide the maximum value 802 by 1
the first time the while loop runs, the next time it is divided by 10, and the
last time it is divided by 100. When using floor division "//", any number
beyond the decimal point are disregarded, and an integer is returned.
On line 11, it is decided where to put a value in the radixArray based on its
radix, or digit in focus. For example, the second time the outer while loop
runs exp will be 10. Value 170 divided by 10 will be 17. The "%10" operation
divides by 10 and returns what is left. In this case 17 is divided by 10 one
time, and 7 is left. So value 170 is placed in index 7 in the radixArray.
REMOVE ADS
Example
A Radix Sort algorithm that uses Bubble Sort:
def bubbleSort(arr):
n = len(arr)
for i in range(n):
for j in range(0, n - i - 1):
if arr[j] > arr[j + 1]:
arr[j], arr[j + 1] = arr[j + 1], arr[j]
def radixSortWithBubbleSort(arr):
max_val = max(arr)
exp = 1
i = 0
for bucket in radixList:
for num in bucket:
arr[i] = num
i += 1
exp *= 10
radixSortWithBubbleSort(mylist)
print(mylist)
See different possible time complexities for Radix Sort in the image below.
Merge Sort
The Merge Sort algorithm is a divide-and-conquer algorithm that sorts an
array by first breaking it down into smaller arrays, and then building the
array back together the correct way so that it is sorted.
Sort
Divide: The algorithm starts with breaking up the array into smaller and
smaller pieces until one such sub-array only consists of one element.
Conquer: The algorithm merges the small pieces of the array back together
by putting the lowest values first, resulting in a sorted array.
The breaking down and building up of the array to sort the array is done
recursively.
In the animation above, each time the bars are pushed down represents a
recursive call, splitting the array into smaller pieces. When the bars are lifted
up, it means that two sub-arrays have been merged together.
The Merge Sort algorithm can be described like this:
How it works:
1. Divide the unsorted array into two sub-arrays, half the size of the original.
2. Continue to divide the sub-arrays as long as the current piece of the array
has more than one element.
3. Merge two sub-arrays together by always putting the lowest value first.
4. Keep merging until there are no sub-arrays left.
Take a look at the drawing below to see how Merge Sort works from a
different perspective. As you can see, the array is split into smaller and
smaller pieces until it is merged back together. And as the merging happens,
values from each sub-array are compared so that the lowest value comes
first.
REMOVE ADS
[ 12, 8, 9, 3, 11, 5, 4]
[ 12, 8, 9] [ 3, 11, 5, 4]
[ 12] [ 8, 9] [ 3, 11, 5, 4]
[ 12] [ 8] [ 9] [ 3, 11, 5, 4]
Step 2: The splitting of the first sub-array is finished, and now it is time to
merge. 8 and 9 are the first two elements to be merged. 8 is the lowest
value, so that comes before 9 in the first merged sub-array.
[ 12] [ 8, 9] [ 3, 11, 5, 4]
Step 3: The next sub-arrays to be merged is [ 12] and [ 8, 9]. Values in both
arrays are compared from the start. 8 is lower than 12, so 8 comes first, and
9 is also lower than 12.
[ 8, 9, 12] [ 3, 11, 5, 4]
[ 8, 9, 12] [ 3, 11, 5, 4]
[ 8, 9, 12] [ 3, 11] [ 5, 4]
[ 8, 9, 12] [ 3] [ 11] [ 5, 4]
Step 5: 3 and 11 are merged back together in the same order as they are
shown because 3 is lower than 11.
[ 8, 9, 12] [ 3, 11] [ 5, 4]
Step 6: Sub-array with values 5 and 4 is split, then merged so that 4 comes
before 5.
[ 8, 9, 12] [ 3, 11] [ 5] [ 4]
[ 8, 9, 12] [ 3, 11] [ 4, 5]
Step 7: The two sub-arrays on the right are merged. Comparisons are done
to create elements in the new merged array:
1. 3 is lower than 4
2. 4 is lower than 11
3. 5 is lower than 11
4. 11 is the last remaining value
[ 8, 9, 12] [ 3, 4, 5, 11]
Step 8: The two last remaining sub-arrays are merged. Let's look at how the
comparisons are done in more detail to create the new merged and finished
sorted array:
3 is lower than 8:
Sort
[
12
,
8
,
9
,
3
,
11
,
5
,
4
]
def mergeSort(arr):
if len(arr) <= 1:
return arr
mid = len(arr) // 2
leftHalf = arr[:mid]
rightHalf = arr[mid:]
sortedLeft = mergeSort(leftHalf)
sortedRight = mergeSort(rightHalf)
[Link](left[i:])
[Link](right[j:])
return result
On line 7, arr[mid:] takes all values from the array, starting at the value on
index "mid" and all the next values.
On lines 26-27, the first part of the merging is done. At this this point the
values of the two sub-arrays are compared, and either the left sub-array or
the right sub-array is empty, so the result array can just be filled with the
remaining values from either the left or the right sub-array. These lines can
be swapped, and the result will be the same.
But Merge Sort can also be implemented without the use of recursion, so
that there is no function calling itself.
Take a look at the Merge Sort implementation below, that does not use
recursion:
Example
A Merge sort without recursion
[Link](left[i:])
[Link](right[j:])
return result
def mergeSort(arr):
step = 1 # Starting with sub-arrays of length 1
length = len(arr)
return arr
You might notice that the merge functions are exactly the same in the two
Merge Sort implementations above, but in the implementation right above
here the while loop inside the mergeSort function is used to replace the
recursion. The while loop does the splitting and merging of the array in
place, and that makes the code a bit harder to understand.
To put it simply, the while loop inside the mergeSort function uses short step
lengths to sort tiny pieces (sub-arrays) of the initial array using the merge
function. Then the step length is increased to merge and sort larger pieces of
the array until the whole array is sorted.
REMOVE ADS
The image below shows the time complexity for Merge Sort.
Merge Sort performs almost the same every time because the array is split,
and merged using comparison, both if the array is already sorted or not.
Python MySQL
MySQL Database
To be able to experiment with the code examples in this tutorial, you should
have MySQL installed on your computer.
Navigate your command line to the location of PIP, and type the following:
C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>python -m pip install mysql-connector-python
demo_mysql_test.py:
import [Link]
Create Connection
Start by creating a connection to the database.
demo_mysql_connection.py:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword"
)
print(mydb)
Now you can start querying the database using SQL statements.
Creating a Database
To create a database in MySQL, use the "CREATE DATABASE" statement:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword"
)
mycursor = [Link]()
[Link]("CREATE DATABASE mydatabase")
If the above code was executed with no errors, you have successfully
created a database.
Example
Return a list of your system's databases:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword"
)
mycursor = [Link]()
[Link]("SHOW DATABASES")
for x in mycursor:
print(x)
Or you can try to access the database when making the connection:
Example
Try connecting to the database "mydatabase":
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
Creating a Table
To create a table in MySQL, use the "CREATE TABLE" statement.
Make sure you define the name of the database when you create the
connection
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
If the above code was executed with no errors, you have now successfully
created a table.
Example
Return a list of your system's databases:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link]("SHOW TABLES")
for x in mycursor:
print(x)
REMOVE ADS
Primary Key
When creating a table, you should also create a column with a unique key for
each record.
We use the statement "INT AUTO_INCREMENT PRIMARY KEY" which will insert
a unique number for each record. Starting at 1, and increased by one for
each record.
Example
Create primary key when creating the table:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
Example
Create primary key on an existing table:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link]()
Example
Fill the "customers" table with data:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link](sql, val)
[Link]()
REMOVE ADS
Get Inserted ID
You can get the id of the row you just inserted by asking the cursor object.
Note: If you insert more than one row, the id of the last inserted row is
returned.
Example
Insert one row, and return the ID:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link]()
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
for x in myresult:
print(x)
Note: We use the fetchall() method, which fetches all rows from the last
executed statement.
Selecting Columns
To select only some of the columns in a table, use the "SELECT" statement
followed by the column name(s):
Example
Select only the name and address columns:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
myresult = [Link]()
for x in myresult:
print(x)
REMOVE ADS
Using the fetchone() Method
If you are only interested in one row, you can use the fetchone() method.
The fetchone() method will return the first row of the result:
Example
Fetch only one row:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
myresult = [Link]()
print(myresult)
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link](sql)
myresult = [Link]()
for x in myresult:
print(x)
Wildcard Characters
You can also select the records that starts, includes, or ends with a given
letter or phrase.
Example
Select records where the address contains the word "way":
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link](sql)
myresult = [Link]()
for x in myresult:
print(x)
REMOVE ADS
Example
Escape query values by using the placholder %s method:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link](sql, adr)
myresult = [Link]()
for x in myresult:
print(x)
Python MySQL Order By
The ORDER BY keyword sorts the result ascending by default. To sort the
result in descending order, use the DESC keyword.
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link](sql)
myresult = [Link]()
for x in myresult:
print(x)
ORDER BY DESC
Use the DESC keyword to sort the result in a descending order.
Example
Sort the result reverse alphabetically by name:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link](sql)
myresult = [Link]()
for x in myresult:
print(x)
Delete Record
You can delete records from an existing table by using the "DELETE FROM"
statement:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link](sql)
[Link]()
REMOVE ADS
Example
Escape values by using the placeholder %s method:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link](sql, adr)
[Link]()
Delete a Table
You can delete an existing table by using the "DROP TABLE" statement:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link](sql)
Drop Only if Exist
If the table you want to delete is already deleted, or for any other reason
does not exist, you can use the IF EXISTS keyword to avoid getting an error.
Example
Delete the table "customers" if it exists:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link](sql)
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link](sql)
[Link]()
Example
Escape values by using the placeholder %s method:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
[Link](sql, val)
[Link]()
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
myresult = [Link]()
for x in myresult:
print(x)
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
myresult = [Link]()
for x in myresult:
print(x)
products
{ id: 154, name: 'Chocolate Heaven' },
{ id: 155, name: 'Tasty Lemons' },
{ id: 156, name: 'Vanilla Dreams' }
These two tables can be combined by using users' fav field and
products' id field.
Example
Join users and products to see the name of the users favorite product:
import [Link]
mydb = [Link](
host="localhost",
user="yourusername",
password="yourpassword",
database="mydatabase"
)
mycursor = [Link]()
sql = "SELECT \
[Link] AS user, \
[Link] AS favorite \
FROM users \
INNER JOIN products ON [Link] = [Link]"
[Link](sql)
myresult = [Link]()
for x in myresult:
print(x)
Note: You can use JOIN instead of INNER JOIN. They will both give you the
same result.
LEFT JOIN
In the example above, Hannah, and Michael were excluded from the result,
that is because INNER JOIN only shows the records where there is a match.
If you want to show all users, even if they do not have a favorite product, use
the LEFT JOIN statement:
Example
Select all users and their favorite product:
sql = "SELECT \
[Link] AS user, \
[Link] AS favorite \
FROM users \
LEFT JOIN products ON [Link] = [Link]"
RIGHT JOIN
If you want to return all products, and the users who have them as their
favorite, even if no user have them as their favorite, use the RIGHT JOIN
statement:
Example
Select all products, and the user(s) who have them as their favorite:
sql = "SELECT \
[Link] AS user, \
[Link] AS favorite \
FROM users \
RIGHT JOIN products ON [Link] = [Link]"
Note: Hannah and Michael, who have no favorite product, are not included
in the result.
Python MongoDB
Python can be used in database applications.
MongoDB
MongoDB stores data in JSON-like documents, which makes the database
very flexible and scalable.
To be able to experiment with the code examples in this tutorial, you will
need access to a MongoDB database.
PyMongo
Python needs a MongoDB driver to access the MongoDB database.
Navigate your command line to the location of PIP, and type the following:
C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>python -m pip install pymongo
Test PyMongo
To test if the installation was successful, or if you already have "pymongo"
installed, create a Python page with the following content:
demo_mongodb_test.py:
import pymongo
If the above code was executed with no errors, "pymongo" is installed and
ready to be used.
MongoDB will create the database if it does not exist, and make a connection
to it.
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
MongoDB waits until you have created a collection (table), with at least one
document (record) before it actually creates the database (and collection).
You can check if a database exist by listing all databases in you system:
Example
Return a list of your system's databases:
print(myclient.list_database_names())
Example
Check if "mydatabase" exists:
dblist = myclient.list_database_names()
if "mydatabase" in dblist:
print("The database exists.")
Creating a Collection
To create a collection in MongoDB, use database object and specify the
name of the collection you want to create.
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
Example
Return a list of all collections in your database:
print(mydb.list_collection_names())
Example
Check if the "customers" collection exists:
collist = mydb.list_collection_names()
if "customers" in collist:
print("The collection exists.")
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
x = mycol.insert_one(mydict)
Example
Insert another record in the "customers" collection, and return the value of
the _id field:
x = mycol.insert_one(mydict)
print(x.inserted_id)
If you do not specify an _id field, then MongoDB will add one for you and
assign a unique id for each document.
Example
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
mylist = [
{ "name": "Amy", "address": "Apple st 652"},
{ "name": "Hannah", "address": "Mountain 21"},
{ "name": "Michael", "address": "Valley 345"},
{ "name": "Sandy", "address": "Ocean blvd 2"},
{ "name": "Betty", "address": "Green Grass 1"},
{ "name": "Richard", "address": "Sky st 331"},
{ "name": "Susan", "address": "One way 98"},
{ "name": "Vicky", "address": "Yellow Garden 2"},
{ "name": "Ben", "address": "Park Lane 38"},
{ "name": "William", "address": "Central st 954"},
{ "name": "Chuck", "address": "Main Road 989"},
{ "name": "Viola", "address": "Sideway 1633"}
]
x = mycol.insert_many(mylist)
Remember that the values has to be unique. Two documents cannot have
the same _id.
Example
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
mylist = [
{ "_id": 1, "name": "John", "address": "Highway 37"},
{ "_id": 2, "name": "Peter", "address": "Lowstreet 27"},
{ "_id": 3, "name": "Amy", "address": "Apple st 652"},
{ "_id": 4, "name": "Hannah", "address": "Mountain 21"},
{ "_id": 5, "name": "Michael", "address": "Valley 345"},
{ "_id": 6, "name": "Sandy", "address": "Ocean blvd 2"},
{ "_id": 7, "name": "Betty", "address": "Green Grass 1"},
{ "_id": 8, "name": "Richard", "address": "Sky st 331"},
{ "_id": 9, "name": "Susan", "address": "One way 98"},
{ "_id": 10, "name": "Vicky", "address": "Yellow Garden 2"},
{ "_id": 11, "name": "Ben", "address": "Park Lane 38"},
{ "_id": 12, "name": "William", "address": "Central st 954"},
{ "_id": 13, "name": "Chuck", "address": "Main Road 989"},
{ "_id": 14, "name": "Viola", "address": "Sideway 1633"}
]
x = mycol.insert_many(mylist)
Just like the SELECT statement is used to find data in a table in a MySQL
database.
Find One
To select data from a collection in MongoDB, we can use
the find_one() method.
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
x = mycol.find_one()
print(x)
Find All
To select data from a table in MongoDB, we can also use the find() method.
The first parameter of the find() method is a query object. In this example
we use an empty query object, which selects all documents in the collection.
No parameters in the find() method gives you the same result as SELECT
* in MySQL.
Example
Return all documents in the "customers" collection, and print each
document:
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
for x in [Link]():
print(x)
Return Only Some Fields
The second parameter of the find() method is an object describing which
fields to include in the result.
This parameter is optional, and if omitted, all fields will be included in the
result.
Example
Return only the names and addresses, not the _ids:
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
Example
This example will exclude "address" from the result:
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
Example
You get an error if you specify both 0 and 1 values in the same object
(except if one of the fields is the _id field):
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
The first argument of the find() method is a query object, and is used to
limit the search.
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
mydoc = [Link](myquery)
for x in mydoc:
print(x)
Advanced Query
To make advanced queries you can use modifiers as values in the query
object.
E.g. to find the documents where the "address" field starts with the letter "S"
or higher (alphabetically), use the greater than modifier: {"$gt": "S"}:
Example
Find documents where the address starts with the letter "S" or higher:
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
mydoc = [Link](myquery)
for x in mydoc:
print(x)
To find only the documents where the "address" field starts with the letter
"S", use the regular expression {"$regex": "^S"}:
Example
Find documents where the address starts with the letter "S":
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
mydoc = [Link](myquery)
for x in mydoc:
print(x)
The sort() method takes one parameter for "fieldname" and one parameter
for "direction" (ascending is the default direction).
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
mydoc = [Link]().sort("name")
for x in mydoc:
print(x)
Sort Descending
Use the value -1 as the second parameter to sort descending.
sort("name", 1) #ascending
sort("name", -1) #descending
Example
Sort the result reverse alphabetically by name:
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
for x in mydoc:
print(x)
Note: If the query finds more than one document, only the first occurrence
is deleted.
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
mycol.delete_one(myquery)
Example
Delete all documents were the address starts with the letter S:
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
x = mycol.delete_many(myquery)
Example
Delete all documents in the "customers" collection:
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
x = mycol.delete_many({})
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
[Link]()
The drop() method returns true if the collection was dropped successfully,
and false if the collection does not exist.
Note: If the query finds more than one record, only the first occurrence is
updated.
The second parameter is an object defining the new values of the document.
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
mycol.update_one(myquery, newvalues)
Update Many
To update all documents that meets the criteria of the query, use
the update_many() method.
Example
Update all documents where the address starts with the letter "S":
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
x = mycol.update_many(myquery, newvalues)
The limit() method takes one parameter, a number defining how many
documents to return.
Consider you have a "customers" collection:
Example
Limit the result to only return 5 documents:
import pymongo
myclient = [Link]("mongodb://localhost:27017/")
mydb = myclient["mydatabase"]
mycol = mydb["customers"]
myresult = [Link]().limit(5)
delattr() Deletes the specified attribute (property or method) from the specifie
divmod() Returns the quotient and the remainder when argument1 is divided b
hasattr() Returns True if the specified object has the specified attribute (prope
map() Returns the specified iterator with the specified function applied to e
Python has a set of built-in methods that you can use on strings.
Note: All string methods returns new values. They do not change the
original string.
Method Description
endswith() Returns true if the string ends with the specified value
expandtabs() Sets the tab size of the string
find() Searches the string for a specified value and returns the position of w
index() Searches the string for a specified value and returns the position of w
isalpha() Returns True if all characters in the string are in the alphabet
isascii() Returns True if all characters in the string are ascii characters
islower() Returns True if all characters in the string are lower case
isupper() Returns True if all characters in the string are upper case
partition() Returns a tuple where the string is parted into three parts
rfind() Searches the string for a specified value and returns the last position
rindex() Searches the string for a specified value and returns the last position
rpartition() Returns a tuple where the string is parted into three parts
rsplit() Splits the string at the specified separator, and returns a list
split() Splits the string at the specified separator, and returns a list
swapcase() Swaps cases, lower case becomes upper case and vice versa
zfill() Fills the string with a specified number of 0 values at the beginning
Note: All string methods returns new values. They do not change the
original string.
Python has a set of built-in methods that you can use on lists/arrays.
Method Description
extend() Add the elements of a list (or any iterable), to the end of the current lis
index() Returns the index of the first element with the specified value
Note: Python does not have built-in support for Arrays, but Python Lists can
be used instead.
Python Dictionary Methods
Python has a set of built-in methods that you can use on dictionaries.
Method Description
items() Returns a list containing a tuple for each key value pair
Python has two built-in methods that you can use on tuples.
Method Description
index() Searches the tuple for a specified value and returns the position
Python has a set of built-in methods that you can use on sets.
intersection_update() &= Removes the items in this set that are not
Method Description
close() Closes the file
fileno() Returns a number that represents the stream, from the operating syste
seekable() Returns whether the file allows us to change the file position
tell() Returns the current file position
Python Keywords
Python has a set of keywords that are reserved words that cannot be used as
variable names, function names, or any other identifiers:
Keyword Description
as To create an alias
or A logical operator
Built-in Exceptions
The table below shows built-in exceptions that are usually raised in Python:
Exception Description
EOFError Raised when the input() method hits an "end of file" con
Feature Description
Global Variables Global variables are variables that belongs to the global
Specify a Variable Type How to specify a certain data type for a variable
Strings are Arrays Strings in Python are arrays of bytes representing Unicod
Identity Operators Identity operators are used to see if two objects are in fa
Loop Through List Items How to loop through the items in a list
Check if List Item Exists How to check if a specified item is present in a list
Check if Tuple Item Exists How to check if a specified item is present in a tuple
Tuple With One Item How to create a tuple with only one item
The pass Keyword in If Use the pass keyword inside empty if statements
While Continue How to stop the current iteration and continue wit the ne
For Continue How to stop the current iteration and continue wit the ne
Looping Through a range How to loop through a range of values
For pass Use the pass keyword inside empty for loops
Function Recursion Functions that can call itself is called recursive functions
Why Use Lambda Functions Learn when to use a lambda function or not
What is an Array Arrays are variables that can hold more than one value
The Class __init__() Function The __init__() function is executed when the class is initia
Object Methods Methods in objects are functions that belongs to the obje
super Function The super() function make the child class inherit the par
Using the dir() Function List all variable names and function names in a module
The strftime Method How to format a date object into a readable string
Date Format Codes The datetime module has a set of legal format codes
Format JSON How to format JSON output with indentations and line br
RegEx Match Object The Match Object is an object containing information abo
This page lists the built-in modules that ship with the Python 3.13 Standard
Library.
These modules are available without extra installation (some are platform-
dependent).
A
Module Description
B
Module Description
bisect Maintain sorted lists; insert and search with binary search
builtins Access Python's built-in objects like len, range, and excep
C
Module Description
REMOVE ADS
D
Module Description
dataclasses Decorator and helpers for classes that store data (auto-ge
etc.).
F
Module Description
faulthandler Dump Python tracebacks on a crash or on demand (helps
fileinput Loop over lines from stdin or a list of files as a single strea
G
Module Description
H
Module Description
I
Module Description
K
Module Description
L
Module Description
M
Module Description
marshal Read and write Python values in a binary format (for .pyc
N
Module Description
P
Module Description
Q
Module Description
R
Module Description
S
Module Description
T
Module Description
typing Type hints and typing helpers for static analysis and tooli
U
Module Description
unicodedata Access the Unicode Character Database (properties, norm
V
Module Description
W
Module Description
wsgiref WSGI utilities and simple reference server for web apps.
X
Module Description
Z
Module Description
Python has a built-in module that you can use to make random numbers.
Method Description
getstate() Returns the current internal state of the random number generator
setstate() Restores the internal state of the random number generator
choices() Returns a list with a random selection from the given sequence
betavariate() Returns a random float number between 0 and 1 based on the Beta dis
gammavariate Returns a random float number based on the Gamma distribution (use
()
gauss() Returns a random float number based on the Gaussian distribution (us
normalvariate( Returns a random float number based on the normal distribution (used
)
vonmisesvariat Returns a random float number based on the von Mises distribution (us
e()
paretovariate() Returns a random float number based on the Pareto distribution (used
weibullvariate( Returns a random float number based on the Weibull distribution (used
)
Python Requests Module
import requests
x = [Link]('[Link]
print([Link])
The HTTP request returns a Response Object with all the response data
(content, encoding, status, etc).
C:\Users\Your Name\AppData\Local\Programs\Python\Python36-32\
Scripts>pip install requests
Syntax
[Link](params)
Methods
Method Description
post(url, data, json, args) Sends a POST request to the specified url
request(method, url, args) Sends a request of the specified method to the specifi
Statistics Methods
Method Description
Math Methods
Method Description
math.expm1() Returns Ex - 1
[Link]() Returns the absolute value of a number
Math Constants
Constant Description
The methods in this module accepts int, float, and complex numbers. It even
accepts Python objects that has a __complex__() or __float__() method.
The methods in this module almost always return a complex number. If the
return value can be expressed as a real number, the return value has an
imaginary part of 0.
cMath Methods
Method Description
REMOVE ADS
cMath Constants
Constant Description
Example Explained
First we have a List that contains duplicates:
Create a dictionary, using the List items as keys. This will automatically
remove any duplicates because dictionaries cannot have duplicate keys.
Create a Dictionary
mylist = ["a", "b", "a", "c", "c"]
mylist = list( [Link](mylist) )
print(mylist)
Now we have a List without any duplicates, and it has the same order as the
original List.
Create a Function
If you like to have a function where you can send your lists, and get them
back without duplicates, you can create a function and insert the code from
the example above.
Example
def my_function(x):
return list([Link](x))
print(mylist)
Example Explained
Create a function that takes a List as an argument.
Create a Function
def my_function(x):
return list([Link](x))
print(mylist)
Create a Dictionary
def my_function(x):
return list( [Link](x) )
print(mylist)
print(mylist)
Return List
def my_function(x):
return list([Link](x))
print(mylist)
print(mylist)
print(mylist)
The fastest (and easiest?) way is to use a slice that steps backwards, -1.
Example Explained
We have a string, "Hello World", which we want to reverse:
Create a slice that starts at the end of the string, and moves backwards.
In this particular example, the slice statement [::-1] means start at the end
of the string and end at position 0, move with the step -1, negative one,
which means one step backwards.
Create a Function
If you like to have a function where you can send your strings, and return
them backwards, you can create a function and insert the code from the
example above.
Example
def my_function(x):
return x[::-1]
print(mytxt)
Example Explained
Create a function that takes a String as an argument.
Create a Function
def my_function(x):
return x[::-1]
print(mytxt)
Slice the string starting at the end of the string and move backwards.
print(mytxt)
Return the backward String
print(mytxt )
print(mytxt)
print(mytxt)
Example
x = input("Type a number: ")
y = input("Type another number: ")
Python Examples
Python Syntax
Python Variables
Python Numbers
Python Casting
REMOVE ADS
Python Strings
Python Operators
Python Lists
Python Tuples
Python Sets
Python Dictionaries
Python Functions
Python Lambda
Python Arrays
Python Iterators
Python Modules
Python Dates
Python Math
Python JSON
Python RegEx
Python PIP
Python MySQL
Python MongoDB