0% found this document useful (0 votes)
5 views4 pages

01 Dataintegration

The document details the process of analyzing Titanic passenger data using Pandas and Seaborn libraries in Python. It includes steps for reading the data, creating specific DataFrames, performing various types of joins (inner, left, right, outer), and concatenating DataFrames. Additionally, it calculates statistical measures such as mean, median, mode, and mid-range for numerical columns like age and fare.

Uploaded by

kamirkarsana04
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views4 pages

01 Dataintegration

The document details the process of analyzing Titanic passenger data using Pandas and Seaborn libraries in Python. It includes steps for reading the data, creating specific DataFrames, performing various types of joins (inner, left, right, outer), and concatenating DataFrames. Additionally, it calculates statistical measures such as mean, median, mode, and mid-range for numerical columns like age and fare.

Uploaded by

kamirkarsana04
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

In [29]: import pandas as pd

import seaborn as sns

# Read the Data with Pandas


df = pd.read_csv("[Link]")
df

Out[29]: PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket

Braund,
A/5
0 1 0 3 Mr. Owen male 22.0 1 0 7
21171
Harris

Cumings,
Mrs. John
Bradley
1 2 1 1 female 38.0 1 0 PC 17599 71
(Florence
Briggs
Th...

Heikkinen,
STON/O2.
2 3 1 3 Miss. female 26.0 0 0 7
3101282
Laina

Futrelle,
Mrs.
Jacques
3 4 1 1 female 35.0 1 0 113803 53
Heath
(Lily May
Peel)

Allen, Mr.
4 5 0 3 William male 35.0 0 0 373450 8
Henry

... ... ... ... ... ... ... ... ... ...

Montvila,
886 887 0 2 Rev. male 27.0 0 0 211536 13
Juozas

Graham,
Miss.
887 888 1 1 female 19.0 0 0 112053 30
Margaret
Edith

Johnston,
Miss.
W./C.
888 889 0 3 Catherine female NaN 1 2 23
6607
Helen
"Carrie"

Behr, Mr.
889 890 1 1 Karl male 26.0 0 0 111369 30
Howell

Dooley,
890 891 0 3 male 32.0 0 0 370376 7
Mr. Patrick

891 rows × 12 columns

In [31]: # 1. Create df1: Identification and Outcome


# Selecting specifically from the columns you listed
df1 = df[['PassengerId', 'Name', 'Age', 'Survived']].copy()

# 2. Create df2: Socio-Economic and Travel Data


df2 = df[['PassengerId', 'Pclass', 'Sex', 'Fare']].copy()

print("DataFrames successfully created!")


display([Link]())
display([Link]())
DataFrames successfully created!
PassengerId Name Age Survived

0 1 Braund, Mr. Owen Harris 22.0 0

1 2 Cumings, Mrs. John Bradley (Florence Briggs Th... 38.0 1

2 3 Heikkinen, Miss. Laina 26.0 1

3 4 Futrelle, Mrs. Jacques Heath (Lily May Peel) 35.0 1

4 5 Allen, Mr. William Henry 35.0 0

PassengerId Pclass Sex Fare

0 1 3 male 7.2500

1 2 1 female 71.2833

2 3 3 female 7.9250

3 4 1 female 53.1000

4 5 3 male 8.0500

In [41]: # Inner Join on 'PassengerId'


merged_inner = [Link](df1, df2, on='PassengerId', how='inner')
print("\nInner Join on 'PassengerId':\n")
merged_inner.head()

Inner Join on 'PassengerId':

Out[41]: PassengerId Name Age Survived Pclass Sex Fare

0 1 Braund, Mr. Owen Harris 22.0 0 3 male 7.2500

Cumings, Mrs. John Bradley


1 2 38.0 1 1 female 71.2833
(Florence Briggs Th...

2 3 Heikkinen, Miss. Laina 26.0 1 3 female 7.9250

Futrelle, Mrs. Jacques Heath


3 4 35.0 1 1 female 53.1000
(Lily May Peel)

4 5 Allen, Mr. William Henry 35.0 0 3 male 8.0500

In [43]: # Left Join on 'PassengerId'


merged_left = [Link](df1, df2, on='PassengerId', how='left')
print("\nLeft Join on 'PassengerId':\n")
merged_left.head()

Left Join on 'PassengerId':

Out[43]: PassengerId Name Age Survived Pclass Sex Fare

0 1 Braund, Mr. Owen Harris 22.0 0 3 male 7.2500

Cumings, Mrs. John Bradley


1 2 38.0 1 1 female 71.2833
(Florence Briggs Th...

2 3 Heikkinen, Miss. Laina 26.0 1 3 female 7.9250

Futrelle, Mrs. Jacques Heath


3 4 35.0 1 1 female 53.1000
(Lily May Peel)

4 5 Allen, Mr. William Henry 35.0 0 3 male 8.0500

In [45]: # Right Join on 'passengPassengerIder_id'


merged_right = [Link](df1, df2, on='PassengerId', how='right')
print("\nRight Join on 'PassengerId':\n")
merged_right.head()

Right Join on 'PassengerId':

Out[45]: PassengerId Name Age Survived Pclass Sex Fare

0 1 Braund, Mr. Owen Harris 22.0 0 3 male 7.2500

Cumings, Mrs. John Bradley


1 2 38.0 1 1 female 71.2833
(Florence Briggs Th...

2 3 Heikkinen, Miss. Laina 26.0 1 3 female 7.9250

Futrelle, Mrs. Jacques Heath


3 4 35.0 1 1 female 53.1000
(Lily May Peel)

4 5 Allen, Mr. William Henry 35.0 0 3 male 8.0500

In [47]: # Outer Join on 'passenger_id'


merged_outer = [Link](df1, df2, on='PassengerId', how='outer')
print("\nOuter Join on 'PassengerId':\n")
merged_outer.head()

Outer Join on 'PassengerId':

Out[47]: PassengerId Name Age Survived Pclass Sex Fare

0 1 Braund, Mr. Owen Harris 22.0 0 3 male 7.2500

Cumings, Mrs. John Bradley


1 2 38.0 1 1 female 71.2833
(Florence Briggs Th...

2 3 Heikkinen, Miss. Laina 26.0 1 3 female 7.9250

Futrelle, Mrs. Jacques Heath


3 4 35.0 1 1 female 53.1000
(Lily May Peel)

4 5 Allen, Mr. William Henry 35.0 0 3 male 8.0500

In [49]: # Concatenate without keys (default behavior)


concatenated = [Link]([df1, df2])
print("\nConcatenation along rows:\n")
[Link]()

Concatenation along rows:

Out[49]: PassengerId Name Age Survived Pclass Sex Fare

0 1 Braund, Mr. Owen Harris 22.0 0.0 NaN NaN NaN

Cumings, Mrs. John Bradley (Florence


1 2 38.0 1.0 NaN NaN NaN
Briggs Th...

2 3 Heikkinen, Miss. Laina 26.0 1.0 NaN NaN NaN

Futrelle, Mrs. Jacques Heath (Lily May


3 4 35.0 1.0 NaN NaN NaN
Peel)

4 5 Allen, Mr. William Henry 35.0 0.0 NaN NaN NaN

In [51]: # Concatenate with specific keys


concatenated_with_keys = [Link]([df1, df2], keys=['df1', 'df2'])
print("\nConcatenation with keys:\n")
concatenated_with_keys.head()

Concatenation with keys:


Out[51]: PassengerId Name Age Survived Pclass Sex Fare

df1 0 1 Braund, Mr. Owen Harris 22.0 0.0 NaN NaN NaN

Cumings, Mrs. John Bradley


1 2 38.0 1.0 NaN NaN NaN
(Florence Briggs Th...

2 3 Heikkinen, Miss. Laina 26.0 1.0 NaN NaN NaN

Futrelle, Mrs. Jacques Heath


3 4 35.0 1.0 NaN NaN NaN
(Lily May Peel)

4 5 Allen, Mr. William Henry 35.0 0.0 NaN NaN NaN

In [53]: # Concatenate along columns (axis=1)


concatenated_axis1 = [Link]([df1, df2], axis=1)
print("\nConcatenation along columns (axis=1):\n")
concatenated_axis1.head()

Concatenation along columns (axis=1):

Out[53]: PassengerId Name Age Survived PassengerId Pclass Sex Fare

Braund, Mr.
0 1 22.0 0 1 3 male 7.2500
Owen Harris

Cumings, Mrs.
John Bradley
1 2 38.0 1 2 1 female 71.2833
(Florence
Briggs Th...

Heikkinen,
2 3 26.0 1 3 3 female 7.9250
Miss. Laina

Futrelle, Mrs.
3 4 Jacques Heath 35.0 1 4 1 female 53.1000
(Lily May Peel)

Allen, Mr.
4 5 35.0 0 5 3 male 8.0500
William Henry

In [73]: # Calculate mean, median, and mode for numerical columns


print("\n----------- Mean -----------\n", titanic[['age', 'fare']].mean())
print("\n----------- Median -----------\n", titanic[['age', 'fare']].median())
print("\n----------- Mode -----------\n", titanic[['age', 'fare']].mode())

----------- Mean -----------


age 29.699118
fare 32.204208
dtype: float64

----------- Median -----------


age 28.0000
fare 14.4542
dtype: float64

----------- Mode -----------


age fare
0 24.0 8.05

In [67]: # Calculate mid-range (average of max and min for each column)
mid_range = titanic[['age', 'fare']].apply(lambda x: ([Link]() + [Link]()) / 2)
print("\n----------- Mid-Range -----------\n", mid_range)

----------- Mid-Range -----------


age 40.2100
fare 256.1646
dtype: float64

You might also like