Project Example2
Project Example2
Authors
Research structure
1. Executive summary
2. Research questions
3. Dataset process
4. First analysis
5. Second analysis
5.1 Cluster analysis
5.2 Hierarchical method
5.3 Non Hierarchical method
K-Means with 7 clusters
K-Means with 6 clusters
K-Means with 5 clusters
Results
Interpretation and characterization of clusters
6. CONCLUSIONS
Business Recommendations
1. Executive summary
This project is a Marketing research case study related to a Rental Airbnb reviews
company.
The goal of our analysis is to understand and highlight possible hidden relationships
within customer satisfaction if we can cluster our clients based on the Factors we
determine with our analysis and further categorize the clusters with qualitative variables
from the dataset.
We aim to point out what drives customer high ratings and discover hidden Factors that
can lead to an improvement in customer satisfaction in order to make recommendations
for new business opportunities, based on customer preferences.
For the purpose of our research, we selected the following 10 Variables, for a total of 111
observations, available on the dedicated dataset:
1. City
2. Typology of Airbnb
3. Price
4. Overall rating
5. Review_scores_accuracy
6. Review_scores_cleanliness
7. Review_scores_checkin
8. Review_scores_communication
9. Review_scores_location
10. Review_scores_value
2. Research questions
3. Dataset
Please access the dedicated data, used for this project, by clicking here.
4. First analysis
For the purpose of this analysis, we decided to run a Factor analysis to understand if there
are any correlations between the following variables listed below.
The goal of the Factor analysis is to reduce the number of variables into a smaller
number of Factors and to determine any hidden Factors related to Satisfaction.
Considered variables:
Overall rating
Review_scores_accuracy
Review_scores_cleanliness
Review_scores_checkin
Review_scores_communication
Review_scores_location
Review_scores_value
The first step of our analysis, to understand if the variables are suitable for our research,
was checking Kaiser's Measure of Sampling Adequacy (KMO).
The KMO indicates the proportion of variance that our variables retain when performing
the Factor Analysis.
Generally speaking, high values (close to 1.0) indicate that a Factor analysis may be useful
as it means that variables are homogeneous.
However, in a marketing context, a value equal or greater than 0.5 is a good indicator that
a factor analysis will be useful with our data:
Upon running the analysis, including all the original variables mentioned above, we
noticed that the KMO < 0.5, a reason to remove the variable “ Overall rating” as it showed
the lowest variance.
According to this, we run the analysis again, removing the variable “ Overall rating”:
The second step of our analysis focuses on determining the number of Factors to correlate
with the original variables, using the following methods.
Type of Methods
Our goal is to choose a number of Factors that together can explain a fair portion of the
variance of the original data. According to this, find below the list of methods considered:
a. Kaiser method
The goal of this method is to determine the number of Factors whose variance is
greater than 1. To do this, we analyzed the following table available below, checking
the values under the “Eigenvalue column”.
b. Pearson method
The goal of this method is to determine the number of factors whose cumulative
variance is lower than 0.8.
However, in a marketing context, we consider a cumulative variance between 0
and 0.5 acceptable.
According to this, we analyzed the table available below, checking the values
under the “Cumulative column”.
c. Scree plot
The goal of this method is to determine the number of Factors
in a graphical way, checking where the Eigenvalue reduces significantly its slope.
To do this, we analyzed the Scree plot table available below, checking where the
Eigenvalue creates the shape of an “elbow”.
3. Step: Results
Please find below the results of our analysis, according to the following picture:
b. Pearson method
As per Pearsons’s method, we decided to choose three Factors too as they have a
cumulative variance of 0.59.
c. Scree Plot
As per the Scree Plot method definition, we should consider a total number of
Factors by checking where the Eigenvalue creates the shape of an “elbow”.
In this case, it’s evident to select three or five Factors, however, since we only have
6 original variables, choosing three Factors would be relevant to reduce
significantly the total amount of variables analyzed.
According to all three methods applied, we get to the conclusion that a total number of
three Factors would be suitable for our research.
After analyzing the table above we can see that there is a significant amount of variance
in most original variables, especially “review_scores_accuracy” and
“review_scores_cleanliness”, however, the variable “review_scores_location” has less than
0.50 variance explained. This is not ideal and may affect the results, nevertheless, this
variable was retained as we believe that performing a Factor Analysis on the remaining
five variables would not paint a complete picture about customer satisfaction on our
observations.
Another value we had to take into consideration is the Root Mean Square Off-Diagonal
Residuals: the difference between correlation created and the real existing between
original variables.
Note: considering that this value should be close to 0.05 we noticed that Root Mean
Square Off-Diagonal Residuals is very high and this may affect the significance of our
analysis.
Upon analyzing the table, we pointed out that Factor 1 highlights a relation between the
first three variables, Factor 2 shows a relation between the last three variables and Factor
3 also shows a relation between the last two variables. These loadings make
interpreting the results more difficult.
For this reason, we decided to run a Factor analysis with four Factors to see if we can have
a better interpretation.
After analyzing the table below we noticed an overall higher amount of variance in the
majority of variables, especially “review_scores_accuracy”, “review_scores_cleanliness”
and “review_scores_location”, all very close to 1.
As mentioned previously, this is ideal and will give better results.
Considering that this value should be close to 0.05 we noticed that Root Mean Square
Off-Diagonal Residuals is lower and this will improve the significance of our analysis.
Upon analyzing the table we pointed out that Factor 1 highlights a relation between the
first three variables, Factor 2 shows a relation between the last three variables, Factor 3
shows a relation between the third last and second the last variable, and Factor 4 will
reflect only the last variable.
4. step: Results
The amount of information (variance) retained per original variable is higher, with all
original variables retaining more than 50% (Table 8).
Also, the Root Mean Square Off-Diagonal Residuals was better in comparison as a
three-factor analysis, decreasing the overall amount to 1.4 does not make for a good
solution. Besides the Root Mean Square and final communality improved compared with
the 3 factors, the loadings are hard to interpret because the same variable is significant in
more than one factor and it does not make a lot of sense to decrease from six original
variables to four factors.
For this reason, we agreed to carry on our analysis focusing on three Factors.
The purpose of using the Orthogonal Varimax method is to maximize the variance of
loading in each Factor, producing, for example, Factors composed by high and low
loadings:
As we can observe in the loadings table above, the results are extremely similar to the
previous rotation. This reinforces the same conclusion that we reached regarding the
Orthogonal Varimax rotation that we performed previously regarding what information
the new three Factors contain.
The Oblique Promax rotation allows Factors to be correlated. It is considered the more
realistic rotation because usually, in the context of marketing problems, the constructs are
correlated.
Considering the Promax rotation approach, this rotation starts with a Varimax prerotation
and, in a second step, the orthogonality is relaxed allowing the Factors to correlate
between each other.
Table 13 shows similar loadings compared with the Orthogonal rotations ( same variables
by factor and no improvements on loadings values).
This means that there is no significant difference between rotations types and all of
them are suitable to run the analysis.
However, we decided to use the Oblique rotation as we are under a marketing research
problem and we assume that variables could be correlated.
According to the results shown in the table, we highlighted the variables that underline
the most significant loading for each retained Factor.
Please note that we consider significant any Factor loadings equal or larger than 0.5.
Factor 1:
We confirm that the following variables “Review_scores_communication” and
“Review_scores_value” are highly positively correlated showing communality between
customers leaving a high review in communication and customers leaving a high review
related to the value score variable too.
Factor 2:
We confirm that the following variables “Review_scores_accuracy” and
“Review_scores_location” are highly correlated showing communality between
customers leaving a high review in accuracy and customers leaving a low review related
to the location score variable too.
Factor 3:
We confirm that the following variables “Review_scores_cleanliness” and
“Review_scores_checkin” are highly correlated showing communality between
customers leaving a high review in cleanliness and customers leaving a low review related
to the check-in score variable too.
After our results, we are able to answer the research problem “Could we find some
hidden Factors of customer satisfaction?”. We argue that there are three hidden
Factors of customer satisfaction, identified as Perceived Quality, Reliability, and
Maintenance Service.
5. Second analysis
For the purpose of this research, we decided to run a Cluster analysis to combine
observations into similar groups. The goal of our analysis is to cluster observations and
categorize them using the qualitative variables listed below.
Our goal is to identify different clusters and name them by customer profile. In addition,
we aim to find homogeneous clusters regarding satisfaction perceived based on the
factors previously identified.
Finally, when it comes to the city, price, and typology of Airbnb we want to highlight
possible trends related to the clusters identified.
We decided to run the cluster analysis, using the Factors selected in the factor analysis
performed before with standardized data, in order to work with a lower number of
variables and eliminate problems of multicollinearity as factors composed of a larger
number of variables will have more weight in forming clusters when using original data.
Choose a method
The advantage of using this method is that we do not need to choose the number of
clusters a priori, something we have to do when it comes to the K-Means method
(Non-Hierarchical). The Hierarchical methods include the following methods listed below.
Considering the analysis purpose, we will analyze the following criteria to understand how
many clusters each method will determine within the Hierarchical method:
● R-squared: This value is a measure of the proportion of the total variance that is
retained in each of the solutions. In other words, it tells us how much Clusters are
different between themself (but similar inside).
● Cubic Clustering Criterium (CCC): this criterion helps to find a solution that
maximizes the criterion; according to this higher values indicate better clustering
and we would accept values larger than 3.
● Semipartial R-Square: we aim to find a low result that measures the loss of
homogeneity by joining two clusters (reduction in R-squared when combining two
clusters).
● Pseudo F Statistics: we look for a solution with a large value for this statistic.
Type of Methods
Based on the Loadings analysis the threshold in the Semipartial R-Square and Pseudo F
Statistics is 9 clusters (Table 14), the threshold in CCC is 4 clusters and the threshold in
R-Square is 5 clusters. Regarding the Dendogram, using the code provided in class, it
showed a horizontal representation that made the interpretation difficult, so we chose to
use only the table above as an indication of how many clusters to consider. Because we
have 2 groups of criteria supporting different results (Semipartial R-Square + Pseudo F
Statistics as a group and CCC + R-Square as another group) we consider using a
maximum of 9 clusters and a minimum of 4 clusters through this method.
Based on the analysis of the Loadings the threshold in the Semipartial R-Square,
R-Square, and Pseudo F Statistics are 4 clusters (Table 15) and the threshold in CCC is 9.
Regarding the Dendogram, using the code provided in class, it showed a horizontal
representation that made interpretation difficult, so we chose to use only the table above
as an indication of how many clusters to consider. Because a majority of criteria (3 out of
4) indicate that we use 4 clusters we would consider this to be the number of clusters to
carry on for the next step of our analysis.
● Average linkage (average of groups) This criterion measures the average of the
distance between all pairs of observations:
Based on the analysis of the Loadings the threshold in the Semipartial R-Square is
composed of 8 clusters, while for R-Square Pseudo-F Statistics and CCC are 7 clusters
(Table 16). Regarding the Dendogram, we determine 9 clusters.
Because a majority of criteria (3 out of 4) indicate that we use 7 clusters we would consider
this to be the number of clusters to carry on for the next step of our analysis.
● Centroid method
The centroid method is a criterion that measures the distance between the mean vectors
(centroids) of the clusters but it is less sensitive to outliers compared with other methods
and it could not work well when clusters have very different sizes. We ran this method
and here are the results:
The threshold in the Semipartial is 8 clusters, in R-Square is 7 clusters, and in CCC and
Pseudo F Statistics are 6 clusters (Table 18) indicating the existence between 8 and 6
clusters. The Dendrogram (Table 19) shows results hard to interpret, so we decided not to
use this criterion to identify the number of clusters. Taking all the criteria results, we
would consider using a maximum of 8 clusters and a minimum of 6 clusters.
● Ward method
The Ward method tends to cluster with few observations and produce clusters with a
similar number of observations and is also very sensitive to outliers. We ran this method
and here are the results:
The threshold in the SemiPartial R_Square is 8 clusters, while in R-Square is 5 clusters, and
the threshold in CCC and Pseudo F Statistics is 4 clusters (Table 20). Regarding the
Dendogram, it shows the existence of 4 clusters (Table 21). Because a majority of criteria (3
out of 5) indicate that we use 4 clusters we would consider this to be the number of
clusters to carry on for the next step of our analysis.
According to the results determined on the Hierarchical method, we are going to try
testing the Non-hierarchical method by starting with 7 clusters and reducing the number
of them until we find the most suitable solution, trying with 6 and 5 clusters too.
The reason why we decided to start with 7 clusters is that 7 is the maximum common
number of clusters supported by the simple, average, and centroid method.
The main advantage of the K-Means method is the possibility to gather better results and
assign observations to different clusters, however, we need to decide on an Nº of clusters a
priori of running the analysis.
● Frequency
This value tells us how many observations we have inside each cluster as our goal
is to create clusters with a homogenous number of observations.
Results appear hard to analyze as 7 clusters have half of the observations
compared with cluster 1.
● RMSSTD
The standard deviation shows a low homogeneity amongst the clusters, with
almost all clusters above 0.5, with clusters 2 and 6 above 0.6, except for clusters 3
and 7 showing the most homogenous results close to 0.50.
● R-squared
The R-Squared shows very good results with results between 0 and 1.
In addition, we can notice that factor 3, within the highest result of 0.71, is the most
heterogeneous factor, and clusters are not very similar.
● Frequency
We noticed an overall good distribution of observations with 6 clusters
although the frequency of clusters 2 and 4 is lower compared with the rest of the
clusters that seem homogenous.
● RMSSTD
The standard deviation shows a low homogeneity amongst the clusters, with all
clusters above 0.55, and with clusters 5 and 6 very close to 0.65, compared with the
rest of the clusters that are more homogeneous.
● R-squared
The R-Squared shows very good results with an overall of 0.63.
We notice that in Factor 2 clusters are more similar.
● Frequency
With 5 clusters we have a good distribution of observations among clusters, with a
minimum of 16 and a maximum of 26.
● RMSSTD
The standard deviation shows a low homogeneity amongst the clusters, with
almost all clusters above 0.6, with clusters 2 and 4 very close to 0.7, except for
cluster 1 that has 0.58 (it is pretty close) which is the most homogeneous.
● R-squared
The R-Squared shows very good results with an overall of 0.6. It also shows that the
factors are more different regarding Factor 1 (highest R-Squared with 0.64) which
means that it is the most discriminant one as well.
Results
We obtain better results in terms of the frequency with the lowest amount of clusters,
although we achieve a better R-squared result with a higher amount of clusters.
For the purpose of our analysis, we decided to choose 5 clusters because we get a much
better frequency and the R-squared does not decrease a lot.
Based on the results of the cluster mean table (Table 28) we focused on naming each
cluster defining the profiling.
According to the value highlighted above please find below the list of our cluster profiling:
● Cluster 1:
People who do not care about Perceived Quality but love Location and Check-In.
Although we were unable to categorize this cluster by a specific preferred city destination
we notice that this cluster would pay the average highest value for the requirements
desired and they tend to usually prefer Airbnb facilities with 2 bedrooms.
We have labeled them Corporate Guests.
● Cluster 2:
People who do not care about Perceived Quality but love Description Accuracy. New York,
Paris, and Rio de Janeiro are the most commonly selected destinations and overall they
would accept to pay an average high price, usually preferring Airbnb with 1 bedroom only.
We have labeled them Backpacker Guests.
● Cluster 3:
People who love Perceived Quality and Check-In and they are attentive to prices, showing
an average that is way below compared to other clusters. We pointed out Paris as a
destination trend.
We have labeled them Romantic Guests.
● Cluster 4:
People who love Perceived Quality, Accuracy, and Cleanliness, the reason why they seem
less price-sensitive compared with other clusters. They tend to book Airbnb facilities with
around 2 bedrooms and we did not notice any predominant preference when it comes to
city destinations.
We have labeled them Affluent Guests.
● Cluster 5:
People who love Location and Cleanliness, although it is the most price-sensitive group of
guests, and they usually prefer facilities with one bedroom only. Paris seems to be the
most popular destination for them.
We have labeled them Family Guests.
To increase customer satisfaction, AirBnB hosts need to focus their service on the 3
hidden factors discovered: Quality Perceived, Reliability, and Maintenance Service.
Good communication with the tenant, defining a good perceived price, the house
location, and cleanliness are 4 important attributes to achieve good ratings.
However, hosts should improve descriptions, in terms of accuracy, and check-in time as
those aspects could negatively impact the ratings of reliability and maintenance service.
Moreover, it is crucial to provide accurate site descriptions as guests could be
disappointed by discovering that their holiday facility is very far from the city center.
Also, having a faster cleanliness process by not losing quality, is an important aspect in
order to increase the efficiency in the check-in service.
According to the factors determined, we were able to categorize our guests into 5 clusters
based on their preferences.
If you want to invest or improve your rental house management bear in mind that:
PRICE
Corporate, Affluent Guests, and Backpacker Guests are the less price-sensitive clusters,
willing to pay more for a rental.
PLACEMENT
Bangkok, New York, Istanbul, Paris, Rio de Janeiro, Rome, and Sydney are cities where all
types of customers travel and rent an Airbnb. While Paris, New York, Bangkok, and Rio are
the ones with a higher number of rentals, it is interesting to point out Paris as the most
preferred destination by Backpacker, Romantic Guests, and Families. We may
recommend Paris to be considered as a location for a possible new business investment,
focusing on Backpackers as a core target due to their insensitivity to price.
PRODUCT
Corporate, Romantic, and Affluent Guests are looking for a bigger Airbnb typology,
while if your Airbnb is smaller you should focus on Backpackers and Family guests.
It is true that Family Guests usually have a higher number of people traveling, however,
be aware they are very sensitive to price, looking for the cheapest solutions. This is
PROMOTION
Promote your Airbnb according to the customer type which is more suitable for your
rental. For example, as Backpacker Guests usually prefer solo trips, booking one bedroom
only, We may focus on this group to run a targeted price promotion campaign in order to
acquire new customers and drive retention, due to their insensitivity to price.
Also, we recommend taking into consideration the reliability aspect: for example by
providing detailed house descriptions and accurate information about the facility location.