A STATISCAL ANALYSIS OF THE CORRELATION BETWEEN
WIN PROBABLITY AND OUTCOMES OF THE PREMIRE
LEAGUE 2023-24 SEASON
The 2023 Premier League season offered an intriguing opportunity to evaluate the
relationship between pre-season probabilities of winning and actual points earned by
teams. Probability forecasts often serve as predictive tools to gauge team performance,
yet their accuracy and association with real outcomes merit statistical examination. This
study employs various statistical tools to analyze this relationship, focusing on
correlation analysis and the chi-square test while addressing key limitations, such as the
presence of outliers and data distribution.
Why Raw Data is Not Appropriate for Direct Use
The direct use of raw probabilities and points is inappropriate for this analysis due to
two primary reasons:
The Influence of Outliers
Manchester City represents a significant outlier, with an overwhelming probability of
winning (65.4%) compared to other teams, whose probabilities mostly fall below 10%.
Similarly, Manchester City scored 91 points in the league, far exceeding most teams.
Such outliers can disproportionately skew statistical analyses, particularly linear
correlation measures like Pearson’s correlation, leading to biased interpretations
(Howell, 2012).
Non-Normal Data Distribution
The dataset exhibits a non-normal distribution, as probabilities are heavily clustered at
lower ranges and points are concentrated below 70. Statistical tests like Pearson’s
correlation assume normality, and deviations from this assumption reduce the reliability
of the results (Field, 2013). These factors necessitate alternative methods and
adjustments, such as rank-based correlation metrics and visual outlier detection (Siegel
& Castellan, 1988).
Statistical Methodologies
Spearman vs. Pearson Correlation
Correlation analysis was employed to determine the relationship between probabilities
and points, using both Spearman and Pearson correlation coefficients. Spearman’s
correlation is particularly robust to outliers and does not assume a linear relationship
(Siegel & Castellan, 1988). Pearson’s correlation measures linear relationships but is
highly sensitive to outliers, making it less reliable for this dataset (Howell, 2012).
Outlier Identification and Treatment
To address the influence of outliers, visual analysis using box-and-whisker plots was
conducted, revealing Manchester City as a significant outlier in both probabilities and
points. Removing this outlier allowed for recalculating correlations, offering a clearer
view of underlying trends (Barnett & Lewis, 1994).
Chi-Square Test of Independence
The chi-square test was applied to examine the relationship between team rankings
(based on points earned) and match outcomes (win-draw-loss records). This test
evaluates whether observed distributions differ significantly from expected distributions
under the assumption of independence (Agresti, 2013).
Results
Spearman and Pearson Correlation Results
The correlation analysis yielded the following results:
Correlation Method Including Manchester City Excluding Manchester City
Spearman 0.86 (strong) 0.83 (strong)
Pearson 0.50 (moderate) 0.67 (strong)
The Spearman correlation remained strong regardless of whether Manchester City was
included, confirming a robust positive relationship between probabilities and points.
However, Pearson's correlation improved significantly upon excluding Manchester City,
highlighting the outlier's influence. This finding underscores the importance of using
rank-based metrics like Spearman’s for such datasets (Siegel & Castellan, 1988).
Correlation analysis
1
0.9
0.8
0.7
0.6
0.5
0.4
0.3
0.2
0.1
0
0.8 1 1.2 1.4 1.6 1.8 2 2.2
Spearman Pearson
Chi-Square Test Results
The chi-square test examined the association between team rankings and match
outcomes. The calculated chi-square statistic was 29.40, with a p-value of 0.021 and 16
degrees of freedom. The low p-value indicates a statistically significant association,
suggesting that higher-ranked teams (e.g., Manchester City, Arsenal, Liverpool) were
more likely to achieve favorable outcomes (wins) than lower-ranked teams (Agresti,
2013).
Discussion
The analysis reveals a strong positive correlation between pre-season probabilities and
points earned, affirming the predictive validity of probabilities in evaluating team
performance (Lane, 2014). The chi-square results further emphasize the influence of
team rankings on match outcomes. Higher-ranked teams consistently outperform lower-
ranked teams, reflecting the inherent disparities in team quality and resources (Dixon &
Coles, 1997). However, the presence of outliers like Manchester City demonstrates the
need for robust statistical methods that account for such anomalies.
Limitations
Outlier Impact
Although outlier management improved the results, the exclusion of Manchester City
may oversimplify the analysis by ignoring the team's legitimate dominance. Future
studies could explore methods to integrate outliers more effectively without excluding
them entirely (Barnett & Lewis, 1994).
Data Scope
The analysis is limited to the 2023 season and does not account for historical trends or
other influencing factors, such as injuries, transfers, or managerial changes. A broader
dataset could provide a more comprehensive view.
Statistical Assumptions
While Spearman’s correlation mitigates some limitations, the chi-square test assumes
independence among match outcomes, which may not fully capture the complexities of
team interactions and strategies (Agresti, 2013).
Conclusion
This study demonstrates a strong positive relationship between pre-season probabilities
and points earned in the 2023 Premier League season. Spearman’s rank correlation
emerged as the most reliable measure, effectively addressing issues of non-normality
and outliers. The chi-square test further highlighted the significant association between
team rankings and match outcomes, underscoring the predictive power of probabilities.
The findings reinforce the utility of statistical tools in sports analysis, while also
highlighting the challenges posed by outliers and data distribution. Future research
could expand on these insights by incorporating additional variables and exploring
alternative statistical models to enhance the accuracy and applicability of predictions
(Lane, 2014).
References:
[Link]
opta-predictor
[Link]
[Link]
2023-2024-odds-title-winner-epl/e6p6xxtp4xymrigpmclwzxqc
Agresti, A. (2013). Categorical Data Analysis (3rd ed.). Wiley-Interscience.
Barnett, V., & Lewis, T. (1994). Outliers in Statistical Data (3rd ed.). Wiley.
Dixon, M. J., & Coles, S. (1997). Modelling Association Football Scores: A
Probabilistic Analysis of the 1994–1995 Season of the English Football League. Journal
of the Royal Statistical Society: Series D (The Statistician), 46(2), 265-280.
Field, A. (2013). Discovering Statistics Using IBM SPSS Statistics (4th ed.). Sage.
Howell, D. C. (2012). Statistical Methods for Psychology (8th ed.). Cengage
Learning.
Lane, D. (2014). Advanced Sports Analytics: Predictive Models for Football and
Other Sports. Springer.
Siegel, S., & Castellan, N. J. (1988). Nonparametric Statistics for the Behavioral
Sciences (2nd ed.). McGraw-Hill.