FULL ARCHIVE
320+ Data Science posts
580+ pages
[Link]
[Link]
Table of Contents
The Must-Know Categorisa3on of Discrimina3ve Models ................................12
Where Did The Regulariza3on Term Originate From? ......................................18
How to Create The Elegant Moving Bubbles Chart in Python? .........................22
Gradient Checkpoin3ng: Save 50-60% Memory When Training a Neural
Network ..........................................................................................................24
Gaussian Mixture Models: The Flexible Twin of KMeans ..................................28
Why Correla3on (and Other Summary Sta3s3cs) Can Be Misleading ...............33
MissForest: A Be[er Alterna3ve To Zero (or Mean) Imputa3on .......................35
A Visual and Intui3ve Guide to The Bias-Variance Problem ..............................39
The Most Under-appreciated Technique To Speed-up Python ...........................41
The Overlooked Limita3ons of Grid Search and Random Search ......................44
An Intui3ve Guide to Genera3ve and Discrimina3ve Models in Machine
Learning ..........................................................................................................48
Feature Scaling is NOT Always Necessary ........................................................55
Why Sigmoid in Logis3c Regression? ...............................................................58
Build Elegant Data Apps With The Coolest Mito-Streamlit Integra3on .............62
A Simple and Intui3ve Guide to Understanding Precision and Recall ................64
Skimpy: A Richer Alterna3ve to Pandas' Describe Method ...............................69
A Common Misconception About Model Reproducibility ................................71
The Biggest Limita3on Of Pearson Correla3on Which Many Overlook .............76
Gigasheet: Effortlessly Analyse Upto 1 Billion Rows Without Any Code ............78
Why Mean Squared Error (MSE)? ....................................................................82
A More Robust and Underrated Alterna3ve To Random Forests ......................90
The Most Overlooked Problem With Impu3ng Missing Values Using Zero (or
Mean) .............................................................................................................93
A Visual Guide to Joint, Marginal and Condi3onal Probabili3es.......................95
Jupyter Notebook 7: Possibly One Of The Best Updates To Jupyter Ever ...........96
How to Find Op3mal Epsilon Value For DBSCAN Clustering? ............................97
Why R-squared is a Flawed Regression Metric .................................................99
75 Key Terms That All Data Scien3sts Remember By Heart.............................102
1
[Link]
The Limita3on of Sta3c Embeddings Which Made Them Obsolete .................109
An Overlooked Technique To Improve KMeans Run-3me ...............................116
The Most Underrated Skill in Training Linear Models .....................................119
Poisson Regression: The Robust Extension of Linear Regression .....................125
The Biggest Mistake ML Folks Make When Using Mul3ple Embedding Models
.....................................................................................................................126
Probability and Likelihood Are Not Meant To Be Used Interchangeably .........129
SummaryTools: A Richer Alterna3ve To Pandas' Describe Method. ................135
40 NumPy Methods That Data Scien3sts Use 95% of the Time .......................136
An Overly Simplified Guide To Understanding How Neural Networks Handle
Linearly Inseparable Data..............................................................................138
2 Mathema3cal Proofs of Ordinary Least Squares .........................................145
A Common Misconcep3on About Log Transforma3on ...................................146
Raincloud Plots: The Hidden Gem of Data Visualisa3on .................................149
7 Must-know Techniques For Encoding Categorical Feature ...........................153
Automated EDA Tools That Let You Avoid Manual EDA Tasks .........................154
The Limita3on Of Silhoue[e Score Which Is Ojen Ignored By Many ..............156
9 Must-Know Methods To Test Data Normality ..............................................159
A Visual Guide to Popular Cross Valida3on Techniques ..................................163
Decision Trees ALWAYS Overfit. Here's A Lesser-Known Technique To Prevent It.
.....................................................................................................................167
Evaluate Clustering Performance Without Ground Truth Labels .....................169
One-Minute Guide To Becoming a Polars-savvy Data Scien3st .......................172
The Most Common Misconcep3on About Con3nuous Probability Distribu3ons
.....................................................................................................................174
Don't Overuse Sca[er, Line and Bar Plots. Try These Four Elegant Alterna3ves.
.....................................................................................................................175
CNN Explainer: Interac3vely Visualize a Convolu3onal Neural Network .........178
Sankey Diagrams: An Underrated Gem of Data Visualiza3on ........................180
A Common Misconcep3on About Feature Scaling and Standardiza3on ..........181
7 Elegant Usages of Underscore in Python .....................................................184
Random Forest May Not Need An Explicit Valida3on Set For Evalua3on ........185
Declu[er Your Jupyter Notebook Using Interac3ve Controls ..........................188
Avoid Using Pandas' Apply() Method At All Times .........................................190