Module-5 Statistical Physics for Computing
Statistical physics is a branch of physics that evolved from a foundation of statistical mechanics, which
uses methods of probability theory and statistics, particularly the mathematical tools for dealing with
large populations and approximations, in solving physical problems.
Descriptive statistics: The term “descriptive statistics” refers to summarizing and organizing the
characteristics of a data set. A data set is a collection of responses or observations from a sample or
entire population.
In quantitative research, after collecting data, the first step of statistical analysis is to describe
characteristics of the responses, such as the average of one variable (e.g., age), or the relation between
two variables (e.g., age and creativity).
Descriptive statistics comprises three main categories – Frequency Distribution, Measures of Central
Tendency, and Measures of Variability.
Inferential Statistics is a method that allows us to use information collected from a sample to make
decisions, predictions, or inferences from a population. The major inferential statistics are based on
statistical models such as Analysis of Variance, chi-square test, student’s t distribution, regression
analysis, etc.
Methods of inferential statistics:
Estimation of parameters
Testing of hypothesis
A variable is any property, characteristic, number, or quantity that increases or decreases over time or
can take on different values in different situations.
Variables in a programming language are names given to computer memory locations in order to store
data in a program. This data can be known or unknown based on the assignment of value to the
variables.
Discrete Variable
A discrete variable is a mathematical term used to describe variable that can only take on a finite
number of values.
Examples
*The number of accidents in the twelve months.
*The number of mobile cards sold in a store within seven days.
*The number of patients admitted to a hospital over a specified period.
Continuous Variable
A continuous variable may take on an infinite number of intermediate values along a specified interval.
Examples are:
*The sugar level in the human body;
*Blood pressure reading;
*Temperature;
*Height or weight of the human body;
*Rate of bank interest;
Poisson Distribution
A Poisson distribution is a discrete probability distribution. It gives the probability of an event
happening a certain number of times (k) within a given interval of time or space. A Poisson distribution
measures how many times an event is likely to occur within an “x” period of time.
Probability mass function
A Poisson distribution is a discrete probability distribution, meaning that it gives the probability of a
discrete (i.e., countable) outcome. For Poisson distributions, the discrete outcome is the number of times
an event occurs, represented by k. The Poisson distribution has only one parameter, λ (lambda), which is
the mean number of events.
A discrete Radom variable X is said to have a Poisson distribution, with parameter, if it has a
probability, Mass Function is given by
𝜆𝐾 𝑒 −𝜆
𝑓(𝑘, 𝜆) = 𝑃(𝑋 = 𝐾) =
𝐾!
Here k is the number of occurrences,
e is Euler’s Number,
! Is the factorial function.
The positive real number λ is equal to the expected value of X and also to its Variance.
Example of probability for Poisson distributions
On a particular river, overflow floods occur once every 100 years on average. Calculate the probability
of k = 0, 1, 2, 3, 4, 5, or 6 overflow floods in a 100-year interval, assuming the Poisson model is
appropriate.
Because the average event rate is one overflow flood per 100 years, λ = 1
𝜆𝐾 𝑒 −𝜆
𝑓(𝐾, 𝜆) = 𝑃(𝑋 = 𝐾) =
𝐾!
𝜆𝐾 𝑒 −𝜆 1𝐾 𝑒 −1
P (K overflow floods in 100 years) = 𝐾! = 𝐾!
𝜆𝐾 𝑒 −𝜆 10 𝑒 −1 𝑒 −1
P (K=0 overflow floods in 100 years) = = = = 0.368
𝐾! 0! 1
𝜆𝐾 𝑒 −𝜆 11 𝑒 −1 𝑒 −1
P (K=1 overflow floods in 100 years) = = = = 0.368
𝐾! 1! 1
𝜆𝐾 𝑒 −𝜆 12 𝑒 −1 𝑒 −1
P (K=2 overflow floods in 100 years) = = = = 0.184
𝐾! 2! 2
Proton decay is a rare type of radioactive decay of nuclei containing excess protons, in which a proton is
simply ejected from the nucleus. The mechanism of the decay process is very similar to alpha decay.
Proton decay is also a quantum tunneling process.
Modeling the Probability for Proton Decay
The probability of observing a proton decay can be estimated from the nature of particle decay and the
application of Poisson Statistics. The number of protons N can be modeled by the decay equation
𝑁 = 𝑁0 𝑒 −𝜆𝑡
Where:
N0: is the initial quantity of the element
λ: is the radioactive decay constant
t: is time
N(t): is the quantity of the element remaining after time t.
Here 𝜆 = 1/𝑡 = 10−33 / 𝑦𝑒𝑎𝑟 is the probability that any given proton will decay in a year.
Since the decay constant λ is so small, the exponential can be represented by the first two terms of the
Exponential Series.
𝑒 −𝜆𝑡 = 1 − 𝜆𝑡
𝑁 = 𝑁0 (1 − 𝜆𝑡)
Most recently the experiment on proton decay has been done by Super Kamiokande, Japan which
started observation in 1996. It is a large water Cherenkov detector which is the most sensitive detector
in the world used to examine proton decay with the huge source with 7.5×1033 protons
For one year of observation, the number of expected proton decays is then
𝑁0 − 𝑁 = 𝑁0 (1 − 𝜆𝑡)
= (7.5 × 1033 𝑝𝑟𝑜𝑡𝑜𝑛𝑠)(10−33 / 𝑦𝑒𝑎𝑟)(1 𝑦𝑒𝑎𝑟)
𝑁0 − 𝑁 = 7.5
Proton decay has not been detected experimentally till now probably because of fact that the event is
extremely rare.
Assuming that λ = 3 observed decays per year is mean, then the Poisson distribution function tells us
that the probability for zero observations of decay is
𝜆𝐾 𝑒 −𝜆 30 𝑒 −3
𝑃(𝐾) = = = 0.05
𝐾! 0!
This low probability for a null result suggests that the proposed lifetime of 1033 years is too short.
Normal Distribution: The bell curve is a normal probability distribution of variables plotted on the
graph and is like a bell shape where the highest or top point of the curve represents the most probable
event out of all the series data.
CHARACTERISTICS
1. The Normal Curve is Symmetrical: The normal probability curve is symmetrical around its
vertical axis called ordinate. The symmetry about the ordinate at the central point of the curve
implies that the size, shape, and slope of the curve on one side of the curve is identical to that of
the other. In other words, the left and right halves of the middle central point are mirror images,
as shown in the figure given here.
2. The Normal Curve is Unimodel: Since there is only one maximum point in the curve, thus the
normal probability curve is unimodal, i.e. it has only one mode.
3. The Normal Curve is Bilateral: The 50% area of the curve lies to the left side of the maximum
central ordinate and 50% of the area lies to the right side. Hence the curve is bilateral.
4. The Normal Curve is a mathematical model in behavioral Sciences: This curve is used as a
measurement scale. The measurement unit of this scale is ± 1σ (the unit standard deviation).
Standard Deviations: The standard normal distribution is a normal probability distribution that has a
mean of 0 and a standard deviation of 1.
The Standard Deviation is a measure of how spread out numbers are 68% of values are within 1 standard
deviation of the mean. 95% of values are within 2 standard deviations of the mean. 99.7%of values are
within 3 standard deviations of the mean
Monte-Carlo Method
Monte Carlo Simulation, also known as the Monte Carlo Method or a multiple probability simulation, is
a mathematical technique, which is used to estimate the possible outcomes of an uncertain event.
The Monte Carlo Method was invented by John von Neumann and Stanislaw Ulam during World War II
to improve decision-making under uncertain conditions. It was named after a well-known casino town,
called Monaco.
The statistical method of understanding complex physical or mathematical systems by using randomly
generated numbers as input into those systems to generate a range of solutions.
How to use Monte Carlo methods
1. Set up the predictive model, identifying both the dependent variable to be predicted and the
independent variables that will drive the prediction.
2. Specify probability distributions of the independent variables.
3. Run simulations repeatedly, generating random values of the independent variables. Do this until
enough results are gathered to make up a representative sample of the nearly infinite number of
possible combinations.
Estimation of Pi
The idea is to simulate random (x, y) points in a 2-D plane with the domain as a square of side 2r
units centered on (0,0).
Imagine a circle inside the same domain with the same radius r and inscribed into the square. We
then calculate the ratio of the number of points that lay inside the circle and the total number of
generated points. Refer to the image below:
We know that the area of the circle πr2, while that of square 4r2. The ratio of these two areas is as
follows:
𝑎𝑟𝑒𝑎 𝑜𝑓 𝑡ℎ𝑒 𝑐𝑖𝑟𝑐𝑙𝑒 𝜋r2
= 4𝑟 2
𝑎𝑟𝑒𝑜𝑓 𝑡ℎ𝑒 𝑠𝑞𝑢𝑎𝑟𝑒
𝜋
=4
𝑛𝑜 𝑜𝑓 𝑝𝑜𝑖𝑛𝑡𝑠 𝑔𝑒𝑛𝑒𝑟𝑎𝑡𝑒𝑑 𝑖𝑛𝑠𝑖𝑑𝑒 𝑡ℎ𝑒 𝑐𝑖𝑟𝑐𝑙𝑒 𝜋
Now for a very large number of generated points =4
𝑛𝑜 𝑜𝑓 𝑝𝑜𝑖𝑛𝑡𝑠 𝑔𝑒𝑛𝑒𝑟𝑎𝑡𝑒𝑑 𝑖𝑛𝑠𝑖𝑑𝑒 𝑡ℎ𝑒 𝑠𝑞𝑢𝑎𝑟𝑒
𝑛𝑜 𝑜𝑓 𝑝𝑜𝑖𝑛𝑡𝑠 𝑔𝑒𝑛𝑒𝑟𝑎𝑡𝑒𝑑 𝑖𝑛𝑠𝑖𝑑𝑒 𝑡ℎ𝑒 𝑐𝑖𝑟𝑐𝑙𝑒
4 ∗ 𝑛𝑜 𝑜𝑓 𝑝𝑜𝑖𝑛𝑡𝑠 𝑔𝑒𝑛𝑒𝑟𝑎𝑡𝑒𝑑 𝑖𝑛𝑠𝑖𝑑𝑒 𝑡ℎ𝑒 𝑠𝑞𝑢𝑎𝑟𝑒 = π
Thus, the title is “Estimating the value of Pi” and not “Calculating the value of Pi”. Below is the
algorithm for the method:
The Algorithm
1. Initialize circle points, square points and interval to 0.
2. Generate random point x.
3. Generate random point y.
4. Calculate d = x*x + y*y.
5. If d <= 1, increment circle_points.
6. Increment square_points.
7. Increment interval.
8. If increment < NO_OF_ITERATIONS, repeat from 2.
9. Calculate pi = 4*(circle_points/square_points).
.
References
1. Information Theory and Coding –K Giridhar.
[Link]://[Link]/estimating-value-pi-using-monte-carlo/
3. Lecture Notes on the Gaussian Distribution Hairong Qi
4. Adamas University web sitePROTON DECAY…..A CANDIDATE FOR FINITE AGE OF THE
UNIVERSE? Posted By Aparajita Bhattacharya July 2, 2020.
5. Bhandari, P. (2023, January 09). Descriptive Statistics | Definitions, Types, Examples. Scribbr.
Retrieved March 17, 2023, from [Link]
6. Statistical Physics: Berkely Physics Course, Volume 5, F. Reif, McGraw Hill.