standard deviation
The standard deviation is a measure of the amount of variation or dispersion in a set of values. It quantifies how much individual data points in a dataset deviate from the mean (average) of the dataset. A smaller standard deviation indicates that the data points tend to be close to the mean, while a larger standard deviation indicates that the data points are spread out over a wider range of values.
Definition
For a dataset with ( n ) values
and mean
, the standard deviation
is defined as:
![]()
where:
is the mean of the dataset:
![Rendered by QuickLaTeX.com [ \bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_i ]](https://i0.wp.com/science.awjunaid.com/wp-content/ql-cache/quicklatex.com-a9ea3392c74a6b9a17e4e0086adb9e9a_l3.png?resize=115%2C22&ssl=1)
are the individual data points.- ( n ) is the number of data points.
Sample vs. Population Standard Deviation
-
Population Standard Deviation: Used when the dataset includes the entire population. The formula divides by ( n ):
![Rendered by QuickLaTeX.com [ \sigma = \sqrt{\frac{1}{n} \sum_{i=1}^{n} (x_i - \bar{x})^2} ]](https://i0.wp.com/science.awjunaid.com/wp-content/ql-cache/quicklatex.com-56ec232d3f0ef7a6f37a76f252ebaa11_l3.png?resize=184%2C32&ssl=1)
-
Sample Standard Deviation: Used when the dataset is a sample of a larger population. The formula divides by ( n-1 ) (Bessel’s correction) to correct for the bias in the estimation of the population variance:
![Rendered by QuickLaTeX.com [ s = \sqrt{\frac{1}{n-1} \sum_{i=1}^{n} (x_i - \bar{x})^2} ]](https://i0.wp.com/science.awjunaid.com/wp-content/ql-cache/quicklatex.com-67a4039f2e5ecfe08bf5bbb854307beb_l3.png?resize=199%2C32&ssl=1)
where
denotes the sample standard deviation.
Properties
-
Non-Negativity:
- Standard deviation is always non-negative because it is the square root of a non-negative value.
-
Same Units as Data:
- Standard deviation is expressed in the same units as the data, making it interpretable in the context of the original measurements.
-
Sensitivity to Outliers:
- Standard deviation is sensitive to outliers. Extreme values can significantly increase the standard deviation, indicating a larger spread in the data.
-
Relation to Variance:
- The standard deviation is the square root of the variance. Variance is another measure of spread, defined as:
and
![Rendered by QuickLaTeX.com [ \text{Standard Deviation} = \sqrt{\text{Variance}} ]](https://i0.wp.com/science.awjunaid.com/wp-content/ql-cache/quicklatex.com-a8faf62ffb399327c0b4460e9e1a2df4_l3.png?resize=263%2C21&ssl=1)
- The standard deviation is the square root of the variance. Variance is another measure of spread, defined as:
Example Calculation
Consider a dataset:
.
-
Calculate the Mean:
![Rendered by QuickLaTeX.com [ \bar{x} = \frac{5 + 7 + 3 + 9 + 6}{5} = \frac{30}{5} = 6 ]](https://i0.wp.com/science.awjunaid.com/wp-content/ql-cache/quicklatex.com-4315c115ef2028fd9fefdb1a59896ad8_l3.png?resize=195%2C22&ssl=1)
-
Calculate the Variance:
![Rendered by QuickLaTeX.com [ = \frac{20}{5} = 4 ]](https://i0.wp.com/science.awjunaid.com/wp-content/ql-cache/quicklatex.com-482553382c1b8f4495a11f22222180a0_l3.png?resize=75%2C22&ssl=1)
-
Calculate the Standard Deviation:
![Rendered by QuickLaTeX.com [ \text{Standard Deviation} = \sqrt{4} = 2 ]](https://i0.wp.com/science.awjunaid.com/wp-content/ql-cache/quicklatex.com-b33f2ae817a827b07000bbc9954e7828_l3.png?resize=238%2C21&ssl=1)
Applications
-
Descriptive Statistics:
- Standard deviation provides a measure of the spread or variability in a dataset, complementing the mean.
-
Quality Control:
- In manufacturing and quality control, standard deviation is used to monitor the consistency of products and processes.
-
Finance:
- In finance, standard deviation is used to measure the volatility or risk of investment returns.
-
Psychometrics:
- In psychometrics and educational testing, standard deviation is used to interpret test scores and assess the spread of scores within a population.
Summary
The standard deviation is a fundamental statistical measure that quantifies the amount of variation or dispersion in a dataset. It is used in various fields to understand the spread of data, assess consistency, and make informed decisions based on data variability.
Discover more from Science blog by awjunaid
Subscribe to get the latest posts sent to your email.
