Probability & Distributions
Probability is essential to statistics because almost every conclusion drawn from data carries some uncertainty. A sample never perfectly represents the population it came from, so any estimate or test result could differ from the true value by chance.
Probability measures how large that chance is, which allows a researcher to state not only what the data suggest but also how confident they can be in that conclusion. Without it, there would be no way to distinguish a genuine effect from random variation.
Basic Probability
Probability measures how likely an event is to occur, expressed as a value between 0 and 1. A probability of 0 means the event is impossible, and a probability of 1 means it is certain.
When all outcomes are equally likely, the probability of an event is the number of outcomes that produce it divided by the total number of possible outcomes.
For example, a standard die has six equally likely outcomes, so the probability of rolling a 4 is 1/6, or about 0.17.
A few basic rules follow from this definition:
The probabilities of all possible outcomes always sum to 1, so the probability that an event does not occur is 1 minus the probability that it does: the probability of not rolling a 4 is 5/6.
Two events are independent when the outcome of one does not affect the other, as with two separate coin flips.
For independent events, the probability that both occur is the product of their individual probabilities, so the probability of flipping heads twice in a row is 0.5 × 0.5 = 0.25.
Probability Distributions
A single probability describes one event, such as the probability of rolling a 4. A probability distribution extends this idea by assigning a probability to every possible outcome at once, using the same basic rules.
Example: Rolling two dice and adding their values is a simple example of how basic probabilities form a distribution.
Each die has six equally likely faces, so there are 6 × 6 = 36 equally likely combinations, each with a probability of 1/36.
The possible totals range from 2 to 12, but they are not equally likely, because some totals can be produced by more combinations than others.
A total of 2 occurs only one way (1 and 1), so its probability is 1/36, or about 0.03.
A total of 7 occurs six ways (1 and 6, 2 and 5, 3 and 4, 4 and 3, 5 and 2, 6 and 1), so its probability is 6/36, or about 0.17.
Calculating the probability of every total in this way produces the full distribution.


The distribution peaks at 7, the most likely total, and falls symmetrically towards 2 and 12, the least likely totals. All eleven probabilities sum to 1, confirming that the distribution accounts for every possible outcome.
Central Limit Theorum
The central limit theorem describes what happens when you add or average many random values: the result tends to form a bell-shaped curve, known as the normal distribution.
The dice example shows this clearly. With one die, every number from 1 to 6 is equally likely, so the graph is flat. With two dice, the totals form a triangle that peaks at 7. With ten dice, the totals form a smooth bell curve, with most results near the middle and very few at the extremes.
The same pattern applies to averages. If you take a large enough sample of almost anything and calculate its average, that average will follow a bell curve, even if the original data do not. This is why the normal distribution appears so often in statistics.


Discrete vs Continuous Distributions
Probability distributions fall into two types, depending on the kind of data they describe.
Discrete distributions describe outcomes that can be counted, such as the roll of a die or the number of customers who make a purchase. Each outcome has its own probability, so discrete distributions are drawn as separate bars, and the probabilities of all the bars add up to 1. Such as rolling a dice.
Continuous distributions describe measurements that can take any value within a range, such as height or time. Because there are no gaps between possible values, they are drawn as a smooth curve, and probability is measured as the area under the curve across a range of values. Such as height and weight.


Connect with Me!
LinkedIn: caroline-rennier
Email: caroline@rennier.com