Statistics

Statistics is a broad field that draws on applied mathematics to identify patterns in data and to quantify how much confidence those patterns should be treated with. Its methods can be organised by the question each answers: how data should be collected, what the collected data look like, how uncertainty can be modelled, what can be concluded about a wider population, and how variables relate to one another.

While I am not going to cover all statistical aspects, I aim to give a general understanding of the field and describe some foundational techniques.

Study Design

  • Study design concerns how data is gathered, which determines what conclusions the data can later support.

  • In a randomised experiment, the researcher selects participants and assigns them by chance to a treatment group or a control group.

  • The treatment group receives the intervention, while the control group does not. Because randomisation balances other influences across the two groups, any difference in outcomes can be attributed to the intervention.

  • An observational study, by contrast, records data without intervention and can establish association but not causation. For example, a survey may find that coffee drinkers report higher productivity, but without random assignment it cannot rule out that productive people simply drink more coffee.

Descriptive Statistics

  • Descriptive statistics summarises the data at hand.

  • Measures of central tendency, such as the mean and median, describe a typical value, while measures of dispersion, such as the standard deviation, describe how widely values spread around it.

  • Two classes may share a mean exam score of 70, yet one may have scores concentrated between 65 and 75 while the other ranges from 40 to 100. In other words, no single measure characterises a distribution fully.

Probability

  • Probability provides the mathematical language for uncertainty, assigning each event a value between 0, impossible, and 1, certain.

  • Probability distributions, such as the normal and binomial distributions, formalise how the outcomes of a random process are expected to vary: a fair coin has a probability of 0.5 of landing heads, and the binomial distribution specifies how likely 7 heads in 10 flips would be.

  • Crucially, probability is the foundation on which all inferential methods are built.

Statistical inference

  • Statistical inference uses a sample to draw conclusions about the population it came from. Estimation produces a value for an unknown quantity along with a confidence interval, a range expressing the precision of that estimate.

  • Hypothesis testing assesses whether an observed effect could plausibly have arisen by chance: a null hypothesis of no effect is assumed, and a p-value gives the probability of a result at least as extreme under that assumption.

    • A p-value of 0.03 suggests such a result would occur only 3% of the time if no true effect existed, although it says nothing about the size or practical importance of the effect.

  • These methods belong to the frequentist approach, which defines probability as how often an outcome would occur if the same study were repeated many times.

    • For example, a 95% confidence interval means that if the study were repeated 100 times, about 95 of the resulting intervals would contain the true value.

  • The alternative, the Bayesian approach, treats probability as a belief that is updated as new information arrives.

    • For example, the probability of rain today might start at 0.3. Once we learn that it rained yesterday, that probability is updated to account for the new information, written as P(rain today | rain yesterday).

Connect with Me!

LinkedIn: caroline-rennier

Email: caroline@rennier.com