Introductory Statistics Explained:
Introduction
©2024 Jeremy Balka
Version 1.11 S24 Draft. Chapter 1: Introduction
“Keep the company of those who seek the truth—run from those who have found
it.”
Vaclav Havel
1
2
1
DESCRIPTIVE STATISTICS
Introduction
The field of statistics is often misunderstood. The study of statistics is not about
learning numerical tidbits and factoids, or learning how to manipulate data to
show what we would like it to show. The often quoted “Lies, damned lies, and
statistics” is typically used dismissively of the field, and is not without an element
of truth, but it is a little like calling a screwdriver a weapon. Sure, a screwdriver
can be used as a weapon, but that is far from its fundamental purpose. When
practiced honestly, the field of statistics is about collecting and analyzing data in
an attempt to learn about an unknown truth. The fact that many people abuse
statistics makes a study of the fundamentals of statistical thought all the more
important.
In statistics, we use data to help us answer questions like:
• Can post-menopausal women lower their heart attack risk by undergoing
hormone replacement therapy?
• Is there a relationship between income and happiness?
• What variables impact length of stay in hospital after a Caesarean section
birth?
To answer these sorts of questions, we will need to find or collect appropriate data.
We must be careful in the planning and data collection process, as sometimes the
data a researcher collects is not appropriate for answering the questions of interest.
(And some questions are, of course, extremely difficult or impossible to answer.)
Once appropriate data has been collected, we summarize and illustrate it with plots
and numerical summaries. Then—ideally—we use the data in the most effective
way possible to address our questions of interest. This will involve some number
crunching, typically carried out by software, but there is much more to it than
that.
Later in the text we will learn to answer research questions like the ones posed
above using statistical inference techniques, but we will first explore the basics
of descriptive statistics.
2
Descriptive Statistics
In descriptive statistics, plots and numerical summaries are used to describe a data
set.
Example 2.1 .
2
3
INFERENTIAL STATISTICS
Coastal Newfoundland is summertime home to
hundreds of thousands of Atlantic puffins, and
puffin viewing is a major tourism draw. As part
of a broader study, a surface count of puffins on
North Bird Island (near Bonavista, NL) was collected once each afternoon for 49 days during a
breeding season.
The counts are illustrated with a histogram and boxplot in Figure 1.12
Surface Count
10
5
0
1000
2000
3000
Surface Count
(a) Histogram of surface counts.
4000
Maximum
3000
2000
1000
0
0
Frequency
15
4000
75th percentile
Median
25th percentile
Minimum
(b) Boxplot of surface counts.
Figure 1: Histogram and an annotated boxplot of Atlantic puffin surface counts
on North Bird Island.
Plots give a helpful visual summary of the data, and numerical summary statistics,
such as the mean, median, and standard deviation are also useful. We will investigate descriptive statistics in greater detail in Chapter ??. But the main purpose
of this text is to introduce statistical inference concepts and methods.
3
Inferential Statistics
The most interesting statistical problems involve investigating relationships between variables. Let’s look at some examples of the types of problems we will
encounter.
Example 3.1 A study from the era of leaded gasoline investigated whether traffic
police officers in Cairo had higher levels of lead in their blood than officers that
worked in suburban Cairo offices.3 The researchers drew random samples of 126
1
Data collected by JB. I’m opening with this example because this place and the puffins are close
to my heart.
2
Boxplots will be discussed in detail in Section ??. In their simplest form, they are plots of the
minimum, 25th percentile, median, 75th percentile, and the maximum. If there are outliers,
they are plotted individually. Boxplots are most useful for comparing two or more groups.
3
3
INFERENTIAL STATISTICS
Cairo traffic officers and 50 officers from the suburbs, and measured their blood
lead levels (µg/dL).
The boxplots in Figure 2 illustrate the data. Boxplots are very useful for comparing the distribution of a quantitative variable between groups.
Blood Lead Level (μg /dL)
50
40
30
20
10
0
Cairo Traffic
Suburbs
Figure 2: Boxplots of blood lead concentration in 126 Cairo traffic officers and 50
officers from the suburbs.
The boxplots show a difference in the distributions; it appears as though the distribution of lead in the blood of Cairo traffic officers is shifted higher than that of
officers from the suburbs. In other words, it appears as though the traffic officers
tend to have a higher blood lead level. Table 1 shows some summary statistics.4
Number of observations
Mean
Standard deviation
Cairo
126
29.2
7.5
Suburbs
50
18.2
5.8
Table 1: Summary statistics for the blood lead level data.
In scenarios like this there are often two main points of interest:
1. Estimating the difference in mean blood lead levels between the two groups.
2. Investigating whether there is strong evidence that the observed difference in
blood lead levels is a real effect, and not simply due to random variability.5
3
Kamal, A., Eldamaty, S., and Faris, R. (1991). Blood level of Cairo traffic policemen. Science
of the Total Environment, 105:165–170. The data used in this text is based on the summary
statistics and histograms in that study.
4
The standard deviation is a measure of the variability of the data. We will discuss it in detail
in Section ??.
5
We sometimes use the phrase statistical significance to indicate strong evidence of a real effect.
This term has a precise, technical meaning that we will discuss in detail later.
4
3
INFERENTIAL STATISTICS
Later in this text we will learn about confidence intervals and hypothesis tests;
these are statistical inference methods that will allow us to formally address these
major points of interest.
Example 3.2 Can self-control be restored during intoxication? Researchers investigated this in a designed experiment involving 44 male undergraduate student
volunteers.6 The volunteers were randomly assigned to one of 4 treatment groups
(11 to each group):
1. An alcohol group, receiving drinks containing a total of 0.62 mg/kg alcohol.
(Group A)
2. A group receiving drinks with the same alcohol content as Group A, but also
containing 4.4 mg/kg of caffeine. (Group AC)
3. A group receiving drinks with the same alcohol content as Group A, but also
receiving a monetary reward for success on the task. (Group AR)
4. A group told they would receive alcoholic drinks, but instead given a placebo
(a drink containing a few drops of alcohol on the surface, and misted to give
a strong alcoholic scent). (Group P)
After consuming the drinks and resting for a few minutes, the participants carried
out a word stem completion task involving “controlled (effortful) memory processes.” Figure 3 shows the boxplots for the four treatment groups. Higher scores
are indicative of greater self-control.
Score on Task
0.6
0.4
0.2
0.0
-0.2
-0.4
A
AC
AR
P
Treatment Group
Figure 3: Boxplots of word stem completion task scores.
The plot seems to show systematic differences between the groups in their task
scores. Is there strong evidence these are real effects? Does this data give evidence
6
Grattan-Miscio, K. and Vogel-Sprott, M. (2005). Alcohol, intentional control, and inappropriate
behavior: Regulation by caffeine or an incentive. Experimental and Clinical Psychopharmacology,
13:48–55.
5
3
INFERENTIAL STATISTICS
that a negative impact of alcohol on self-control can be overcome with the use of
caffeine or a monetary reward? In Chapter ?? we will use a statistical inference
method called one-way ANOVA to help answer these questions.
Shell Thickness (mm)
Example 3.3 A study investigated a possible relationship between eggshell thickness and a variety of environmental contaminants in brown pelican eggs.7 It was
suspected that higher levels of contaminants would result in thinner eggshells. One
contaminant was DDT, measured in parts per million of the yolk lipid. Figure 4
shows a scatterplot of shell thickness vs. DDT in a sample of 65 brown pelican
eggs from Anacapa Island, California.
0.5
0.4
0.3
0.2
0.1
0
500
1000
1500
2000
2500
3000
DDT (ppm)
Figure 4: Shell thickness (mm) vs. DDT (ppm) for 65 brown pelican eggs.
There appears to be a decreasing trend. Does this data give strong evidence of
a relationship between DDT contamination and eggshell thickness? Can we come
up with a statistical model to estimate the relationship? In Chapter ?? we will
use regression analysis to help address these types of questions.
Example 3.4 A study investigated various aspects of oxygen uptake and resting
energy expenditure in overweight subjects.8 One part of the study explored the
relationship between resting energy expenditure and total organ weight (estimated
via an MRI). Figure 5 illustrates the results for 57 females and 14 males.
Is there a difference in resting energy expenditure between males and females after
accounting for organ weight? This is a common type of research question, and
one of the most meaningful things a statistical analysis can do for us is allow us
to explore relationships between variables after adjusting for the effects of other
variables. For example, we may wish to:
7
Risebrough, R. (1972). Effects of environmental pollutants upon animals other than man. In
Proceedings of the Sixth Berkeley Symposium on Mathematical Statistics and Probability.
8
Pourhassan et al. (2015). Relationship between submaximal oxygen uptake, detailed body composition, and resting energy expenditure in overweight subjects. American Journal of Human
Biology, 27:397–406.
6
INFERENTIAL STATISTICS
Females
Males
6
5
3
4
REE (kJ/min)
7
8
3
3.0
3.5
4.0
4.5
5.0
5.5
6.0
Organ weight (kg)
Figure 5: Resting energy expenditure versus total organ weight for 71 individuals.
• Estimate differences in post-operative outcomes between smokers and nonsmokers, after adjusting for age and various health indicators.
• Investigate whether minorities tend to be underpaid in an occupation, after adjusting for variables such as education, experience, and performance
reviews.
We often use multiple regression analysis to build statistical models that can
address these sorts of questions.
In the world of statistics—and our world at large for that matter—we rarely know
anything with certainty. Our statistical interpretations and conclusions will involve
a measure of the reliability of our estimates. We will make statements like we can
be 95% confident that the true difference in means lies between 4 and 10, or, the
probability of seeing the observed difference, if in reality the new drug has no effect,
is less than 0.001. So probability plays an important role in inferential statistics,
and we will study probability in Chapter ??.
Before we get to our study of formal statistical inference, we will first need to
discuss a number of related topics, including descriptive statistics, probability, and
fundamental concepts related to how we obtain data to address our questions of
interest, and how those data collection methods impact our interpretations and
conclusions. The following chapter discusses some of these fundamental concepts
in data collection.
7
REFERENCES
References
Grattan-Miscio, K. and Vogel-Sprott, M. (2005). Alcohol, intentional control, and
inappropriate behavior: Regulation by caffeine or an incentive. Experimental
and Clinical Psychopharmacology, 13:48–55.
Kamal, A., Eldamaty, S., and Faris, R. (1991). Blood level of Cairo traffic policemen. Science of the Total Environment, 105:165–170.
Pourhassan et al. (2015). Relationship between submaximal oxygen uptake, detailed body composition, and resting energy expenditure in overweight subjects.
American Journal of Human Biology, 27:397–406.
Risebrough, R. (1972). Effects of environmental pollutants upon animals other
than man. In Proceedings of the Sixth Berkeley Symposium on Mathematical
Statistics and Probability.
8