Chapter One
What is Statistics?
1
What is Statistics?
“Statistics is a way to get information from data.”
2
What is Statistics?
“Statistics is a way to get information from data”
Statistics
Data
Information
Statistics is a tool for creating new understanding from a set of numbers.
Definitions: Oxford English Dictionary
3
Key Statistical Concepts
Population
— a population is the group of all items of interest to
a statistics practitioner.
— frequently very large; sometimes infinite.
E.g., All 5 million Florida voters, per Example 12.5
Sample
— A sample is a set of data drawn from the
population.
— Potentially very large, but less than the population.
E.g., a sample of 765 voters exit poll on election day.
4
Key Statistical Concepts
Parameter
— A descriptive measure of a population.
Statistic
— A descriptive measure of a sample.
5
Key Statistical Concepts
Population
Sample
Subset
Parameter
Statistic
Populations have Parameters,
Samples have Statistics.
6
Descriptive Statistics
…are methods of organizing, summarizing, and presenting
data in a convenient and informative way. These methods
include:
Graphical Techniques (Chapters 2 & 3), and
Numerical Techniques (Chapter 4).
The actual method used depends on what information we
would like to extract. Are we interested in…
• measure(s) of central location? and/or
• measure(s) of variability (dispersion)?
Descriptive Statistics helps to answer these questions…
7
Inferential Statistics
Descriptive Statistics describe the data set that’s being
analyzed, but doesn’t allow us to draw any conclusions or
make any interferences about the data. Hence we need
another branch of statistics: inferential statistics.
Inferential statistics is also a set of methods, but it is used to
draw conclusions or inferences about characteristics of
populations based on data from a sample.
8
Statistical Inference
Statistical inference is the process of making an estimate,
prediction, or decision about a population based on a sample.
Population
Sample
Inference
Statistic
Parameter
What can we infer about a Population’s Parameters
based on a Sample’s Statistics?
9
Statistical Inference
We use statistics to make inferences about parameters.
Therefore, we can make an estimate, prediction, or decision
about a population based on sample data.
Thus, we can apply what we know about a sample to the
larger population from which it was drawn!
10
Statistical Inference
Rationale:
• Large populations make investigating each member impractical
and expensive.
• Easier and cheaper to take a sample and make estimates about the
population from the sample.
However:
Such conclusions and estimates are not always going to be correct.
For this reason, we build into the statistical inference “measures of
reliability”, namely confidence level and significance level.
11
Confidence & Significance Levels
The confidence level is the proportion of times that an
estimating procedure will be correct.
E.g. a confidence level of 95% means that, estimates based on this
form of statistical inference will be correct 95% of the time.
When the purpose of the statistical inference is to draw a
conclusion about a population, the significance level
measures how frequently the conclusion will be wrong in the
long run.
E.g. a 5% significance level means that, in the long run, this type of
conclusion will be wrong 5% of the time.
12
Confidence & Significance Levels
If we use α (Greek letter “alpha”) to represent significance,
then our confidence level is 1 - α.
This relationship can also be stated as:
Confidence Level
+ Significance Level
=1
13
Confidence & Significance Levels
Consider a statement from polling data you may hear about
in the news:
“This poll is considered accurate within 3.4
percentage points, 19 times out of 20.”
In this case, our confidence level is 95% (19/20 = 0.95),
while our significance level is 5%.
14
Exercise #1
• A researcher at University of Toronto wants to estimate
the average number of credits earned by students enrolled
last semester at York University. She randomly selects
750 students from last semester and finds that they
averaged 13.75 credits per student. The population of
interest to the researcher is:
a. all York students
b. all University of Toronto students
c. all York students enrolled last semester
d. the 750 York students selected at random
15
Exercise #2
• A company has developed a new power cell and wants to
estimate its average lifetime. A random sample of 650
power cells is tested and the average lifetime of this
sample is found to be 315 hours. The 315 hours is the
value of a:
a. parameter.
b. statistic.
c. sample.
d. population.
16
Chapter Two
Graphical
Descriptive Techniques 1
17
Introduction & Re-cap…
Descriptive statistics involves arranging, summarizing, and
presenting a set of data in such a way that useful information
is produced.
Statistics
Data
Information
Its methods make use of graphical techniques and numerical
descriptive measures (such as averages) to summarize and
present the data.
18
Populations & Samples
Population
Sample
Subset
The graphical & tabular methods presented here apply to both entire
populations and samples drawn from populations.
19
Definitions…
A variable is some characteristic of a population or sample.
E.g. student grades.
Typically denoted with a capital letter: X, Y, Z…
The values of the variable are the range of possible values
for a variable.
E.g. student marks (0..100)
Data are the observed values of a variable.
E.g. student marks: {67, 74, 71, 83, 93, 55, 48}
20
Types of Data & Information
Data (at least for purposes of Statistics) fall into three main
groups:
Interval Data
Nominal Data
Ordinal Data
21
Interval Data…
Interval data
• Real numbers, i.e. heights, weights, prices, etc.
• Also referred to as quantitative or numerical.
Arithmetic operations can be performed on Interval Data,
thus its meaningful to talk about 2*Height, or Price + $1, and
so on.
22
Nominal Data…
Nominal Data
• The values of nominal data are categories.
E.g. responses to questions about marital status, coded
as:
Single = 1, Married = 2, Divorced = 3, Widowed = 4
These data are categorical in nature; arithmetic operations
don’t make any sense (e.g. does Widowed ÷ 2 = Married?!)
Nominal data are also called qualitative or categorical.
23
Ordinal Data…
Ordinal Data appear to be categorical in nature, but their
values have an order; a ranking to them:
E.g. College course rating system:
poor = 1, fair = 2, good = 3, very good = 4, excellent = 5
While its still not meaningful to do arithmetic on this data
(e.g. does 2*fair = very good?!), we can say things like:
excellent > poor or fair < very good
That is, order is maintained no matter what numeric values
are assigned to each category.
24
Calculations for Types of Data
As mentioned above,
• All calculations are permitted on interval data.
• Only calculations involving a ranking process are allowed for
ordinal data.
• No calculations are allowed for nominal data, save counting the
number of observations in each category.
This lends itself to the following “hierarchy of data”…
25
Hierarchy of Data…
Interval
Values are real numbers.
All calculations are valid.
Data may be treated as ordinal or nominal.
Ordinal
Values must represent the ranked order of the data.
Calculations based on an ordering process are valid.
Data may be treated as nominal but not as interval.
Nominal
Values are the arbitrary numbers that represent categories.
Only calculations based on the frequencies of occurrence are valid.
Data may not be treated as ordinal or interval.
26
Graphical & Tabular Techniques for Nominal Data…
The only allowable calculation on nominal data is to count
the frequency of each value of the variable.
We can summarize the data in a table that presents the
categories and their counts called a frequency distribution.
A relative frequency distribution lists the categories and the
proportion with which each occurs.
27
Example 2.1 Work Status in the GSS 2012 Survey
[GSS2012*] In Chapter 1 we briefly introduced the General Social Survey.
In the 2012 survey respondents were asked the following question.
Last week were you working full time, part time, going to school, keeping
house, or what? The responses were
1. Working full time
2. Working part time
3. Temporarily not working
4. Unemployed, laid off
5. Retired
6. School
7. Keeping house
8. Other
The responses were recorded using the codes 1, 2, 3, 4, 5, 6, 7, and 8,
respectively.
28
Frequency and Relative Frequency Distributions
Work Status
Code
Working full-time
1
Working part-time
2
Temporarily not working 3
Unemployed, laid off
4
Retired
5
School
6
Keeping house
7
Other
8
Frequency
912
226
40
104
357
70
210
54
Relative Frequency (%)
46.2
11.5
2.0
5.3
18.1
3.5
10.6
2.7
29
Nominal Data (Frequency)
Bar Charts are often used to display frequencies…
30
Nominal Data (Relative Frequency)
Pie Charts show relative frequencies…
31
Nominal Data
It all the same information,
(based on the same data).
Just different presentation.
32
Describing the Relationship between Two Nominal Variables
To describe the relationship between two
nominal variables, we must remember that we
are permitted only to determine the frequency
of the values.
As a first step we need to produce a crossclassification table, which lists the frequency
of each combination of the values of the two
variables
33
Example 2.4 Newspaper Readership Survey
In a major North American city there are four competing
newspapers: the Globe and Mail (G&M), Post, Sun, and
Star. To help design advertising campaigns, the advertising
managers of the newspapers need to know which segments
of the newspaper market are reading their papers. A survey
was conducted to analyze the relationship between
newspapers read and occupation. A sample of newspaper
readers was asked to report which newspaper they read:
Globe and Mail (1) Post (2), Star (3), Sun (4), and to
indicate whether they were blue-collar worker (1), whitecollar worker (2), or professional (3). The responses are
stored in Xm02-04 using the codes. Some of the data are
listed here.
34
Example 2.4
Reader
Occupation
Newspaper
1
2
2
2
1
4
3
2
1
.
.
.
.
.
.
.
.
352
3
2
353
1
3
354
2
3
Determine whether the two nominal variables are related.
35
Cross-Classification Table of Frequencies
Occupation
Blue collar
White collar
Professional
Total
G&M
27
29
33
89
Newspaper
Post
Star
18
38
43
21
51
22
112
81
Sun
37
15
20
72
Total
120
108
126
354
36
Row Relative Frequencies
Occupation
Blue collar
White collar
Professional
Total
G&M
.23
.27
.26
.25
Newspaper
Post
Star
.15
.32
.40
.19
.40
.17
.32
.23
Sun
.31
.14
.16
.20
Total
1.00
1.00
1.00
1.00
37
Graphing the Relationship between 2 Nominal Variables
60
Post
50
Post
Star Sun
40
30
20
G&M
Post
G&M
G&M
Star Sun
Star
Sun
10
0
Blue collar
White collar
Professional
Occupation
38
INTERPRET
If the two variables are unrelated, the patterns exhibited in
the bar charts should be approximately the same. If some
relationship exists, then some bar charts will differ from
others.
The graphs tell us the same story as did the table. The shapes
of the bar charts for occupations 2 and 3 (White-collar and
Professional) are very similar. Both differ considerably from
the bar chart for occupation 1 (Blue-collar).
39
Chapter Three
Graphical
Descriptive Techniques II
40
Example 3.1
The game of bridge is played all over the world. There are
two versions. There is rubber bridge, which is usually
played in private and often for money. The second, more
popular version is duplicate bridge, which is played in
clubs and tournaments around the world. The American
Contract Bridge League is the organization that runs
duplicate bridge. Anyone who has played in club games
will notice that a great majority of players are seniors.
(Xm03-01)
41
Example 3.1
To gain useful information we need to know how the
ages are distributed between 16 and 99. Are there many
old players with few young ones? Are the ages somewhat
similar or do they vary considerably? To help answer
these questions and others like them, we will construct a
frequency distribution from which a histogram can be
drawn.
We create a frequency distribution for interval data by
counting the number of observations that fall into each of
a series of intervals, called classes that cover the
complete range of observations.
42
Example 3.1
We have chosen nine classes defined in such a way that each
observation falls into one and only one class. These classes are
defined as follows:
Classes
Ages that are more than 10 and less than or equal to 20
Ages that are more than 20 but less than or equal to 30
Ages that are more than 30 but less than or equal to 40
Ages that are more than 40 but less than or equal to 50
Ages that are more than 50 but less than or equal to 60
Ages that are more than 60 but less than or equal to 70
Ages that are more than 70 but less than or equal to 80
Ages that are more than 80 but less than or equal to 90
Ages that are more than 90 but less than or equal to 100
43
Example 3.1
44
(36+27+12+6=81)÷200 = 40.5%
i.e., over 40% of the bride players are 60 or more.
The histogram gives us a clear view of the way the ages are distributed. As
expected about 40% of the sample are older than 60. About a sixth are in their
teens and twenties. There appears to be a good supply of younger players.
However, there is an unexpected gap of players in the 40 to 50 range. This
group may be individuals who are working and have little time for bridge.
Building a Histogram…
1) Collect the Data
2) Create a frequency distribution for the data…
How?
a) Determine the number of classes to use…
How?
Refer to table 3.2:
With 200 observations,
we should have
between 7 & 10
classes…
Alternative, we could use Sturges’ formula:
Number of class intervals = 1 + 3.3 log (n)
46
Building a Histogram…
1) Collect the Data
2) Create a frequency distribution for the data…
How?
a) Determine the number of classes to use. [9]
b) Determine how large to make each class…
How?
Look at the range of the data, that is,
Range = Largest Observation – Smallest
Observation
Range = 99 – 16 = 83
Then each class width becomes:
Range ÷ (# classes) = 83 ÷ 9 ≈ 9.22
47
Building a Histogram…
CLASS LIMITS
FREQUENCY
10-20*
6
20-30
27
30-40
30
40-50
16
50-60
40
60-70
36
70-80
27
80-90
12
90-100
6
TOTAL
200
* Each class represents ages that are more
than the lower limit and less than or equal
to the upper limit.
48
Building a Histogram…
CLASS LIMITS
10-20*
20-30
30-40
40-50
50-60
60-70
70-80
80-90
90-100
TOTAL
FREQUENCY
6
27
30
16
40
36
27
12
6
200
49
Shapes of Histograms…
Variable
Frequency
Frequency
Frequency
Symmetry
A histogram is said to be symmetric if, when we draw a
vertical line down the center of the histogram, the two sides
are identical in shape and size:
Variable
Variable
50
Shapes of Histograms…
Frequency
Frequency
Skewness
A skewed histogram is one with a long tail extending to
either the right or the left:
Variable
Positively Skewed
(Right Skewed)
Variable
Negatively Skewed
(Left Skewed)
51
Shapes of Histograms…
Modality
A unimodal histogram is one with a single peak, while a
bimodal histogram is one with two peaks:
Bimodal
Frequency
Frequency
Unimodal
Variable
Variable
A modal class is the class with
the largest number of observations
52
Shapes of Histograms…
Many statistical techniques
require that the population
be bell shaped.
Drawing the histogram
helps verify the shape of
the population in question.
Frequency
Bell Shape
A special type of symmetric unimodal histogram is one that
is bell shaped:
Variable
Bell Shaped
53
Histogram Comparison…
Compare & contrast the following histograms based on data
from Examples 3.3 and 3.4:
unimodal vs. bimodal
The two courses, Business Statistics
and Mathematical Statistics have very
different histograms…
54
Relative Frequencies…
From Example 3.1, we had 6 observations in our first class
( ages from 10 to 20). Thus, the relative frequency for
this class is 6 ÷ 200 (the total # of bridge players) = 0.03
(or 3%)
CLASS LIMITS
10-20*
20-30
30-40
40-50
50-60
60-70
70-80
80-90
90-100
TOTAL
FREQUENCY
6
27
30
16
40
36
27
12
6
200
RELATIVE FREQUENCY
0.03
0.14
0.15
0.08
0.20
0.18
0.14
0.06
0.03
1.00
55
Describing Time Series Data
Observations measured at the same point in time are called
cross-sectional data.
Observations measured at successive points in time are
called time-series data.
Time-series data graphed on a line chart, which plots the
value of the variable on the vertical axis against the time
periods on the horizontal axis.
56
Example 3.5
We recorded the monthly average retail price of gasoline
since 1976.
Xm03-05
Draw a line chart to describe these data and briefly describe
the results.
57
Example 3.5
Average price of gasoline
450
Line Chart
400
350
300
250
200
150
100
50
0
1
25 49 73 97 121 145 169 193 217 241 265 289 313 337 361 385
Month
58
Example 3.6 Price of Gasoline in 1982-84 Constant
Dollars
Xm03-06
Remove the effect of inflation in Example 3.5 to
determine whether gasoline prices are higher than they have
been in the past after removing the effect of inflation.
59
Ajusted price of gasoline
Example 3.6
200
180
160
140
120
100
80
60
40
20
0
1
25 49 73 97 121 145 169 193 217 241 265 289 313 337 361 385
Month
60
Example 3.6
Using constant 1982-1984 dollars, we can see that the average
price of a gallon of gasoline hit its peak in the middle of 2008
(month 390). From there it dropped rapidly and in late 2009 it
was about equal to the adjusted price in 1976.
61
Graphing the Relationship Between Two Interval Variables…
Moving from nominal data to interval data, we are
frequently interested in how two interval variables are
related.
To explore this relationship, we employ a scatter diagram,
which plots two variables against one another.
The independent variable is labeled X and is usually placed
on the horizontal axis, while the other, dependent variable,
Y, is mapped to the vertical axis.
62
Example 3.7
A real estate agent wanted to know to what extent the selling
price of a home is related to its size. To acquire this
information he took a sample of 12 homes that had recently
sold, recording the price in thousands of dollars and the size
in hundreds of square feet. These data are listed in the
accompanying table. Use a graphical technique to describe
the relationship between size and price. Xm03-07
Size 2354 1807 2637 2024 2241 1489 3377 2825 2302 2068 2715 1833
Price 315 229 355 261 234 216 308 306 289 204 265 195
63
Example 3.7
It appears that in fact there is a relationship, that is, the
greater the house size the greater the selling price…
Price $1,000s)
Scatter Diagram
400
350
300
250
200
150
100
50
0
0
500
1000
1500
2000
2500
3000
3500
4000
House size
64
Patterns of Scatter Diagrams…
Linearity and Direction are two concepts we are interested in
Positive Linear Relationship
Negative Linear Relationship
Weak or Non-Linear Relationship
65
Summary …
Interval
Data
Histogram
Frequency and
Relative Frequency
Tables, Bar and Pie
Charts
Scatter Diagram
Cross-classification
Table, Bar Charts
Single Set of
Data
Relationship
Between
Two Variables
Nominal
Data
66
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )