KABARAK
UNIVERSITY
Of Sciences, Engineering and Technology
SCHOOL ----------------------------------------------------------P.O Private Bag – 20157, KABARAK, KENYA
Course code MATH 424 Course Name NONPARAMETRIC METHODS 3 Credit Hour3 CF
COURSE OUTLINE
Lecturer; RAGAMA PHILIP
Cell phone; 0723235132
Email; peragama55@yahoo.co.uk, pragama@kabarak.ac.ke
Course description
The general purpose of nonparametric methods is to introduce students to the nontraditional methods of hypothesis testing in statistics.
Course Objectives
At the end of this course, the learner should be able to:
1. Use Goodness of fit tests and 2
2. Discuss test of independence contingency table analysis and associated tables and 2 .
3. Use the Rank correlation test, Kalmogorov-Smirnov test.
4. Compare two populations, sign test, run test, median test, wilcoxon test and Mann
Whitney u-test.
Expected Learning outcomes
At the end of this course the learner should be able to:
1. Explain differences between parametric and non-parametric methods.
2. Apply point estimations to single samples using non-parametric methods.
Page 1 of 62
3. Compare two populations, sign test, run test, median test, wilcoxon test and Mann
Whitney u-test
4. Use the Rank correlation test, Kalmogorov-Smirnov test to compare associations in two
variables
5. Discuss test of independence contingency table analysis and 2 of goodness of fit test.
6. Introduction to binary data analyses
Learning approaches
E.g Lectures, discussions in class and groups, short presentations, individual exercises and
structured activities.
Weeks
Week 1
Dates
Detailed course content
Topic:
Introduction: comparison between parametric and non parametric
methods
o Advantages and disadvantages of both nonparametric and
parametric methods
o
Introduction to Sign test compared to t-test
o How to test hypothesis using Sign test for small and large
samples
o
Week 2
The paired-sample sign test
o A rank sum test: Mann- Whitney U-test
o Hypothesis concerning U-test for two-tailed test and onetailed test
Week 3
o Hypothesis for U-test for small samples and large samples
o Hypothesis test for comparing two populations- Wilcoxon
signed rank test
Week 4
o Kolmogorov-Smirnov test (K-S) test
o Hypothesis test using K-S test on uniform, Poisson and
normal distributions
o
Run test for randomness for small and large samples
o Two -sample median test for small and large samples
Page 2 of 62
Week 5
CAT 1
Week 6
K-samples median test
K-sample median test for small and large samples
Week 7
Kruskal-Wallis test (H-test)
Friedman's ANOVA test
Week 8
Measures of association: Spearman‟s rank correlation coefficient
Week 9
Pearson‟s Chi-square statistics
McNemar's test
Likelihood-ratio chi-square (G2)
Week 10
Odd ratio statistics
Mantel-Haenszel chi-square
Fisher‟s exact test
CAT 2
Week 11
Week 12
Phi coefficient
Cramer‟s V
Gamma (γ)
Kendall‟s W test
Stuart‟s tau-c
Somers‟ D (C|R)
Week 13
Order statistics
Week 14
Order statistics
Week 15
Exams
1 Course assessment
Continuous assessment tests (CATS)
30
Final examination
70
Total
100
Core Reading Texts
1. Gupta S.C. and Kapoor V.K. (1974) Fundamentals of Applied Statistics. 3rd Ed. Sultan
Chand and Sons. Montgomery, DC.
2. Hogg R.V. and Graig A.T. (1995) Introduction to Mathematical Statistics. Prentice-Hall
of India.
Page 3 of 62
3. Mood A.M. and Graybill B. (1974) Introduction to Theory of Statistics. McGraw-Hill.
E Books (At least 3)
1.
2.
3.
Journals (At least 2)
Tables you need for this course
Critical values of t-distribution Critical value of Chi-square
distribution
F- test
Binomial test p=0.5 table
Mann- Whitney table
Critical value of Wilkoxon Tdistribution
Kurtosis measures,g2
Proportion of normal curve Ƶtable
Proportional of the Binomial
Distribution for p=q=0.5
Fisher‟s exact table
Critical values of C for the
sign test
Critical values of MannWhitney U-distribution
Page 4 of 62
Critical values of Friedman‟s
2 test
Critical value of dmax for the
Kolmogorov- Smirnov test
Kruskul-Wallis H-distribution
table
1.0 Introduction:
Parametric method or linear models assumes the following model
Yi 0 1 xi i
i =1,2……..n
With assumption
εi~NID (0,σ2)
i)
i s are normally distributed
ii)
i s has constant variance
iii)
i s are independent
iv)
i s are independent of x‟s
v)
There is linearity (additivity),
a
x
i 1
i
0
Non parametric methods
When to Use Nonparametric Tests
> When the dependent variable is nominal
(What are ordinal, nominal, interval, and ratio scales of measurement?)
> Used when either the dependent or independent variable is ordinal
> Used when the sample size is small
> Used when underlying population is not normal
Limitations of Nonparametric Tests
> Cannot easily use confidence intervals or effect sizes
> Have less statistical power than parametric tests
> Nominal and ordinal data provide less information
> More likely to commit type II error
•
Review: What is type I error? Type II error?
Disadvantages
-Due to re-ordering of the data, we waste information & hence they are less efficient
Page 5 of 62
Non parametric to be considered will be based on order of statistics. We shall concentrate mostly
on continuous random variables
-Based on cumulative distribution function
-Point estimation, interval estimation and testing hypothesis
-Population quintile‟s (Median, quartiles, mode, deciles and percentiles).
Tolerance Limits
-Similarities of differences of tolerance limits and confidence limits
-Test Homogeneity of two populations
Advantages of non-parametries
1) They are distribution free (No assumption are made about parent population)
2) Simple to understand and easy to compute
3) We apply them when sample sizes are small
4) Non-parametric tests are applicable to all types of data- qualitative like counts (nominal
scaling), data in rank form (ordinal scaling) as well as data with interval scale or ratio
scaling.
5) Since we have minimal assumption, they are normally easily met
Example where we use non-parametric compared parametric
Situation
One sample
Non-parametric method
Wilcoxon signed rank test
Two independent samples
Two paired samples
One sample, two
quantitative variables
Wilcoxon sum rank test
Wilcoxon signed rank test
Correlation coefficient of
Spearman‟s
Differences between independent groups
Parametric
Nonparametric
Wald-Wolfowitz runs test
t-test for independent samples
Mann-Whitney U test
Page 6 of 62
Parametric methods
Z statistic ( t test)
Z statistic for two independent
samples (t test)
Z-paired statistic (t-paired test)
Correlation coefficient of Pearson
Differences between independent groups
Parametric
Two samples- compare mean of
value of some variables of interest t-test for independent
samples
Analysis of variance
(ANOVA/MANOVA)
Multiple groups
Nonparametric
Wald-Wolfowitz runs test
Mann-Whitney U test
Kolmogorov-Smirnov two
sample test
Kruskal-Wallis test
Median Test
Differences between dependent groups
Parametric
Nonparametric
t-test for dependent
Sign test
Compare two variables measured in
samples
the same sample
Wilcoxon‟s matched pairs test
If more than two variables are
Repeated measures
Friedman‟s two way analysis
measured in the same sample
ANOVA
of variance
Cochran Q
Relationships between variables
Parametric
Two variables of interest are categorical
Correlation coefficient
Nonparametric
Spearman R
Kendall Tau
Coefficient Gamma
Chi square
Phi coefficient
Fisher exact test
Kendall coefficient of concordance
Sample Characteristics
Level of
Measurement
1 sample
Independent
2
Categorical or
Nominal
Rank or
Ordinal
Parametric
(interval & Ratio)
Z-test
or t test
or
binomial
Mann Whitney
U
2
t test
between
2 Sample
K Sample (i.e., >2
Dependent
Independent
MacNarmar's
X2
Wilcoxon
Matched
Pairs Signed
Ranks
t test within
groups
Page 7 of 62
Correlation
2
Dependent
Cochran's
Q
Kruskal Wallis
H
Friendman's
ANOVA
Spearman's
rho
1 way ANOVA
between
1-way
ANOVA
Pearson's r
groups
groups
(within or
repeated
measures)
Factorial (2 way)
ANOVA
(Plonskey, 2001)
Sign Test
Simplest of all non-parametric tests. The name is derived from the fact that it is based on
direction (+) or (-) of a pair of observations.
It is an alternative to one –sample t-test for H 0 : ~ ~0 ( ~ = median)
In sign test first replace each sample that exceeds ~ with „+‟ and each sample that is less
than ~ with “-“; the sample that is equal to ~ assign it „0‟ and will be discarded from
the calculations. We then test H₀ that the number of „+‟ sign is a value of random
variable having a binomial distribution with parameters n (Number of positives or
negatives) and =½
The two sided alternative H 1 : ~ ~0 becomes ≠ ½ and the one-sided alternative ˂½
or ˃½
To perform a sign test when the sample size is very small, we refer directly to table of
Binomial probability. Otherwise for large samples, we use normal approximation to
Binomial distribution.
If S is the number of times the “+”signs occur then S is Binomial distribution with p=½
or θ=½
The critical value for a two –sided alternative at =0.05 can conveniently be found by
the expression
K
n 1
0.98 n
2
H0 is rejected if S≤K for the sign test; otherwise fail to reject H0 if S>K
Types of the Sign Test
sign test can be two types
i)
The one-sample sign test
ii)
The paired sample sign test
Page 8 of 62
One-Sample Sign test
H 0 : 0 versus H1 : 0 or 0
Example:
Breaking strength of a cotton ribbon (in lbs)
163
165
160
189
161
176
158
151
163
139
172
165
148
148
166
172
Use the sign to test the hypothesis H 0 : μ=160 versus H1: >160 at 𝞪=0.05
169
187
162
173
Solution
a) Use the test statistic x. The observed number of signs.
b) Replace each value exceeding 160 with a “+” sign and each value less than 160 with “-“ value
while values that are equal to 160 are assigned “0”
we have + + 0 + + + - - + + - + + - + + + + + +
So that n=19 and X=15; Using Binomial table we find P(X≥15) =0.0095 for =½
c) Since the P= 0.0095 is less than 0.05 we reject H 0 and conclude that the mean breaking
strength of the given kind of ribbon exceeds 160 lbs
Example: Amount of Sulphur Oxides emitted by a larger industry in 40 days
17
24
19
19
15
20
23
10
20
17
28
23
29
6
19
18
19
24
16
31
18
14
22
13
22
15
24
20
Using the sign test to test Ho: =21.5 Against H 0 : <21.5 at 𝞪=0.01
25
23
17
17
27
24
20
24
9
26
13
14
Solution:
a) Ho: 21.5 ; against H1: <21.5 at 𝞪=0.01 No of positive signs=16, No of –ve=24
- - - + - - + + + - + - - - + - - + + + - + + - - ++ - - - - - + - + - - - + Reject Ho at Ƶ≤ =-2.33
Where Z
x n
n 1
, =½
,n=Number of „+‟
Since n=40, X=16 we get n =40●½=20
n (1 ) 40 0.5 0.5 3.16
Page 9 of 62
Ƶ=
=-1.26; since Ƶ = -1.26>Ƶ0.01=-2.33 we cannot reject H 0
Paired-sample Sign test
This comes as a result of a need to compare periods (before and after) or successive /failure in
national schools exams. Impact study of introduced of a new technology and a few years after we
introduce the technology. Mother response Vs daughter response. This presents a perfect data for
comparing using pairwise comparison
Example
Consider 15 students competing in typing
Students
Words/min
A
81
+21
d i WPM 60
B
76
+16
C
53
-7
D
71
+11
E
66
+6
F
59
-1
G
88
+28
H
73
+13
I
80
+20
J
66
+6
K
58
-2
L
70
+10
M
60
0
N
56
-4
0
55
-5
H 0 : d 60 Versus H 1 : d 60
Since the alternative H 0 is one-sided, we obtain the probability of our results from tables and
compare with 𝞪=0.005
Results:
N○ of “+” signs =9
No of “–“signs =5
Number of 0‟s
= 1 (which should be discarded)
Total sample size 14
K=
K=
-0.98√
-0.98√
=6.5-3.67= 2.83
Since S˃K the null hypothesis is not rejected
Example:
Use the Sign test to test if there is a difference between the number of days until collection of an
account receivable before and after a new collection policy use
Before
30 28 34 35 40 42 33 38 34 45 28 27 25 41 36
After
32 29 33 32 37 43 40 41 37 44 27 33 30 38 36
1st- 2nd = di - + + + - + + - + 0
Page 10 of 62
H0: 1 2 0 versus H1: 1 2 0
No of “+” signs =6
No of “-” signs =8
No of zeros
=1
14
K=
-0.98 √
=
-0.98 √
=6.5 =3.67 =2.83
Since S˃K H0 the null hypothesis is not rejected.
For large samples, n> 25 for Sign test, the normal approximation to the binomial distribution
may be used correcting for continuity.
Since P=0.5 then μ= np = 0.5n
Std =
1
n npq
2
The actual value of Z can be computed using the formula
Ƶ=
x np
x= No of + signs
np(1 p)
Example:
Daily production of cement in tanks for 30 days
11.5
10
11.2
10
9.3
10.7
11.3
10.4
10.8
11.9
12.4
9.6
H 0 : 11.2 versus H1 : 11.2 at
12.3
11.4
10.5
=0.05
11.1
12.3
11.6
Solution:
+-0-+----- --+ - +++-+- -++--+---+
No of “+” signs =11
No of “–“ signs =18
No of zero signs =1
Total
29 (less zero)
Page 11 of 62
10.2
11.4
8.3
9.6
10.2
9.3
8.7
11.6
10.4
9.3
9.5
11.5
X= 11;
n = 29
p= ½
Z
11 29 0.5
1.46
29 0.5 0.5
-ᵶ0.05=- 1.645
Calculated Ƶ tabulated Ƶ0.05
H 0 is rejected and we conclude that the production is less than 11.2 tons
Rank Methods for comparing two populations
Suppose we have two populations with distributions F and G which were normal with common
variances from two independent samples, X 1 , X 2 ,... X n1 and Y1 , Y2 ,...Yn 2 one from each
population. We suppose that F and G are continuous
The Wilcoxon Statistics
We test H 0 : F G versus H 1 : F G
We could consider the one-tailed test like
P(Y>t) ≥ P(X>t) or equivalently G(t)≤F(t) for all t but G≠F
For parametric we could have tested as
H 0 : 2 1
Assign ranks R1, R2, …, Rn, where n= n1 n2 of X 1 , X 2 ,... X n1 and Y1 , Y2 ,...Yn 2
Wilcoxon paired-sample test (Wilcoxon, 1945; Wilcoxon and Wilcox, 1964) is a non-parametric analogue
to paired-sample t-test, just as the Mann-Whitney test is a nonparametric procedure is to two-sample
t-test. It is the one referred to Wilcoxon with a suffix “rank sum” or “signed rank”. When the
differences d j ' s are assumed to be normally distributed then the parametric t-test is valid but when one
is not sure of this fact then, the nonparametric will apply.
The test criteria is T
n( n 1)
T
2
For a two-tailed test we reject H0 if either T- or T+ is less or equal to the critical T ( 2), n that is read from
the tables.
Page 12 of 62
Example
A sample of 10 deers were measured the average hind- and fore-legs. Test the hypothesis of equal lengths
of pair of legs
H 0 : Hindlegs Forelegs versus H1 : Hindlegs Forelegs
Deer (j)
Hind leg lengths
(cm) X 1 j
Fore-leg lengths
(cm) X 2 j
D j X1 j X 2 j
142
140
144
144
142
146
149
150
142
148
138
136
147
139
143
141
143
145
136
146
4
4
-3
5
-1
5
6
5
6
2
1
2
3
4
5
6
7
8
9
10
Difference
Rank of
Dj
Signed Rank of
Dj
4.5
4.5
3
7
1
7
9.5
7
9.5
2
4.5
4.5
-3
7
-1
7
9.5
7
9.5
2
versus
n=10; df=n-1=9
t=
H0
7
T
n( n 1)
T
2
0.005<P
T+ =4.5+4.5+7+7+9.5+7+9.5+2=51
T- =3+1=4
T0.05( 2),10 8 Since T_ 3 T0.05( 2),10 8 ; H0 is rejected 0.01 P(T_ orT 4) 0.05
References
1.- Last JM. A dictionary of epidemiology. New York, 4ª ed. Oxford University Press,
2001:173.
2.- Kirkwood BR. Essentials of medical ststistics. Oxford, Blackwell Science, 1988: 1-4.
3.- Altman DG. Practical statistics for medical research. Boca Ratón, Chapman & Hall/ CRC;
1991: 1-9.
Page 13 of 62
A Rank Sum Test : The Man Whitney U- Test
-The sign test for comparing two population distributions ignores the actual magnitude of the
paired observations and hence discards information that would be useful in detecting a departure
from H0.
-A statistical test that partially circumvents this loss by utilizing the relative magnitudes of the
observations was proposed by Mann and Whitney and is equivalent to the t-test proposed
independently by Wilcoxon.
-Rank sum test is a whole family of tests. Order the two random samples n1 and n2 to form R1
and R2 so that we get
n1 n2
n1 n2 n1 n2 1
2
This assists us to find R2 if we know R1 and vice versa
The test was to replace the two-sample t-test the statistic is
U 1 n1n2
U 2 n1n2
n1 n1 1
R1 …………………………………………….(i)
2
n2 n2 1
2
R2 ………………………………….……….(ii)
Where n1and n2 are the size of samples 1&2 and R1 & R2 are the rank sums of the corresponding samples
1 & 2. If both n1 and n2 are less than 10 (some suggest 8) special tables must be used.
The calculated, U is compared to the two-tailed value of U (2),n1 .n2 found from statistical tables. The table
assumes that n1˃n2. Should that not be the case we just reverse n2 to n1 as; U (2) n2, n1, as the critical
value. For a two-tailed test we must first compute both U1 or U2 and the larger of the two is compared to
the critical value U (2) ,n1 n2
If equation (i) was used to calculate U1 then we can compute quickly
U2 =n1n2 –U1
Similarly U1= n1 n2 –U2
If U1 and U2 is as greater than U (2), n1,n2, Ho is rejected at -level of significance
If U is smaller than the critical value H 0 can be related to the standard normal curve by the
statistics
Z
n1n2
2
n1 n2 n1 n2 / 12
U
Page 14 of 62
When n1 &
distribution
n 2 >8 it is reasonable to assume that U 1 and U 2
We estimate
be are random variable having approximately normal
n1n2
n n n n2 1
2
, variances 1 2 1
2
12
Proof:
Since we have no ties and
R1
n1 & n 2 are positive integers then
is a random variable with mean
and variances
2
n1 n1 n2 1
2
n1n2 n1 n2 1
12
n1 (n1 1)
the mean and variances of r.v corresponding to U 1 are
2
Since
U 1 R1
U
n1 n1 n2 1 n1 n1 1 n1n2
2
2
2
and variances
2
n1n2 n1 n2 1
12
Page 15 of 62
Review
Done Sign test
One sample sign test for
H 0 : 0 against H 1 : 0 or 0
Used test criteria K
n 1
0.98 n
2
Reject Ho if S K for the Sign test, where S= total “+” signs
Consider the small samples
Where np or n(1-p)<5, p=0.5 and n=sample size. We use binomial distribution. Let X be the
random variable representing the number of “+” signs
Example:
H 0 : p ½, H1 : p =½
We compute binomial value P(X≥No of „+‟ signs. n, p) is more than
Otherwise reject H0
- level then accept;
Example: A medical representative visited at random 9 clinics and observed the waiting time to
see a doctor: 16, 30, 17, 22, 31, 11, 19, 18, and 8. Check whether the waiting time is 15 minutes
at = 0.05?
Solution:
H 0 : 15 against H 1 : 15 using sign test at = 0.05
n=9, P=0.5, np =n(1-p) = 4.5
Sign are: + + + + + - + + Solution; No of „+‟ sign =7
No of „-„ sign =2
Let X be a random variable for representing „+‟ sign
P(X≥7, n=9, p=0.5) = 1- P(X≤6, 9,p=0.05)=1-0.9102=0.0898
So P(X≥7 ) is more than the significance level of =0.05, we reject H 0 and conclude that
waiting time is more than 15 minutes.
Page 16 of 62
Example 2
A Salesman visited at random 8 cities and got the following orders
5, 6, 4, 8, 2, 4, 9, 1,
Check H 0 : 7 or p= ½ against H1 : 7 or p<½ use sign test at = 0.05
Solution:
So „+‟ if observation is ˃7 and „-„ otherwise we have - - - + - - + -
+ Sign =2
- Sign= 6
Let X be a random variable representing „+‟ sign
Pr (X≤ 2, n=8, P=0.5) =0.1445
Since p(X≤ 2) ˃ =0.05 we accept H0
Two One-Sam ple Sign Test for small samples
1
1
versus H 0 : p Sample size n, np as well as n(1-p)˂5. H 0 is tested using
2
2
binomial distribution with P=½ where X is a random variable representing the „+‟ sign. The
Objective is to check whether population is to be accepted at at both tails of distribution.
H0 : p
If Pr (X≥ No of „+‟ sign, n, p) is more than
then accept H 0 otherwise reject H 0
Example
Annual income in 100 dollars of 9 randomly selected CEO‟s of companies are 12, 16, 20, 15, 14,
8, 11, 17, 13
Check H 0 : 10 or P=½ against H1 : 10 or P ½
Using the Sign test at =0.10
Solution
Sample size n=9, np=n(1-p)= 4.5˂ 5. We use binomial =0.10 and ⁄ =0.05.Assign „+‟ to
samples ˃10 and „-„ of less than 10
We have: + + + + + - + + +
So „+‟ signs =8
„-„ signs = 1
Page 17 of 62
Pr (X≥8, n=9, p=0.5) = 1-P(X≤ 7, n=9, p= 0.05)
=1-0.9805 =0.0195
Pr X 8
2
0.05
This value falls in rejection region. We reject Ho and conclude that CEO‟s earn more than 1000
dollars.
One-Tailed One- Sample sign tests for large samples
Test 1: H0 : p=½ against H1:P˃½
Test 2 : H0 : p= ½ against H1 : p˂½
A random sample of size n, where np & n(1-p) is at test 5 is taken and the H0 is tested using
normal approximation to binomial distribution with p=½ where X is the random variable
representing the number of signs‟+‟ signs
If X~B(X, n, p), then for large value of n (n˃30)., one can approximate the binomial distribution
with normal distribution X~N(np, np(1-p)), where μ=np and σ²=np(1-P) of normal distribution.
The formulae for Z-Statistics is
Z
X np
np(1 p)
Example.
Specific weight drug is 100mg. A sample of 36 drugs were taken their weight as follows
101
101
95
108
96
103
101
103
105
99
100
99
102
100
102
108
102
95
98
101
109
103
102
109
106
92
97
99
108
105
104
103
99
108
Verify H 0 : 100 or p=½ against H 0 : 100 or p>½ using sign test at =0.05
Solution:
Assign „+‟ to values ˃100 of „-„ ˂ 100 & 0 otherwise
+ + - + -+ + + + - 0 - + 0 + + + + - - - + + + + + + - - - + + + + - +
No of “+” signs=23
No of “-“ signs=11
No of zeroes= 2
Total
35 (less zero)
Since n˃ 30 we use normal approximation to binomial distribution
Page 18 of 62
101
91
mean=np=35x0.5=17.5
Variance, 2 np(1 p ) 35x0.5 x0.5 8.75
Then X~ N np, np(1 p)
the Ƶ statistic is computed as
Ƶ=
(
√
)
=
√
=2
The value of Z at =0.05 is 1.64. Calculated Ƶ0 =2.0 ˃Ƶ (1.64) We reject H 0 & conclude the
weight of the drugs ˃100mg
Tests 2: H 0 :p=½ & H 1 : p˂ ½
If the computed Ƶ value ˃- Z then accept Ho otherwise reject H 0
Example:
A personal manager of a firm is recruiting trainees. It uses performance index of the scale (0-10
scale) that follows non-normal distribution. The manager feels that the index should be more
than 8.He takes the sample of 40 trainees under the following indices.
5
6
7
2
6
9
10
9
9
5
8
9
4
2
7
8
2
4
9
5
5
5
3
4
5
8
6
7
7
6
7
9
7
10
6
8
9
7
2
10
Test H 0 : 8 or p=½ against H1 : 8 or p≤ ½ using sign test at =0.01
Solution:
Sample size n=40, =0.01.
Assign „+‟ to value ˃8 and „-‟ to value ˂ 8, 0 otherwise
--+------+-+----0-+--+0-+----+-++0---+0+
The no of „+‟ sign=11
no of „-‟ sign =25
no of 0 sign=4
mean=np= 36x0.5=18
Variance=np(1-p) =36x0.5x0.5=9
SD= 9 3
then X~N(18,9)
Page 19 of 62
The Ƶ statistic is comprised as
Ƶ=
√
(
)
=
11 18
= -2.333
3
The value of Z at =0.05 is =-2.33 Calculated Zo= -2.333˂(-2.333) We fail to reject H0 &
conclude test performance index of trainee is 8
Assignment
A special diet is fed to adult turkeys to see if they will gain weight. The before and after weights (in kg)
are given below:
Before 30
26
31
32
34
35
27
28
30
38
37
39
31
30
32
33
33
37
31
27
32
37
38
40
After
a) Use the paired sample sign test at 0.05 to see if there is weight gain.
b) Do a pairwise t-test on the same data. at 0.05
c) Do Spearman rank correlation on the same data at 0.05
Kolmogorov- Smirnov test K-S
K-S is similar to 2 test. The K-S is more powerful for small samples whereas the 2 – test is
suited for large samples.
Ho: The given set of observation does not follow the assumed distribution
H1: The given set of observation does not follow the assumed distribution.
In K-S test, the cumulative distributions of random variables are compared with corresponding
theoretical cumulative distribution and their absolute differences are calculated. The maximum
absolute differences is treated as the calculated statistics of K-S test
Dcal=max|OFi-EFi| where OFi = Observed cumulative probability for the ith value of the random
value X; EFi is expected cumulative test for ith value of random value X from the theoretical
distribution
If Dcal> Theoretical value of D then reject H 0 ; otherwise fail to reject H 0
Example 1: A sales manager of a Cement factory feels the quarterly demand for cement in tons
follow uniform distribution. Observed frequencies of demand values are given below.
Page 20 of 62
Demand (tons)
Observed
Frequency
2
3
1
4
2
2
4
2
4
1
30
31
32
33
34
35
36
37
38
39
Check whether the given data follows uniform distribution using K-S test at =0.05
Solution
H 0 : The given data follows uniform distribution
H1 : Does not follow uniform distribution
frequency=25
Significance level = 0.05
No of random variables X=10
Expected frequency of each Xi = = 2.5
Expected probability for each Xi =
Serial
Demand
value
= 0.10
Observed
Freq (Oi)
Observed
Prob
Expected
Prob
Observed
Cum prob
(OFi)
Expected
Cum prob
(EFi)
Di=
|OFi-EFi|
2
3
1
4
2
2
4
2
4
1
0.08
0.12
0.04
0.16
0.08
0.08
0.16
0.08
0.16
0.04
0.10
0.10
0.10
0.10
0.10
0.10
0.10
0.10
0.10
0.10
0.08
0.20
0.24
0.40
0.48
0.56
0.72
0.80
0.96
1.0
0.10
0.20
0.30
0.40
0.50
0.60
0.70
0.80
0.90
1.00
0.02
0.00
0.06
0.00
0.02
0.04
0.02
0.00
0.06
0.00
(Xi )
1
2
3
4
5
6
7
8
9
10
30
31
32
33
34
35
36
37
38
39
The maximum value for Dcal=0.06 where Xi is equal to 32 or 38. The theoretical value of Dn for
the sample size (n) of 25 at = 0.05 is 0.27. Since Dcal (0.06) < tabulated D25(0.27) , we fail to
reject Ho and conclude that the data follows uniform distribution
Page 21 of 62
Experiment 2. The arrival rate of loan customers at a loan counter of a bank appears to follow
Poisson distribution. The observed frequencies of arrival rates are given below
Observed frequency of arrival rates
Period (i) Arrival arte
(Xi )
1
0
2
1
3
2
4
3
5
4
6
5
7
6
Observed
Freq (Oi )
2
6
18
12
8
3
1
X i Oi
0
6
36
36
32
15
6
131
Check whether the given data follows Poisson distribution using Kolmogorov-Smirnov test
Solution:
H0: The given data follows Poisson distribution
H1: The given data does not follow Poisson distribution
The total observed frequency=50
No of random variable Xi=7
7
Mean arrival rate λ=
xo
i 1
7
i i
o
i 1
=
=2.62
i
The probability distribution for Poisson distribution is P( X )
x e
x!
x=0, 1,2……6. We
calculate K-S test
Summary of calculations of K-S test
Period
Arrival
rates ( X i )
Observed
Freq (Oi)
Observed
Prob
Expected
Prob
1
2
3
4
5
6
7
0
1
2
3
4
5
6
2
6
18
12
8
3
1
0.04
0.12
0.36
0.24
0.16
0.16
0.02
0.073
0.191
0.250
0.218
0.143
0.075
0.033
Page 22 of 62
Observed Expected
Cum prob Cum prob
(OFi)
(EFi)
0.04
0.073
0.16
0.264
0.52
0.514
0.76
0.732
0.92
0.875
0.98
0.950
1.00
0.983
Di=
|OFi-EFi|
0.033
0.104
0.006
0.028
0.045
0.030
0.017
The maximum value of Dcal =0.104 when Xi is 1. The theoretical value for Dn for the sample size
(n=50) at =0.05 is computed from the formulae D50 =
= =0.192
√
√
Since Dcal (0.104) < tabulated D50(0.192) . We fail to reject Ho & conclude that the data follows
Poisson distribution
Example3: The demand of a product in Ksh appears to follow a normal distribution. The
observed frequencies of the demand rate are given below
Observed frequency of demand values
Serial No Demand Observed mi
(i)
(Xi )
Freq (Oi )
1
0-5
4
2.5
2
5-10
9
7.5
3
10-15
15
12.5
4
15-20
20
17.5
5
20-25
18
22.5
6
25-30
7
27.5
7
30-35
5
32.5
Check whether given data follows normal distribution using K-S test at =0.10
Solution
H0: The given data follows a normal distribution
H1: The given data does not follow a normal distribution
The total number of observation n=78. No of classes, Xi=7
mi oi
=17.6282; Variance, 2 =57.033; Std (σ)=7.552; Significance level
oi
X
,=0.10. The formula for the standard normal statistic is Z 0 i
, i= 1, 2……., 7
Mean demand, =
The calculations for D values for K-S test follows
Demand
0-5
5-10
10-15
15-20
20-25
25-30
30-35
Mid points
(Xi )
Observed
Freq (Oi)
Observed
Prob
Z
values
2.5
7.5
12.5
17.5
22.5
27.5
32.5
4
9
15
20
18
7
5
0.0513
0.1154
0.1925
0.2564
0.2318
0.0817
0.0641
-2.00
-1.34
-0.68
-0.02
0.65
1.31
1.97
Observed
Cum prob
(OFi)
0.0513
0.1667
0.3590
0.6154
0.8462
0.9359
1.00
Expected
Cum prob
(EFi)
0.0228
0.0901
0.2483
0.4920
0.7422
0.9047
0.9756
Di=
|OFi-EFi|
0.0285
0.0766
0.1107
0.1234
0.1040
0.0310
0.0244
The maximum value of Dcal =0.1234. The theoretical value of Dn for the sample size (n=78) at
𝞪=0.10 is computed using the following formulae
Page 23 of 62
D78 =
√
=
√
=0.1381
Dcal (0.1234) < tabulated D78(0.1381) .We fail to reject Ho & conclude that the data follows
normal distribution
Assignment
The following frequency distribution gives the number of cars arriving at a certain T junction of
a road in 60 seconds intervals.
No. of cars
Observed
Frequency
0
7
1
23
2
49
3
63
4
26
5
20
6
15
7
7
8
5
9
3
10
1
Use Kolmogorov–Smirnov test of goodness of fit to check whether the Poisson distribution
provides a suitable distribution? Use 0.025
Run test for Randomness
Consider the arrival of customers at a branch office of the telephone department for payment of
bills. The sequence of arrival for male and female is as follows
MMFFF MMMFFFMFMMMMFF
No of males =10
No of females=9; No of runs, r=8
We need to test randomness of occurrence of runs at a given significance level 𝞪. Let n1 be
frequency of occurrence of a particular symbol and n2 the frequency of occurrence of another
symbol value to the number of runs. If n1 as well as n2 are less than 20, then the sample is treated
as a small sample. Hypothesis for the run test of randomness is
H0: The occurrence of the runs in a given streams of symbols is random
H1: The occurrence of the runs in a given streams of symbols is not random
Runs test for randomness of small samples
When n1+n2 <20
Observed number of runs =r
If r≤ critical value of more than or equal to the larger critical value at 𝞪 then reject Ho otherwise
fail to reject H0
Page 24 of 62
Example (we use the data given above)
No of males n1=10
No of females n2=9
No of runs r=8
Significance level 𝞪=0.05
n1+n2=19<20
small sample
H0: The occurrence of the runs of a given stream of symbols (M,F) is random
H1: The occurrence of the runs of (M,F) are not random
From table smaller critical values & large critical value for the given combination of n1(10) or
n2(9) at 𝞪=0.05 are 5 & 16 respectively
Since calculated number of runs rcal (8) is not less than (5) as well as greater than 16 we fail to
reject Ho and conclude that the occurrence of (F/.M) are random
Run Test of Randomness of large samples
Should n1 or n2 or both have more than 20, then the sample is treated as a large sample. Under
such a situation one can approximate the sampling distribution of r to normal distribution with
the following mean and variance
Mean of r,
2n1n2
1
n1 n2
Variance of r, 2
2n1n2 2n1n2 n1 n2
n1 n2 n1 n2 1
2
The standard normal statistics to test the significance of r is Z
r
This is a two-tailed test hence Z /2 or Z /2 define the left critical and right critical values of r
.If the calculated ᵶ -values is between Z /2 and Z /2 accept H0; otherwise reject H0
Example: Large sample
The marketing manager is analyzing the outcomes of different potential quotations from
customers. The outcome is winning (W) or loosing (L). The sequences of outcomes of 40
different quotations are given below. Check whether different events of winning or losing the
orders is random at =0.05
Page 25 of 62
WWLLWWWWWWLLWWWLWWWLLWWLLWWLLLW LLWWW
L LWW
Solution
Frequency of occurrence of W, n1 =24
Frequency of occurrence of L, n2 =16
No of runs
r=17
Significance level =0.05;
=0.025
Since n1>20 the sample problem is from large samples
H0: The occurrence of the runs of the given streams of symbols (W,L) is random
H1: (W,L) is not random n1n2
Mean of r, =
Variance of r,
2n1n2
2 x24 x16
1
1 20.2
n1 n2
24 26
2=
2n1n2 2n1n2 n1 n2
n1 n2 n1 n2 1
2
=√
2 x 24 x16 x 2 x 24 x16 24 16
24 16 24 16 1
2
=8.96
=2.993
The formulae for the standard normal statistic to test the significance of r is
ᵶ=
r
=
=-1.069
The value of Z 0.025 1.96 & Z 0.025 1.96
Since the calculated Z 0 1.069 of r falls in the acceptance region we fail to reject Ho and
conclude that W& L are random
References
a) "Runs Test for Detecting Non-randomness".
b) Sample 33092: Wald–Wolfowitz (or runs) test for randomness
c) Magel, RC; Wibowo, SH (1997). "Comparing the Powers of the Wald–Wolfowitz and
Kolmogorov–Smirnov Tests". Biometrical Journal. 39 (6): 665–675.
doi:10.1002/bimj.4710390605.
d) Barton, DE; David, FN (1957). "Multiple runs". Biometrika. 44 (1–2): 168–178.
doi:10.1093/biomet/44.1-2.168.
e) Sprent P, Smeeton NC (2007) Applied Nonparametric Statistical Methods, pp. 217–219. Boca
Raton: Chapman & Hall/ CRC.
Page 26 of 62
f) Alhakim, A; Hooper, W (2008). "A non-parametric test for several independent samples". Journal
of Nonparametric Statistics. 20 (3): 253–261. CiteSeerX 10.1.1.568.6110.
doi:10.1080/10485250801976741.
Assignment
1. An economic entomologist rates the annual incidence of damage by a certain beetle as
mild (M) or heavy (H). For a 27 year period he records the following: H M M M H H M
M H M H H H M M H H H H M M H H M M M M. Test the null hypothesis that the
incidence of heavy damage occurs randomly over the years.
2. The winning numbers for the Kenya sweepstake drawing for the month of January are
listed below:
457 605 348 927 463 300 620 261 614 098 467 961 957 870 262 571 633 448
187 462 565 180 050 004 530 350
Classifying each number as odd or even, test for randomness at 0.05 .
TWO-SAMPLE MEDIAN TEST
The different Sign test for two samples as previously presented are applied if the two samples
have the same size. The sign test cannot be applied when two independent samples with different
sizes are taken from two population .In this case use median test.
The objective of the median test is to check whether the two independent samples are from two
different populations with the same median test
H0: The two independent samples are from the population with the same median
H1: The two independent samples are not from the population sample with same median
Procedure for Median Test
Step 1: Pool the observation of the two samples and order them in ascending order
Step 2: Find the median of the combined observation
Step 3: Form frequency table for pooled observation. Let n1 be the size of 1st sample, n2 be the
size of 2nd sample & N=n1+n2 & be the median of the pooled observation
Page 27 of 62
Frequency of the pooled observations
Sample 1
Sample 2
Above median
p
q
Below median
r
s
Column total
n1
n2
Step 4: If pooled sample is small, go to step 5 otherwise go to step 6
Row total
p+q
r+s
N= n1 n2
Step 5 : 5,1 Compute the probability which is to be compared with the given significance level
using the formulae
n1 n2
p q
n !/ (n1 p)!n2 !/ (n2 q)!
P= 1
n1 n2 n1 n2 !/ n1 n2 p q !
pq
5.2 If calculated value of P in step 5.1 is more than 𝞪- level accept Ho; otherwise reject Ho go to
step 7
Step 6.1: Compute 2 statistic using the formulae
2
N
N | ps qr |
2
2
p q r s p r q s
N n1 n2
6.2 If the calculated 2 is less than tabulated 2 Accept Ho; otherwise reject Ho and go to step 7.
Step 7: End
Example: -Small samples
In a computer company, the increment of monthly sales revenue (ksh) in a randomly selected
sales region against two different pure selling strategies (increased warranty period and increased
price discount) are shown below. Check whether two samples are from the population with the
same median at 𝞪=0.05
Sales
Increased
Increased price
region warranty period
discount
1
20
23
2
25
27
3
41
30
4
35
42
5
60
56
6
40
46
Page 28 of 62
7
Solution
28
-
The combination of hypothesis for the two samples median test is
H 0 : The incremental monthly sales revenue due to two selling strategies has the same median
H1: The incremental monthly sales revenue due to two selling strategies do not have the same
median.
Size of the 1st sample n1=7
Size of the 2nd sample n2=6
Size of the pooled observation ( n1 n2 )=13
Level of significance 𝞪=0.05
Ordered pooled observation in ascending are 20, 23, 25, 27, 28, 30, 35, 40, 41, 42, 48, 56, 60
35 median
The classification of the pooled observation is given below
Above median
Below median
Column total
sample 1
p=3
r=4
p+r=7
sample 2
q=3
s=3
q+s=6
Row total
p+q=6
r+s=7
N= n1 n2 =13
76
3 3
7!/ 4!*6!/ 3!
0.408 ; Calculated value of P(0.408) >𝞪=0.05; we fail to reject H 0
P=
13!/ 6!
13
6
and conclude that the two samples have the same median
Example –Large Samples
In a chemical company, the incremental weekly productions in tons using two catalysts are given
below. Check whether two samples are from the population with the same median at 𝞪=0.05.
Week
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17
25 30 46 40 65 45 45 33 35 41 51 64 66 77 82 96 68
28 32 42 47 62 60 53 29 34 54 57 63 70 80 78 93 Solution. The combination of hypotheses for the two-samples median test is
Catalyst X
Catalyst Y
H 0 : The incremental weekly production of chemical company using two catalysts has the same
median
Page 29 of 62
H1: They have different medians
Size of the 1st sample n1 =17
Size of the 2nd sample n 2 =16
Size of the combine samples n1 n2 =33
Level of significance = 𝞪=0.05
The pooled observation in increasing order are: 25, 28, 29, 30, 32, 33, 34, 35, 40, 41, 42, 45, 46,
47, 51, 53, 54, 57, 62, 63, 64, 65, 66, 68, 70, 77, 78, 80, 82, 93, 96,
The median 53
The classification of the pooled observation is given below
sample 1
Above median
p=7
Below median
r=10
Column total
p+r=17
Same n1+n2 is large, the χ2 statistic is used.
2
sample 2
q=9
s=7
q+s=16
Row total
p+q=16
r+s=17
N= n1 n2 =33
2
N
33
N |ps qr|
33* |7 x7 10 x9|
2
2
2
0.0303
17x16x17x16
p q r s p r q s
02.05 (1) 3.841=3.841>Calculated 2 (0.0303). we fail to reject Ho and conclude that the two
samples have same median
K-SAMPLES TESTS
There may be situations in which simultaneously more than two samples (K-samples) are to be
checked whether they are from identical populations. The following tests can be used for testing
such K-samples.
K-samples median test, Kruskal –Willis test
K-samples median test
They are two types
a) K-samples median test for small samples
b) K-samples median test for large samples
Page 30 of 62
The objective of K-samples median test is to check whether the K-independent samples are from
K-different population with the same median. The combination of hypothesis for the K-samples
median test is as follows
H0: K- Independent samples are from the population with same median
H1: K- Independent samples are from the population with different medians.
Procedure for K-samples median test
Step1: Pool the observation of K samples & order them in ascending order
Step 2: Find the median of the combined observations
Step3: Form frequency tables for pooled observations. Let K be the number of samples; n j be the
sizes of the sample. j =1, 2,…..,k: N be the size of the pooled observation (n1+n2+……..+nk=N).
And be the median of the pooled observation.
Above median
1
p1
Samples
2
3 ………….k
p2 p3
Below median
q1
q2
q3
Row total
p1 + p2 +…+ p k
q1 + q 2 +….+ q k
n3 ---------- n k N n1 n2 .... nk
Step 4: If the pooled sample is small, go to step 5 .if it is large, go to step 6
Column total
n1
n2
Step 5: 5.1: Compute the probabilities which is to be compared with the probability of
significance level 𝞪 using the formulae
n1 n2 nk
...
p
p
p
n !/ (n1 p1 )!n2 !/ (n2 p21 )!....nk !( nk pk )!
P 1 2 k 1
n1 n2 ... nk
n1 n2 ... nk !/ ( p1 p2 ... pk )!
p1 p2 ... pk
5.2. If calculated value of P in 5.1 is more than 𝞪-level accept Ho, otherwise reject H0 go to step
7
Step 6. 6.1 Compute 2 statistic
6.2 If calculated 2 is less than tabulated 2 (K-1) accept Ho , otherwise reject Ho and go to
step 7
Step 7 : End
Example In a company, the monthly percentage of absenteeism of employees in the first quarter
of the last 5 years is given below. Check whether there is significance difference between the
monthly percentage absenteeism of the employees using the median test at 𝞪=0.05
Page 31 of 62
Percent absenteeism of employees
January February March
10
18
13
15
25
21
8
16
26
20
17
14
19
11
7
Solution
Combination of hypothesis for the three samples median test is
H 0 : There is no significance between the percentage absenteeism of employees in different
months
H1: There is significance difference between the percentage absenteeism of employees in
different months
Size of 1stsample, n1=5
Size of 2nd sample, n2=5
Size of 3rd sample, n3=5
Size of pooled samples ( n1 n2 n3 )=15
Level of significance 𝞪=0.05
Ordered pooled samples are 7, 8, 10, 11, 13, 14, 15, 16, 17, 18, 19, 20, 21, 25, 26
The median of the pooled observations 16 .
The classification of the pooled observation
Above median
Below median
Column total
Samples
2
3
p2 3 p3 2
q1 3 q2 2 q3 3
5
5
5
Row total
1
p1 2
7
8
N= n1 n2 n3 =15
5 5 5
2 3 2
5!/ 3!*5!/ 2!*5!/ 3!
P
0.155
15!/ 7!
15
7
P(0.155)>𝞪=0.05. Hence accept Ho and conclude that there is no significance difference
between absenteeism of the employees in the three months
Page 32 of 62
Example: In a survey, the age of respondents in three different regions are summarized. Check
whether there is a significant difference between the ages of the respondents in the three different
regions using median test at 𝞪=0.05
Age of respondents
Region 1 Region2 Region3
35
45
38
42
55
26
28
32
43
31
47
30
44
53
60
50
46
39
33
41
54
52
34
57
56
49
27
Solution:
The combination of hypothesis for the three samples median test is as follows
Ho: There is no significant difference between the ages of respondents in three regions
Hi: There is significant difference between the ages of respondents in 3 regions
Size of 1st sample n1=9
Size of 2nd sample, n2=9
Size of 3rd sample, n3=
Level of significance 𝞪=0.05
The pooled observation in ordered format are: 27, 24, 30, 31, 32, 33, 34, 35, 36, 38, 39, 41, 42,
43, 44, 45, 46, 47, 49, 50, 52, 53, 54, 55, 56, 57, 60; Median=43
The classification of the pooled observations
Samples
1
2
3
p1 4 p2 6 p3 3
q1 5 q2 3 q3 6
9
9
9
Above median
Below median
Column total
Row total
13
14
N= n1 n2 n3 =27
Since (n1+n2+n3) is large we calculate statistic is to be used for this problem
2
2
3
2
i 1 j 1
O
ij
Eij
Eij
2
2.0769 , where Eij
Rowtotal * columntotal
N
Page 33 of 62
02.05 (2) 5.991; Calculated 2 (2.0769) is less than tabulated 2 (5.991); we accept H0 and
conclude that the ages of respondents in 3 regions are the same
Kruskal-Wallis Test (H-test)
These methods can be used when the data cannot be measured on a quantitative scale, or when
the numerical scale of measurement is arbitrarily set by the researcher, or when the parametric
assumptions such as normality or constant variance are seriously violated.
Kruskal-Wallis test similar to the K-samples median test in which the objective is to check
whether the K independent samples have the same median, hence identical populations. This is
called H-test. This is a non-parametric test which is an alternative to the single factor one-way
ANOVA
Ho: K –independent samples are from K identical populations
H1: K –Independent samples are not from K identical populations
The observations of all K samples are pooled together and then ranks are established. The
smallest value will have rank 1, and the largest N. In case of any ties then the respective ranks
are averaged.
H follows 2 distribution with (K-1) df. The formulae for H-statistic of Kruskal -wills test is
R 2j
12
H
3( N 1)
N ( N 1) j 1 n j
where Rj is the sum of ranks of the sample j; nj is the size of the sample j,
j=1,2,………,K and N=n1+n2+,…….+nk
2
The calculated Ho value is to be compared against tabulated 0.05
k 1
Rule for Decision Making
2
If H>tabulated 0.05
k 1 reject Ho otherwise accept H0
Example:
Production volume of units assembled by 3 different operators (1, 2, 3) during 9 shifts are given
below. Check whether there is significant difference between production volume of units
assembled by 3 operators using Kruskal-Wallis test at 𝞪=0.05
Production volume of Units assembled
Shift No Operator 1 Operator 2 Operator 3
1
29
30
26
2
34
21
36
Page 34 of 62
3
4
5
6
7
8
9
34
20
32
45
42
24
35
23
25
44
37
34
19
38
41
48
27
39
28
46
15
Solution: No of shifts worked by 3 operators =9 each
N=n1+n2+n3=27
H0: There is no significance between production vols of units assembled by 3 operators
H1: There is significant difference between the production vols of units assembled by 3 operators
Rule for decision making. If calculated H> 02.05 (k 1) , reject H0; otherwise accept H0
We now rank the observations polled together
Sample No
3
2
1
2
2
1
2
3
3
3
1
2
1
Observation
15
19
20
21
23
24
25
26
27
28
29
30
32
Rank
1
2
3
4
5
6
7
8
9
10
11
12
13
Sample No
1
1
2
1
3
2
2
3
3
1
2
1
3
3
Observation
34
34
34
35
36
37
38
39
41
42
44
45
46
48
Observation 34 is repeated 3 times hence revise rank is
Rank
14}
15}15
16}
17
18
19
20
21
22
23
24
25
26
27
(
Ranks for sample 1= (3, 6, 11, 13, 15, 15, 17, 23, 25)=R1=128
Ranks for sample 2 = (2, 4, 5, 7, 12, 15, 19, 20, 24,)=R2=108
Ranks for sample 3= (1, 8, 9, 10, 18, 21, 22, 26, 27,)=R3=142
H= 2 (3 1) => 2 ( 2) =5.991
Page 35 of 62
)
=15
H=
3 R2
1282 1082 1422
12
12
j
=
3
(
N
1
)
3(27 1) 1.03
27 27 1 9
9
9
N ( N 1) j 1 n j
Since H(1.03)< 2 ( 2) =5.991
We accept Ho and conclude that there is no significant difference between the production
volumes under 3 operators
Differences between several related groups: Friedman's ANOVA
Friedman's ANOVA is the non-parametric analogue to a repeated measure ANOVA where the
same subjects have been subjected to various conditions.
Example here: Testing the effect of a new diet called 'Andikins diet' on n=10 women. Their
weight (in kg) was tested 3 times: Start; Month 1; Month 2
Would they lose weight in the course of the diet? Subject's weight on each of the 3 dates is listed
in a separate column. Then ranks for the 3 dates are determined and listed in separate columns.
Then, the ranks are summed up for each Condition (Ri)
Diet data with Ranks
Person
Start
Month 1
(ranks)
(Ranks)
1
63.75
65.38
81.34
1
2
2
62.98
66.24
69.31
1
2
3
65.98
67.7
77.89
1
2
4
107.27
102.72
91.33
3
2
5
66.58
69.45
72.87
1
2
6
120.46
119.96
114.26
3
2
7
62.01
66.09
68.01
1
2
8
71.87
73.62
55.43
2
3
9
83.01
75.81
71.63
3
2
10
76.62
67.66
68.6
3
1
R1=19
R2=20
From the sum of ranks for each group, the test statistic Fr is derived:
Fr
start
Month 1
Month2
Month2
(Ranks)
3
3
3
1
3
1
3
1
1
2
R3=21
12
12
Ri2 3N (k 1) =
192 202 212 (3 *10)(3 1)
Nk (k 1) i 1
10 * 3(3 1)
= 12/120 (361+400+441) – 120=0.1 (1202) – 120=120.2 - 120 = 0.2
Page 36 of 62
The F-Statistics is called Chi-Square, here. It has df=2, ((k-1), where k is the number of groups).
The statistics is not significantly different
Assignment
Given below is the amount of food in grams eaten by male hooded rats following 0, 24 and 72 hours of
food deprivation. Test the hypothesis that the data provide sufficient evidence to indicate a difference in
the effects of three levels of food deprivation by Friedman test.
subject Hours of food deprivation
0
24
72
1
3.5
5.9
13.9
2
3.7
8.1
12.6
3
1.6
8.1
8.1
4
2.5
8.6
6.8
5
2.8
8.1
14.3
6
2.0
5.9
4.2
7
5.9
9.5
14.5
8
2.5
7.9
7.9
MATHEMATICAL PROPERTIES OF THE CHI-SQUARE DISTRIBUTIONS
A normal distribution can be transformed to the standard normal distribution
Zi
X i
where Zi is a value from the standard normal distribution, Xi is an observation from
the normal distribution to be transformed , and μ and σ re the mean and standard deviation
respectively of the distribution
If we square the Zi variates we get
X
Z i then Z² follows a Chi-square distribution
2
2
i
APPLICATION OF CHI-SQUARE DISTRIBUTION
1) To test if the hypothetical value of the population variance is the same
2) To test for the goodness of fit
3) To test for the independence of attributes
4) To test the homogeneity of independent estimates of the population variance
5) To combine various probabilities obtained from independent experiments to a single test
of significance
Page 37 of 62
6) To test the homogeneity of independence estimates of the population correlation
coefficient
CHI-SQUARE TEST FOR INDEPENDENCE
A research question that frequently arises is whether two variables are associated. A consumer
protection official may be interested in knowing whether price is associated with the quality of
small household appliances. A school nutritionist may want to know whether the nutritional
status of students is associated with the academic performance, If there is no association between
two variables, we say they are independent. If two variables are associated, knowledge of one is
helpful in predicting the value of the other.
We use the chi-square test of independence to decide whether two variables in a population are
independent.
Assumption:
(i) The data consist of a simple random sample of size n from some population of interest
(ii) The observations in the sample may be cross-classified according to two criteria, so that
each observation belongs to one and only one category of each criterion. The criteria are
the variables of interest in a given situation
(iii)The variables may be inherently categorical, or they may be quantitative variables whose
measurements are capable of being classified into mutually exclusive numerical
categories
The data may be displayed in a contingency table as follows in which the observed number nij of
the subject characterized by one category of each criterion is placed in the cell formed by the
intersection of the ith row and jth column. The cell entries are referred to as observed cell
frequencies, and they are designated Oij -that is Oij = nij . The observed frequency Oij represents
the joint occurrence in the sample subjects of the ith category of the first criterion of classification
with the jth category of the second.
Page 38 of 62
Table 5.1 Contingency table for Chi-square test of independence
First
criterion of
classification
1
Second criterion of classification
category
l
….
…
n1 j
1
n11
2
n12
2
n21
n22
.
I
ni 2
…
nij
…
nic
.
R
.
n i1
.
nr1
nr 2
…
nrj
…
nrc
.
ni
.
nr
Total
n1
n2
…
n j
…
nc
n
…
n2 j
…
c
n1c
Total
n1
n2 c
n2
Hypothesis
H0: The two criteria of classification are independent
H1: The two criteria of classification are not independent
Test statistics
The probability that a subject will belong in cell ij is calculated as follows:
n n j
P(Subject belongs in cell ij)= i
n n
contingency table , we calculate
n n j
Eij n i
n n
r
c
2
i 1 j 1
o
ij
. To estimate the expected frequency for cell ij of the
Eij
2
Eij
. When the differences between observed and expected frequencies are
large, 2 is large, when there is close agreement between them, 2 is small
Example: The following table gives a sample of married women, the level of education and adjustment
in married life scores.
Level of
Adjustment scores
Total
education
Very low
Low
High
Very high
College
24
97
62
58
241
High school 22
28
30
41
121
Primary
32
10
11
20
73
Total
78
135
103
119
435
Test the hypothesis that the higher level of education, the greater is the degree of adjustment in marriage
Page 39 of 62
ni n j
n n
Solution: we use the following formula to calculate the expected cell frequencies Eij n
level of education adjusted score Observed Expected
(O-E)²
College
very low
24
43.21379 369.1698
High school
very low
22
21.69655 0.092081
Primary
very low
32
13.08966 357.6011
College
low
97
74.7931 493.1463
High school
low
28
37.55172 91.23543
Primary
low
10
22.65517 160.1534
College
High
62
57.06437 24.36047
High school
High
30
28.65057 1.820949
Primary
High
11
17.28506 39.50195
College
Very high
58
65.92874 62.86485
High school
Very high
41
33.10115 62.39184
Primary
Very high
20
19.97011 0.000893
2
(O
i 1
i
Ei ) 2
E
2
(Oi Ei ) 2
E
8.542871
0.004244
27.31937
6.593472
2.429594
7.069175
0.426895
0.063557
2.285323
0.953527
1.884884
4.47E-05
57.57296
57.57296 we reject H0 since calculated > 2 0.05 (6) 12.59
CH-SQUARE TEST FOR POPULATION VARIANCE IS THE SAME
H0 : 2 02 versus H1 : 2 02
2
n
yi
SS
1 n 2 i 1
2
Test statistics 0 2 2 yi
0 0 i 1
n
2
0
n 1 s 2
02
follows chi-square distribution with (n-1) df. By comparing calculated 02 with tabulated
2 n 1 , we reject H0 if calculated 02 > 2 (n 1)
2
2
Remarks:
i)
ii)
The above test is valid if the population was drawn from a normal population
Should the sample size , n be large (n≥30) we can use Fisher‟s approximation,
2 2 ~ N
i.e. Z= 2 2
2n 1,1
2n 1 ~N(0, 1) and we can apply Normal Test
Page 40 of 62
Examples
1. It is believed that the precision of the instrument is no more than 0.16. Write down the null
hypothesis and alternative hypothesis to test this belief at α=0.01. The measurement of the
instrument on the same subjects gives the following results: 2.5, 2.3, 2.4, 2.3, 2.5, 2.7, 2.5, 2.6,
2.6, 2.7 and 2.5
Solution:
xi
xi x
x x
2.5
2.3
2.4
2.3
2.5
2.7
2.5
2.6
2.6
2.7
2.5
27.6
-0.01
-0.21
-0.11
-0.21
-0.01
0.19
-0.01
0.09
0.09
0.19
-0.01
0.0001
0.0441
0.0121
0.0441
0.0001
0.0361
0.0001
0.0081
0.0081
0.0361
0.0001
0.1891
x x
11
x
2
i
i 1
2
i
=
1 n
xi 27.6 /11 2.51
n i 1
Under the null hypothesis H0 : 02 0.16 the test statistics
11
02
n 1 s
2
2
0
Tabulated
2
0.01
10 xi x
i 1
2
0
2
10 x0.1891
11.81875 ; which follows chi-square with 10 df.
0.16
2
(10) 23.21 . We fail to reject H since 02 0.01
(10) and conclude that the data is
0
consistence with the hypothesis that the precision of the instrument is no more than 0.16.
2. Example: Test H 0 : 10 given that s=15 for the sample size n=50 from a normal population
Solution: H 0 : 10 , n=50, s=15
02
ns 2
02
50 x152
112.5
102
Since n is large, using the test statistic is
Page 41 of 62
Z= 2 2
2n 1 = 225 99 15 9.95 5.05
Since Z 3 , it is significant at all levels of significance and hence H0 is rejected and we conclude that
10
3. Example: The theory predicts the performance of beans in an experiment in the four groups A, B,
C, D should be in the ratio 9:3:3:1. In an experiment among 1600 beans, the numbers in the four
groups were 882, 313, 287 and 118. Does the experimental result support the theory?
Solution:
H 0 : The theory supports the experimental results
Total number of beans=882+313+287+118=1600; These are to be distributed in the ratio 9:3:3:1
E(882)=
9
3
3
x1600 900 ; E(313)= x1600 300 ; E(287)= x1600 300 ;
16
16
16
E(118)=
1
x1600 100 ;
16
Oi Ei 2 882 900 2 313 300 2 287 300 2 118 100 2
=4.7266
Ei
900
300
300
100
i 1
4
2
0
2
Df=(4-1)=3 and tabulated 20.05 (3) 7.815 . Since calculated 0 is less than the tabulated value,
we fail to reject H0 and conclude that the experiment supports the theory.
4. Example: A survey of 320 families with 5 children each revealed the following distributions:
No of boys
5 4
3
2 1 0
No of girls
0 1
2
3 4 5
No of families 14 56 110 88 40 12
Is this result consistent with the null hypothesis that male and female are equally probable?
Solution:
Let the H0: the data is consistent with equal probability for male and female births. P=the probability of male
birth=½=q
P(r)=probability of “r” male births in a family of 5.
5
5 r 5 r 5 1
= p q
r
r 2
Page 42 of 62
5
5
5 1
=10x r
r 2
The frequency of r male births is given by f(r)=N.p(r)=320x
(*)
Substituting r=0, 1, 2, 3, 4 successively in 9*), we get the expected frequencies as follows
5
5
5
=50; f(2)=10x =100, f(3)=10x =100
1
2
3
f(0)=10x1=10, f(1)=10
5
5
=50 and f(5)=10x =10
4
5
f(4)=10x
Calculation of Chi-square
Observed freq (Oi)
Expected freq (Ei)
Oi Ei
14
56
110
88
40
12
320
10
50
100
100
50
10
320
16
36
100
144
100
4
Oi Ei
2
2
/E
1.6
0.72
1
1.44
2
0.4
7.16
Oi Ei 2
7.16 df=6-1=5df
E
i 1
2
Tabulated 20.05 (5) 11.07 So the calculated chi-square is less than tabulated values.
5. Example: Fit a Poisson distribution to the following data and test the goodness of fit.
x
f
0
275
1
72
2
30
3
7
4
5
fi x
Solution: We estimate the mean of the distribution; x
i 0 fi
6
5
2
6
1
189 0.482
392
i
In order to fit the Poisson distribution we take the parameter, λ as the mean, x =0.482. The frequency of r successes
is given by the Poisson law as
Now f(0)=392 e
0.482
f (r ) Np(r ) 392
e0.482 0.482
r!
392 xanti log 0.482log e
=392 x antilog[-0.482xlog2.7183] {⸪ e=2.7183}
=392 x antilog[-0.482 x 0.4343]
=392 x antilog[-0.2093]
Page 43 of 62
r
, r=0, 1, 2, …, 6
=392 x antlog 1.7907 392 x0.6176
=242.1
f(1)=λ x f(0)=0.482x 242.1=116.69, f(2)=
f(3)=
f(5)=
3
5
xf (1) 0.241x116.69 28.12
2
xf (2)
0.482
0.482
x 28.12 4.518 , f(4)= xf (3)
x 4.518 0.544
3
4
4
xf (4)
0.482
0.482
x0.544 0.052 , f(6)= xf (5)
x0.052 0.004
5
6
6
Hence the theoretical Poisson frequencies corrected to one decimal place are as given below:
x
Expected frequency
0
242.1
1
116.7
2
28.1
3
4.5
4
0.5
5
0.1
6
0
Total
392
Observed freq(O)
275
72
30
Expected freq (E)
242.1
116.7
28.1
(O-E)
32.9
44.7
1.9
(O-E)²
1082.41
1998.09
3.61
(O-E)²/E
4.471
17.121
0.128
7
5
15
2
1
4.5
0.5
5.1
0.1
0
9.9
98.01
19.217
392
392
40.937
Oi Ei
40.937 , df=7-1-1-1-3=2df
E
i 1
2
2
Tabulated 0.05 (2) 5.99 So the calculated chi-square is greater than tabulated values and hence we
conclude that Poisson distribution is not a good fit to the data.
2
Exercise
1. Below is part of the study designated to identify possible subcultural, ethnic, and family
influences on mobility after training. Are the data showing sufficient evidence for us to reject the
H0 that opportunity level is independent of labour force mobility? What is the P-value?
Opportunity levels for underclass whites and labour force mobility
Opportunities Labour force mobility
Low
High
Low
45
19
High
6
43
Total
51
62
Total
64
49
113
Page 44 of 62
2. The following figures show the distribution of digits in the numbers chosen at random from a
telephone directory
digit
Freq
0
1026
1
1107
2
997
3
966
4
1075
5
933
6
1107
7
972
8
964
9
853
total
10000
Test the hypothesis whether the digits occur equally frequently in the directory
3. The number of airplane accidents in an Indian airport per day in a week is reported below. Test
whether the accidents are uniformly distributed in a day
Days
Mon Tues Wed Thurs Friday Sat
Accidents 14
18
15
20
16
14
2x2 CONTINGENCY TABLE for a 2x2 tables
a
b
a+b
c
d
c+d
a+c b+d N=a+b+c+d
Prove that chi-square test of independence gives
2
N (ad bc) 2
, where N=a+b+c+d
a c b d a bc d
Solution:
Under the hypothesis of independence of attributes
E (a)
E (d )
2
a b a c ;
N
E (b)
a b b d ;
N
E (c )
c d a c
N
d b d c
N
[a E (a)]2 [b E (b)]2 [c E (c)]2 [d E (d )]2
……………………………(*)
E (a)
E (b)
E (c )
E (d )
where a-E(a)= a
a b a c a(a b c d ) (a 2 ac ab bc) ad bc
N
N
Similarly, we get
b E (b)
ad bc
c E (c ) ;
N
d E (d )
ad bc
N
Page 45 of 62
N
Similarly in (*), we get
2
ad bc2
N
1
1
1
1
E (a) E (b) E (c) E (d )
2
ad bc
N
1
1
1
1
a b a c a b b d a c a c d b d c
ad bc2
N
bd ac
bd ac
a b a c b d a c d c b d
cd ab
(ad bc) 2
a ba c b d c d
N (ad bc) 2
a c b d a bc d
YATE’S CORRECTION
In a 2x2 contingency table, the number of degree of freedom (2-1)(2-1)=1. If any one of the
theoretical cell frequency is less than 5, then the use of pooled method for 2 -test results with
2 with 0 df. In this case we apply a correction due to F.Yates (1934) known as Yate‟s
Correction for continuity
For a 2x2 contingency table we have
N (ad bc) 2
2
a c b d a bc d
a
b
c
d
According to Yate‟s correction , we subtract (or add) ½ from a and d add (subtract)½ to b and c
so that the marginal totals are not distributed at all. Thus, the correction value of 2 is given as
c2
N (a ½)(d ½) (b ½)(c ½)
a cb d a bc d
2
Page 46 of 62
N
2
Numerator N ad bc ½(a b c d) N | ad bc |
2
2
2
N
N | ad - bc | -
2
c2
a c b d a b c d
Exercise
a) Two lotions are used for getting relief by a group of patients. Test the hypothesis that the
proportion of persons experiencing relief is the same in both lotions at 0.05 .
Lotion 1
Relief No relief
Lotion 2 Relief
11
6
No relief 10
24
i) Examine whether Yates‟s correction modifies the conclusion or not
ii) Compare results from McNemar‟s test and Yate‟s correction for continuity
iii) Derive Yate‟s correction factor for continuity formula
MCNEMAR TEST FOR TWO RELATED SAMPLES
ASSUMPITION
The data consists of N randomly selected subjects (or items) or pairs of subjects, depending on whether
subjects acts as their own controls or whether experimental subjects are paired with a matched control
The measurement scale is nominal with four categories. When subjects are their own control, they are
independent of each other. When matched pairs are used, the pairs are independent, but observation on
each pair are related
Hypotheses
A. We may wish to test the null hypothesis that the proportion of items or subjects with the
characteristic of interest is the same under two conditions (or treatments). We let p 1 be the
proportion with the characteristic of interest under the other condition
H 0 : p1 p2 or p1 p2 0
H1 : p1 p2 or p1 p2 0
Page 47 of 62
B. If we want to see whether we can conclude that the proportion with characteristic of interest
under condition 1 is higher than under condition 2, we have a one-sided test, and the hypotheses
are
C. H 0 : p1 p2 or H1 : p1 p2
D. H 0 : p1 p2 or H1 : p1 p2
Test statistics
p1
A B
AC
and p2
N
N
Where the characteristic of interest is a Yes response
The difference between sample proportions is
p1 p2
A B AC B C
N
N
N
The null hypothesis is that the expectation of (B-C)/N is zero. McNemar (T40) shows that an
appropriate test statistics is
Z=
B C
when H 0 : p1 p2 is true, z is distributed approximately as the standard normal
BC
variate, provided that B+C is at least 10
Decision Rule
A. Reject H0 at α level of significance if the computed z is equal to or greater than the z to the right of
which lies α/2 of the area under the standard normal curve (two-sided test)
B. Reject H0 at α level of significance if the computed z is greater than or equal to the tabulated value
that has α of the area to the right.
C. Reject H0 at α level of significance if the computed z is less than or equal to the tabulated z that has α
of the area to its left.
The McNemar test has been recommended for use in certain before- and after experiments when the
experimenter is interested in the number of subjects who respond differently after they are exposed to
some intervening condition or treatment
Example: 85 matched patients treated for Hodgkin‟s disease with a nonpatient sibling of the same
sex and within five years of age of the patient. The question to be answered is whether there is a
differential rate of tonsillectomies in the two groups
Page 48 of 62
Hypothesis: The characteristic of interest is a history of a tonsillectomy in a given subject. The
appropriate null hypothesis is that the proportion of tonsillectomies in the population of Hodgkin‟s
disease patients is the same as in the population of matched siblings at α=0.05
History of tonsillectomy in patients with Hodgkin‟s disease and in matched controls
Tonsillectomy in matched controls
Yes
No
Total
Tonsillectomy
Yes
26
15
41
in patients
No
7
37
44
Total
33
52
85
H 0 : p1 p2 or H1 : p1 p2
Test Statistics
Z=
15 7
1.71
15 7
Since 1.71 <Z=1.96, we cannot reject H0. Since we have a two-sided test, the p value is
2(0.0436)=0.0872
Exercise
Respond of Rabbits to light is under study. Individual ticks were placed first in an arena that was 1 inch in
diameter, then in arena 2 inches in diameter. The response of interest was whether the tick left the arena
on the side toward the light or away from the light. Can we conclude from these data that the size of the
arena affects the response of rabbit ticks to light by McNemar at α=0.05? What is the P-value?
Response of rabbit ticks to light
Left toward light from1-inch arena
Yes
No
Total
Left toward light from2-inches arena Yes
8
9
17
No
5
8
13
Total 13
17
30
KENDALL’S TAU
Like the spearman‟s rank correlation, Kendall‟s ˆ is based on the ranks of observations, and can assume
values between -1 and +1. Despite this similarity when we compute them using similar data they give
different values. The main difference between the r and ˆ , ˆ provides an unbiased estimator of a
population parameter, while the sample statistic r does not provide an estimate of a population coefficient
of rank correlation.
The parameter estimated by ˆ may be defined as the probability of concordance minus the probability of
discordance.
Page 49 of 62
The observation pairs X i , Yi and X j , Y j are said to be concordant if the difference between X i and
X j is in the same direction as the difference Yi and Yj . In other words, if either X i > X j and Yi > Yj or
X i < X j and Yi < Yj we have concordance. The observation X i , Yi and X j , Y j are said to be
discordant if the directions of the differences are not the same. If
X
i
X j and/or Yi Y j , the
observation pairs are neither concordant nor discordant.
The objective when we use Kendall‟s ˆ for inferential purposes is to test the null hypothesis that X and Y
are independent (which implies
0 ) against one of the following alternatives 0, 0 or 0 .
We may interpret 0 to indicate a direct association between X and Y, and
0 to mean that X and Y
are inversely associated.
Assumptions
A. The data consist of a random sample of n observation pairs (X, Y) of numeric or nonnumeric
observations. Each pair of observations represents two measurements taken on the same unit of
association.
B. The data are measured on at least an ordinal scale, so that we can rank each X observation in
relation to all other observed X‟s and each y observation in relation to all other observed Y‟s.
Hypothesis
A. (Two-sided)
H0: X and Y are independent
H1 : 0,
B. (One-sided)
H0: X and Y are independent versus H1 : 0
C. (one-sided)
H0: X and Y are independent versus H1 : 0
Test statistics
The test statistics, which is also a measure of association in the sample, is given by
ˆ
S
n(n 1) / 2
When n is the number of (X, Y) observations (or ranks). To obtain S, and consequently ˆ , we
proceed as follows:
1) Arrange the observations X i , Yi in a column according to the magnitude of the X‟s , with
the smallest X first, the second smallest second, and so on. Then we say that the X‟s are in
natural order.
2) Compare each Y value, one at a time, with each Y value appearing below it. In making these
comparisons, we say that a pair of Y values (a Y being compared to the Y below it) is a
Page 50 of 62
natural order if the Y below is larger than the Y above. We say that a pair of Y values is in
reverse natural order if the y below is smaller than the Y above
3)
Let P be the number of pairs of natural order and Q the number of pairs in the reverse natural
order.
4) S=P-Q; that is, S in the above equation is equal to the difference between P and Q
n
2
A total of n(n 1) / 2 possible comparison of Y values can be made in this manner. If all
the Y pairs are in natural order, then P=n(n-1)/2, Q=0, S=[n(n-1)/2]=-0=n(n-1)/2, and we have
ˆ
n n 1
n n 1
1
Indicating perfect direct correlation between the rankings of X and Y. On the other hand, if all the
Y pairs are in reverse natural order, we have P=0, Q=n(n-1)/2, S=0-[n(n-1)/2]=-n(n-1)/2, and
ˆ
n n 1
1 , indicating a perfect inverse correlation between the X and Y rankings.
n n 1
Thus ˆ cannot be greater than +1 or less than -1. We can think of ˆ as a relative measure of the
extent of the disagreement between the observed order of the Y observations and the two
orderings that represent a perfect correlation between the X and Y rankings. If the number of Y
pairs that are in natural order exceeds the number in reverse natural order, we have a direct
correlation between the X and Y rankings, and ˆ is positive. If the number of Y pairs that are in
reverse natural order exceeds the number in natural order, we have an inverse correlation between
the X and Y rankings, and ˆ is negative. The strength of the correlation is indicated by the
magnitude of the absolute value of ˆ .
Decision Rule
The decision rules for the three sets of hypotheses are as follows
A. (Two-sided). Reject H0 at α level of significance if the computed value of ˆ is either positive
or larger than the ˆ * entry for n and α/2, or negative and smaller than the negative of the ˆ *
entry for n and α/2.
B. (One-sided): Reject H0 at the α level of significance if the computed value of ˆ is positive
and larger than the ˆ * entry for n and α
C. (One-sided): Reject H0 at the α level of significance if the computed value of ˆ is smaller
than the negative of the ˆ * entry for n and α
Page 51 of 62
Example
Using the data below we wish to compute ˆ to check whether there is sufficient evidence to
conclude that benchmark achievement and management rating are directly related
Territory rankings based on Benchmark achievement and performance rating
Benchmark
achievement
(X)
2
9
7
23
5
17
16
25
4
10
20
15
8
Territory
1
2
3
4
5
6
7
8
9
10
11
12
13
Management
rating (Y)
4
2
20
17
5
7
6
24
3
21
18
9
8
Territory
14
15
16
17
18
19
20
21
22
23
24
25
Benchmark
achievement
(X)
11
1
21
14
3
13
18
22
19
24
6
12
Management
rating (Y)
10
1
14
15
11
13
19
25
16
23
22
12
Hypotheses
H0: Benchmark achievement and management rating are independent
H1: Benchmark achievement and management rating are directly related ( 0 )
Test statistic
We first arrange the data so that the X ranks are in natural order. The number of Y pairs in natural
and reverse natural order with respect to each Y is shown in the table below.
From the table we compute S=P-Q=218-82=136 so that
F
136
136
0.45
25 24 / 2 300
Decision
With n=25, from tables reveals that we can reject H0 at α=0.005 level since ˆ =0.45 is larger than
ˆ =0.367. we can conclude that there is a direct relationship between benchmark achievement
and management ranking in the sampled population
Arrangement of data for computing ˆ
Y pairs in
Y pairs in reverse
(X,Y) rankings
natural order
natural order
(1, 1)
24
0
Page 52 of 62
(2,4)
(3,11)
(4, 3)
(5,5)
(6, 22)
(7, 20)
(8, 8)
(9, 2)
(10, 21)
(11, 10)
(12, 12)
(13, 13)
(14, 15)
(15,9)
(16, 6)
(17,7)
(18, 19)
(19, 16)
(20, 18)
(21, 14)
(22, 25)
(23, 17)
(24, 23)
(25, 24)
21
14
20
19
3
4
14
16
3
11
10
9
7
8
9
8
3
5
3
4
0
2
1
0
P=218
2
8
1
1
16
14
3
0
12
3
3
3
4
2
0
0
4
1
2
0
3
0
0
0
Q=82
PHI COEFFIENT
The phi coefficient was designed for use with dichotomous variables- variables that can assume only one
of the two possible mutually exclusive values, like gender (male, female); products (defective, nondefective), marital status (married, not married)
The phi coefficient.
ad bc
a b c d a c b d
The Phi coefficient may assume values -1 and +1. The phi coefficient is related to the ch-square statistic.
The relationship is expressed by
2 2 / n
To determine whether a computed value of Ф is significant, we may convert it to chi-square by
2 n 2
Page 53 of 62
We compare the Chi-square with tabulated 2 1
Example
A sample of 125 employees classified by gender and sexual harassment on the job
Gender Sexual harassment
Yes
No
Total
Male
15
35
80
Female
50
25
75
Total
65
60
125
ad bc
a b c d a c b d
15 x 25 35 x50
0.3595
50 x75 x65 x60
We now have a measure of the strength of association between gender and sexual harassment experience
of 125 workers
To test for significance we first use the equation
2 n 2 125 0.3595 16.16 >3.841
2
YULE’S Q
When measuring the strength of the association between two dichotomous variables, some researchers
prefer a statistic that was introduced and called Q by Yule (1900)
Q
ad bc
ad bc
Like the phi coefficient, Q may assume any value between -1 and +1, inclusive
Using the example above
Q
ad bc 15 x 25 35 x50
0.647
ad bc 15 x 25 35 x50
The following two measures of association are for use with r x c contingency tables, i.e. tables in which
there are two categorical variables and one or both have more than two categories.
Page 54 of 62
The CRAMER STATISTICS
A statistic suggested by Cramér provides an appropriate measure of the strength of association between
two categorical variables yielding data that may be displayed in a contingency table of any size. When the
contingency table has two rows and two columns, the Cramér coefficient yields a value that is identical,
except for a possible difference in sign, to the contingency coefficient. The Cramér coefficient is defined
by
C
2
n t 1
Where χ² is a chi-square statistic computed by
2
n
Oi Ei
i 1
Ei
2
, n=sample size and t=number of either
rows or columns in the contingency table, whichever is smaller.
Example: A survey was conducted among homeowners in a certain state. The response on how satisfied
on their place of residence is considered
We wish to use the Cramér statistic to measure the strength of the association between place of residence
and level of satisfaction with community of residence. We compute χ²=53.178. n=230, number of rows=3
is smaller than the number of columns, t=3-1=2
Place of residence
Level of satisfaction with community of residence
Very satisfied
Satisfied
Unsatisfied
Very unsatisfied
Rural
30
16
10
5
Suburban
40
20
15
10
Urban
10
15
20
40
C
2
n t 1
53.178
0.34
230(2)
To test C for significance, we compare χ² in the equation with tabulated χ²(r-1)(c-1). The significance of
C depends on the significance of χ². Value of C ranges 0≤C≤1. Value of C=0 will mean there is no
association between the two variables. If r=c then C=1, implies perfect correlation between the two
variables
Page 55 of 62
Exercise
The study of the views of the criminal justice system mentioned in the problems follows
Fear of punishment is the best way to discourage criminal acts
Number of criminal justice
courses completed
Agree
Disagree
Total
None
17
3
20
1-3
18
9
27
4+
21
19
40
Total
56
31
87
a) Calculate Cramér C and interpret the results
b) Is the result significant at α=0.05
MODELLING BINARY DATA
Suppose we have a number of individuals, each yielding an observation which takes one of two
possible forms.
Examples:
1. Electronic component: defective / non defective
2. Insect: Survives / dies from exposure to a given dose
3. Patient: obtains / does not obtain relief from symptoms.
Each experimental unit yields a binary observation, usually coded as 0 (failure) and 1 (success).
Data can be grouped or ungrouped.
Example 1:
Germination of seeds:
Batches of 10 seeds exposed to different temperatures, each seed germinates (1) or fails to
germinate (0).
Denote response for i‟th individual by r.v. zi, coding the two possible values of zi 0 and 1.
Write P(zi=1) =pi – success probability; P(zi=0) 1-pi – failure probability; then
E(zi)=pi =0 x P(z=0) + 1 x P(z=1).
Page 56 of 62
Now suppose we have n independent sets of binary observations such that in the i'th set are yi
successes (1‟s) in ni observations with same success probability pi, i=1(1)N. Y is an observation
on a binomial r.v. Yi==Zi B(ni,pi) and Z Bernoulli and P(Z=z) = Pz(1-P)1-z, z=0,1; where
grouped yi B(ni,pi).
How can we assess the dependence of pi on explanatory variables or groupings of the data?
The model that is normally used to do this is the logistic model i.e.
p
logit(pi)= log( i ) ; normally referred to as the linear logistic model for binomial data.
1 pi
Consider two independent samples of binary data, of sizes n1 and n2, and denote the success
probabilities underlying each sample by p1 and p2. If the number of success in the two samples
are y1 and y2, the estimated success probabilities are given by p1
y1
y
and p 2 2 ,
n1
n2
respectively. Since the two estimates are independent of one another, the variance of p1-p2, the
estimated difference between the two success probabilities, is given by
Var (p1-p2)=Var(p1) + Var(p2) =
p1 (1 p1 ) p 2 (1 p 2 )
.
n1
n2
Odds and odds ratio
It is sometimes helpful to describe the chance that a binary response variable leads to a success
in terms of the odds of that event. The odds of a success is defined to be the ratio of the
probability of success to the probability of failure. Thus if p is the true success probability, the
odds of success is
p
. If the observed binary data consist of y successes in n observations,
(1 p )
the odds of a success can be estimated by
p
y
=
.
(1 p ) (n y )
When two sets of data are to be compared, a relative measure of the odds of a success in one set
relative to that in the other is the odds ratio. Suppose that the p1 and p2 are the success
probabilities in these two sets, so that the odds of success in the ith set is
Page 57 of 62
pi
, i=1, 2. The
(1 pi )
ratio of the odds of success in one set of binary data relative to the other is usually denoted by ,
so that
=
p1 (1 p1 )
is the odds ratio.
p 2 (1 p 2 )
When the odds of success in each of the two sets of binary data are identical, is equal to one.
This will happen when the two success probabilities are equal. Values of less than one (1)
suggest that the odds of a success are less in the first set of data than in the second, while an odds
ratio greater than one (1) indicates that the odds of a success are greater in the first set of data.
The odds ratio is a measure of the difference between two success probabilities which can take
any positive value, unlike the difference between two success probabilities, p1-p2, which is
restricted to a range (-1, 1).
In order to estimate the ratio of the odds of a success in one data set relative to another, suppose
that the binary data are arranged as in the following 2 x 2 contingency table.
Number of success Number of failures Total
Data set 1
a
b
a+b
Data set 2
c
d
c+d
The estimate success probabilities in the two data sets are p1
a
c
and p 2
, and so
( a b)
(c d )
the estimated odds ratio,, is given by
=
p1 (1 p1 ) ad
=
.
p 2 (1 p 2 ) bc
The estimate is the ratio of the products of the two pairs of diagonal elements in the above 2 x 2
table, and for this reason is sometimes referred to as the cross-product ratio.
In order to construct a confidence interval for the true odds ratio, we first note that the logarithm
of the estimated odds ratio is better approximated by a normal distribution than the odds ratio
itself, especially when the total number of binary observations is not very large. The approximate
standard error of the estimated log odds ratio, log, can be shown to be given by
Page 58 of 62
s.e. (log) ≈
1 1 1 1
( )
a b c d
An approximate 100(1-)% confidence interval for log ranges from [log -z/2 s.e.log] to
[log +z/2 s.e.log], where z/2 is the upper (100/2)% point of the standard normal distribution.
These confidence limits are then exponentiated to get corresponding interval for itself.
If the approximate standard error of itself is required, this can be found using the results that
s.e.() ≈s.e.(log),
Because an odds ratio of unity will be obtained when the success probabilities in the two sets of
binary data are equal, the null hypothesis that the true odds ratio is equal to unity, H0: =1, can
be tested using the statistics.
Worked example:
Table 1: Number of mice developing lung tumours when exposed or not to cigarette smoke
Group
Tumour Present Tumour absent Total
Treated 21
2
23
Control 19
13
32
For the data above, estimate odds of a tumour occurring in mice exposed to cigarette smoke is
21/(23-21)=10.50, while the corresponding estimate odds of a tumour in the control group is
19/(32-19)=1.46. The estimated ratio of the odds of a tumour occurring in the treated group
relative to the control group is given by
=
21x13
10.50
7.184,
7.184
2 x19
1.462
The interpretation of this odds ratio is that the odds of tumour occurrence in the treated group is
more than seven times that for the control group. A 95 % confidence interval for the true log
odds ratio can be found from the standard error of log. Using the formula above,
s.e.(log)=0.823. A 95% confidence interval for log ranges from [log (7.184) -1.96x0.823 to
[log (7.184)+1.96x0.823], i.e. from 0.359 to 3.585. A corresponding confidence interval for the
Page 59 of 62
true odds ratio is the interval from e0.359 to e3.585 that is from 1.43 to 36.04. This interval does not
include unity, which indicates that the evidence that the odds of tumour is greater amongst mice
exposed to the cigarette smoke is certainly significant at the 5% level.
Exercise 1: (Class)
Two types of trees were planted by a farmer, after two weeks the farmer went to check and
counted those that survived and those that had died. The results are as in table 2 below.
Table 2: Survival of trees planted by a farmer
Tree type
Alive
Dead
Calliandra
52
38
Eucalyptus
24
66
a. Calculate the odds of each tree surviving.
b. Calculate the odds ratio and comment
Measures of Association between disease and exposure
The probability of a disease occurring during a given period of time, also known as the risk of a
disease occurrence, is the number of cases of the disease that occur in a time period, expressed as
a proportion of the population at risk. In studies where a population at risk is monitored to
determine the incidence of the disease, the risk can be determined directly from the proportion of
individuals in the sample who develop the disease during the period in question.
The main purpose of estimating risks is to be able to make comparisons between the risk of
disease at different levels of exposure factor, and thereby assess the association between the
disease and that exposure factor. For this, a measure of the difference in risk of disease
occurrence for one level of an exposure factor relative to another is required. Suppose that a
particular exposure factor has two distinct values or levels, labeled exposed and unexposed. For
example, if the exposure factor were smoking history, a person may be classified as a smoker
(exposed) or non-smoker (unexposed). Denote the risk of a disease occurring during a given time
Page 60 of 62
period by pe for the exposed persons and pu for the unexposed person. The risk of the disease
occurring in an exposed person relative to that for an unexposed person is then
pe
.
pu
This is called the relative risk of disease, and is a measure of the extent to which an individual
who is exposed to a particular factor is more or less likely to develop a diseased than someone
who is unexposed. In particular, an exposed person is times more likely to contract the disease
in a given period of time than unexposed person. When the risk of the disease is similar for the
exposed and the unexposed groups, the relative risk will be close to unity (1). Values of less
than one (1) indicate that an unexpected person is more at risk than an exposed person, while a
value of greater than one (1) indicates that an exposed person is at greater risk.
Example
Table 3: Disease state and exposure level
Diseased Not diseased
Exposed
a
b
Unexposed c
d
The estimated risk of disease for individuals in the exposed and unexposed groups are given by
c
a
pe
and pu
, respectively, and so the relative risk is estimated by
cd
( a b)
p e a (c d )
p u c ( a b)
Exercise 2: (Class)
Two farmers Alex and Maxi took seedlings from a project nursery to plant. They were given
instructions on how to minimize the drying of the seedlings to ensure survival. On the first visit
to the two farms, the researcher was able to assess the number alive at that moment in time, the
results of what the officer observed is in table 4.
Page 61 of 62
Table 4: Farmers and state of trees seedlings
Farmer
State of tree seedlings
Dead
Alive
Total
Alex
52
20
72
Maxi
24
66
90
a. Calculate the risk of giving seedlings to each of the farmers.
b. Calculate the relative risk for the two farmers.
Exercise 3: (Homework to handed in tomorrow)
Table 5: Survival of plum root-stock cuttings
Length of cutting Surviving Not Surviving
Short
107
31
Long
156
84
1. a. Calculate the odds of each type of cutting.
b. Calculate the odds ratio and comment
2. a. Calculate the risk of giving the different types cuttings to farmers.
b. Calculate the relative risk for the types of cuttings.
References:
An introduction to categorical data analysis (1996): Alan Agresti, University of Florida.
Modelling binary data (1999): D. Collet, Department of Applied Statistics, University of
Reading, U.K, Chapman & Hall /CRC, London.
Page 62 of 62
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )