Finite Markov Chains
A complete treatment of the theory of finitf;! ..
Markov chains,suitable as an undergraduate
introduction to probability theory and as a.
reference. Examples from physics,·economjcs
and the life sciences. The central techniques ...
(matrix operations forthe transition matrices
representing the different chains) can be ....
programmed easily. This is particularly impor~
tant in .the application of the theory.
A new editionofa sequel ofthis book, .•..
Denumerable Markov Chains, is published by
Springer-Verlag as Volume 40 of the series
Graduate Texts in Mathematics.
. .!.
ISBN 0-3S7 ~9
ISBN 3-.540-9
Undergraduate Tex~s in Mathematics
Apostol: Introduction to Analytic
Number Theory.
1976. xii, 338 pages. 24 illus.
Halmos: Naive Set Theory.
1974, vii, 104 pages.
Iooss/Joseph: Elementary Stabil
Apostol: Introduction to Analytic
Halmos: Naive Set Theory.
Number Theory.
Armstrong: Basic Topology. 1974, vii, 104 pages. Bifurcation Theory.
1980. xv, 286 pages. 47 illus.
24 illus.
1976. xii, 338 pages. 1983.
xii, 260 pages. 132 illus.
looss/Joseph: Elementary Stability and
Bifurcation Theory. Kemeny/Snell: Finite Markov C
Armstrong: Basic Topology.
Bak/Newman: Complex Analysis.
ix, 224 pages. I! illus.
1976.
iIlus.
1980. xv, 286 pages. 47
1983. xii, 260 pages. 132 illus.
1982. x, 224 pages. 69 illus.
Undergraduate
Analysis
Markov
Chains.
Kemeny/Snell: FiniteLang:
Bak/Newman: Complex
Analysis.
Ranchoff/Wermer:
Linear Algebra
1983.
1976. ix, 224 pages. l!
illus.xiii, 545 pages. 52 illus.
illus. Geometry.
1982. x, 224 pages. 69Through
1983. x, 257 pages. 81 illus. Lang: UndergraduateLax/Burstein/Lax:
Calculus wit
Analysis
Banchoff/Wermer: Linear Algebra
and Computing,
1983. xiii, 545 pages. Applications
52 illus.
Through Geometry. Childs: A Concrete Introduction to
Volume 1.
illus. Algebra.
1983. x, 257 pages. 81Higher
1976. xi,with
513 pages. 170 illus.
Lax/Burstein/Lax: Calculus
1979. xiv, 338 pages. 8 illus. Applications and Computing,
Childs: A Concrete Introduction to
EeCuyer: College Mathematics
Volume 1.
Higher Algebra.
A Programming
Language.
Chung: Elementary Probability
Theory
1976.
xi, 513 pages. 170
illus.
1979. xiv, 338 pages. with
8 illus.
1978. xii, 420 pages. 144 illus.
Stochastic Processes.
1975. xvi, 325 pages. 36 illus.LeCuyer: College Mathematics with
Macki/Strauss: Introduction to
Chung: Elementary Probability Theory
A Programming Language.
Control
with Stochastic Processes.
1978. xii, 420 pages. 144
illus. Theory.
Croom: Basic Concepts of Algebraic
1981. xiii, 168 pages. 68 illus.
36 illus.
1975. xvi, 325 pages. Topology.
1978. x, 177 pages. 4,6 illus. Macki/Strauss: Introduction to Optimal
Malitz: Introduction to Mathem
Control Theory.
Croom: Basic Concepts of Algebraic
Logic:
198!. xiii, 168 pages. 68
iIlus.Set Theory - Computabl
Topology.
Fischer: Intermediate Real Analysis.
Functions - Model Theory.
1978. x, 177 pages, 46 1983.
illus. xiv, 770 pages. 100 illus.
xii, 198 pages. 2 illus.
Malitz: Introduction 1979.
to Mathematical
Set Theory - Computable
Fischer: Intermediate Fleming:
Real Analysis.
Functions of SeveralLogic:
Variables.
Martin: The Foundations of G
Functions - Model Theory.
1983. xiv, 770 pages. 100
illus.edition.
Second
the Non-Euclidean Plane.
1979. xii, 198 pages. 2and
iIlus.
1977. xi, 411 pages. 96 illus.
1975. xvi, 509 pages. 263 illus.
Fleming: Functions of Several Variables.
Martin: The Foundations of Geometry
Second edition.
Foulds: Optimization Techniques:
AnNon-Euclidean
Martin:
and the
Plane,Transformation Geom
1977. xi, 411 pages. 96Introduction.
iIlus.
Introduction
to Symmetry.
ilIus,
1975. xvi, 509 pages. 263
1981. xii, 502 pages. 72 illus.
1982. xii, 237 pages. 209 illus.
Foulds: Optimization Techniques: An
Martin: Transformation Geometry: An
Introduction.
Franklin: Methods of Mathematical
Millman/Parker: Geometry: A
Introduction to Symmetry.
1981. xii, 502 pages. 72Economics.
iIlus.
Linear and Nonlinear
Approach
1982. xii, 237 pages. 209
ilIus. with Models.
Programming. Fixed-Point Theorems.
1981. viii, 355 pages. 259 illus.
Franklin: Methods of 1980.
Mathematical
x, 297 pages. 38 illus. Millman/Parker: Geometry: A Metric
Economics. Linear and Nonlinear
Owen: A First Course in the
Approach with Models.
Programming. Fixed-Point
Theorems.
Mathematical
Foundations of
Halmos:
Finite-Dimensional 1981.
Vectorviii, 355 pages. 259
illus.
Thermodynamics.
1980. x, 297 pages. 38 Spaces.
illus. Second edition.
1974. viii, 200 pages.
in thexiv, 178 pages. 52 illus.
Owen: A First Course1983,
Mathematical Foundations of
Halmos: Finite-Dimensional Vector
continued
Thermodynamics.
Spaces. Second edition.
1974. viii, 200 pages.
1983, xiv. 178 pages. 52 illus.
continued
Halmos: Naive Set Theory.
John G. Kemeny
1974, vii, 104 pages.
J. LaurieIooss/Joseph:
Snell Elementary Stabil
Apostol: Introduction to Analytic
Number Theory.
1976. xii, 338 pages. 24 illus.
Bifurcation Theory.
1980. xv, 286 pages. 47 illus.
Armstrong: Basic Topology.
1983. xii, 260 pages. 132 illus.
Kemeny/Snell: Finite Markov C
1976. ix, 224 pages. I! illus.
Finite IVlarkov Chains
Bak/Newman: Complex Analysis.
1982. x, 224 pages. 69 illus.
Lang: Undergraduate Analysis
1983. xiii, 545 pages. 52 illus.
Ranchoff/Wermer: Linear Algebra
Through Geometry.
1983. x, 257 pages. 81 illus.
Lax/Burstein/Lax: Calculus wit
With a New Appendix
Applications and Computing,
"Generalization
of a Fundamental
Matrix"
Childs: A Concrete Introduction
to
Volume 1.
Higher Algebra.
1979. xiv, 338 pages. 8 illus.
1976. xi, 513 pages. 170 illus.
Chung: Elementary Probability Theory
with Stochastic Processes.
1975. xvi, 325 pages. 36 illus.
EeCuyer: College Mathematics
A Programming Language.
1978. xii, 420 pages. 144 illus.
With 12 Illustrations
Macki/Strauss: Introduction to
Croom: Basic Concepts of Algebraic
Topology.
1978. x, 177 pages. 4,6 illus.
Control Theory.
1981. xiii, 168 pages. 68 illus.
Fischer: Intermediate Real Analysis.
1983. xiv, 770 pages. 100 illus.
Malitz: Introduction to Mathem
Logic: Set Theory - Computabl
Functions - Model Theory.
1979. xii, 198 pages. 2 illus.
Fleming: Functions of Several Variables.
Second edition.
1977. xi, 411 pages. 96 illus.
Martin: The Foundations of G
and the Non-Euclidean Plane.
1975. xvi, 509 pages. 263 illus.
Foulds: Optimization Techniques: An
Introduction.
1981. xii, 502 pages. 72 illus.
Martin: Transformation Geom
Introduction to Symmetry.
1982. xii, 237 pages. 209 illus.
Franklin: Methods of Mathematical
Economics. Linear and Nonlinear
Programming. Fixed-Point Theorems.
1980. x, 297 pages. 38 illus.
Millman/Parker: Geometry: A
Approach with Models.
1981. viii, 355 pages. 259 illus.
Halmos: Finite-Dimensional Vector
Spaces. Second edition.
1974. viii, 200 pages.
Owen: A First Course in the
Mathematical Foundations of
Thermodynamics.
1983, xiv, 178 pages. 52 illus.
Springer -Verlag
New York Berlin Heidelberg rrokyo
continued
3. 4;. Sncll
Department of l\lathcmat
Department of Mathcrnat~cs
Dartmouth College
Dartmouth College
J. G. Kemeny Ilanover, hTH03755 J. L. Snell
Hanover, NH 03755
Department of l\fathcmatics
Departmen t of lITathcmat
U.S.A. ie>
U.S.A.
Dartmouth College
Dartmouth College
Hanover, NH 03755
Hanover, NH 03755
U.S.A.
"c.S.A.
Editorial Board
Editorial Board
F. W. Gehring
P. R. Halmos
Department of Xathematics
Nathematics Department
Unlrersity of Michigan
Indiana University
P.
R.
Halmos
F. W. Gehring Ann Arbor, X I 48104
Bloomington, I X 47405
Mathematics Department
Department of Mathematics
U.S.A.
U.S.A.
Indiana University
University of :Michigan
Bloomington, IN 47405
Ann Arbor, J\U 48104
U.S.A. 60-01, 60510
U.S.A.
AJIS Subject Classifications:
AlliS Subject Classifications:
60-01, 60JI0
Originally published
in 1960 by Tan Sostrand, Princeton, K J .
Kew appendix or~ginallyappeared in Linear A l g c b ~ aand zts Applic
Elsewer
Worth
Holland,
Inc., Princeton,
1981, pp. 193-206.
Reprinted with
Originally published
in 19GO
by Yan
Nostrand,
NJ.
Libraryappeared
of Congress
Cata.loging
inand
Publication
Data vol. 38,
::\ew alJpenclix originally
in Linear
Algebra
its Applications,
Elsevier North Holland,
Inc.,
1981,
Kemcny,
John
G. pp. 193-206. Reprinted with permission.
Finite Xarkov Chains.
Library of Congress Cataloging
in Publication
Data
(Undergraduate
texts in mathematics)
Remeny, John G. Reprint. Originally published : Princeton, N J :
Finite :Markov Chains.
\-an Kostmnd, 1960. With new appendix.
(Undergraduate texts
in mathematics)
1. Marlcov
processes. I. Snell, J. Laurie (James
Reprint. Originally
published:
. 11. Title.NJ:111. Series.
Laurie),
1925- Princeton,
Van Nostrand, 1960.
'With new appendix.
QA274.7K4.5
1983
519.2'33
83-17031
1. Markov processes. 1. Snell, .J. Laurie (James
Laurie), 1925- . II. Title. III. Series.
0 1960, 5J9.2'33
1976 by J. 6. ICemeny,
QA274.7.K45 1983
83-li031 J. L. Snell
All rights reserved. KOpart of this book may be translated or rep
form without written permission from copyright holder or publis
© 1060, 1976 by .J. G. Kemeny, ,T. L. Snell
All rights reserved.Printed
No part
of bound
this book
may
translated6iorSons,
reproduced
in any V
R.beDonnelley
Harrisonburg,
and
by R.
form without written
permission
from copyright
or publisher.
Printed
in the United
States ofholder
America.
Printed and bounel by R. R. Donnelley & Sons, Harrisonburg, VA.
Printed in the United States of America.
ISBX 0-357-90192-2 Spi-inger-VerlagXew Pork Berlin Heidelber
ISBN 3-540-90192-2 Springer-Verlag Berlin Heidelberg New Yor
98765432
ISBN 0-387·90192-2 Springer-Verlag Xew York Berlin Heidelberg Tokyo
ISBN 3-540-90192-2 Springer-Verlag Berlin Heidelberg New York Tokyo
3. 4;. Sncll
Department of Mathcrnat~cs
Dartmouth College
03755
Ilanover, hTHPREFACE
U.S.A.
Department of l\lathcmatics
Dartmouth College
Hanover, NH 03755
U.S.A.
The basic concepts of Markov chains were introduced by A. A. Markov
in 1907. Since that time Markov chain theory has been developed by
Editorial
Board
It is only in very recent times
a number of leading
mathematicians.
that the importance of Markov chain theory to the social and biological
F. W.recognized.
Gehring
R. Halmos
we believe,
sciences has become
This new interestP.has,
Xathematics
Nathematics
produced a real Department
need for a oftreatment,
in English, of
the basicDepartment
ideas
of Michigan
Indiana University
of finite 1\Iarkov Unlrersity
chains.
X I 48104
47405
By restricting Ann
our Arbor,
atteDtion
to finite chains, we Bloomington,
are able toI Xgive
U.S.A.
U.S.A.
quite a complete treatment and in such a way that a minimum amount
of mathematical background is needed. For example, we have written
that Classifications:
it can be used60-01,
in an60510
undergraduate probthe book in suchAJIS
a way
Subject
ability course, as well as a reference book for workers in fields outside of mathematics.
The restrictionOriginally
of this book
to finite
chains
hasSostrand,
made it Princeton,
possible to
published
in 1960
by Tan
KJ.
give simple, closed-form matrix expressions for many quantities usually
c b ~ aand ztsto
Application
Kew
or~ginally
appeared
in Linear
given as series. It
is appendix
shown that
it suffices
for all
types Aofl g problems
Holland,
Inc.,
1981, pp.
193-206.and
Reprinted
types ofWorth
Markov
chains,
namely
absorbing
ergodicwith perm
consider just two Elsewer
chains. A "fundamental matrix" is developed for each type of chain,
Library of Congress Cata.loging in Publication Data
and the other interesting quantities are obtained from the fundamental
Kemcny, John G.
matrix operations.
matrices by elementary
Finite Xarkov Chains.
One of the practical
advantages
of in
this
new treatment of the sub(Undergraduate texts
mathematics)
ject is that these elemenL"ry
matrix operations
can easilyNbe
Reprint. Originally
published : Princeton,
J : programed
for a high-speed \-an
computer.
The1960.
authors
developed a pair of proKostmnd,
With have
new appendix.
grams for the IBM1.704,
one processes.
for each I.type
ofJ.chain,
Marlcov
Snell,
Lauriewhich
(Jameswill find a
11. Title.
111process
. Series. directly from the
Laurie),
1925- . for
number of interesting
quantities
a given
1983 were 519.2'33
transition matrix.QA274.7K4.5
These programs
invaluable in 83-17031
the computation
of examples and in the checking of conjectures for theorems.
A significant feature
the by
new
approach
is J.
that
it makes no use of
0 1960,of1976
J. 6.
ICemeny,
L. Snell
The authors
in each
that theor reprodu
the theory of eigen-values.
All rights reserved.
KOpart found,
of this book
maycase,
be translated
expressions in matrix
fanE are
simpler
than the
corresponding
expresform without
written
permission
from
copyright holder
or publisher.
sions usually given in terms of eigen-values. This is presumably due
Donnelley
Harrisonburg,
Printed
and bound matrices
by R. R. have
to the fact that the
fundamental
direct6i Sons,
probabilistic
in- VA.
Printed
the United States
the in
eigen-values
do not.of America.
terpretations, while
The book falls into three parts. Chapter I is a very brief summary
of prerequisites. Chapters II-VI develop the theory of Markov chains.
of this
theory to problems
a variety
Chapter VII contains
ISBXapplications
0-357-90192-2
Spi-inger-Verlag
Xew PorkinBerlin
Heidelberg To
of fields. A summary
of the symbols
used and ofBerlin
the principal
ISBN 3-540-90192-2
Springer-Verlag
HeidelbergdefiNew York To
nitions and formulas can be found in the appendices together with page
references. Therefore, there is no index, but it is hoped that the dev
tailed table of contents and the appendices will serve a mo
PREFACE
purpose.
I t was not intended that Chapter I be read as a unit. The
tailed table of be
contents
the appendices
willreader
servehas
a the
moreoption
useful
11, and the
of lookin
started and
in Chapter
purpose.
brief summary of any prerequisite topic not familiar to him,
It was not intended that Chapter I be read as a unit. The book can
needs it in a later chapter.'
be started in Chapter II, and the reader has the option of looking up the
The book was designed so that it can be used as a text for a
brief summary of any prerequisite topic not familiar to him, when he
graduate mathematics course. For this reason the proofs wer
needs it in a later chapter.*
out by the most elementary methods possible. The book is
The book was designed so that it can be used as a text for an underfor a one-semester course in Markov chains and their app
graduate mathematics course. For this reason the proofs were carried
Selections from the book (presumably from Chapters 11, 111
bookof isansuitable
out by the most
elementary
possibly
VII) methods
could alsopossible.
be used The
as part
upper-class
for a one-semester course in Markov chains and their applications.
probability theory. For this use, exercises have been given a
book (presumably
from Chapters II, III, IV, and
Selections from oftheChapters
II-VI.
possibly VII) could also be used as part of an upper-class course in
The following system of notation has been used in the boo
probability theory. For this use, exercises have been given at the end
bers are denoted by small italic letters, matrices by capital italic
of Chapters II-VI.
by Greek letters. Functions, sets, and other abstract objects
The following system of notation has been used in the book: Numnoted by boldface letters.
bers are denoted by small italic letters, matrices by capital italics, vectors
The authors gratefully acknowledge support by the Nationa
by Greek letters. Functions, sets, and other abstract obiects are deFoundation to the Dartmouth Mathematics Project. Many of
noted by boldface letters.
nal results in this book were found by the authors while wo
The authors gratefully acknowledge support by the National Science
this project. The authors are also grateful for computing ti
Foundation to the Dartmouth Mathematics Project. Many of the origiavailable by the M.I.T. and Dartmouth Computation Cente
nal results in this book were found by the authors while working on
development of the above-mentioned programs and for the use
this project. The authors are also grateful for computing time made
programs.
available by the M.LT. and Dartmouth Computation Center for the
The authors wish to express their thanks to two research a
development of P.
thePerkins
above-mentioned
programs
and for
the usesuggestions
of these as w
and B. Barnes,
for many
valuable
programs.
their careful reading of the manuscript. Thanks are due to
The authors wish
to express
their
to two
research
Andrews
and Mrs.
H. thanks
Hanchett
for typing
theassistants,
manuscript.
P. Perkins and B. Barnes, for many valuable suggestions as well as for
their careful reading of the manuscript. Thanks are due to Mrs. M.
THEA
Andrews and Mrs. H. Hanchett for typing the manuscript.
Hanover, New Hampshire
vi
THE AUTHORS
* A more detailed treatment of most of these topics may be found in
Hanover, New Hampshire
following books: (If Modern Mathematical Methods and Models, Volum
by the Dartmouth Writing Group, published by the Mathematical Ass
1958.
[Referred
to as
M4.1may
( 2 be
) Introduction
Finile
* A more detailedAmerica,
treatment
of most
of these
topics
found in oneto of
the Mathe
1957. [Referred
Kemeny,
and Thompson,
following books: (1)
ModernSnell,
Mathematical
MethodsPrentice-Hall,
and Models, Volumes
1 and 2, to as
Mathematical
Structu~es,
Mirkil,Association
Snell, and Thompson
by the DartmouthFinite
Writing
Group. published
by by
the Kemeny,
Mathematical
of
Hall, 1959.
to Introduction
as FMS.1 For
prerequisites
in probability
America, 1958. [Referred
to as[Referred
M4.J (2)
to the
Finite
Mathematics,
by
as a treatment
of Markov 195i.
chains [Referred
from a different
view, the
wellThompson.
Kemeny, Snell, and
Prentice-Hall,
to as point
FM.] of (3)
alsoStructures,
wish to consult
Introduction
Probability
Theory and
Its Applicati
Finite Mathematical
by Kemeny,
Mirkil,toSnell,
and Thompson,
PrenticeFeller,
1957.For the prerequisites in probability theory. as
Hall, 1959. [Referred
to Wiley,
as FMS.J
well as a treatment of Markov chains from a different point of view, the reader may
also wish to consult Introduction to Probability Theory and Its Applications, by W.
Feller, Wiley, 195i.
tailed table of contents and the appendices will serve a mo
purpose.
I t was not intended that Chapter I be read as a unit. The
be started in Chapter 11, and the reader has the option of lookin
PREFACE
TO THE
SECOND
PRINTING
brief summary
of any
prerequisite
topic not familiar to him,
needs it in a later chapter.'
bookFinite
was designed
that it can
used as a text for a
When the authorsThe
wrote
.MarkovsoChains,
twobefundamental
graduate mathematics
For this
reason
theused
proofs wer
matrices N for absorbing
chains and Zcourse.
for ergodic
chains
were
out bydescriptive
the most elementary
methods
possible.
to compute the basic
quantities for
Markov
chains. The
The book is
a one-semester
course
Markov
chains and
their app
choice of N was for
natural
but the choice
of Zinwas
less natural.
Z was
Selections
from
the
book
(presumably
from
Chapters
needed to solve equations of the form (I-P)x = f where f is known. 11, 111
VII)
could also
be used
as part ofbyanadding
upper-class
not have
an inverse,
J -P
was modified
Since I-P does possibly
probability theory. For this use, exercises have been given a
the matrix A, all of whose rows are the fixed probability vector a,
of Chapters II-VI.
and the resulting inverse Z = (/ _P+A)-l was used as the fundaThe following system of notation has been used in the boo
mental matrix for ergodic chains. This had the disadvantage bfhaving
bers are denoted by small italic letters, matrices by capital italic
to find a before comput,ing Z.
by Greek letters. Functions, sets, and other abstract objects
vVith the development of pseudo-inverses, it was pointed out by
noted by boldface letters.
C. D. Meyer* that pseudo-im-erses could be used to find basic
The authors gratefully acknowledge support by the Nationa
quantities for ergodic
Markov
chains,
including
the fixedProject.
vector Many
a.
Foundation
to the
Dartmouth
Mathematics
of
Independently, while
teaching
Markov
chains,
Kemeny
noticed
that
nal results in this book were found by the authors while wo
it was not, necessary
to use A The
and authors
that A are
could
replacedforbycomputing
any
this project.
alsobegrateful
ti
matrix B all of whose
rO\VSbyare
same and
vector
(3 whose Computation
components Cente
available
thetheM.I.T.
Dartmouth
sum to 1. The resulting
matrix
(J -P+B)-l serves
in many
ways
development
of Z
the= above-mentioned
programs
and
for the use
the same role as programs.
the fundamental matrix used in the book. (Actually,
it is sufficient to The
assume
thatwish
the to
sum
of the
components
is
authors
express
their
thanks to of
two(3 research
a
non-zero.) The P.
fixed
vector
a is
from
such
a Z bysuggestions
a = (3Z
Perkins
and
B. obtained
Barnes, for
many
valuable
as w
reading
of the manuscript.
are due to
and certain othertheir
basiccareful
quantities
for ergodic
chains, such Thanks
as the mean
Andrews
Hanchett
typing
the manuscript.
first passage times,
have and
the Mrs.
sameH.formula
as for
in the
book.
In any
event, the fundamental matrix used in the book is easily obtained
from any matrix in this new class. Kemeny further showed that thcsc THEA
Hanover,
Hampshire
results are special
cases ofNew
a more
general theorem in linear algebra.
vVe have included Kemeny's paper as an appendix to enable the
readers to benefit* from
simplification
from
themay
usebeoffound in
A morethe
detailed
treatment ofresulting
most of these
topics
this more generalfollowing
class of
fundamenta.l
(If Modernmatrices.
Mathematical Methods and Models, Volum
books:
by the Dartmouth Writing Group, published by the Mathematical Ass
America, 1958. [Referred to as M4.1 ( 2 ) Introduction to Finile Mathe
[Referred
Snell,
andgroup
Thompson,
Prentice-Hall,
* Carl D. :Meyer,Kemeny,
"The role
of the
generalized
in\"orse in1957.
the themy
of to as
FiniteSIAM
Mathematical
Structu~es,
by Kemeny, Mirkil, Snell, and Thompson
finite :Markov chains",
Rev, 17:443-464,
1975.
Hall, 1959. [Referred to as FMS.1 For the prerequisites in probability
well as a treatment of Markov chains from a different point of view, the
also wish to consult Introduction to Probability Theory and Its Applicati
Feller, Wiley, 1957.
vii
TABLE OF CONTENTS
CHAPTER I-PREREQUISITES
SECTION
1.1 Sets
1.2 Statements
1.3 Order Relations
1.4 Communication Relations
1.5 Probability Measures
1.6 Conditional Probability
1.7 Functions on a Possibility Space
1.8 Mean and Variance of a Function
1.9 Stochastic Processes
1.10 Summability of Sequences and Series
1.11 Matrices
1
2
3
5
7
9
10
12
14
18
19
CHAPTER II-BASIC CONCEPTS OF MARKOV CHAINS
2.1
Definition of a Markov Process and a Markov Chain
Examples
2.3 Connection with Matrix Theory
2.4 Classification of States and Chains
2.5 Problems to be Studied
2.2
24
26
32
35
38
CHAPTER III-ABSORBING MARKOV CHAINS
3.1
3.2
3.3
3.4
3.5
Introduction
The Fundamental Matrix
Applications of the Fundamental Matrix
Examples
Extension of Results
jx
43
45
49
55
58
x
TABLE OF CONTENTS
CHAPTER
IV-REGULAR
TABLE
OF CONTENTS
x
MARMOV CHAINS
SECTION
SECTION
4.1
4.2
4.3
4.4
4.5
4.6
4.7
4,8
4.1 Basic Theorems
CHAPTER4.2IV-REGULAR
CHAINS
Law of Large MARKOV
Nulnbers for
Regular Markov Chain
4.3 The Fundamental Matrix for Regular Chains
69
Basic Theorems
4.4 First Passage Times
73
Law of Large Numbers for Regular Markov Chains
4.5 Variance of the First Passage Time
The Fundamental
Matrix
for
Regular
Chains
75
4.6 Limiting Covariance
First Passage
Times
78
4.7 Comparison of Two Examples
Variance of 4.8
the First
Passage Two-State
Time
82
The General
Case
Limiting Covariance
84
Comparison of Two Examples
90
The General Two-State Case
94
CHAPTER V-ERGODIC
MARKOV CHAINS
5.1 Fundamental Matrix
CHAPTER
MARKOV
5.2V-ERGODIC
Examples of Cyclic
ChainsCHAINS
5.3 R e ~ e r s eMarkov Chains
5.1 Fundamental Matrix
5.2 Examples of Cyclic Chains
5.3 Reverse 1Iarkov Chains
105
CHAPTER VI-FURTHER
RESULTS
99
102
6.1 Application of Absorbing Chain Theory to Ergodie C
CHAPTER
VI-FURTHER
RESULTS
6.2 Application
of Ergodic
Cliain Theory to Absorbing
Markov Chains
6.1 Application 6.3
of Absorbing
Chain
Theory to Ergodic Chains
Combining
States
6.2 Application 6.4
of Ergodic
Cnain Theory to Absorbing
Weak Lumpability
11arkov Chains
6.5 Expanding a Markov Chain
6.3 Combining States
6.4 Weak Lumpability
6.5 Expanding a Markov Chain
7.1 Random Walks
CHAPTER VII-APPLICATIONS
OF MARKOV CHAINS
7.2 Applications to Sports
7.3 Ehrenfest Model for Diffusion
7.1 Random Walks
7.4 Applications to Genetics
7.2 Applications7.5
to Sports
Learning Theory
7.3 Ehrenfest .:\lodel
for Diffusionto Mobility Theory
7.6 Applications
7.4 Applications7.7to The
Genetics
Open Leontief Model
7.5 Learning Theory
7.6 Applications to ::VIobility Theory
7.7 The Open Leontief Model
112
117
123
132
140
149
161
167
176
182
191
200
x
TABLE OF CONTENTS
CHAPTER IV-REGULAR
TABLE OF CONTENTS
MARMOV CHAINS
SECTION
xi
4.1 Basic Theorems
I-SUMMARY
OF Nulnbers
BASIC NOTATION
207
4.2 Law of Large
for Regular Markov
Chain
4.3
The
Fundamental
Matrix
for
Regular
Chains
APPENDIX II-BASIC DEFINITIONS
207
4.4 First Passage Times
APPENDIX III-BASIC
QUANTITIES
FOR Time
4.5 Variance
of the First Passage
4.6ABSORBING
Limiting Covariance
CHAINS
208
4.7 Comparison of Two Examples
APPENDIX IV4.8
-:-BASIC
FORMULAS
FORCase
The General
Two-State
ABSORBING CHAINS
209
APPENDIX
APPENDIX
V-BASIC QUANTITIES FOR
209
ERGODIC CHAINS
CHAPTER V-ERGODIC MARKOV CHAINS
APPENDIX VI-BASIC FORMULAS FOR
CHAINS
210
5.1ERGODIC
Fundamental
Matrix
5.2
Examples
of
Cyclic
Chains
APPENDIX VII-SOME BASIC EXAMPLES
210
5.3 R e ~ e r s eMarkov Chains
APPENDIX VIII-GENERALIZATION OF A
21l
FUNDAMENTAL MATRIX
CHAPTER VI-FURTHER
RESULTS
6.1 Application of Absorbing Chain Theory to Ergodie C
6.2 Application of Ergodic Cliain Theory to Absorbing
Markov Chains
6.3 Combining States
6.4 Weak Lumpability
6.5 Expanding a Markov Chain
7.1
7.2
7.3
7.4
7.5
7.6
7.7
Random Walks
Applications to Sports
Ehrenfest Model for Diffusion
Applications to Genetics
Learning Theory
Applications to Mobility Theory
The Open Leontief Model
CHAPTER I
PREREQUISITES
§ 1.1 Sets. By a set a mathematician means an arbitrary but welldefined collection of objects. Sets will be denoted by bold capital
letters. The objects in the 8011ection are called elements.
If A is a set, and B is a set whose elements are some (but not necessarily all) of the elements of A, then we say that B is a .'i'Ubset of
symbolized as B s; A. If the two sets have exactly the same elements,
then \ve say that they are equal, i.e. A = B. Thus A = B if and only
if As; Band B s; A. If B i~ a subset of A and is not equal to A, then we
say that it is a proper subset, ",nd write BcA. If A and B have no
element in common, we say that they are disjoint.
Very frequently we will deal with a given set of objects, and discuss
various subsets of it. The entire set will be called the universe, U.
A particularly interesting subset is the set with no elements, the
empty set E.
Given a set, there are a number of ways of getting new subsets from
old ones. If A and B are both subsets of U, then we define the foHowing operations:
(1) The complement of A, J, has as elements all the elements of U
which are not in A.
(2) The union of A and B, A u B, has as elements all the elements of
A and all the elements of B.
(3) The intersection of A and B, A n R, has as elements all the
elements that A and B have in common.
(4) The difference of A and E, A-B has as elements all the elements
of A that are not in B.
To illustrate these operations, we will list some easily provable
relations between these sets:
~
AnB=EnA
AuB=AnB
-
~
A-B=An'B
-
-
AnB=AuB
AuB=BuA
A uE = A
2
FINITE MARKOV CHAINS
If A,,A2,. . . , A, are subsets of U, and every element of U
Aj, thenCHAINS
we say t h a t A = {AI, CHAP.
Az,. I. . , A,.)
and only
one set
FINITE
MARKOV
tition of U .
If A l , A 2 , • • . , AT
are subsets
U, andaevery
element
U is
in one we
If wc
wish to ofspecify
set by
listingofits
elements,
,
and only one set
A
then
we
say
that
A
=
{Al'
A
...
,
Ar}
is
a par- the s
j
z,
elements inside curly brackets. Thus, for example,
tition of U.
first five positive integers is {I, 2 , 3, 4, 5 ) . The set ( 1 , 3, 5) i
If we wish to
specify
a set
its elements,
we write
thefive-el
subset
of it.
(21, which
is also a subset
of the
Thebysetlisting
elements insideiscurly
brackets.
Thus,
for
example,
the
set
of
the
called a unit set, since i t has only one element.
first five positive integers
is {l, 2,
4, 5}.
3, 5}
a proper
I n the course
of 3,this
bookThe
we set
will{I,
have
t oisdeal
with both
subset of it. The
set {2},
which
is also
a subset
five-element
infinite
sets,
i.e. with
sets
havingofa the
finite
number or set,
a n infini
is called a unit set,
since it hasThe
onlyonly
one infinite
element.sets t h a t are used repeatedl
of elements.
In the courseset
of of
thisintegers
book we
deal
with simple
both finite
and of thi
(1,will
2, 3,have
. . .} toand
certain
subsets
infinite sets, i.e. with
having
a finite
number
infinite
For sets
a more
detailed
account
of tor
h ean
theory
of number
sets see FM C
of elements. The
only Chapter
infinite 1I.t
sets that are used repeatedly are the
or FMS
set of integers {I, 2, 3, ...} and certain simple subsets of this set.
$ 1.2 account
Statements.
We areofconcerned
a process
w
For a more detailed
of the theory
sets see FMwith
Chapter
II
frequently
be a scientific experiment or a game of chance.
or FMS Chapter II. t
a number of different possible outcomes, and we will consid
§ 1.2 Statements.
We about
are concerned
with a process which will
statements
the outcome.
frequently be a scientific
of chance.
are
of aallgame
logically
possibleThere
outcomes.
We formexperiment
the set U or
Th
a number of different
possible
outcomes,
and
we
will
consider
various
be so chosen t h a t we are assured t h a t exactly one of these
statements about
the outcome.
place.
The set U is called the possibility space. If p is any
We form theabout
set Uthe
of all
logically
possible
must
outcome,
then
i t will outcomes.
(in general)These
be true
accordin
be so chosen that
we
are
assured
that
exactly
one
of
these
will
possibilities, and false according to others. The settake
P of all po
place. The setwhich
U is called
possibility
If pthe
is any
p truespace.
is called
truthstatement
set of p. Thu
wouldthemake
about the outcome,
then it
will (in
true
according
to some
statement
about
thegeneral)
outcomebewe
assign
a subset
of U as a
false
according
to
others.
The
set
P
of
all
possibilities
possibilities, andThe
choice of U for a given experiment is not unique. For
p true
is called
truthweset may
of p. analyze
Thus to the
eachpossib
which would make
for two
tosses
of athecoin
statement about
outcome
we TT)
assign
a=
subset
of U2N).
as a truth
U =the
{NM,
WT, TW,
or U
(ON, IH,
I n theset.
first cas
The choice of Uthe
foroutcome
a given of
experiment
notinunique.
Foronly
example,
each toss is
and
the second
the numbe
for two tosseswhich
of aturn
coinup.we(For
maya more
analyze
the possibilities
detailed
discussion of as
this co
U = {HH, HT, TH,
TT}
or
U
=
{OH,
lH,
2H}.
In
the
first
case we give
FRI Chapter II or FMS Chapter 11.)
the outcome of each
tosstwo
ELndstatements
in the second
only
the number
of heads
Given
p and
q having
the same
subject m
which turn up.the(For
a U),
morewedetailed
this of
concept
seenew s
same
have a discussion
number ofofways
forming
FM Chapter II from
or FMS
Chapter
II.)
them. (We will assume that the statements have
Given two statements
truth sets :)p and q having the same subject matter (i.e.
the same U), we have a number of ways of forming new statements
The
statenlent
p (read
"not p")
is true
if and
only if
from them. (We(1)
will
assume
that -the
statements
have
rand
Q as
P
as
truth
set.
Hence
i
t
has
truth sets:)
(2) The statement pV q (read "p or q") is true if either p
(1) The statement q~isp true
(reador"not
p") is
true iift has
and only if p is false.
both.
Hence
Hence it has
P
as
truth
set.
t FM =Kemeny, Snell, and Thompson, Introduction Lo Finite Mathema
(2) The statement
p\f N.J.,
q (read
"p or q")Inc.,
is true
wood Cliffs,
Prentice-Hall,
1957. if either p is true or
Mirkil,
Snell,
Finite .?ltrthematicnl Struc
q is true orFMS=Kemeny,
both. Hence
it has
P UandQThompson,
as truth set.
wood Cliffs, N.J., Prentice-Hall, Inc., 1959.
t FM=Kemeny, Snell, and Thompson. Introduction to Finite Mathematics, Engle·
2
wood Cliffs, N .•J., Prentice-Hall, Inc., 19:)7.
FMS= Kemeny, Mirkil, Snell. and Thompson, Finite 3l1'nthematica! Structures, Englewood Cliffs. N.J .. Prentice-Hall, Inc., 19M),
2
FINITE MARKOV CHAINS
If A,,A2,. . . , A, are subsets of U, and every element of U
PREREQCISITES
and only one
set Aj, then we say t h a t A = {AI, Az,.3. . , A,.)
tition
of
U
.
(3) The statement p Aq (read "p and q") is true if both p and q are
If itwchas
wish
set by listing its elements, we
true. Hence
P (\toQspecify
as trutha set.
elements inside curly brackets. Thus, for example, the s
Two special kinds
of positive
statements
are among
principal
first five
integers
is {I, 2the
, 3, 4,
5 ) . Theconcerns
set ( 1 , 3, 5) i
of logic. A statement
trueset
for(21,
each
logically
outcome,
subset ofthat
it. isThe
which
is alsopossible
a subset
of the five-el
that is, a statement
having
as since
its truth
set,
is said
to be logically
i t has
only
one element.
is called
a unitU set,
true (such a statement
sometimes
called
tautology).
I n theis course
of this
booka we
will have tAo statement
deal with both
that is false forinfinite
each logically
possible
a statement
sets, i.e. with
setsoutcome,
having a that
finiteisnumber
or a n infini
having E as its of
truth
set, is logically
falseinfinite
or self-contradictory.
elements.
The only
sets t h a t are used repeatedl
Two statements
said to be
they
have simple
the same
truth of thi
set are
of integers
(1, equivalent
2, 3, . . .} if
and
certain
subsets
set. That meansFor
thata one
true if and
only of
if the
moreis detailed
account
t h e other
theoryis oftrue.
sets see FM C
The statements
PI, P2,
... , Pk1I.t
are inconsistent if the intersection of
or FMS
Chapter
their truth sets is empty, i.e., PI (\ P2 (\ ... (\ Pk=E. Otherwise
Statements.
We are concerned
with a then
process w
they are said to be$ 1.2
consistent.
If the statements
are inconsi.s~ent,
frequently
bethey
a scientific
experiment
a game
they cannot all be true. If
are consistent,
thenor
they
couldofallchance.
be
a number of different possible outcomes, and we will consid
true.
statements
about
The statements
PI, pz, ...
, Pk the
areoutcome.
said to form a complete set of
of all logically
possible
outcomes.
We
form
the
set
Th
alternati'z:es if for every element of UUexactly
one of them
is true.
This
so chosen tof
h aany
t wetwo
are truth
assured
a t empty,
exactly and
one the
of these
means that thebeintersection
setst his
is called
possibility
Theisset
If pset
is any
union of all theplace.
truth sets
U. U Thus
the the
truth
sets of aspace.
complete
about
the
outcome,
then
i
t
will
(in
be true accordin
of alternatiyes form a partition of U. A completegeneral)
set of alternatives
and false according
to others.
The
set P of all po
provides a newpossibilities,
way (and normally
a less detailed
way) of
analyzing
which would make p true is called the truth set of p. Thu
the possible outcomes.
statement about the outcome we assign a subset of U as a
§ l.3 Order The
relations.
"VeU will
some
simple ideas
the For
for aneed
given
experiment
is notfrom
unique.
choice of
theory of orderfor
relations.
A complete
treatment
of this
theorythe
willpossib
two tosses
of a coin
we may
analyze
be found in lH4,
II,WT,
unitTW,
2.tTT)We
take IH,
only2N).
It few concepts
U =Vol.
{NM,
or will
U = (ON,
I n the first cas
from that treatment.
the outcome of each toss and in the second only the numbe
Let R be a relation
between
a specified
which turn
up. two
(Forobjects
a more(selected
detailed from
discussion
of this co
set U). \Ve denote
by
aRb
the
fact
that
a
holds
the
relation
R to b.
FRI Chapter II or FMS Chapter 11.)
Some special properties
of such
relations
are of
interestthe
to same
us. subject m
Given two
statements
p and
q having
the sameThe
U), relation
we haveR aisnumber
of ifways
of forming
1.3.1 DEFINITION.
reflexive
xRx'holds
for allnew s
from them. (We will assume that the statements have
x in U.
truth setsT:)hr: relation R is symmetric if whenever xRy
1.3.2 DEFI:\ITION.
statenlent
holds, then yRx (1)
alsoThe
holds,
for all x, - yp in(read
U. "not p") is true if and only if
P as truth
1.3.3 DEFINITION.Hence
Thei t has
rdation
R is set.
transitive if whenever
(2)then
The
statement
pVforq all
(read
or U.
q") is true if either p
xRy AyRz holds,
xRz
also holds,
x, y,"pz in
q
is
true
or
both.
Hence
i
t
has
1.3.4 DEFINITION. A relation that is reflexive, symmetric, and
FM =Kemeny, relation.
Snell, and Thompson, Introduction Lo Finite Mathema
transitive is ant equivaJc·,ce
SEC. 3
wood Cliffs, N.J., Prentice-Hall, Inc., 1957.
FMS=Kemeny,
and Thompson,
Finite .?ltrthematicnl
The fundamental
property Mirkil,
of anSnell,
equivalence
relation
is that it Struc
Cliffs,More
N.J., Prentice-Hall,
1959.
partitions the wood
set U.
specifically, Inc.,
let us
suppose that R is an
t M'=Modern Mathematical }..teli,ods and },{odels, by the DArtmouth Writing Group.
!\-IFlthf'mAtiC'fll As:"'o('i8tion of America.! l05S.
FINITE MARKOV CHAINS
4
equivalence relation defined on U. . We put elements of U in
in suchFINITE
a manner
that two
elements a and b are
in the
sam
CHAP.
I
MARKOV
CHAINS
aRb. It can be shown that the resulting classes are well de
equivalence relation
defined
on U. , giving
We putuselements
of Uofinto
mutually
exclusive,
a partition
U. classes
These classe
in such a manner
that twoclasses
elements
equ.ivalence
of R.a and b are in the same class if
aRb. It can be shown
that theletresulting
classes
are"xwell
defined
For example,
xRy express
that
is the
same and
height as
mutually exclusive,
a partition
U. These
classes are
the divi
U is agiving
set ofus
human
beings.of Then
the resulting
partition
equ.ivalence classes
of
R.
people according to their heights. Two men are in the sam
For example, lence
let xRy
express
the are
same
y," where
class
if andthat
only"xif isthey
theheight
same as
height.
4
U is a set of human beings. Then the resulting partition divides these
1.3.5
relation
T the
i s said
be consistent
people according to
theirDEFINITION.
heights. TwoA men
are in
samet oequivaR if, height.
given that xRy, then if xTz hold
equivalence
lence class if and only
if they relation
are the same
yTz, and if zTx holds so does zTy.
1.3.5 DEFINITION. A relation T is said to be consistent with the
1.3.6R DEFINITION.
relation
re$exive
and tra
equivalence relation
if, given that A
xRy,
then ifthat
xTz i sholds
so does
known
a
s
a
weak
ordering
relation.
yTz, and if zTx holds so does zTy.
A weak
can be used
order the
1.3.6 DEFINITION.
A ordering
relation relation
that is reflexive
and to
transitive
is eleme
Given
a
weak
ordering
T,
and
given
any
two
elements
a an
known as a weak ordering relation.
there are four possibilities: (1) aTb AbTa; then the two elem
A weak ordering relation can be used to order the elements of U.
"alike" according to T. (2) aTb A- (bTa) ; then a is "ahea
Given a weak ordering T, and given any two elements a and b of U,
(3) -(aTb) AbTa; then b is "ahead." (4) ~ ( a T bA-(bTa);
)
there are four possibilities:
aTb /\ bTa;
thenobjects.
the two elements are
are unable to(1)
compare
the two
"alike" according For
to T.
(2) aTb
1\ ~ (bTa);
thenthat
a is"I"ahead"
b. as w
example,
if xTy
expresses
like x a tofleast
then
b
is
"ahead."
(4)
-(aTb)I\-(bTa);
then
we
(3) -(aTb)l\bTa;
then the four cases correspond to "I like them equally," "I p
are unable to compare
"I preferthe
y,"two
aildobjects.
"I cannot choose," respectively.
For example, ifThe
xTyrelation
expresses
that "I
likeacts
x atasleast
as well as y,"
of being
alike
an equivalence
relation.
then the four cases
correspond
to
"I
like
them
equally,"
"Ithen
prefer
it can be shown that if T is a weak ordering,
thex,"
relation
"I prefer y," and
"I cannot
choose,"
respectively.
expresses
that
xTy AyTx
is an equivalence relation consisten
The relation of
being
alike acts
equivalence
relation.
Thus
T serves
bothastoanclassify
and to
order. Indeed,
Consistency a
it can be shownthat
that equivalent
if T is a weak
ordering,
then
the
relation
elements of U have the same xRy
placethat
in the or
expresses that xTy
l\yTx
is an equivalence
consistent
T. weak
For
example,
if we chooserelation
"is a t least
as tall"with
as our
Thus T serves both
to
classify
and
to
order.
Consistency
assures
this determines the equivalence relation "is the sameus
height,,
that equivalentconsistent
elements of
U have
the same
place in the ordering.
with
the original
relation.
For example, if we choose "is at least as tall" as our weak ordering,
1.3.7
DEFINITION.
a weak
ordering,
then the
this determines the
equivalence
relationi f"isTthei ssame
height,"
which is
xTy
AyTx irelation.
s the equivalence relation determined by it.
consistent with the
original
1.3.8 If
DEFINITION.
i f Tordering,
i s a weak
and the eq
1.3.7 DEFINITION.
T is a weak
thenordering,
the relation
determined
by i t i s the byidentity
relation ( x = y ) the
xTy /\yTx is the relation
equivalence
rela"tion determined
it.
partial ordering.
1.3.8 DEFINITION. If T is a weak ordering, and the equivalence
The significance
a partial
ordering
by it is the of
identity
relation
(X""isy)that
thennoTtwo
is distinct
a
relation determined
are alike according to it. One simple way of getting a partia
partial ordering.
is as follows: Let T be a weak ordering defined on U.
Defi
The significance
of a partial
no two distinct
relation
T* on ordering
the set isofthat
equivalence
classeselements
by saying th
are alike according
to ifit.every
Oneelenlent
simple way
getting
a partial Tordering
holds
of u of
bears
the relation
to every elem
is as follows: Let T be a weak ordering defined on U. Define a new
relation T* on the set of equivalence classes by saying that uT*v
holds if every element of u bears the relation T to every element of v.
4
FINITE MARKOV CHAINS
equivalence relation defined on U. . We put elements of U in
in such a PREREQUISITES
manner that two elements a and b are in the
5 sam
aRb. It can be shown that the resulting classes are well de
This is a partialmutually
ordering exclusive,
of the equivalence
and we
call These
it the classe
giving usclasses,
a partition
of U.
partial orderingequ.ivalence
induced byclasses
T.
of R.
For example,
let xRy
thata"xminimal
is the same
height as
1.3.9 DEFINITION.
An element
a ofexpress
U is called
element
is afor
setall
of human
Thenelement
the resulting
partition
if aTx impliesUxTa
x E U. beings.
If a minimal
is unique,
we divi
people
according
to
their
heights.
Two
men
are
in
the sam
call it a minimum.
lence class if and only if they are the same height.
We can define "maximal element" and "maximum" similarly. If
1.3.5 DEFINITION. A relation T i s said t o be consistent
U is a finite set, then it is easily shown that for any weak ordering there
equivalence relation R if, given that xRy, then if xTz hold
must be at least one minimal element. However, this minimal element
yTz, and if zTx holds so does zTy.
need not be unique. Similarly, the weak ordering must have a maxi1.3.6
DEFINITION.
A relation that i s re$exive and tra
mal element, but not
necessarily
a maximum.
known a s a weak ordering relation.
§ 1.4 Communication
An important
of order
A weakrelations.
ordering relation
can beapplication
used to order
the eleme
relations is theGiven
studya of
communication
networks.
Lettwo
us. elements
suppose a an
weak
ordering T, and
given any
that r individuals
areare
connected
through a(1)
complex
network.
Each
there
four possibilities:
aTb AbTa;
then the
two elem
individual can pass
a message
on to
the
"alike"
according
to aT.subset
(2) of
aTb
A-individuals.
(bTa) ; then This
a is "ahea
we will call direct
contact.AbTa;
These
messages
may be(4)relayed,
(3) -(aTb)
then
b is "ahead."
~ ( a T band
A-(bTa);
)
relayed again, etc.
This will
indirectthe
contact.
It will not be assumed
are unable
to be
compare
two objects.
that a member can
contact
himself
directly.
Letthat
aTb "I
express
the as w
For
example,
if xTy
expresses
like xthat
a t least
individual a canthen
contact
b (directly
or indirectly)
or like
thatthem
a=b.equally,"
It is "I p
the four
cases correspond
to "I
easy to verify that
T is ay,"
'weak
the set of
individuals. It
"I prefer
aild ordering
"I cannotofchoose,"
respectively.
determines the equivalence
relation
xTyalike
,~yTx,
be read relation.
as
The relation
of being
actswhich
as anmay
equivalence
"x and y can communicate
withthat
each
other,
or x=y."
it can be shown
if T
is a weak
ordering, then the relation
This equivalence
relation
usedis to
the individuals.
expresses
thatmay
xTy be
AyTx
an classify
equivalence
relation consisten
Two men will beThus
in theTsame
equivalence
class if and
they to
canorder.
communicate,
serves
both to classify
Consistency a
that is, if each can
the other
one. ofThe
ordering
thatcontact
equivalent
elements
U induced
have thepartial
same place
in the or
T* has a very intuitive
meaning:ifThe
uT*v
if all
members
For example,
we relation
choose "is
a t holds
least as
tall"
as our weak
of the class u this
can determines
contact all the
members
of therelation
class v, "is
but
con·height,,
equivalence
thenot
same
= v. Thus
the the
partial
ordering
shows us the possible
versely unless uconsistent
with
original
relation.
flow of information.
DEFINITION.
i f ofT the
i s partial
a weakordering
ordering,
then the
In particular, u1.3.7
is a maximal
element
if its
xTycontacted
AyTx i s the
determined
members cannot be
byequivalence
members relation
of any other
class, by
andit.u
1.3.8if its
DEFINITION.
i f T contact
i s a weak
ordering,
and the eq
members cannot
members
of other
is a minimal element
relation
determined
i t i s the initiators,
identity relation
x = y ) the
classes. Thus the
maximal
sets areby message
while (the
partial ordering.
minimal sets are terminals
for messages. (See M4 VoL II, Cnit 2.)
It is interesting
study a given
equivalence
two
Thetosignificance
of a partial
orderingclass.
is thatAny
no two
distinct
members of such
classaccording
CiLl1 communicate with each other.
Hence
areaalike
to it. One simple way of getting
a partia
any member can
contact
anyLet
other
member.
But ho;wdefined
long does
it Defi
is as
follows:
T be
a weak ordering
on U.
take to contactrelation
other members?
As
a
unit
of
time
we
will
take
the
T* on the set of equivalence classes by saying th
time needed to holds
send aif mesEage
from anyone
member
to any Tmember
every elenlent
of u bears
the relation
to every elem
he can contact directly. We call this one step. We will assume that
member i sends out a message, and we will be interested to know where
t.he message could possihly be after n steps.
SEC. 4
F I N I T E MARKOV CHAINS
6
Let Nij be the set of n such that a message starting from
MARKOV
CHAINS
CHAP.We
I will
can beFINITE
in member
j's hands
a t the end of n steps.
sider Nil, the possible times a t which a message can retu
Let N i } be theoriginator.
set of n such
that
a message
from member i
It is
clear
that if astarting
E Nii and b E Nii, then a + b
can be in member
./S hands at the end of n steps.
vVe
will be
first
COIlall the message can return in a steps and can
sent
out aga
sider Nit, the possible
timesafter
at which
message
canset
return
to its unde
received back
b more asteps.
So the
Nif is closed
originator. It The
is clear
that if anumber-theoretic
E Nii and b E Nii , then a+b E Nii after
following
result will be useful. Its
all the message given
can return
in
a
steps
can be sent out again and be
a t the end of theand
section.
6
received back after b more steps. So the set Nii is closed under addition.
1.4.1 THEOREM.
A set
that isi s clo
The following number-theoretic
result
willof bepositive
useful. integers
Its proof
addition
contains all but a finite number of multiples of i
given at the end of
the section.
common divisor.
104.1 THEOREM. A set of positive integers that is closed under
If the
greatest
common
of the elements
of Nii is d
of multiples
of its grwtest
addition contains
all but
a finite
numberdivisor
dl, i t is clear that the elements of Nii are all multiples of
common divisor.
Theorem 1.4.1 tells us in addition that all sufficiently high
If the greatest
divisor
of common
d z are in the
set. of the elements of Nii is designated
d i , it, is clear that
the each
elements
of Nit
all multiples
of d imember
But in i
.
Since
member
canare
contact
every other
Theorem 1.4.1 tells us in addition that all sufficiently
high
multiples
lence class, the Nil are non-empty. We next prove that f
of d; are in the set.
in the same equivalence class, di = d j = d , and that the el
Since each member can contact every other member in its equh-aa given Nil are congruent to each other modulo d (their diff
lence class, the Ni! are non-empty. \Ve next prove th2,t for i and j
multiple of d). Suppose that a E Nil, b E Nij, and c E Nji.
in the same equivalence class, di=el;=d, and that the elements of
First of all, member i can contact himself by sending a m
a given Nij are congruent to each other modulo d (their difference is a
member j and getting a message back. Hence a i c E Nti.
multiple of d). sage
Suppose
Nih b E Nij, and C E X j1:.
couldthat
also agoE to
member j , come back to member j , an
First of all, member i can contact himself by sending a message to
to member i. This could be done in a + kdj + c steps, w
member j and getting
a message
~~ii.
Theofmesd j mustaTe
be aE multiple
di. But
sufficiently
large. back.
Hence Hence
sage could also the
go to
member
j,
come
back
to
member
j,
and
then go of d
same way we can prove that di is a multiple
to member i. d This
be done in a + kd j + C steps, where k is
i = d j =could
d.
sufficiently large. OrHence
But
j must be a multiple of d i .
again,dthe
message could go to member
j ininb exactly
steps, and th
the same way member
we can i.prove
that
eli
is
a
multiple
of
d j . bHence
Hence b c E Nii. Hence a + c and
c are bot
di=d j =d.
bv" d ,, and thus we see that a = b (mod d ) . Thus the elem
Or again, the message could go to member j in b steps, and then back to
given Nij are congruent to each other
d . We can t
member i. Hence b + c E N ii . Hence a + c and b + c are both divisible
duce numbers t i j , with O < t i j < d , so that any element of
by d, and thus we see that a b (mod d). Thus the elements of a
gruent to tij, modulo d . I t is also easy to see that Nii conta
given N ij are congruent to each other modulo d. \Ve can thus introa finite. number of the numbers tij kd.
duce numbers lij, with 0"; til < d, so that any element of N tj is conIn particular we see that tit = 0 in each case, and hence
gruent to tij, modulo d. It is also easy to see that N tj contains all but
(mod d ) . Also tij + ti, =tirn (mod d ) . From this it is easily
a finite numbertij=O
of theisnumbers
lij + led. relation. Let us call such an e
an equivalence
In particularclass
we see
that
lit. = 0 in each case, and hence tij + tji == 0
a cyclic class.
(mod d). Also tl}Since
+ tjm ==ttjtimt j(mod
From
thissee
it that
is easily
m= ttmd).(mod
d ) , we
tij =seen
tim ifthat
and only
til = 0 is an equivalence relation.
Let us call such an equivalence
hence if and only if members j and m are in the same cy
class a cyclic class.
Let n be any integer. If n r ti; (mod d ) , then the message
Since lij + tjm from
== tim member
(mod d), iwe
that
tim if and only if tim = 0,
cansee
only
be lij
in=this
one cyclic class after n ste
hence if and only if members j and rn are in the same cyclic class.
Let n be any integer. If n == lij (mod d), then the message originating
from member i can only be in this Olle cyclic class after n steps. From
+
+
mocha
+
+
F I N I T E MARKOV CHAINS
6
Let Nij be the set of n such that a message starting from
can be in member j's hands a t the end of n steps. We will
PREREQUISITES
SEC. 5
7
retu
sider Nil, the possible times a t which a message can
originator.
It is there
clear that
if a E Nii
and b classes,
E Nii, then
this it immediately
follows that
are exactly
d cyclic
anda + b
all the
message
can return
in a steps
and can
sent of
out aga
that the message
moves
cyclically
from class
to class,
withbecycle
backseen
after
b more
So the
sethas
Nif is
closed unde
length d. It isreceived
also easily
that
aftersteps.
sufficient
time
elapsed,
number-theoretic
result
willappropriate
be useful. Its
it can be in the The
handsfollowing
of any member
of the one cyclic
class
given a t the end of the section.
for n.
While this description
of an equivalence
of the communication
1.4.1 THEOREM.
A set class
of positive
integers that i s clo
network holds inaddition
complete
generality,
cycle number
degenerates
when of i
contains
all butthe
a finite
of multiples
d = 1. In this case
there
is
a
single
"cyclic
class,"
and
after
sufficient
common divisor.
time has elapsed the message can be in the hands of any member at
If the greatest common divisor of the elements of Nii is d
any time.
is clear
thatthat
theifelements
of Nii
areequivalence
all multiples of
In particular,dl,
it isi tworth
noting
any member
of the
Theorem
tells us
in addition
that isallimmediately
sufficiently high
class can contact
himsdf1.4.1
directly,
then
d = 1. This
of d zthat
are d
inisthe
set.
seen from the fact
a divisor
of any time in which a member
Since
contact
can contact himself,
andeach
heremember
d has to can
divide
I every
. other
' member in i
We
next
proveitsthat f
lence
class,
the
Nil
are
non-empty.
The number-theoretic result, § 1.4.1, is of such interest that
in the
proof will be given
here.same equivalence class, di = d j = d , and that the el
giventhat
Nil are
congruent
each other
modulo
d (their
First, of all wea note
if the
greatest to
common
divisor
d of the
set diff
multiple
of d).
Suppose by
thatd, aand
E Nil,
b E Nij,
c E Nji.
is not 1, then we
can divide
all elements
reduce
the and
problem
of all,
member
can contact
to the case d = 1. First
Hence
it suffices
to itreat
this case.himself
Hereby
wesending
have a m
and getting
a message
back.
Hence
i c E Nti.
member
a set of numbers
whosej greatest
common
divisor
is 1, and
wea must
j , comebyback
to member j , an
sage could
to memberHence,
have a finite subset
withalso
thisgoproperty.
a well-known
i. This could
beaznz
done
in a++
kdj +ofc the
steps, w
member
result, there is to
a linear
combination,
alnl +
+ ...
aknk
d j must
be a multiple
Hence
elements (with sufficiently
positive or large.
negative
integers
a;) which
is equaloftodi.l. But
thethesame
way and
we all
canthe
prove
that terms
di is aseparately,
multiple of d
If we collect all
positive
negative
d i = dthe
j = dset
. is closed under addition, we note that there
and remember that
message
could
go to
member
j inbeing
b steps,
must be elements Or
m again,
and n the
in the
set, such
that
m-n=
1 (m
theand th
member
i. Hence
b +the
c E Nii.
Hence
a + c and terms).
b + c are bot
sum of the positive
terms,
and -n
sum of
the negative
(mod d )q. ~Thus
thus number,
we see that
a = bprecisely
bv" d ,, andlarge
Let q be any sufficiently
or more
n(n - the
1). elem
We can t
given Nij are
congruent
to each
mocha d . Then
We can write q=an+b,
where
a~(n-l)
and other
o,,;b";(n-l).
t i j ,hence
with qOmust
< t i j <be
d , in
so the
thatset.
any element of
duce
numbers
we see that q = (a
- b)n
+ bm, and
gruent to tij, modulo d . I t is also easy to see that Nii conta
§ 1.5 Probability
measures.
In the
making
a probability
tij + kd. analysis of an
a finite.
number of
numbers
experiment there In
are particular
two basic we
steps.
First,
a
set
of each
logical
possibilisee that tit = 0 in
case,
and hence
ties is chosen. (mod
This problem
was
discussed
in
§
1.2.
Second
a probad ) . Also tij + ti, =tirn (mod d ) . From this
it is easily
bility measure is
assigned.
The way that
this second
is carried
tij=O
is an equivalence
relation.
Let step
us call
such an e
out will be discussed
this class.
section. We consider first a finite possiclass aincyclic
bility space. (ForSince
a more
Chapter
or only
ttj detailed
t j m= ttm discussion
(mod d ) , wesee
seeFM
that
tij = timIV
if and
FMS Chapter III.)
hence if and only if members j and m are in the same cy
anyU integer.
If n,rar}
ti; be
(mod
, then
thepossimessage
Let n beLet
1.5.1 DEFINITION.
= {aI, az, ...
a setd )of
logical
from member
i canfor
only
in this one
classtoafter
measure
U be
is obtained
by cyclic
assigning
each n ste
bilities. A probability
element aj a positive number w(aj), called a weight, in such a way that
the weights assigned have sum 1. The measure of a subset A of U,
denoted by m(A), is the sum of the weights assigned to elements of A.
+
FINITE MARKOV CHAINS
8
8
1.5.2 THEOREM. A probability measure m assigned to a
set UFINITE
has the following
MARKOVproperties:
CHAINS
CHAP. I
bsetmeasure
B of U,mO<m(
1.5.2 THEOREM. A probability'
assigned to a possibility
set U has the foUowing properties:
re disjoint subsets of U,then
(1) For any subset P of U, 0 ~ m(P) ~ I.
(2) IfPandQare disjointsubsetsofU, thenm(P u Q)=m(P)+m(Q).
(3) For any subsets P and Q of U, m(PuQ)=m(P}+m(Q)m(P n Q). 1.5.3 ~ E F I N I T I O N . Let p be a statement relative to a set U h
h e probability
P. U, Tm(P)=
(4) For any setset
P in
l-m(P). of p relative t o the probability m
i s dejned as m ( P ) .
1.5.3 DEFINITION. Let p be a statement relative to a set U having truth
I n any discussion
wheretothere
is a fixed probability
measu
set P. The probability
of p relative
the probability
measure m
refer simply to the probability of p without mentioning eac
is defined as m(P).
measure. From Theorem 1.5.2 and the relation of the conn
In any discussion
where
there is we
a fixed
measure
we shall
the set
operations,
haveprobability
the following
theorem:
refer simply to the probability of p without mentioning each time the
Let U
be a setof
of the
possibilities
for which
a
1.5.4 THEOREM.
measure. From Theorem
1.5.2 and the
relation
connectives
to
T h e probabilities of statements
hasthe
been
assigned.
the set operations,measure
we haye
following
theorem:
by this measure have the following properties:
1.5.4 THEOREM. Let U be a set of possibilities for which a probability
( 1 ) For a nThe
y statement
p, 0 <of
Pr[p]
< 1. determined
meaSUTe has been assigned.
probabilities
statements
(
2
)
If
p
clnd
q
are
inconsistent
then
Pr[
by this measure have the following properties:
(3) Forp, a0n~y Pr[p]
two ~statements
p and q,
(I) For any statement
1.
Pr[p Ad. then Pr[pV q] = Prep] + Pr(qJ.
(2) If p and q are inconsistent
(
4
)
For any statement
p] = 1- Pr[p].
(3) For any two statements
p and p,
q, Pr[Pr[pVq]=Pr[pl+Pr[q]-
Prep 1\ q]. 1.5.5 EXAMPLE. Given any finite set having s eleme
a Pre
probability
by assigning weight 1
(4) For anydetermine
statement p,
~ p] = 1 - measure
Prep J.
element of U. This measure is called the equiprobable mea
1.5.5 EXAMPLE.
Given
anyr elements,
finite set m(A)
having
8 elements we can
For example, this i
A with
=rjs.
any set
determine a probability
by assigning
weight
l/s outcomes
to each for
sure which measure
would normally
be assigned
to the
element of U. aThis
measure
is
called
the
equiprobable
measure.
For
die. I n this case U = { l , 2 , 3, 4, 5 , 6) and a weight
of ' 1 6
any set A with r elements, m(A) = r/s. For example, this is the meato each.
sure which would normally be assigned to the outcomes for the roll of
1.5.6
EXAMPLE.
a n aexample
of 1!a6 issituation
a die. In this case
U = {I,
2, 3, 4, 5, AS
6} and
weight of
assignedwher
weights would be assigned consider the following: A man
to each.
race between three horses a , b, and c. H e feels t h a t a and
As an of
example
a tsituation
whereas different
1.5.6 EXAMPLE.
same chance
winningofbut
h a t c is twice
likely to win
weights would be assigned consider the following: A man observes a
take the possibility set to be U = {a,b, c) and assign weights
race between three horsesand
a, b,w(c)
and=c.'1% He feels that a and b have the
=
same chance of w(b)
winning
but that c is twice as likely to win as a. We
It set
is occasionally
to extend
the above
take the possibility
to be U = {a,necessary
b, c} and assign
weights
w(a) = concepts
1! 4,
the case of a n experiment with a n infinite sequence of possibl
W(b)=lj4 and W(C)=l/z.
For example, consider the experiment of tossing a coin un
It is occasionally necessary to extend the above concepts to include
the case of an experiment with an infinite sequence of possible outcomes.
For example, consider the experiment of tossing a coin until the first
8
SEC. 6
FINITE MARKOV CHAINS
1.5.2 THEOREM. A probability measure m assigned to a
set U hasPREREQUISITES
the following properties:
9
of U,O<m(
time that a llead turns up. bset
TheB possible
outcomes would be
The above definitions
theorems
apply equally
re disjointand
subsets
of U,then
well to this possibility set. vVe will have an infinite number of weights
assigned but we still must require that they have sum 1. In the
example just mentioned we would assign weights (liz, 1/ 4, lis, . .. ).
These weights form a geometric progression having sum 1.
U={l, 2, 3, ... }.
p be a statement
to a set U h
1.5probability.
.3 ~ E F I N I T I O
. Let happens
§ 1.6 Conditional
ItN often
that a relative
probability
T
h
e
probability
of
p
relative
t
o
the
probability
m
set
P.
me8.sure h8.s been assigned to a set U and then we learn that a certain
i sto
dejned
as m ( PWith
) . this new information we change
statement q relative
U is true.
the possibility set Ito
the discussion
truth set where
Q of q.there
Weiswish
to probability
determine ameasu
n any
a fixed
probability measure
this new
set probability
from our original
measurementioning
m. We eac
referon
simply
to the
of p without
do this by requiring
that elements
of Q should
the same
relative
measure.
From Theorem
1.5.2have
and the
relation
of the conn
weights as they the
hadset
under
the original
assignment
of weights.
operations,
we have
the following
theorem:This
means that our new weights must be the old weights multiplied by a
Let U be awill
set be
of possibilities
for of
which a
1.5.4sum
THEOREM.
constant to give them
1. This constant
the reciprocal
T
h
e
probabilities
of
statements
measure
has
been
assigned.
the sum of the weights of all elements in Q, i.e. I/m(Q). (See FM
this measure
Chapter IV or FMSbyChapter
III.) have the following properties:
( 1 ) For a n y statement p, 0 < Pr[p]< 1.
1.6.1 DEFINITION. Let U ={al' a2, . . . , a r } be a possibility set for
( 2 )been
If passigned,
clnd q are
inconsistent
then Pr[w(aj). Let
which a measure has
determined
by weights
(3) Forto a U
n y (not
two a.statements
p and q, The conq be a statement relative
self-contradiction).
ditional probability measure
given q is a pj'obability measure defined
Pr[p Ad.
on Q the tmth set of (q,4 )dFor
eterrnined
by weights
any statement
p, Pr[- p] = 1- Pr[p].
_
w(aj)
1.5.5 EXAMPLE.
v;(aj) = -Given
- . any finite set having s eleme
m(Q) measure by assigning weight 1
determine a probability
This measure is called the equiprobable mea
elementLet
of U.
1.6.2 DEFINITION.
p and q be two statements relative to a set
example, this i
r elements,
m(A)
=rjs. For
any set A with The
U (q not a self-contradiction).
conditional
probability
of p given q.
sure
which
would
normally
be
assigned
to
the
outcomes for
denoted by Pr[plq] is the probability of p computed from the conditional
a
die.
I
n
this
case
U
=
{
l
,
2
,
3,
4,
5
,
6)
and
a
weight
of ' 1 6
probability measure given q.
to each.
1.6.3 THEOREM. Let P CLnd q be two statements relative to U (q not
1.5.6 Asswne
EXAMPLE.
a n example
of am situation
a self-contradiction).
that aAS
probability
measure
has been wher
A man
weights
would
be
assigned
consider
the
following:
assigned to U. Then
race between three horses a , b, and c. H e feels t h a t a and
Pr[p /\q]
same chance of winning
but t h a t c is twice as likely to win
Pr[plq]
Pr[q]
take the possibility set
to be U = {a,b, c) and assign weights
'1% the measure m.
where Pr[p /\ q]w(b)
and =
Pr[ q]and
are w(c)
found= from
It
is
occasionally
necessary
extend
above
concepts
1.6.4 EXAMPLE. In Example 1.5.6 assumeto that
thethe
man
learns
the
case
of
a
n
experiment
with
a
n
infinite
sequence
of
possibl
that horse b is not going to run. This causes him to consider the new
For
example,
consider
the
experiment
of
tossing
a
coin
un
possibility space Q={a, c}. The new weights which determine the
1/ 4
1/ ')
1/ 4 + 1/2
1/ 4 + 1/2
conditional measure are w(a)=---=I!a and w(c)=-,_'_"_=2/3'
FINITE MARKOV CHAINS
10
CHAP. I
We observe that it is still twice as likely that c will win than it is that
a will win.
1.6.5 DEFINITION. Two statements p and q (neither of which is
a self-contradiction) are independent if PrEp i\ q] = Pr[pJ· Pr[qJ.
It follows from Theorem 1.6.3 that p and q are independent if and
only if Pr[plq]=Pr[p] and Pr[qlp]=Pr[q). Thus to say that p and q
are independent is to say that the knowledge that one is true does not
effect the probability assigned to the other.
1.6.6 EXAMPLE. Consider two tosses of a coin. We describe the
outcomes by U = {HH, HT, TH, TT}. We assign the equiprobable
measure. Let p be the statement "a head turns up on the first toss"
and q the statement "a head turns up on the second toss." Then
Pr[pi\q]=lj4, Pr[p] =Pr[q] = ljz. Thus p and q are independent.
§ 1.7 Functions on a possibility space. Let U = {:1l, a 2, . . • , a r } be a
possibility space. Let f be a function with domain U and range
R = {fl, rz, ... , fs}. That is, f assigns to each element U a unique
element of R. If f assigns fk to aj, we write f(aj)=rk. We write
f = rk for the statement "the value of the function is fk." This is a
statement relative to U, since its truth value is known when the
outcome aj is known. Hence it has a truth set which is a subset of U.
(See FMS Chapters II, III, or M4 Vol. II, Unit 1.)
DEFINITION.
Let f be a function with domain U and range R.
Assume that a measure has been assigned to U. For each rk in R
let w(rk) = Pr[f = rkJ. The weights w(rk) determine a probability
measure on the set R, called the induced measure for f. The weights
are called the induced weights.
1. 7.1
We shall normally indicate the induced measure by giving both the
range values and the weights in the form:
f:
... ,
Thus the induced weight of rk in R is the measure of the truth set
of f=rk in U.
1.7.2 EXAMPLE. In Example lo6.6let f be the function which gives
the number of heads which turn up. The range of f is R = {a, 1, 2}.
The Pr[f=O]=lj4, Pr[f=1]=lj2, and Pr[f=2]=lj4. Hence the range
and induced measure is:
2'
ljJ
SEC. 7
PREREQUISITES
11
1.7.3 DEFINITION. Let U be a possibility space, and f and g be two
functions with domain D, each having as range a set of nllmbers. The
function f + g is the funch:on with domain U which assigns to aj the
number f(aj) + g(aj). The junction f· g is the function with domain U
which assigns to aj the number f(aj)·g(aj). For any number G the
constant function c is the function which assigns the number c to every
element of U.
Let U be a possibility space for which a measure has been assigned.
Then if f and g are two numerical functions with domain U, f + g and
f· g will be functions with domain U, and as such have induced measures.
In general there is no simple connection between the induced measures
of these functions and the induced measure for f and g.
1.7.4 EXAMPLE. In Example 1.6.6 let g be a function having the
value 1 if a head turns up on the first toss and 0 otherwise. Let h be
a function having the vaiue 1 if a head turns up on the second toss
and 0 if a tail turns up. Then the range and induced measures for
g, h, g+h, and g·h are
( 0
1~2}
{1~2 l~J
21
g+h:
{1~4 liz 1/4J
{3~4 I;}
g:
il .
I
<: 2
h:
g·h:
1.7.5 DEFINITIO=". Lei f be a function defined on U. Let p be a
statement relative to U hav'ing truth set P. Assume that a measure m
has been assigned to U. Let f' be the function f considered only on the
set P. Then the induced mwsure for f' calculated from the conditional
measure given p is called the conditional induced measure for f
given p.
1.7.6 DEFI!<ITION. Let f and g be two functions defined on a space 13
for which a probability measure has been assigned. Then f and g are
independent if, for any rIc in the range of f and Sf in the range of g,
the statements f = rIc and g = Sf are independent statements.
An equivalent way to state the condition for independence of two
FINITE MARKOV CHAINS
1%
functions is to say that the induced measure for one func
changed
by the MARKOV
knowledgeCHAINS
of the value of the other.
FINITE
CHAP. I
Throughout
3 1.8
and variance
of a for
function.
functions is to say
that Mean
the induced
measure
one function is not this
shall
assume
that
the
functions
considered
are functions w
changed by the knowledge of the value of the other.
set is a set of numbers. (A detailed discussion of the conc
§ 1.8 Mean and variance of a function. Throughout this section we
duced in this section is given in FMS Chapter 111, or M4 Vol.
shall assume that the functions considered are functions whose range
Let f be a of
function
defined introon a possi
1.8.1 DEFINITION.
set is a set of numbers.
(A detailed discussion
the concepts
U
=
(81,
ag,
.
.
.
,
a,.),
for
which
a
measure
determined
4
duced in this section is given in FMS Chapter III, or M Vol. II, Unit 1.)
w(aj) has been assigned. Then the mean value of f denoted
1.8.1 DEFINITION. Let f be a function defined on a pos8ibility space
U = {aI, a2, ... ,ar }, for which a meas1~re determined by weights
w(aJ) has been assigned. Then the mean value of f denoted by M[f] is
The term expected value i s often used in place of mean v
M[f] = 2:f(aj).w(aJ)'
1.8.2 THEOREM.
Let f be a function defined on U . Assu
J
a probability
measure
definedof on
U , value.
the function f h
The term expected
value is often
used m
in place
mean
measure
1.8.2 THEOREM. Let f be a function defined on U. Assume that for
a probability measure m defined on U, the function f has induced
meaS1(re
12
f:
Then
1.8.3 EXAMPLE.In Example 1.6.6 let f be the numb
2: fj'W(fj).
the definition of mean value we ha
which turn M[f]
up. = From
j
1.8.3 EXAMPLE. In Example 1.6.6 let f be the number of heads
which turn up. From the definition of mean value we have
M[f] = f(HH)· 1/4 + f(HT)· 1/4 +f(TH)· 1/4+ f(TT) .1/ 4
We can also calculate the mean of f by making use of The
= 2.1/4+1.1/4+1.1/4+0.1/4
= 1.The range and induced measure for f is
We can also calculate the mean of f by making use of Theorem 1.8.2.
The range and induced measure for f is
Thus by Theorem 1.8.2,
Thus by Theorem 1.8.4
1.8.2, DEFINITION. Let f be a function dejined on a poss
Let M[f] = m be
UM[f]
for which
a measure has been= assigned.
= 0.1/4+1.1/2+2.1/4
1.
this function. Then the variance of f, denoted by Var[f],
1.8.4 DEFINITION.
Let f be a (f-m)2.
function defined
on a possibility
space
The standard
deviation
denoted
of the function
M[f]
=
m
be
the
mean
of
U for which a measure
has
been
assigned.
Let
the square root of the variance.
this f1mction. Then the variance of f, denoted by Var[f], is the mean
of the f1mclion (f - m)2. The standard deviation denoted by sd[f], is
the sq1~are root of the variance.
FINITE MARKOV CHAINS
1%
functions is to say that the induced measure for one func
changed by the knowledge of the value of the other.
PREREQ,UISITES
13
SEC. 8
3 1.8 Mean and variance of a function. Throughout this
shall Let
assume
the functions
considered
w
1.8.5 THEOREM.
f be that
a function
having mean
value are
m. functions
Then
(A detailed discussion of the conc
set 2is
Va.r[!] =M[f2] -m
• a set of numbers.
duced in this section is given in FMS Chapter 111, or M4 Vol.
1.8.6 EXAMFLE. Let f be the function in Exa.mple 1.8.3. We
found tha.t M[f] = 1.1.8.1
ThusDEFINITION. Let f be a function defined on a possi
U = (81, ag, . . . , a,.), for which a measure determined
Var[f] = (2-1)2.1/
+(1-1)2.1/4+(1-1)2.1/4+(0-1)2.1/
4 of f denoted
w(aj)4has
been assigned. Then the mean value
= 1/2.
An alternative way to compute the variance is to make use of
Theorem 1.8.5. Using
this result
we find
The term
expected
value i s often used in place of mean v
THEOREM. Let f be a function
M[f2]1.8.2
= 4..1/4+1.1/4+1.1/4+0.1/4
= 3/ 2. defined on U . Assu
a probability measure m defined on U , the function f h
Since M[f] = 1, we have
Var(£] = 3/z-1 = 1/ 2.
measure
1.8.7 THEOREM. Let f and g be any two junctions jor which means and
variances have been defined. Then
(1) M[c] = c.
(2) M[f+g] = M[f]+M[gJ.
(3) M[c.f] = e·M[f].
(4) Var[e·f] = c2 • Var[f].
(5) Var[c+f] = Var[f].
(6) Var[e] = O.
Ij f and g are independent
functions In
thenExample 1.6.6 let f be the numb
1.8.3 EXAMPLE.
which turn up. From the definition of mean value we ha
(7) M[f· g] = M[f]- M[g].
(8) Var[f+g] = Var[f]+Var[g].
1.8.8 DEFINITION. Let p be a statement relative to a possibility set U
jor which a measure has been assigned. Let f be a fu.nction with
can also mean
calculate
mean of
making
conditional
andthe
variance
offf by
given
p areuse
theof The
domain U. The We
The of
range
and induced
measure
for f ismeastwe given p.
f comp1ded
from the
conditional
mean and variance
We denote these by M[flp] and Var[flp].
1.8.9 THEOREM. Let PI, pz, ... , pr be a complete set of alternatives
rela1ive to a set U. Let f be a function with domain U. Then
Thus by Theorem 1.8.2,
M[f] =
L M[fJpiJ ·Pr[pj]'
j
Let f be aoffunction
dejined
1.8.4
1.8.10 THEOREM.
If fDEFINITION.
functions
such on
thata poss
1 , f 2 , . . • is a sequence
for some. constant Uc,for which a measure has been assigned. Let M[f] = m be
this function. Then the variance of f, denoted by Var[f],
of the function (f-m)2. The standard deviation denoted
as n -700, then the square root of the variance.
and for any E> 0
Pr[lfn-cl > E] - r 0
asn -+00.
FINITE MARKOV CHAINS
13
14
1 D E F I N I T I O N .Let fi und fz be two functions wit
I
covariance of fl and CHAP.
fz i s defined
T h e n theCHAINS
and FINITE
sd[f] = bi.MARKOV
1.8.11 DEFINITW:'i". Let fl and f2 be two functions with M[f;] = ai
and sd[fiJ = bi . Then the covariance of f1 and f2 is defined by
axd the correlation of 4 and f~ i s
COV[f1, f 2] = M[(fl - a1)(f2 - a2)],
and the correlation of f1 and f2 is
COV[fl' f 2] I n this section we shall brief
5 1.9 CStochastic
[f f] _processes.
orr 1, 2 b b
'
A more complete trea
the concept
of a stochastic
.
l' process.
2
be found in FM Chapter I V or FMS Chapter 111.
§ 1.9 Stochastic processes. In this section we shall briefly describe
We wish t o give a probability measure to describe a n
the concept of a stochastic process. A more complete treatment may
which takes place in stages. The outcome a t the n-th stag
be found in FM Chapter IV or FMS Chapter III.
to depend on the outcomes of the previous stages. It i
INe wish to give a probability measure to describe an experiment
however, t h a t the probability for each possible outconle a t a
which takes place in stages. The outcome at the n-th stage is allowed
stage is known when the outcomes of all previous stages
to depend on the outcomes of the previous stages. It is assumed,
From this knowledge we shall construct a possibility space a
however, that the probability for each possible outcome at a particular
for the over-all experiment.
stage is known when the outcomes of all previous stages are known.
We shall illustrate the construction of the possibility
From this knowiedge we shaH construct a possibility space and measure
measure by a particular example. The general procedure w
for the over-all experiment.
from this.
We shall illustrate the construction of the possibility space and
measure by a particular
The
1.9.1 example.
EXAMPLE.
Wegeneral
choos procedure
a t randomwill
onebeofclear
two co
has heads on both sides.
from this.
Coin A is a fair coin and coin
chosen is tossed. If a tail comes up a die is rolled. If a he
1.9.1 EXAJ\1PLE. We choose at random The
of two
coins
or experim
B.
stage
of Athe
the coin is thrown again. one first
Coin A is a fair coin and coin B has heads on both sides. The coin
choice of a coin. At the second stage, a coin is tossed. A
chosen is tossed. If a tail comes up a die is rolled. If a head turns up
stage a coin is tossed or a die is rolled, depending on the
the coin is thrown again. The first stage of the experiment is the
the first two stages.
choice of a coin. At the second stage, a coin is tossed. At the third
We indicate the possible outcomes of the experiment b
stage a coin is tossed or a die is rolled, depending on the outcome of
shown in Figure 1 - 1 .
the first two stages.
The possibilities for the experiment are t l = (A, H, W ) , t z
We indicate the possible outcomes of the experiment
by a tree as
t3 = ( A ,T,I ) , t4 = ( A ,T, %), etc. Each possibility may b
shown in Figure 1-1.
with a path through the trees. Each path is made up of li
The possibilities for the experiment are h = (A, H, H), tz = (A, H, T),
called branches. I n the tree we have just given, there are
t3 = (A, T, I), t4each
= (A, T, 2), etc.
Each possibility may be identified
having three branches.
with a path through the trees. Each path is made up of line segments
We know the probability for each outcome a t a given sta
called branches. In the tree v,e have just given,
there areifnine
pathsA o
outcome
previous stages are known. For example,
each having three branches.
first stage and T on the second stage, then the probabilit
We know the probability for each outcome at a given stage when the
the third stage is 1 1 6 We assign these known probabil
previous stages arc known. For example, if outcome A occurs 011 the
branches and call them branch probabilities.
first stage and T on
second
thentothe
of ato 1the
for pro
Wethe
next
assignstage,
weights
theprobability
paths equal
the third stage probabilities
is lis. We assigned
assign these
known
probabilities
the F
to the
components
of thetopath.
branches and call them branch probabilities.
We next assign weights to the paths equal to the product of the
probabilities assigned to the components of the path. For example
FINITE MARKOV CHAINS
13
1 D E F I N I T I O N .Let fi und fz be two functions wit
PREREQUISITES
s defined
and sd[f]
= bi. T h e n the covariance of fl and fz i15
SEC. 9
the path t? corresponds to outcome A on the first stage, T on the second,
and 5 on the third. The weight assigned to this path is
axd the correlation of 4 and f~ i s
1/2.1/ 2
.1/ 6 = 1/24.
This procedure assigns a weight to each path of the tree and the sum
of the weights assigned .is 1. The set V of all paths may be considered
a sUitable possibility
space
for the processes.
consideration
of any
statement
I n this
section
we shall brief
5 1.9
Stochastic
whose truth valuethe
depends
outcome process.
of the total
experiment.
A more
complete trea
conceptonofthe
a stochastic
The measure assigned
by in
theFM
path
weights
I V isor the
FMSappropriate
Chapter 11prob1.
be found
Chapter
ability measure.
We wish t o give a probability measure to describe a n
liz
/
which takes place in stages. The outcome a t the n-th stag
It i
to depend on the outcomesw(t)
of the
previous
f l (t)
£2(t) stages.
f3(t)
however, t h a t the probability for each
possible
outconle
at a
II·
II
A
1/8
tt
stage is known when the outcomes of all previous stages
1:
'2 m
From this
wetzshall1/construct
1/2
A a possibility
II
T space a
8
II knowledge
.i
for the over-all experiment.
A
1/z4
We shall illustrate the
ofT the possibility
t3 construction
1
measure by a particular
example. The general procedure w
from this.
A
T
2
1/24
t4
X
7
~a:
1.9.1 EXAMPLE.
We
a t Arandom
T one 3of two co
ts choos
1/24
has heads on both sides.
Coin A is a fair coin and coin
1/6 4If a tail
4
If a he
T rolled.
1/24 up Aa die is
t6 comes
chosenTis tossed.
the coin is thrown again. The first stage of the experim
choice of a coin. At the
is 5tossed. A
1/2
A a coin
T
llz4 stage,
t7 second
:"
stage a coin ~5
is tossed or a die is rolled, depending on the
the first two stages.
A
6
T
6
1/24
is
We indicate the possible outcomes of the experiment b
shownIIin Figure 1II
- 1 . tg
B
1 ;'2
B
H
II
The possibilities for the experiment are t l = (A, H, W ) , t z
FrauHE l~l
t3 = ( A ,T,I ) , t4 = ( A ,T, %), etc. Each possibility may b
with a path through the trees. Each path is made up of li
The above procedure can be carried out for any experiment that
called branches. I n the tree we have just given, there are
takes place in stages. We require only that there be a finite number
each having three branches.
of possible outcomes at each stage and that we know the probabilities
We know the probability for each outcome a t a given sta
for any particularprevious
outcomestages
at theare
j-th
stage, given
the knowledge
of A o
For example,
if outcome
known.
the outcome for the
first
j
1
stages.
For
each
j
we
obtain
a
tree
V
j •
first stage and T on the second stage, then the probabilit
The set of paths ofthe
thisthird
tree stage
serves isas 1a1 6possibility
space
for
any
stateWe assign these known probabil
ment relating to the first j experiments. On this tree we assign a
branches and call them branch probabilities.
measure to the set of
allnext
paths.
"Veweights
first assign
branch
We
assign
to the
pathsprobabilities.
equal to the pro
Then the weight assigned
to a path
is thetoproduct
of all branch
probaprobabilities
assigned
the components
of the
path. F
~
bilities on the path. The tree measures are consistent in the following
sense. A statement whose truth value depends only on the first j
stages may be considered a statement relative to any tree U i for i~j.
FTNJTE MARKOV CHAINS
16
Each of these trees has its own tree measure and the probabi
statement
couldMARKOV
be found from
any one of these CHAP.
measures.
FINI'l'E
CHAINH
I
in every case the same probability would be assigned.
Each of these trees
has itst hown
treehave
measure
of the
Assume
a t we
a treeand
forthe
a nprobability
n stage experiment.
statement coulda be
found from
of these
measures.
and value th
function
with anyone
domain the
set of
paths U, However,
in every case the
same
probability
would
assigned.fl, fz, . . . , f,, are calle
thebefunctions
a t the
j-th
stage. Then
we have a The
tree set
for an
n
stage
experiment.
be
Assume thatfunctions.
a
of functions
fl, fz, . . . , f,Letis fjcalled
a function with domain the set of paths Un and value the outcome
process. (In Markov chain theory i t is convenient to denot
at the j-th stage. Then the functions fl' f2, ... , fn are called outcome
outcome by fo instead of 4.)
of functions
fl' f2,are
...three
,fn isoutcome
called afunctions.
stochastic We
functions. The set
I n our
example there
process. (In Markov chain theory it is convenient to denote the first
cated in Figure 1-1 the value of each function on each path.
outcome by fo instead
Thereofisfl.)
a simple connection between the branch probab
In our example
are functions.
three outcome
We have aindiThe functions.
branch probabilities
t the first
the there
outcome
cated in Figure 1-1 the value of each function on each path.
r[fl
=
rz]
There is a simple connection between the branch probabilities and
a t the second
the outcome functions.
The stage
branch probabilities at the first stage are,
Pr[fz = rijEi = ri]
Pr[fl = Ti]
a t the third stage
at the second stage
Pr[f3 = rr/fi = rj Afl = ri]
etc.
at the third stage
I n our example,
10
etc.
rCf1 = A] = w(tl) + . . . +w(ts) =
In our example,
Pr[fl = A] = W(tl)+ ... +w(ts)= 1/2
P [f
r
2
= T'f = A] = Pr[f2 = T /\fl = A]
I 1
Pr[ f 1 = A]
W(t3)+ ... + wits)
w(td+ ... +w(ts)
1/ 4
1/2 = liz
= -
Pr[fs = 1/\£2 = T /\fl = A]
Pr[f 2 = T /\fl = A]
1.9.2
EXAMPLE.
We
wits)
= 1/24 = shall
1/ 6. often deal with experiments
allow
of stages. For example, in
W(t3)
+ ...a n+arbitrary
wits)
1/number
4
the tosses of a coin, we can envision any number of tosses
1.9.2 EXAMPLE. \Ve shall often deal with experiments where we
for three tosses and the path measure is shown in Fig. 1.2.
allow an arbitrary number of stages. For example, in considering
For any number of tosses we can construct a tree. I t is ev
the tosses of a coin,
we can continuing
envision anythe
number
of tosses. The
tree
to consider
tree indefinitely
to obtain
a
for three tosses infinite
and the paths.
path measure
is
shown
in
Fig.
1.2.
Our procedure for assigning a measure wo
For any number
tosses
can construct
is evenweight
possible
0 to e
thisofcase
bewe
adequate
since ai ttree.
wouldItassign
to consider continuing
indefinitely
obtain
a tree with
We shall the
not, tree
however,
have totoassign
a measure
to the in
infinite paths. This
Our isprocedure
for assigning
a measureabout
would
in
the case because
the statements
thenot
process
t
this case be adequate
since
it would assign weight 0 to every path.
us will depend only on a finite part of the tree, and for any fin
We shall not, however, have to assign a measure to the infinite tree.
This is the case because the statements about the process that interest
us will depend only on a finite part of the tree, and for any finite number
Pr[f3 = I1f2 = T /\fl = A]
FTNJTE MARKOV CHAINS
16
Each of these trees has its own tree measure and the probabi
statement could
be found from any one of these measures.
PREREQUIS1TES
17
in every case the same probability would be assigned.
of stages we haveAssume
a method
a measure.
shall,
howt h a tofweassigning
have a tree
for a n n \Ve
stage
experiment.
ever, consider functions
requires
infinite
U, tree.
and value th
a functionwhose
with definition
domain the
set ofthe
paths
For example,a tin
1.9.2Then
let the
of f fl,
befz,
the
at calle
. . stage
. , f,, are
the value
functions
theExample
j-th stage.
which the firstfunctions.
head occurs.
Then
f
is
defined
for
all
paths
with
at
The set of functions fl, fz, . . . , f, is called a
least one head. process.
This is a(In
SD bset of paths in the infinite tree.
\~T shall
Markov chain theory i t is convenient
to denot
outcome by fo instead of 4.)
w(t) We
I n our example there are three outcome functions.
l/S path.
h each
cated in Figure 1-1 the value of each function on
There is a simple connection between the branch probab
probabilities
a t l/S
the first
112
T
the outcome functions. The branch
tz
1,12
H
r[fl = rz]
a t the second stage
112
H
l/S
Pr[fz = rijEi = ri] ta,.
l/Z
a t the third stage
l/S
Pr[f3 = rr/fi = rj Afl = ri]
t4
etc.
SEC. 9
0
~H
.
1/ 2
1/
'I,
~H
ts
l/S
1/2
T
ts
l/S
~H
t7
lIs
T
ts
l/S
I n our example,
T~H
rCf1 = A] = w(tl) + . . .'I,+w(ts) =
T
1/2
Flm:p.i:.. 1·2
speak of the mean value of such a function when the following conditions are satisfied:
1.9.2 EXAMPLE.We shall often deal with experiments
a n arbitrary
number
of stages.
(a) There is allow
a sequence
of numerical
range
values f1,For
f2, example,
' . . such in
that the truth
valueofofathe
statement
f = rj depends
only onofthe
the tosses
coin,
we can envision
any number
tosses
andofthe
pathand
measure
in Fig. 1.2.
outcomes for
of athree
finitetosses
number
stages
Pr[fis=shown
rjj = 1,
.L
For any number of tosses we canj construct a tree. I t is ev
to
continuing the tree indefinitely to obtain a
(b)
rjPr[f=rj] consider
< co.
j
infinite paths. Our procedure for assigning a measure wo
thishold,
casewebesay
adequate
since
i t would
weight 0 to e
In case (a) and (b)
that f has
a mean
value assign
given by
We shall not, however, have to assign a measure to the in
M[f]
= because
rjPr[fthe
= rj}
This is the
case
statements about the process t
j
us will depend only on a finite part of the tree, and for any fin
.L
2:
\Vhen f has a mean a, we shall say that f has a variance if (f-a)2 has
a mean. If so, Var[f]=M[(f-a)2J.
All properties of means and variances given in § 8 hold for these
FINITE MARKOV CHAINS
18
extended mean values. I n addition we shall need t
FINITE MARKOV CHAINS
CHAP. I
theorem.
18
extended mean values.
addition we
following
Let shall
f l , fi, need
. . . bethe
functions
such that
1.9.3InTHEOREM.
theorem.
each f j is a subset of the same jinite set of numbers. Le
. . Then i f the mean o f s exists,
1.9.3 THEOREM. Let f1' f2' . . . be functions such that the range of
each fj is a subset of the same finite set of numbers. Let s=f1 +f2 +
Then if the mean of sexists,
.
2:
A stochastic
M[s] = process
M[fj].for which the outcome functions al
which are subsets
j of a given finite set is called afcnite stoch
Thus Theorem 1.9.3 states t h a t in a finite stochastic proc
A stochastic process for which the outcome functions all have ranges
of the sum of the functions (if this mean exists) is the sum
which are subsets of a given finite set is called a finite stochastic process.
of the functions.
Thus Theorem 1.9.3 states that in a finite stochastic process the mean
of the sum of the functions
this mean exists)
is the sum
the means
It may oc
of sequences
andofseries.
§ 1.10(if Summability
of the functions.
divergent sequence so, s l , sz, . . . we can form a sequenc
of the terms, and t h a t this new sequence converges.
sequences
andoriginal
serIes. sequence
It may occur
that for a by m
§ 1.10 Summability
we of
say
that the
is summable
divergent sequence averaging
80, 81, 82, . . . we can form a sequence of averages
process. We will be concerned with only two
of the terms, and averaging.
that this new sequence converges. In this case
we say that the original sequence is 8ummable by means of the
n- 1
ayeraging process. We will be concerned
with only two methods of
k)ist for
Let t,= (lln) si and let un =
averaging.
f(:)kn-i(l -
i=O
i=O
of these
is a n average of terms of
< 1.
Let tn = (1 /n) ~~ Slthat
and O<
letk 7),n
== i~Each
(~) kn-i
(1 - k )ISi for some k such
with non-negative coefficients whose sum is 1. If
tl. ta, . . . converges to a limit t , then we say t h a t the origin
that 0< k< 1. EachCesaro-summable
of these is an average
terms
of theu lsequence,
t o t. Ifofthe
sequence
, uz, . . . conver
with non-negative we
coefficients
whose
sum
is
1. If the sequence
say that the original sequence is Euler-summable t o u
tl, t2, ... converges to a limit t, then we say that the original sequence is
For example, consider the sequence 1 , 0 , 1 , 0 , 1 , 0 , .
Cesaro-8ummable to that
t. Ift ,the
Uz, ...
u, odd.
then This
= sequence
1/2 if n is7~1,even,
1/2 converges
+ 1/2n if to
?z is
we say that the original sequence is Euler-summable to u.
verges to 1/2, and hence the original sequence is Ce
:For example, consider the sequence 1,0,1,0,1,0,.... We find
It is easy to verify that lim u n = 112 and henc
to
that tn = 12 if n is even, Yz. + Yzn if n is odd. Thisn+msequence converges to 12, and sequence
hence theis original
sequence is Cesaro-summable
also Euler-summable
to I / * . But the orig
diverges.
to l/ z. It is easy to
verify that lim Un = 1/2 and hence the original
These two summability methods have the following two pr
sequence is also Euler-summable to 1/ 2 . But the original sequence
a sequence converges, then it is summable by each metho
diverges.
(2) If a sequence is summable by both methods, the two sum
These two summability methods have the following two properties: (1) If
same.
a sequence converges, then
it is summable
by each
method to
limit. To
Summability
may also
be applied
to its
a series.
(2) If a sequence is summable
m by both methods, the two sums must be the
same.
series
ak is summable by a given method means that i
Summability may also k=O
be applied to a series. To say that the
'"
series
ak is sum mabIe by a given method means that its sequence of
2
2:
k=O
FINITE MARKOV CHAINS
18
extended mean values. I n addition we shall need t
theorem.
SEC. 11
?REREQUISITES
19
Let
f
l
,
fi,
.
.
.
be
functions
such
that
i 1.9.3 THEOREM.
partial sums 8t = each
ale isf j summable
thatsame
method.
Forof example
is a subset by
of the
jinite set
numbers. Le
. . Then i f the mean o f s exists,
if we apply Cesaro-summability to the partial sums, we obtain
2:
k~O
,,-1 n -
tn=
.
k
'2 --ale·
n
k~O
for which
al
§ 1.11 Matrices. AAstochastic
matrix is aprocess
rectangular
arraythe
ofoutcome
numbers.functions
An
are subsets
of a agiven
is called
r x 8 matrix has r which
rows and
s columns,
totalfinite
of rsset
entries
(or afcnite
com- stoch
Thus Theorem
states
a t especial1y
in a finiteimportant
stochastic proc
ponents). Three special
kinds of 1.9.3
matrices
willt hbe
of
the
sum
of
the
functions
(if
this
mean
exists)
is the sum
in this book. A matrix having the same number of rows as columns
of the functions.
is called a square matrix.
That is, a square matrix is r x r. If r = I,
that is, the matrix consists 0: a single row, then we call it a row vector.
§ 1.10 Summability of sequences and series. It may oc
If s = I, i.e. the matrix has a single column, we call it a column vector.
divergent sequence so, s l , sz, . . . we can form a sequenc
Matrices will be denoted
by capitals
vectors
by small converges.
Greek
of the terms,
and t hand
a t this
new sequence
letters.
we say that the original sequence is summable by m
Let the r x s matrix
A haveprocess.
components
r' x 8' matrix
B two
averaging
We atj,
willand
be the
concerned
with only
have components averaging.
bij • Then we define the following operations and
relations:
n- 1
f(:)kn-i(l -
(1) The matrix kALet
has t,=
components
That
is, a multiplication
and let
un =
k)ist for
(lln) sikatj.
i=O
O
of the matrix by a number means
multiplyingi =each
component
by this number.
Thek <matrix
-A is (-l)A.
that O<
1. Each of these is a n average of terms of
(2) If r = r' and with
8 = s', then the matrix sum A + B has components
non-negative coefficients whose sum is 1. If
alj + bu.
That
addition is carried out componentwise.
tl. is,
ta, . . . converges to a limit t , then we say t h a t the origin
8
Cesaro-summable
o t. toIfhave
the sequence
u l , uz,aikbkj.
. . . conver
(3) If 8=r', we define
the productt AB
components
k~l
we say that the original sequence is Euler-summable
to u
Note that the For
product
of an consider
r x 8 and the
s x t sequence
matrix is 1an
example,
, 0 , r1 x, 0t, 1 , 0 , .
matrix. Thisthat
definition
row This
t , = 1/2also
if napplies
is even,to1/2the+ product
1/2n if ?zofis aodd.
vector and a verges
matrix,too:A,1/2,
or to
a column
vector, is Ce
anda matrix
hence times
the original
sequence
Af3. In the to
former case
of that
a 1 xlim
r and
x 8 henc
It isthe
easyproduct
to verify
u n = an
112 rand
n+m matrix A is
matrix is a I x 8 matrix, or a row vector. If the
But the
orig
alsovector
Euler-summable
to I / *number
.
square, the sequence
resulting isrow
has the same
of
diverges.
components as
0:.
Th;.;s a square matrix may be thought of
These two
summability
havewe
thecan
following
as a transformatio!~l
of row
vectors. methods
Similarly
think two pr
a
sequence
converges,
then
it
is
summable
by
of it as a transformation of column vectors. This will beeach
our metho
If aproduct
sequence
methods, the two sum
principal use (2)
of the
ofisa summable
vector andbya both
matrix.
(4) We say that same.
A:) B (or that A = B) if alj:) bij (or aij = bij ) for all
may also
applied
to a series. To
i and j. That Summability
is, matrix relations
mustbe
hold
componentwise~
m
for all corresponding
series components.
ak is summable by a given method means that i
The r x 8 matrix
(5) Some special matrices.
k=O play an important role.
having all components equal to 0 is denoted by Orxs. The
subscripts are omitted whenever there is no danger of con~
fusion. The r x r matrix having l's as components ali ("on
2:
2
FINITE MARKOV CHAINS
20
20
the main diagonal") and 0's elsewhere is denoted b
CHAP.!
FINITE MARKOV
CHAINS The role that
subscript
is often omitted.
these ma
csn be seen as follows. Let A, I , and 0 be r x r , l
the main diagonal") and O's elsewhere is denoted by IT' The
r-component row vector, and @ an r-component colu
subscript is often omitted. The role that these matrices play
:
Then
can be seen as follows. Let A, I, and 0 be r x r, let a be an
r-component row yector, and f1 an r-component column vector.
Then:
A+O=O+A=A
A+(-A) = (-A)+A = 0
AI = IA = A
al = a
If3
=
f3
AO = OA = 0
0f1 =the
0 matrices 0 and I play somewhat the same
Thus
aO = O. 0 and 1.
numbers
(6) I n analogy to the reciprocal of a number we define
Thus the matrices 0 and I play somewhat the same role as the
of a matrix. The r x r matrix B is said to be the inv
numbers 0 and 1.
r x r matrix A if AB= I . If such an inverse exists, it
(6) In analogy to the reciprocal of a number \ve define the inverse
by A-1. The inverse can be found by solving r2 si
of a matrix. The r x r matrix B is said to be the inverse of the
equations. Of course, these equations may fail
r x r matrix A ifsolution.
AB = I. If such an inverse exists, it is denoted
But when they do have a solution, the
by A-I. The inverse can be found by solving r2 simultaneous
unique, and we can show that AA-I= A-'A =I.
equations. Of course, these equations may fail to have a
they
do haye operations
a solution, on
thematrices,
solutionwhenev
is
solution. But
The when
various
arithmetical
The one majo
defined,
obey
thethat
usual
laws of arithmetic.
we can
show
AA-I=A-IA
=1.
unique, and
to this is that matrix multiplication is not commutativ
The various arithmetical
matrices,
whenever
they are
AB need notoperations
equal BA.on One
important
case where
matrice
defined, obey the
usual
laws
arithmetic.
Thematrix.
one major
is the
case
of of
powers
of a given
Letexception
An be A mu
to this is that itself
matrix
multiplication
is Am
not= Am.
commutative,
i.e.nthat
An for every
and m.
Then An.
n times.
AB need not equal BA. One important case where matrices commute
AO= I .
is the case of powers of a given matrix. Let An be A multiplied by
I t is convenient to introduce row vector 7 , and the colum
itself n times. Then An. Am = Am. An for every nand m. We define
having all components equal to 1. The subscript is aga
AO=I.
when possible. These vectors are convenient for summ
It is convenient to introduce row vector 1)r and theThe
column
vector
product
a ttris a
or rows and columns of matrices.
having all components
equal
to
1. The subscript is again omitted
more precisely a matrix with a single entry) which is the
when possible. components
These vectors
are Similarly
convenient
vectorsA t is
The product
of a.
forfor
$. summing
or rows and columns
of
matrices.
The
product
ag
is
a
number
(or
vector whose i-th component gives the sum of the compon
more precisely ai-th
matrix
with
a
single
entry)
which
is
the
sum
of
the
row of A (or the i-th row sum of A). Similarly ?A
components of column
a. Similarly for 1)(3. The product At is a column
sums of A. We shall denote by E a square mat
vector whose i-th component giyes the sum of the components in the
entries 1. Note that E = 5~
i-th row of A (or the i-th row sum of A). Similarly 1)A gives the
column sums of A. We shall denote by E a square matrix with all
entries 1. Note that E = g1).
FINITE MARKOV CHAINS
20
the main diagonal") and 0's elsewhere is denoted b
21
PREREQUISITES
subscript
is often omitted. The role that these
ma
A,
I
,
and
0
be
r
x
r
, l
csn
be
seen
as
follows.
Let
Let us give some examples of these operations and relations.
r-component row vector, and @ an r-component colu
Then :
(6
SEC. 11
3
\0
-:)
(2o 1) + (-1 0) = (1 1)
-1
0
-2
0
-3
(1,2,3)+ (2,1,0) = (3, 3,3)
(1'2'3)(~
-:)
(5, -1)
Thus the matrices 0 and I play somewhat the same
(~(6) I n analogy to the reciprocal of a number we define
1numbers 0 and 1.
of a matrix. The r x r matrix B is said to be the inv
r x r matrix A if AB= I . If such an inverse exists, it
by
The=inverse
(_3 ) can be found by solving r2 si
( :'0) A-1.l\(~)
1
-1)\1/ Of course, these equations may fail
equations.
solution. But when they do have a solution, the
unique, and we can show that AA-I= A-'A =I.
(J C:)
The various arithmetical operations on matrices, whenev
>
defined, obey the usual laws of arithmetic. The one majo
to this is that matrix multiplication is not commutativ
AB need not equal BA. One important =case
1. where matrice
Let An be A mu
-1is the2 case of
-1powers
2 of1 a 1given matrix.
0 1
itself n times. Then An. Am = Am. An for every n and m.
(2 1)( 1 -1) ( 1 -1)(2 1) (1 0)
1
Therefore,
1
AO= I .
12
I t is convenient
to introduce row vector 7 , and the colum
- 1\ =
having all components equal to 1. The subscript is aga
2} These
\1 vectors are convenient for summ
when possible.
The product
a t is a
or
rows
and
columns
For a square matrix A we introduce of
itsmatrices.
transpose AT.
The ij-th
more
precisely
a
matrix
with
a
single
entry)
which
entry of AT is the ji-th entry of A. We also define the matrix A dg is the
The product
of a.diagonal,
Similarly
which agrees withcomponents
A on the main
butfor
is $.elsewhere.
The A t is
vector
whose
i-th
component
gives
the
sum
of
the
matrix Asq is formed from A. by squaring each entry. This, of course,compon
i-th row
of A).
Similarly ?A
rowsame
of A as(orA the
be the
2.
(But
D2 =sum
Dsq for
a diagonal
will not normally i-th
We shall
denote
square mat
columnwhose
sums only
of A.non-zero
entries
are by
on Ethea main
matrix D, i.e. a matrix
1.
Note
that
E
=
5~
entries
diagonal.) Similarly we define asq for a vector a.
It is often convenient to give a matrix or a vector in terms of its
components. We thus write {a;j} for the matrix whose ij component
°
22
F I N I T E MAEKOV CHAINS
is aij. Similarly we write { a j ) for a row-vector, and {ai) for
vector.FINITE
The following
relations
will illustrate this
notation.
MARKOV
CHAINS
CHAP.!
22
+
+
{all) {bij)and
= {aij
bij),a column
is alj. Similarly we write raj} for' a row-vector,
{aj} for
vector. The following relations will illustrate this
0 =notation.
{O),
&
I= E = {I),
{ad + {bij} = {aii +
bij},
o = {O}, {aij}sq = {a2ti),
) ~ {aji),
t7] = E = {I}, { ~ i j =
2
{aii}sq = {a jj} , 3{m) = ( 3 4 ,
{atilT = {ajt}, {ai){bj} = {atbj).
The last example
that the product of a column vec
3{ai}shows
= {3at},
r
components)
is a matrix (with r x r com
row vector (each
with
{a/}{b j } = {atb j }.
This must be contrasted with t'he product in the reverse ord
The last example
the product
a column
and a
a is a row vecto
is a shows
single that
component.
For ofexample,
if vector
row vector (each
with
r
components)
is
a
matrix
(with
r
x
r
components).
gives the sum of its components. However, Ea gives an r
This must be contrasted
with a for with
each the
row.product in the reverse order, which
is a single component.
a is a row
vector, then
at ent
Ak, with
Suppose For
thatexample,
we have aif sequence
of matrices
gives the sum We
of its
components.
However,
ta
gives
an
r
x
r
matrix
will say that the series Ao+ A1 + A2+ . . . converges if e
with ex: for each row.
of entries converges, i.e, if ~ ( 0+ ) ~ ~+ a(2)fj+ . . . converges
Suppose thati we
a sequence
of matrices
with
a(k)ij.
andhave
j. And
if the sum
of this Ak,
series
of entries
components
is ai
We will say that
the
series
Ao+A1
+A2+
...
converges
if
each
series
i and j , and if A is the matrix with these entries as compon
of entries converges,
if a(O)o
a(1)/j + a(2)/j + ... converges for every
A is+the
sum of the infinite series of matrices.
we sayi.e.that
i and j. And we
if the
sum an
of infinite
this series
is aij,
for each
define
sumofofcomponents
matrices by
forming
the sum
i and j, and if component.
A is the matrix with these entries as components, then
we say that A is the sum of the infinite series of matrices. In brief,
1.11.1sum
THEOREN.
If An
to 0 the
(zerosum
matrix)
n tends
we define an infinite
of matrices
by tends
forming
for as
each
then ( I - A ) has a n inverse, and
component.
m
CA~.
1.11.1 THEOREl\L If An (tends
(zero
I - A )to- 10 =
I + matrix)
A + A Z +asAn3 tends
+ . .to. infinity,
=
then (1 - A) has an inverse, and
k=O
PROOF.
Consider the identity
( I - A ) . ( I + A + A z + . . . +An-l) = I - A n ,
PROOF.
Consider
identity
whichthe
is easily
verified by multiplying out the left side.
By h
we know
that+A2+
the right
tends =to I-An,
I. This matrix has dete
(I-A)·
(1 +A
...side
+An-l)
Hence for sufficiently large n, I - An must have a non-zero det
which is easily But
verified
multiplying
the leftofside.
By hypothesis
the by
determinant
of out
a product
two matrices
is the prod
we know that the
right
side
tends
to
I.
This
matrix
has
determinant
1.
determinants, hence I - A cannot have a zero determin
Hence for sufficiently
large
n,
1
An
must
have
a
non-zero
determinant.
determinant not being equal to zero is a sufficient condi
But the determinant
product
two matrices
the
matrixoftoa have
an of
inverse.
HenceisIthe
- Aproduct
has an of
inverse.
determinants, inverse
hence exists,
I - A cannot
have
a
zero
determinant.
The
we may multiply both sides of the identity by
determinant not bcing equal to zero is a sufficient condition for a
A + A 2I-A
+ . . .has
+An-l
= ( I - A )Since
- l . ( I -this
A*).
matrix to have an inverse.I + Hence
an inverse.
inverse exists, we may mUltiply both sides of the identity by it :
I+A+A2+ ... +An-l = (l-A)-l.(l-An).
F I N I T E MAEKOV CHAINS
22
is aij. Similarly we write { a j ) for a row-vector, and {ai) for
vector. The following relations will illustrate this notation.
SEC. II
PREREQUISITES
23
{all) {bij) = {aij + bij),
But the right side of this new identity clearly tends to (1 _A)-l,
0 = {O),
which completes the -pIOOf.
&I=sequences
E = {I), and series
One can define the summability of matrix
exactly as in § 1.10, applying the averaging method
component
{aij}sqto=each
{a2ti),
of the matrix. Then there is a generalization of{ the
~ i jprevious
=
) ~ {aji), theorem:
If the sequence An is summable to 0 by some averaging method, then
3{m) = ( 3 4 ,
the matrix 1 -A has an inverse, and the series I +A +A2+ ... is
summable by the same method to (I -A)-l. {ai){bj} = {atbj).
The last
showsA that
the product
of a column
vec
1.11.2 DEFINITION.
A example
square matrix
is positive
semi-definite
if
r
components)
is
a
matrix
(with
r
x
r
com
row
vector
(each
with
for any column vector y, yTAy;) 0.
This must be contrasted with t'he product in the reverse ord
l.U.3 THEOREM.
For any
positive semi-definite
matrix
row vecto
is a single
component.
For example,
if a Ais athere
is a matrix B such
thatthe
A =sum
BT B.
gives
of its components. However, Ea gives an r
with a for each row.
Suppose that we have a sequence of matrices Ak, with ent
We will say that the series Ao+ A1 + A2+ . . . converges if e
of entries converges, i.e, if ~ ( 0+ ) ~ ~+ a(2)fj+ . . . converges
i and j. And if the sum of this series of components is ai
i and j , and if A is the matrix with these entries as compon
we say that A is the sum of the infinite series of matrices.
we define an infinite sum of matrices by forming the sum
component.
+
1.11.1 THEOREN. If An tends to 0 (zero matrix) as n tends
then ( I - A ) has a n inverse, and
m
(I-A)-1 = I+A+AZ+A3+
... = C A ~ .
k=O
PROOF.
Consider the identity
( I - A ) . ( I + A + A z + . . . +An-l) = I - A n ,
which is easily verified by multiplying out the left side. By h
we know that the right side tends to I. This matrix has dete
Hence for sufficiently large n, I - An must have a non-zero det
But the determinant of a product of two matrices is the prod
determinants, hence I - A cannot have a zero determin
determinant not being equal to zero is a sufficient condi
matrix to have an inverse. Hence I - A has an inverse.
inverse exists, we may multiply both sides of the identity by
I + A + A 2 + . . . +An-l = ( I - A ) - l . ( I - A * ) .
CHAPTER I
CHAPTER II
ASIC CONGE
BASIC CONCEPTS OF MARKOV CHAINS
§ 2.1 Definition o f a Markov process and a
arkov
recall that for a finite stochastic process we have a tre
measure and a sequence of outcome functions f n , n =
§ 2.1 Definition of a Markov process and a Markov chain. vVe
The domain of fn is the tree T n and the range is the set U
recall that for a finite stochastic process we have a tree and a tree
outcomes for the n-th experiment. The value of fn is sj if
measure and a sequence of outcome functions fn, n = 0, 1,2, ....
of the n-th experiment is sj (see $ 1.9). I n the followin
The domain of fn is the tree Tn and the range is the set Un of possible
whenever a conditional probability P r [ q / p ]occurs, it is a
outcomes for the n-th experiment. The value of fn is Sj jf the outcome
p is not logically false. The reader may find it convenie
of the n-th experiment is 5j (see § 1.9). In the following definitions,
to time to refer to the summary of basic notations and
whenever a conditional probability Pr[qlp] occurs, it is assumed that
the end of the book.
p is not logically false. The reader may find it convenient from time
A finite stochastic process is an independent process if
to time to refer to the summary of basic notations and quantities at
(I) For any statement p whose truth value depends only on
the end of the book.
before the
A finite stochastic process
is n-th,
an independent process if
Br[&depends
= sj lp] =
Pr[fn
(I) For any statement p whose truth value
only
on =
thesjJ.o1ttcomcs
before the n-th,
For such a process the knowledge of the outcome of a
experiment
affect
Pr[fn does
= Sj !p]not= Pr[f
Sj]. predictions for the next
n = our
For a Markov process we weaken this to allow the know
For such a process the knowledge of the outcome of any preceding
immediate past to influence these predictions.
experiment does not affect our predictions for the next experiment.
2.1.1 weDEFINITION.
processofi sthe
a jn
For a Markov process
weaken this A
to finite
allow Markov
the knowledge
suchthese
that predictions.
immediate past toprocess
influence
(11) For
whose truth
depends
only o
2.1.1 DEFINITION.
A any
finitestatement
Markovp process
is avalue
finite
stochastic
before the n - st,
pmcess such that
(II) For any statement p whose truth value depends only on the outcomes
before the n- st,
Pr[fn We
=sj!<f
Pr[fn =s/fn=s£J.
n- 1 =Si)
shall
referI\p]to =condition
II 1as
the Markov prop
Markov process, knowing the outcome of the last experi
neglect any other information we have about the past
We shall refer to condition II as the Markov property. For a
the future. I t is important to realize that this is the ca
Markov process, knowing the outcome of the last24experiment we can
neglect any other information we have about the past in predicting
the future. It is important to realize that this is the case only if we
24
SEC. 1
BASIC CONCEPTS OF MARKOV CHAINS
25
know exactly 'the outcome of the last experiment.
For example,
if
CHAPTER
I
we know only that the outcome of the last experiment· was either 8,
or 8A: then knowledge of the truth value of a statement p relating to
earlier experiments may affect our future predictions.
ASIC CONGE
2.1.2 DEFINITION. The n-th step transition probabilities for a
Markov process, denoted by pij(n) are
pjj(n) = Pr[fn=sjlfn-1=Si].
2.1.3 DEFINITION. §A2.1
finite
Markov chain
is a finiteprocess
Markovand
process
Definition
o f a Markov
a arkov
do
not
depend
on
n. have
In a tre
such that the transition
probabilities
Pij(n)
recall that for a finite stochastic process we
this case they are measure
denoted byand
Pii. The elements of U are called states.
a sequence of outcome functions f n , n =
The The
domain
of fn ismatrix
the tree
and the chain
rangeisisthe
the set U
2.1.4 DEFINITION.
transition
forTan Markov
outcomes
the n-th
experiment.
valueis ofthe
fn is sj if
matrix P with entries
Pii'forThe
initial
probability The
vector
the n-th
vector 1TO = {p/o)} =of{Pr[fo
= sJ]}.experiment is sj (see $ 1.9). I n the followin
whenever a conditional probability P r [ q / p ]occurs, it is a
For a Markov chain we may visualize a process which moves from
p
is not
logically
false. The
readerIfmay
find
it convenie
state to state. It starts
in 8j
with probability
p(O)j.
at any
time
it
to
time
to
refer
to
the
summary
of
basic
notations
and
is in state St, then it moves on the next "step" to Sj with probability Pli'
the
end
of
the
book.
The initial probabilities are thought of as giving the probabilities for
A finite stochastic process is an independent process if
the various possible starting states. The initial probability vector and
For anydetermine
statement pthe
whose
truth chain
value depends
the transition matrix (I)
completely
Markov
process, only on
buildthe
then-th,
entire tree measure. Thus, given
since they are sufficient tobefore
any probability vector 1T0 and any probability
there
Br[& =matrix
sj lp] = P,
Pr[fn
= sjJ.is a
unique Markov chain (except possibly for renaming the states) which
For such a process the knowledge of the outcome of a
will have the 1T0 as initial probability vector and P as transition matrix.
experiment does not affect our predictions for the next
In most of our discussions we will consider a fixed transition matrix
For a Markov process we weaken this to allow the know
P, but we will wish to vary the initial vector 1T. The tree measure
immediate past to influence these predictions.
assigned will depend on the initial vector 1T that is chosen. Hence if
2.1.1 toDEFINITION.
Markov
p is any statement relative
the tree, or fAis finite
a function
withprocess
domaini s a jn
process
such
the tree, Prep], M[f] , and
Var[f]
allthat
depend on 1T. We indicate this by
writing Prn[p], Mn[f] and
The special
casetruth
where
1T has a 1
(11)Var~[fJ.
For any statement
p whose
value
depends only o
in the i-th component (process
is
started
in
state
5i) is denoted Prt[p],
before the n - st,
Mi[f], Vari[f].
We shall give several examples of Markov chains in the next section.
We conclude this section with a few brief remarks about the Markov
property.
We shall refer to condition II as the Markov prop
It can be easily proved
the Markov
is equivalent
to experi
Markovthat
process,
knowingproperty
the outcome
of the last
the following property
more
symmetric
with
respect
to
time.
neglect any other information we have about the past
future.whose
I t is
important
to realize
this is the ca
(II') Let p be any the
statement
truth
value depends
only that
on outcomes
whose truth value
after the n-th experiment and q be any statement 24
depends only on outcomes before the n-th experiment. Then
Prep j\qlfn=siJ = Pr[plf,,=sj] .Pr[qlf,,=sj].
26
FINITE MARKOV CHAINS
This condition says essentially that, given the presen
This II
more
and future
areMARKOV
independent
of each other. CHAP.
FINITE
CHAINS
definition suggests in turn that a Markov process should
This condition
says process
essentially
given the
present,order.
the past
Markov
if i tthat,
is observed
in reverse
That t
and future aretrue
independent
of
each
other.
This
more
symmetric
is seen from the following theorem. (We shall not
definition suggests in turn that a Markov process should remain a
theorem.)
Markov process if it is observed in reverse order. That the latter is
Given a(We
Markov
let pthis
be an
THEOREM.
true is seen from2.1.5
the following
theorem.
shall process
not prove
whose truth value depends only on experiments after the n-th
theorem.)
Then
2.1.5 THEOREM. Given a Markov proce88 let p be any 8tatement
whose truth value depends only on experiment8 after the n-th experiment.
Then
Since a Markov process observed in reverse order remain
process, it mighti\p]
be =suspected
the same is true for
f ,,+l=Sj].
Pr[f,,=sJi(f"+l=SI)
Prlfll =sji that
chain. This would be the case if the "backward transit
Since a Markov
processp*$j(n)
observed
in reverse
order
a Markov of
bilities,"
= Pr[fn
= sjlfn+l
=st], remains
were independent
process, it might
be suspected
the same
is true
: for a Markov
probabilities
maythat
be found
as follows
chain. This would be the case if the "backward transition probabilities," p*lj(n)=Pr[f,,=sjif"+l=s,], were independent of n. These
probabilities may be found as follows:
26
* ()
p tj n =
Pr[fn=sj i\fn+ 1 =51]
Pr[fn + 1 =sj-]Pr[fn+l = s¥n = Sj]' Pr[f" = 5,]
Pr[fn+l = St]
pwPr[f,,=sJJ
These
transition probabilities would be independent of n
probability
of being. in a particular state a t time n was indep
Pr[f,,+l=St]
This is certainly not the case in general. For example, if
probabilities
be independent of n only if the
These transition
81 with probability 1, then the probabili
is started
in statewould
probability of being
a particular
stateisat
timeThus,
n was in
independent
of n. sl]
thereinon
the next step
pll.
general, Pr[fo=
This is certainlyThus
not the
case
in
general.
For
example,
if
the
system
a Markov chain looked a t in reverse order will be
is started in state
81 with probability 1, then the probability that it is
process,
but in general its transition probabilities will depe
there on the next
step
is Pl!.
Thus,
general,
Pr[fO=sl]#Pr[f1=srJ.
and
hence
i t will
notin be
a Markov
chain. We will ret
Thus a Markov chain looked at in reverse order will be a Markov
problem in 3 5 . 3 .
process, but in general its transition probabilities will depend on time
this section
shalltogive
and hence it will92.2
not Examples.
be a MarkovIn chain.
We willwereturn
this sev
examples of Markov chains which will be used in futur
problem in § 5.3.
illustrative purposes. The first five examples relate to what
§2.2 Examples.
this section
we shall
give several
simple
calledIna "random
walk."
We imagine
a particle
which
examples of Markov
chains
which
will
be
used
in
future
work
straight line in unit steps. Each step is one unit for
to the
illustrative purposes.
The first
to what
normally q
probability
p orfive
oneexamples
unit to relate
the left
with isprobability
called a "random
walk."
We one
imagine
particle points
which which
moves are
in called
a
until
i t reaches
of twoa extreme
straight line in points."
unit steps.
Each
step
is
one
unit
to
the
right
with
The possibilities for its behavior a t these point
probability p or one unit to the left with probability q. It moves
until it reaches one of two extreme poin"ts which are called "boundary
points." The possibilities for its behavior at these points determine
FINITE MARKOV CHAINS
26
This condition says essentially that, given the
presen
27
and future are independent of each other. This more
definition
suggests
in turn
a Markov
process should
several different kinds
of Markoy
chains.
Thethat
states
are the possible
Markov
process
i t is observed
reverse
order.
positions. We take
the ca.se
of 5ifstates,
states 51inand
55 being
the That t
the following
theorem. (We shall not
"boundary" states,true
and is52,seen
83, S4from
the "interior
states."
theorem.)
51
52
S3
54
55
SEC. 2
BASIC CONCEPTS OF :>lARKOV CHAINS
THEOREM.Given a Markov process let p be an
whose truth value depends only on experiments after the n-th
Then EXAMPLE 1
2.1.5
Assume that if the process reaches state 81 or 85 it remains there
from that time on. In this case the transition matrix is given by
Since a Markov process observed in reverse order remain
process, it might
be suspected that the same is true for
81 82 83 84 S5
chain. This would be the case if the "backward transit
0 0 0 0
bilities,"S1 p*$j(n)=
Pr[fn= sjlfn+l=st], were independent of
probabilities
may0 bePfound
as follows :
0
S2
P
S2,
q
0
p
S,1
0
q
0
Ss
0
0
0
0
(1)
EXAMPLE 2
Assume now that the particle is "reflected" when it reaches a
probabilities
would
be independent
of n
boundary point and These
returnstransition
to the point
from which
it came.
Thus if
a particular
a t time
was indep
probability
of being
it ever hits 51 it 'goes
on the next
step in
back
to 82. Ifstate
it hits
85 it n
goes
This is
the case
in general.
For example, if
on the next step back
to certainly
84.
The not
matrix
of transition
probabilities
is started in state 81 with probability 1, then the probabili
becomes in this case
there on the next step is pll. Thus, in general, Pr[fo= sl]
Thus a Markov
looked
a t in reverse order will be
81 82chain
S3 84
85
process, but in general
its
transition
probabilities will depe
0 0
51
and hence i t will not be a Markov chain. We will ret
problemS:2in 3 5 . 3 0. P 0
j)
P
q 0In P this section we shall (2)
S3
92.2
Examples.
give sev
examples
of
Markov
chains
which
will
be
used
in futur
0 q 0
Sel
illustrative purposes. The first five examples relate to what
0 walk."
0
called as:,"random
We imagine a particle which
straight line in unit steps. Each step is one unit to the
probability p or one unit to the left with probability q
EXAMPLE 3
until i t reaches
one of two extreme points which are called
points."we The
possibilities
for its behavior
a t these
assume
that whenever
the particle
hits point
As a third possibility
one of the boundary states, it goes directly to the center state S3. We
may think of this as the process of Example 1 started at state S3 and
28
FINITE MARKOV CHAINS
repeated each time the boundary is reached.
28
repeated each time the boundary IS reached.
The transitio
CHAP. II
FINITE MARKOV CHAINS
The transition matrix is
S1 S2 Sa S4 S5
S1
0
1
0
0
S2
0
P
0
0
0
P = S3
0
q
0
p
S4
0
0
q
0 P
S5
0
0
1
0
(3)
0EXAMPLE 4
Assume now t h a t once a boundary state is reached
stays a t thisEXAMPLE
state with 4probability
and moves to the othe
state with probability 11% I n this case the transition matr
Assume now that once a boundary state is reached the particle
sz other
s3 s4boundary
s5
stays at this state with probability 1/ 2and movess1to the
state with probability 1/ 2. In this case the transition matrix is
51
52 S3 S4
S5
S1
1/ 2 0
0
0
1/ 2
S2
0
P
0
0
P = Sa
0
q
0
p
0
S4
0
0
q
0
p
Ss
1/2 0
0
0
1/ 2
(4)
EXAMPLE
5
As the final choice for the behavior a t the boundary, le
that when the
particle reaches one boundary i t moves dir
EXAMPLE 5
other. The transition matrix is
As the final choice for the behavior at the boundary, let us assume
that when the particle reaches one boundary it moves directly to the
other. The transition matrix is
S2 S3 S4
Sl
0
0
0
S2
0
P
0
0
0
P = S3
0
q
0
p
S4
0
0
q
0 P
S5
1
0
0
0
(5)
0
We next consider a modified version of the random w
process is inEXAMPLE
one of the6 three interior states, i t has equal
of moving right, moving left, or staying in its present sta
We next consider a modified version of the random walk. If the
process is in one of the three interior states, it has equal probability
of moving right, moving left, or staying in its present state. If it is
28
FINITE MARKOV CHAINS
repeated each time the boundary is reached.
SEC. 2
BASIC CONCEPTS OF MARKOV CHAINS
The transitio
29
on the boundary, it cannot stay, but has equal probability of moving
to any of the four other states. The transition matrix is:
Sl
52
Sa
S4
85
Sl
0
1/4
1/4
1/4
1/4
52
1/3
1/3
1/ 3
0
0
P = Sa
0
lla
1/3 1/3
0
S4
0
0
1/ 3
1/
E3
XAMPLE 4
1/ 3
(6)
Assume
t once
1/now
1/4 a 0boundary state is reached
55
4 1/4t h a1/4
stays a t this state with probability
and moves to the othe
state with probability
EXAMPI.E 711% I n this case the transition matr
sz take~as
s3 s4 s5
A sequence of digits is generated at random.s1 We
states
the following: 81 if a 0 occurs, S2 if a 1 or 2 occurs, S3 if a 3, 4, 5, or 6
occurs, S4 if a 7 or 8 occurs, S5 if a {) occurs. This process is an independent trials process, but we shall see that Markov chain theory gives
us information even about this special case. The transition matrix is
51
82
83
54
S5
51
.1
.2
.4
.2
.1
S2
.1
.2
.4
.2
.1
EXAMPLE 5
.1 .2 .4 .2 .1
(7)
As the final choice for the behavior a t the boundary, le
.2 .4 reaches
.2 .1 one boundary i t moves dir
.1 particle
S4
that when
the
other. S5 The.1transition
.2 .4 matrix
.2 .1 is
P = S3
EXAMPLE
8
According to Finite Mathematics (Chapter V, Section 8), in the Land
of Oz they never have two nice days in a row. If they have a nice day
they are just as likely to have snow as rain the next day. If they
have snow (or rain) they have an even chance of having the same the
next day. If there is a change from snow or rain, only half of the
time is this a change to a nice day. We form a three-state Markov
cha·in with states R, N, and S for rain, nice, and snow, respectively.
The transition matrix is then
We next consider a modified version of the random w
process is in one of the three interior states, i t has equal
of moving right, moving left, or staying in its present sta
(8)
30
FINITE MARKOV CHAINS
I n our previous examples t h e Markov property clearly
this case
i t could
only be
regarded as a n approximation
FINITE
MARKOV
CHAINS
CHAP. II
knowledge of the weather the last two days, for example,
In our previous
Markov property
clearlythe
held.
In o
us toexamples
differentthe
predictions
than knowing
weather
this case it could
only day.
be regarded
as to
animprove
approximation
since the is
previous
One way
this approximation
knowledge of the
weather
the lastfor
twotwo
days,
for example,
lead wo
states
the weather
successive
days. might
The states
us to different NN,
predictions
knowing
the SR,
weather
only transition
on the p
NR, NS, than
RN, RR,
RS, SN,
SS. New
previous day. would
One way
this approximation
is would
to takestill
as be
haveto toimprove
be estimated.
A single step
Rtates the weather
twoNR,
successive
days. weThe
states
would
be R
for example,
could
move
onlythen
to states
thatfor
from
NN, NR, NS, RN,
RR, RS, SN,
SR, kind
SS. i New
transition
probabilities
I n examples
of this
t is possible
to improve
the app
would have to be
estimated.
A
single
step
would
still
be
so o
still using the Markov chain theory, but a tone
theday,
expense
that from NR, for
westates.
could move only to states RN, RR, RS.
theexample,
number of
In examples of this kind it is possible to improve the approximation,
still using the Markov chain theory,. but at the expense of increasing
the number of states.
EXAMPLE 9
30
An urn contains two unpainted balls. At a sequence
ball is chosen EXAMPLE
a t random,9 painted either red or black, and
If the ball was unpainted, the choice of color is made
An urn contains
unpainted
balls. is At
a sequence
times
a
If i t two
is painted,
its color
changed.
We of
form
a Marko
ball is chosen at
random,
painted
either
red
or
black,
and
put
back.
taking as a state three numbers (x, y, z) where z is the
If the ball wasunpainted
unpainted,
they choice
of color
is made
balls,
the number
of red
balls, at
andrandom.
z the numb
If it is painted,balls.
its color
is
changed.
We
form
a
Markov
chain by
The transition matrix is then
taking as a state three numbers (x, y, z) where x is the number of
unpainted balls, y the number of red balls, and z the number of black
balls. The transition matrix is then
(0,1,1)
(0,2,0)
(0,0,2)
(2,0,0)
(1,1,0)
(1,0,1)
o
liz
1! 2
0
0
0
(0,2,0)
0
0
0
0
(0,0,2)
0
0
0
0
°
(2,0,0)
0
1/2
1/2
(1,1,0)
1/4
0
0
liz
(1,0,1 )
0
°
°
1/4
0
1/ 2
0
10
(0,1,1)
0
EXAMPLE
0
(9)
Assume that a student going to a certain college has e
probability p EXAMPLE
of flunking10out, a probability q of having t o
year, and a probability r of passing on to the next year.
Assume thatMarkov
a student
going
to aascertain
college flunked
has eachout,
year
a
chain,
taking
states sl-has
sz-has
probability p of
flunking
out, se-is
a probability
of having
to repeat ss-is
the a
SO--is
a senior,
a junior,q s5-is
a sophomore,
year, and a probability r of passing on to the next year. \Ve form a
Markov chain, taking as states s1-has flunked out, s2-has graduated,
s3-is a senior, 84-is a junior, s5-is a sophomore, 56-is a, freshman.
30
SEC. 2
FINITE MARKOV CHAINS
I n our previous examples t h e Markov property clearly
_B_A~SI_C~C~O~N~C~E~P_T~S_O~F~1=fA=R=K~O_V~CH~A~IN=T~S
~31
this case i t could only be regarded ______
as a n approximation
knowledge
The transition matrix
is then of the weather the last two days, for example,
us to different predictions than knowing the weather o
previous day. One way to improve this approximation is
states the weather for two successive days. The states wo
NN, NR, NS, RN, RR, RS, SN, SR, SS. New transition p
would have to be estimated. A single step would still be
to states R
that from NR, for example, we could move only(10)
I n examples of this kind i t is possible to improve the app
still using the Markov chain theory, but a t the expense o
the number of states.
EXAMPLE
EXA}lPLE
9
11
An urn contains two unpainted balls.
At a sequence
A man is playing
slot-machines.
Thepainted
first machine
pays
balltwo
is chosen
a t random,
either red
or off
black, and
with probability c, If
thethe
second
he loses,
plays
ball with
was probability
unpainted, d.the Ifchoice
of he
color
is made
the same machine again;
he wins, its
he color
switches
to the other
If i t isifpainted,
is changed.
Wemachine.
form a Marko
Let s, be the state oftaking
playing
i-th machine.
The transition
asthe
a state
three numbers
(x, y, z) matrix
where isz is the
unpainted balls, y the number of red balls, and z the numb
balls. The transition matrix is then
s~
( 1- c
d
C .).
(11)
I-d
As c and d take on all permissible values (0 ~ c ~ 1, 0,;; d ,;; I) we get all
2 x 2 Markov chains.
EXAMPLE
12
Consider the special two-state Markov chain (Example 11) with
transition matrix
p
EXAMPLE
10
(l1a)
Assume that a student going to a certain college has e
of flunking out, a probability q of having t o
(This can be calledprobability
Example 1]p a.)
passingchain
on toasthe
next year.
year,
and we
a probability
From this Markov
chain
form a newr of
Markov
follows.
chain,
states
flunked out,
A state in the new Markov
chain will
be a taking
pair of as
states
in sl-has
the old chain.
Thatsz-has
SO--isS)82,
a senior,
se-is The
a junior,
s5-is isa sophomore,
is, the states are SIS),
S2S), 52S2.
new chain
in state SiSj ss-is a
on the n-th step if the old chain was in state Si on the n-th step and sf
on the (n+ l)-th step.
F I N I T E MARKOV CHAINS
32
The transition matrix for the new chain (Example 12) i
FINITE MARKOV CHAINS
CHAP. II
32
The transition matrix for the new chain (Example 12) is
S151 S152
SIS 1
p=
51S2
c~'
5251
1/2
0
0
1/ 4
82 52
'~)
(12)
1/2
0
1/2
5251
We shall see in 3 6.5 t h a t the study of this new chain gi
1/ 4the 3/original
0
0about
5252 information
detailed
process than could
4
directly from the two-state chain.
We shall see in § 6.5 that the study of this new chain gives us more
Connection
withprocess
matrix than
theory.
5 2.3about
section w
detailed information
the original
couldI nbethis
obtained
the
connection
between
Markov
chain
theory and matrix t
directly from the two-state chain.
shall start with t h e general finite Markov process and the
§ 2.3 Connection
with matrix
In this
section we shall show
our results
to the theory.
finite Markov
chain.
the connection between Markov chain theory and matrix theory. \Ve
2.3.1
THEOREM.
Let fn process
be the outcome
function
clt time
shall start with the
general
finite Markov
and then
specialize
processchain.
lcith trnnsition probabilities pt,(n), t h ~ n
our results to the Mnrkov
finite :'IIarkov
2.3.1 THEOREM. Let fn be the outcome fu.nction at time n for (£ ,finite
jl,farkov processwilh transition probabilities ]Jij(n), then
PROOF. The statement f n = s, is a statement relative to
=5v]
=
Pr[fn-we
1 =5,,]pur(n).
To Pr[fn
find its
probability,
add the weights of all paths in i
" paths which end in outcome s,. Thus i
T h a t is, all possible
PROOF.
The isstatement
= Sv is a statement
a possiblefnsequence
of states relative to the tree Tn.
To find its probability, we add the weights of all paths in its truth set.
That is, all possible paths which end in outcome sv. Thus if.}, k, ... , U
is a possible sequence of states
L:
Pr[fn = sv]
L: Pr[fo =
t.Jn-l =
('.fn = sv].
By the Markov property this is
= L: Pr[fo=sj/\'" /\fn-l=5u]·Pr[fn=5vlfo=sj/\'·· /\fn--l=Su].
u.
=
Sj /\ ' ••
Su
j,k, ... , U
j, k •... ,
By the IVlarkov property this is
If in this last sum we keep u fixed and sum over the remai
P(fo = 5j t, ... /\fn - 1 = su]puv(n).
we obtain
j, k, ... , U
L:
2
If in this last sum we keep u fixedW and
over Pr[f,-1
the remaining
indices
f n =sum
sv] =
= s,]p,,(n).
I1
we obtain
This completes the proof.
Pr[fn = sv]
We can write the result of this theorem in matrix for
This completes the proof.
\Ve can write the result of this theorem in matrix form.
Let '/Tn
F I N I T E MARKOV CHAINS
32
SEC. 3
The transition matrix for the new chain (Example 12) i
BASIC CONCEPTS OF MARKOV CHAINS
33
be a row vector which gives the induced measure for the outcome
function f n . That is
71"
= {p(n)l, p(n)z, ... , p(n)r},
is the probability that the process
where p(n)j=Pr[fn=sil. Thus.
will after n steps be in state Sj. The vector 710 is the initial probabilities
vector. Let P(n) be the matrix with entries Pij(n). Then the result
We be
shall
see inin3 the
6.5 form
t h a t the study of this new chain gi
of Theorem 2.3.1 may
written
detailed information about the original process than could
71 n = 71 n - l ' P(n)
directly from
the two-state chain.
for n;? 1.
By successive application of this result we have
5 2.3 Connection with matrix theory. I n this section w
71n connection
= 71o·P(1)·P(2)
. ...Markov
. P(n). chain theory and matrix t
the
between
shall
start
with
t
h
e
general
finite Markov
process
In the case of a Markov chain process, all the F\n)'s
are the same
andand the
our results
to the finite
Markov chain.
we obtain the following
fundamental
theorem.
'
Let fnmeasure
be the outcome
2.3.2 THEOREM. 2.3.1
Let -r;-nTHEOREM.
be the induced
for thefunction
outcomeclt time
Mnrkov
process
lcith
trnnsition
probabilities
pt,(n),
710 t h ~ n
f1mction fn for a finite 1fJarkov chain with initial probability vector
and transition matrix P. Then
f n = s, is a statement relative to
This theorem showsPROOF.
that theThe
key statement
to the study
of the induced measures
To
find
its
probability,
we
the
of of
all the
paths in i
for the outcome functions of a finite Markov add
chain
is weights
the study
T h a t is,matrix.
all possible
paths
which
in outcome
s,. Thus i
powers of the transition
The
entries
ofend
these
powers have
is a possible
sequence of
states
themseh'es an interesting
probabilistic
interpretation.
To see this,
Path
Weights
1/4
1/ 8
By the Markov property this is
If in this last sum we keep u fixed and sum over the remai
we obtain
W f n = sv] =
2 Pr[f,-1 = s,]p,,(n).
I1
This completes the proof.
1/
We can write the result of this theorem in 8matrix for
FIC1.an; 2·1
take as initial vector 770 the vector with 1 in the i-th comp
0 otherwise.
by Theorem
B uII
t noP*
FINITE Then
MARKOV
CHAINS3.2, a,=noPn.CHAP.
-------------------------row of the matrix Pn. Thus the i - t h row of the n-th po
take as initial vector
7T0 the
vector
withthe1 probability
in the i-th component
transition
matrix
gives
of being in and
each of t
otherwise. Then
by
Theorem
3.2,
TTn
=
TTOP".
But
7T
the i-th
OPn isstarted
states under the assumption t h a t t h e process
in stat
row of the matrixI npn.
Thus the
i-thusrow
of the
n-ththepower
of the
Example
1, let
assume
that
process
starts i
transition matrix
gives
the probability ofWe
being
each
the various
770 = ( 0 , 0 , 1 , 0, 0).
caninfind
theofinduced
measures
Then
states under thefor
assumption
that
the
process
started
in
state
Si.
the first three outcome functions by constructing
a tre
In Example measure
1, let usforassume
the
process starts
in tree
state
53.
the firstthat
three
experiments.
This
is given
in
Then TTO={O, 0,1,0,
O}. this
We tree
can find
measures
(seecompute
§ 1.7)
From
and the
treeinduced
measure
we easily
th
for the first three
outcome
by constructing
a treeareand tree
measures
forfunctions
the functions
f1, 4, f3. They
measure for the first three experiments. This tree is given in Figure 2-1.
From this tree and tree measure we easily compute the induced
measures for the functions fl' f2, f3. They are
34
°
7Tl = {a, Ih, 0, liz, O}
By Theorem
these
induced measures should also be
TT2 = {l/4,2.3.2
0, l/z,
0, 1/4}
P, P4,21/, and
row in 7T3
the= matrices
{l/4, 1/4,0,1/
4}. P3, since the starting sta
These matrices are
By Theorem 2.3_2 these induced measures should also be the third
row in the matrices P, p2, and P3, since the starting state was S3.
These matrices are
° °
1/2
°
1/ 2
°
0 liz
°
0
°
° 1/4
liz 1/ 4
° 1/4°
liz
1/4
l! 2
P=
0
0
0
0
liz
0
0
0
1
0
0
0
0
0
0
0
p2 =
0
1/4
0
0
0
1/4
1/2
°
0
° 1/ 8
5/ 8
1/4
°
1! 4 1/4
1/4 1/ 4
°
1/4
I! 8
0
0
0
0
p3 =
0
0
0
0
0
0
5! 8
take as initial vector 770 the vector with 1 in the i-th comp
0 otherwise. Then by Theorem 3.2, a,=noPn. B u t noP*
SEC. 4
BASIC
OF Pn.
MARKOV
row CONCEPTS
of the matrix
Thus CHAINS
the i - t h row of the35n-th po
transition
matrix
gives
the
probability
of being
in each of t
We thus see that these matrices furnish us several tree
measures
sim ultaneously. states under the assumption t h a t t h e process started in stat
I n Example 1, let us assume that the process starts i
§ 2.4 Classification of states and chains. We wish to classify the
Then 770 = ( 0 , 0 , 1 , 0, 0). We can find the induced measures
states of a Markov
according
to whether
it is possible
to go a tre
for chain
the first
three outcome
functions
by constructing
from a given state to another given state. This problem is exactly
measure for the first three experiments. This tree is given in
like the one treated From
in § 1.4.thisIftree
,':e interpret
to meanwethat
the proand treeiTj
measure
easily
compute th
cess can go from state
Si to stctte Sj (not necessarily in one step), then
measures for the functions f1, 4, f3. They are
all the results of that section are applIcable.
In particular, the states are divided into equivalence classes. Two
states are in the same equivalence class if they "communicate," i.e. if
one can go from either state ~.) the other one. The resulting partial
ordering shows us the
possible directions
which measures
the process
can also be
By Theorem
2.3.2 theseininduced
should
proceed.
row in the matrices P, P 2 , and P3, since the starting sta
The minimal elements
of theare
partial ordering are of particular
These matrices
interest.
2.4.1 DEFINITION. The minimal elements of the partial ordering of
equivalence classes are called ergodic sets. The remaining elements
are calle.d transient sets. ThE elements of a transient set are called
transient states. The ei£lnents of an ergodic set are called ergodic
(or non-transient) states.
Since every finite partial ordering must have at least one minimal
element, there must be at least one ergodic set for every Markov chain.
However, there need be no transient set. The latter will occur if the
entire chain consists of a single ergodic set, or if there are several
ergodic sets which do not communicate with others.
From the results of § 1.4 we see that if a process leaves a transient
set it can never return to this set, v;hile if it once enters an ergodic set,
it can never leave it. In particular, if an ergodic set contains only
one element, then we ha,-e a state which once entered cannot be left.
Such a state is called absorbing. Since from such a state we cannot
go to another state, the following theorem characterizes absorbing
states.
2.4.2
THEOREM.
A state Si is a))sorbing if and only if Pii = 1.
It is convenient to use Gur classification to arrive at a canonical
form for the transition matrix. We renumber the states as follows:
The elements of a given equivalence class will receive consecutive
numbers. The minimal sets will come first, then sets that are one
level above the minimal sets, then sets two levels above the minimal
sets, etc. This will assure us that we can go from a given state to
another in thc same class, or to a state in an earlier class, but not to
a state in a later class. If the equivalence classes arranged as here
36
described are ul, uz, . . . , uk, then our matrix will appear as f
CHAP.
MARKOV
(whereFIKITE
k is taken
as 5 , for CHAINS
the sake of 'illustration)
: II
-------------------
described are Ul, U2, . . . , Uk, then our matrix will appear as follows
(where k is taken as 5, for the sake of'illustration):
Ul:
U2 :
U3 :
/
PI
P2
R2
Rs
0
P3
R4
P4 I
Here
the
Pi
represent
transition
within a given equiv
I matrices
pU5 :
Rs
class. The region 0 consists entirely of 0's. The matrix Ri w
entirely 0 if Pi is an ergodic set, but will have non-zero ele
Here the Pi represent
otherwise.transition matrices within a given equivalence
consists
The happens
matrix Rl
class. The regionI n0 this
form ientirely
t is easyof
t o O's.
see what
as will
P is be
raised to p
entirely 0 if Pi
is
an
ergodic
set,
but
will
have
non-zero
elements
Each power will be a matrix of the same form ; in Pn we still have
otherwise.
in the upper region, and we simply have P*a in the diagonal re
In this form it is easy to see what happens as P is raised to powers.
This shows that a given equivalence class can be studied in iso
Each power will be a matrix of the same form; in pn we still have zeros
by treating the submatrix Pi. This will be considered in detail
in the upper region, and we simply have P"j in the diagonal regions.
We can also apply the subdivision of an equivalence class cons
This shows that a given equiv<Llence class can be studied in isolation,
in the previous chapter. We saw there that each equivalence
by treating the submatrix Pi. This will be considered in detail later.
can be partitioned into cyclic classes. If there is only one
We can also apply
the subdivision
of the
an equivalence
class, then
we say that
equivalenceclass
classconsidered
is regular, otherw
in the previoussay
chapter.
\Ve
saw
there
that
each
equivalence
class
that i t is cyclic.
can be partitioned
intoequivalence
cyclic classes.
is then
only after
one cyclic
If an
class If
is there
regular,
sufficient tim
class, then we say that the equivalence class is regular, otherwise we
elapsed the process can be in any state of the class, no matter
say that it is cyclic.
of the equivalent states it started in (see $ 1.4). This means th
If an equivalence class is regular, then after sufficient time has
sufficiently high powers of its Pi must be positive (i.e. have
elapsed the process can be in any state of the class, no matter which
positive entries). If the equivalence class is cyclic, then no po
of the equivalent
started in (see § 1.4). This means that all
Pt states
can be itpositive.
of classification
its Pi must of
be states
positive
sufficiently high Prom
powersthis
we (i.e.
can have
arriveonly
a t a classif
positive entries). If the equivalence class is cyclic, then no power of
of Markov chains. We have noted that there must be an ergod
Pi can be positive.
but there need be no transient set. This will lead to our pr
From this classification of states we can arrive at a classification
subdivision. Within this we can subdivide according to the n
of Markov chains.
We have
notedsets.
that there must be an ergodic set,
and type
of ergodic
but there need be no transient set. This will lead to our primary
subdivision. Within
thisWithout
we canTransient
subdivideSets
according to the number
I. Chains
and type of ergodic
sets.
If such a chain has more than one ergodic set, then there is
U4 :
I
1___
' _
lutely
no interaction
between these sets.
1. Chains Without
Transient
Sets
Hence we have two o
unrelated Markov chains lumped together. These chains m
If such a chain has more than one ergodic set, then there is abso·
studied separately, and hence without loss of generality we
lutely no interaction between these sets. Hence we have two or more
unrelated Markov chains lumped together. These chains may be
studied separately, and hence without loss of generality we may
SEC. 4
described are ul, uz, . . . , uk, then our matrix will appear as f
(where
k isCONCEPTS
taken as 5 , OF
for the
sake of CHAINS
'illustration) :
BASIC
MARKOV
37
a.ssume tha.t the entire chain is a single ergodic set.
of a single ergodic set is called <in ergodic chain.
A chain consisting
I-A. The ergodic set is regular. In this case the chain is called a
regular Markov chain. As Vie see from previous considerations, all
sufficiently high powers of P must be positive in this case. Thus no
matter where the process starts, after sufficient lapse of time it
could be in any state.
I-B. The ergodic set is cyclic. In this case the chain is called a
cyclic Markov
a ehain
has a matrices
period d,within
and its
statesequiv
Herechain.
the Pi Such
represent
transition
a given
are subdivided
into
d cyclic sets (d>l). For a given starting
class. The region 0 consists entirely of 0's. The matrix Ri w
position entirely
it will move
through
cyclic
sets
a definite
order, ele
is an the
ergodic
set,
butinwill
have non-zero
0 if Pi
returningotherwise.
to the set of the starting state after d steps. We also
know that Iafter
elapsed,
the process
be in to p
n thissufficient
form i t istime
easyhas
t o see
what happens
as Pcan"
is raised
any stateEach
of the
cyclic
set
appropriate
for
the
moment.
power will be a matrix of the same form ; in Pn we still have
the Transient
upper region,
II. Chain8 in
With
Sets and we simply have P*a in the diagonal re
This shows that a given equivalence class can be studied in iso
In such a chain the process moves towards the ergodic sets. As will
by treating the submatrix Pi. This will be considered in detail
be seen in the next chapter, the probability that the process is in an
We can also apply the subdivision of an equivalence class cons
ergodic set tends to 1 ; and it cannot escape from an ergodic set once it
in the previous chapter. We saw there that each equivalence
enters it. Hence it is fruitful to classify such chains by their ergodic scts.
can be partitioned into cyclic classes. If there is only one
II-A. Allclass,
ergodic
are that
unit the
sets.equivalence
Such a chain
called an
otherw
thensets
we say
class isis regular,
absorbingsay
chain.
case the process is eventually trapped in a
that iIn
t isthis
cyclic.
single (absorbing)
state. This class
type of
alsoafter
be characterIf an equivalence
is process
regular,can
then
sufficient tim
ized by the
fact that
all the ergodic
states
absorbing
states.
elapsed
the process
can be in
any are
state
of the class,
no matter
of
the
equivalent
states
it
started
in
(see
$
1.4).
This
means th
II-B. All ergodic sets are regular, but not all are unit sets.
sufficiently high powers of its Pi must be positive (i.e. have
II-C. Allpositive
ergodic entries).
sets are cyclic.
If the equivalence class is cyclic, then no po
Pt can
positive.
II-D. There
are be
both
cyclic and regular ergodic sets.
Prom this classification of states we can arrive a t a classif
Naturally,
in each chains.
of these We
classes
can that
further
classify
chains
of Markov
havewe
noted
there
must be
an ergod
according to
how
many
ergodic
sets
there
are.
Of
particular
interest
but there need be no transient set. This will lead to our pr
is the question
whether there
arethis
one we
or more
ergodic sets.
subdivision.
Within
can subdivide
according to the n
We can illustrate
of these sets.
types except II-D by the random walk
and type all
of ergodic
examples.
I. Chains
Without
Transient
For Example
1: The
states
81 and Sets
85 are absorbing states.
The
states S2, 83, If
S4 such
are transient
is possible
to go set,
between
a chain states.
has moreItthan
one ergodic
then any
there is
two of these
states.
Hence theybetween
form a single
set. we
\Ve
have
lutely
no interaction
these transient
sets. Hence
have
two o
an absorbing
Markov Markov
chain-that
is, case
II-A. together. These chains m
unrelated
chains
lumped
For Example
:2: In
this exa;nple
is possible
to go
from
state we
studied
separately,
and ithence
without
loss
of any
generality
to any other state. Hence there are no transient states and there is a
single ergodic set. Thus we have an ergodic chain. It is possible to
return to a state only in an even number of steps. Thus the period
38
FINITE MARKOV CHAINS
of the states is 2. The two cyclic sets are {sl, s3, ssj a
This isFINITE
type I-B.
:MARKOV CHAINS
CHAP. II
For Example 3 : Again we can go from any state t o any o
of the states isHence
2. The
two cyclic
setsergodic
are {51,chain.
S3, S5} It
andis {S2'
S4}' t o
we again
have an
possible
This is type I-B.
state s3 from ss in either two or three steps. Hence th
For Examplecommon
3: Againdivisor
we cand go
state.
= 1,from
and any
the state
periodtoisany
1. other
This is
type IHence we again For
haveExample
an ergodic
It is {sl,
possible
to ergodic
return set
to whi
4 : I nchain.
this example
ss) is an
state 83 from regular.
S3 in either two or three steps.
Hence
the greatest
transient
set. This i
The set {sz,sa,s4}is the single
common divisor d=
and the 5period
1. have
Thisa issingle
type ergodic
I-A.
For1,Example
: Hereiswe
set {sl, sg
For Example period
4 : In this example {51, 55} is an ergodic set which is clearly
2. The set (s2, s3, s4) is again a transient set. This i
regular. The set {S2' 53, S4} is the single transient set. This is type II-B.
be studied.
us consider
For Example 5:$2.5
Here Problems
v.. e have atosingle
ergodic Let
set {51,
55} whichour
hasvario
chains,
and
ask
what
types
of
problems
we
would
like
period 2. The set {S2' S3, S4} is again a transient set. This is type II-C.t o an
following chapters.
§ 2.5 ProblemsFirst
to beofstudied.
Letwish
us consider
variousMarkov
types ofchai
all we may
to studyour
a regular
chains, and askawhat
types
of
problems
we
would
like
to
answer
in the
chain the process keeps moving through all the
states,
following chapters.
where i t starts. Some of the questions of interest are :
First of all we may wish to study a regular Markov chain. In such
(1) If
a chain
startsthrough
in st, what
is the
probability
after n
a chain the process
keeps
moving
all the
Rtates,
no matter
will beofinthe
sj ?questions of interest are:
where it starts. i t Some
(2) Can we predict the average number of times that
(I) If a chain
St. what is the probability after n steps that
is starts
in si ? inAnd
if so, how does this depend on where the pro
it will be in Sj? (3) We may wish t o consider t h e process as i t goes fro
(2) Can we predict
ofoftimes
that theof process
What is the
the average
mean andnumber
variance
the number
steps need
is in Si? And if
so,
how
does
this
depend
on
where
the
process
starts'1
are the mean and variance of the number
of states
passed
(3) We may the
wish
to consider
it goesthrough
from 5is*?
to Sf.
probability
t h athe
t theprocess
processaspasses
What is the mean(4)
andWe
variance
of the
steps needed?
may wish
to number
study aofcertain
subset of ,"Vhat
states, a
are the mean and
variance
of
the
number
of
states
passed?
What does
is t
the process only when i t is in these states. How
the probabilityour
that
the process
passes
through
Sk ~
? These
previous
results
questions
are treated in Chap
(4) We may wish to study a certain subset of states, and observe
t o states.
study a How
cyclicdoes
chain.
the sa
the process only Next
whenwe
it may
is inwish
these
thisHere
modify
questionsThese
are ofquestions
interest are
as for
a regular
chain.IV.Naturally
our previous results?
treated
in Chapter
chain is easier to study; and we will find that, once w
N ext we mayanswers
wish tofor
study
a cyclic
chain.
kindsthe
of co
regular
chains,
i t isHere
not the
hardsame
t o find
questions are of
interestforasallforergodic
a reg 11lar
chain.This
Naturally,
answers
chains.
extensiona ofregular
regular c
chain is easiertotothe
study;
we willchains
find that,
once out
we inhave
the V.
theoryand
of ergodic
is carried
Chapter
answers for regular
chains,
it
is
not
hard.
to
find
the
corresponding
Next we may wish to consider a Markov chain with trans
answers for all ergodic chains. This extension of regular chain theory
There are two kinds of questions to be asked here. One w
to the theory of
ergodic
chains
is carried
out ini tChapter
Y.ergodic set, whi
the
behavior
of the
chain before
enters an
Next we maykind
wishwill
to consider
a
Markov
chain
with
transient
states.set.
apply after the chain has entered an ergodic
There are two questions
kinds of questions
to be asked
One considered
will concernabov
are no different
fromhere.
the ones
the behavior ofchain
the chain
before
it enters
ergodic
set,
while
enters
a n ergodic
set an
i t can
never
leave
it,the
andother
hence th
kind will applyofafter
the
chain
has
entered
an
ergodic
set.
The
latter of
states outside the set is irrelevant. Thus questions
questions are no
the ones
considered above.
a
kinddifferent
can be from
answered
by considering
a chaii Once
consisting
chain enters anergodic
ergodic set,
set it
can
never
leave
it,
and
hence
the
existence
i.e. a n ergodic chain.
38
of states outside the set is irrelevant. Thus questions of the second
kind can be answered by considering a chain consisting of a single
ergodic set, i.e. an ergodic chain.
FINITE MARKOV CHAINS
38
SEC. 5
of the states is 2. The two cyclic sets are {sl, s3, ssj a
BASIC
CONCEPTS
39
This is
type I-B. OF ?Y1ARKOV CHAINS
For Example
3 : Again
we can gooffrom
state
o any o
The really new questions
concern
the behavior
the any
chain
up tto
weanagain
have
ergodic chain.
is possible t o
the moment thatHence
it enters
ergodic
set. anHowever,
for theseItquestions
s3 from
either two
steps. them
Hence th
the nature of thestate
ergodic
statesssis in
irrelevant,
and or
wethree
may make
common
divisor
d
=
1,
and
the
period
is
1.
This
is
type Iall into absorbing states if we wish. More generally, if we wish to
For Example
: I n of
thistransient
examplestates,
{sl, ss) is
ergodic
set whi
study the process while
it is in a4 set
weanmay
make
regular.
{sz,
sa,s4}
is
the
single
transient
set.
The
set
all other states absorbing. This modified process will serve to find allThis i
For Example
: Here
havequestions
a single ergodic
the answers we desire.
Hence 5the
onlywenew
concern set
the{sl, sg
period 2. cha.in.
The set (s2, s3, s4) is again a transient set. This i
behavior of an absorbing
Some of the questions
that are to
of be
interest
a transient
$2.5 Problems
studied.concerning
Let us consider
our vario
state Si are:
chains, and ask what types of problems we would like t o an
following
chapters.
(1) The probability
of entering
a given ergodic set, starting from s,.
First
of
all we
maynumber
wish toofstudy
regular
(2) The mean and variance
of the
timesa that
the Markov
process chai
a chain
process
moving
through
all theonstates,
is in Si before entering
an the
ergodic
set, keeps
and how
this number
depends
where i t starts. Some of the questions of interest are :
the starting position.
(3) The mean and
the number
of steps
beforeafter n
(1) variance
If a chainofstarts
in st, what
is theneeded
probability
entering an ergodic
set be
starting
i t will
in sj ? at Si.
(4) The mean number
passed
(2) Canofwestates
predict
the before
averageentering
numberanof ergodic
times that
set, starting at St.
is in si ? And if so, how does this depend on where the pro
may
wish t ochains,
consider
e process
as i t goes fro
Chapter III will (3)
dealWe
with
absorbing
andt hall
these questions
What
is the
and variance
the numberquestions
of steps need
will be answered.
Thus
we mean
will :find
the mostofinteresting
are the
mean
and variance
of the III,
number
of states
passed
about finite Markov
chains
answered
in Chapters
IV, and
V.
the probability t h a t the process passes through s*?
(4) We
may wish
to study
Exercises
for Chapter
II a certain subset of states, a
the process only when i t is in these states. How does t
For § 2.1
? These questions are treated in Chap
our previous results
1. Five points are marked. on a circle. A process moves from a given
Next we may
wish t o study
a cyclic
chain. Here
point to one of its neighbors,
with probability
liz for
each neighbor.
Findthe sa
questions
are of interest
for a regular chain. Naturally
the transition matrix
of the resulting
Markovas
chain.
easier
to A
study;
we with
will probability
find that, 2/once
w
2. Three tanks chain
fight a isduel.
Tank
hits itsand
target
3,
tank B with probability
1/2,for
andregular
tank C with
probahility
1/3. hard
Shots
fired
answers
chains,
i t is not
t oare
find
the co
once afor
tank
is hit it chains.
is out of This
action.extension
As a state
we
simultaneously, and
answers
all ergodic
of regular
c
choose the set of tanks still in action. If on each step each tank fires at its
V.
to
the
theory
of
ergodic
chains
is
carried
out
in
Chapter
strongest opponent., verify that the following transition matrix is correct:
Next we may wish to consider a Markov chain with trans
B
C
AC BC toABC
ThereE are ~'1..
two kinds
of questions
be asked here. One w
the
behavior
of
the
chain
before
i
t
enters
E
0
0
0
0
0 \ an ergodic set, whi
0
kind will apply after the chain has entered an ergodic set.
A
0
0
0 different
0
0 the 0ones considered abov
questions
are no
from
0 a n ergodic
B
chain0 enters
0 set0i t can
0 never
0 leave it, and hence th
of states
outside
Thus questions of
() the set is
0
C
0
0 irrelevant.
0
0
kind9' can be answered by considering a chaii consisting
AC
1/ 9 2/ 9chain.
0
-i9 set, i.e. 0a n ergodic
0
ergodic
Be
ABC
1/6
0
1/3
I! 6
0
1/ 3
0
0
0
0
4/ 9
2/ 9
2/ 9
1/ 9
40
FINITE MARKOV CHAINS
CH
3. Modify the transition matrix in the previous exercise, assumin
fires a t B,'B a t 6 , and CHAP.
G a t A.II
when all FINITE
tanks are MARKOV
in action, ACHAINS
4. We carry out a sequence of experiments as follows : At first a fa
3. Modify the
transitionThen,
matrix
in the previous
assuming
that
n- 1 exercise,
comes out
heads, we
toss a fai
if experiment
is tossed.
when all tanks ifare
action,
fires we
at B,'B
C, and
C athas
A. probability l / n of com
it in
comes
outAtails,
toss at
a coin
which
What are
the transition
4. We carryheads.
out a sequence
of experiments
as probabilities?
follows: At firstWhat
a fair kind
coin of pro
thisif
? experiment n-l comes out heads, we toss a fair coin;
is tossed. Then,
if it comes out tails, we toss a coin which has probability
For $ 2.2 lin of (loming up
heads. What are the transition probabilities 1 What kind of process is
this 1
5. Modify Example 1 by assuming that when the process reache
Form the new transition matrix.
goes on the next step
Fort o§state
2.2 sz.
6 . Modify the process described in Example 2 by assuming that wh
5. Modify Example
1 by assuming
when
thenext
process
81 on
it the thi
sl it staysthat
there
for the
two reaches
steps and
process reaches
goes on the next
step to
to state
moves
state 82.
sz. Form
Show the
thatnew
thetransition
resulting matrix.
process is not a Markov
(with
the five
given in
states).
6. Modify the
process
described
Example 2 by assuming that when the
process reaches 817.itIstays
there 6for
the that
nextwe
twocan
steps
and
the third
n Exercise
show
treat
theonprocess
as astep
Markov ch
moves to stateallowing
82.
Show
that number
the resulting
process
is not
a Markov
chain matrix
Write
down
the transition
a larger
of states.
(with the five given states).
8. Modify the transition matrix of Example 7, assuming that the d
7. In Exercise
6 show
that to
webe
can
treat thethan
process
a Markov
twice
as likely
generated
anyas
other
digit. chain, by
allowing a larger number of states. Write down the transition matrix.
9. Modify Example 7, assuming that the same digit is never ge
8. Modify the
transition
matrix
of Example
7, assuming
thatlikely
the digit
0 is
twice
in a row,
but otherwise
digits
are equally
to occur.
twice as likely to 10.
be generated
than
any other
I n Example
8 allow
onlydigit.
two states: Nice and not nice. Sho
9. Modify Example
7, assuming
that the
sameand
digit
never
generated
the process
is still a Markov
chain,
findisits
transition
mat,rix.
twice in a row, but otherwise digits are equally likely to occur.
10. In Example 8 allow only two states: NiceFor
and
not nice. Show that
$ 2.3
the process is still a Markov chain, and find its transition matrix. .
11. I n Example I l a compute P2, P4, Pa,and P16, and write the
the trend, and interpret your results.
as decimal fractions.ForNote
§ 2.3
12. Show that, no matter how Example 7 is started, the probabili
11. In Example
p2,states
P4, p8,
and1 step
p16, agree
and write
beinglla
in compute
each of the
after
with the
the entries
common row
as decimal fractions.
Note
the trend,
and
interpret
your results.
What
are
the probabilities
after n steps?
transition
matrix.
12. Show that,13.
no Assume
matter how
the probabilities
thatExample
Example 78isis started,
started with
initial vectorfor
vo= (2/5,
being in each of
thenstates
step
What1 is
n,?agree with the common row for the
l , nz. after
Find
transition matrix.14.What
are the is
probabilities
steps?of Oz. What kind of w
The weather
nice today after
in then Land
13. Assume is
that
Example
8 is
started
vector 170= e/5, 1/ 5 , 2/5).
most
likely to
occur
daywith
afterinitial
tomorrow?
Find 7T1, 7T2. What
15. isI n17n?
Example 11, assume that c = V 2 and d=1I4. The man ra
14. The weather
is nice
today in to
theplay
Land
of Oz.
What
kind
of weather
What
is the
probability
that he p
chooses
the machine
first.
is most likely to
occurmachine
day after
better
(a)tomorrow!
on the second play, (b) on the third play, and (c
fourth
15. In Example
11,play?
assume that c=l/z and d=1/4. The man randomly
chooses the machine
first. 2 assume
What isthat
the the
probability
he plays
the s3. Co
16. to
I n play
Example
process isthat
started
in state
better machinea (a)
thetree
second
play,for
(b)the
on first
the third
and (c) onUse
thethis to
treeonand
measure
three play,
experiments.
fourth play? induced measure for the first three outcome functions. Verify th
and P3.
results
agree with
foundinfrom
16. In Example
2 assume
that the probabilities
process is started
statePS3., P2,Construct
a tree and tree measure for the first three experiments. Use this to find the
induced measure for the first three outcome functions.
For $ 2.4 Verify that your
results agree with the probabilities found from P, pz, and P3.
17. For the following M a r k o ~chain, give a complete classificatio
states and put the transition
For § 2.4 matrix in canonical form.
40
17. For the following Markov chain, give a complete classification of the
states and put the transition matrix in canonical form.
FINITE MARKOV CHAINS
40
CH
3. Modify the transition matrix in the previous exercise, assumin
A fires a t CHAINS
B,'B a t 6 , and G a t A.
when
all tanks
are in action,
BASIC
CONCEPTS
OF MA.RKOV
41
4. We carry out a sequence of experiments as follows : At first a fa
is tossed. Then, if experiment n55 1 comes out heads, we toss a fai
if it comes out tails, we toss a coin which has probability l / n of com
0
0
0
1
o
o
0
What kind of pro
heads. What are the transition probabilities?
this ?
000
o o o 1
o 0
o oFor o$ 2.2 0
SEC. 5
5. Modify
1 : assuming
1/2 Example
0
0 1 by
o othat0when the process reache
i2
goes on the next step t o state
sz. Form the new transition matrix.
0the
0process
0 described
o 1 in Example
o 0 2 by assuming that wh
6 . Modify
sl 0
it stays there
process reaches
00
o ofor theo next two steps and on the thi
moves to state sz. Show that the resulting process is not a Markov
liz states).
0
o
o l/Z 0
(with the ofive given
7. I n Exercise 6 show that we can treat the process as a Markov ch
18. A Markov
chain ahas
thenumber
following
transition
matrix,
non· zero matrix
Write
downwith
the transition
allowing
larger
of states.
entries marked by x. Give a complete classification of the states and put the
theform,
transition matrix of Example 7, assuming that the d
transition matrix8.inModify
canonical
twice as likely to be generated than any other digit.
84 7, 85
56
87
88 same
89 digit is never ge
9. Modify Example
assuming
that
the
0
o are
0 equally
o likely
0' to occur.
twice in a row, but o
otherwise
digits
8o allow
x
0 onlyo two xstates:
x10. I nx Example
o Nice0 and not nice. Sho
the process is still a Markov chain, and find its transition mat,rix.
o
o
0
o 0 o x o 0
x
0
o o 0
o For0 $ 2.3o x
0
x
o P2, 0P4, Pa,
o and0 P16, and write the
o11. I noExample
Iol a compute
your results.
as o
decimal
o fractions.
x
o Note
0 thex trend,
0 andointerpret
0
no
matter
7 is 0
started, the probabili
o that,
x
o12. Show
o
0 how
o Example
0
o
being in each of the states after 1 step agree with the common row
0
x probabilities
0
x
0
o x matrix.
oWhat0 are the
S8
n steps?
after
transition
13.
Assume
that
Example
8
is
started
with
initial
vector vo= (2/5,
o
o
0
x
0
o
0
o
x
59
Find n l , nz. What is n,?
19. Cbssify the14.
following
chainsis as
ergodic
of thekind of w
The weather
nice
todayorinabsorbing.
the Land ofWhich
Oz. What
ergodic chains is most
regular?
likely to occur day after tomorrow?
15. I n Example 11, assume that c = V 2 and d=1I4. The man ra
chooses the machine to play first. What is the probability that he p
(a)
better machine (a) on the second play, (b) on the third play, and (c
fourth play?
16. I n Example 2 assume that the process is started in state s3. Co
a tree and tree measure for the first three experiments. Use this to
o 0 functions. Verify th
induced measure for the first three outcome
results agree
with
the
probabilities
found
1/from
3 1/ 3P , P2, and P3.
o .
liZ)
(c)
o
1/ 3
For $ 2.4
1/3
1/3 1h
17. For the following M a r k o ~chain, give a complete classificatio
states and put the transition matrix in canonical form.
(e)
P =
(~
0
1/2 1/2
001 )
20. In Example 9 classify the states. Put the transition matr
canonicalFINITE
form. What
type ofCHAINS
chain is .this?
MARKOV
CHAP. II
21. For an ergodic chain the i-th state is made absorbing by replacin
20. In Example
9 classify
the states.
Putbythe
transition
in the in
i-th compo
i-th row
in the transition
matrix
a row
with a 1matrix
canonical form.Prove
What
type
chain ischain
this? is absorbing.
that
theof
resulting
21. For an ergodic
the i-th
made absorbing
bydreplacing
theresulting
22. Ichain
n Example
11,state
giveisconditions
on c and
so that the
i-th row in theistransition matrix by a row with a 1 in the i-th component.
Prove that the resulting chain
is absorbing.
(a) ergodic
(b) regular (c) cyclic ( 6 ) absorbing
22. In Example 11, give conditions on c and d so that the resulting chain
is
For the entire chapter
42
(a) ergodic (b) regular (c) cyclic (d) absorbing
23. In a certain state a voter is allowed to change his party affiliatio
primary elections)
by chapter
abstaining from the primary for one year.
For theonly
entire
sl indicate that a man votes Democratic, sz that he votes Republican,
23. In a certain
a voter in
is allowed
changeExperience
his party affiliation
that state
he abstains,
t,he giventoyear.
shows that(for
a Democr
primary elections)
only112bythe
abstaining
from
the primary
for one
year. Letwill absta
abstain
time in the
following
primary,
a Republican
81 indicate thattime,
a man
votesa Democratic,s2
that he for
votes
Republican,
andlikely
S3
while
voter who abstained
a year
is equally
to vo
that he abstains,
in the
given
shows
a Democrat
will
either
party
in year.
the nextExperience
election. [We
willthat
refer
to this as Example
13.
abstain l/z the time
in thethe
following
primary,
(a) Find
transition
matrix.a Republican will abstain 1/4
time, while a voter
for a that
year ais man
equally
vote for this yea
(b) who
Find abstained
the probability
wholikely
votes to
Democratic
either party in the next
election.
willfrom
refernow.
to this as Example 13.]
abstain
three[We
years
(a) Find the transition
matrix.
(c) Classify
the states.
(b) Find the probability
thatyear
a man of
who
Democratic
year will112 Repub
thevotes
population
votes this
Democratic,
(d) I n a given
abstain three years
from
now. What proportions do you expect in the next pr
the rest
abstain.
(c) Classify the states.
election ?
(d) In a given year
of the population
votesisDemocratic,
Republican,
24. A1/4
sequence
of experiments
performed, l}Z
in each
of which two fai
the rest abstain.
do you
the next
sl indicate that
twoexpect
headsin
come
up, szprimary
that a head and
Letproportions
are tossed.What
election 1come up, and s3 that two tails turn up. [We will refer to this as Examp
24. A sequence(a)
of experiments
is performed,
in each of which two fair coins
Find the transition
matrix.
are tossed. Let 81
heads
up, S2 toss,
that awhat
headisand
tail
(b)indicate
If two that
headstwo
turn
up come
on a given
thea probability
come up. and S3 that two
tails
turn up.
[We will
refer
to ?this as Example 14.]
heads
turning
up three
tosses
later
(a) Find the transition
matrix.
(c) Classify
the states.
(h) If two heads turn up on a given toss, what is the probability of two
heads turning up three tosses later 1
(c) Classify the states.
20. In Example 9 classify the states. Put the transition matr
canonical form. What type of chain is .this?
21. For an ergodic chain the i-th state is made absorbing by replacin
i-th row in the transition matrix by a row with a 1 in the i-th compo
Prove that the resulting chain is absorbing.
22. I n Example 11, give conditions on c and d so that the resulting
is
(a)CHAPTER
ergodic (b) regular
III (c) cyclic ( 6 ) absorbing
For the entire chapter
23. In a certain state a voter is allowed to change his party affiliatio
ABSORBING
MARKOV
primary
elections) only
by abstainingCHAINS
from the primary for one year.
sl indicate that a man votes Democratic, sz that he votes Republican,
that he abstains, in t,he given year. Experience shows that a Democr
abstain 112 the time in the following primary, a Republican will absta
time, while a voter who abstained for a year is equally likely to vo
§ 3.1 Introduction.
recall
the basic
definitions
relevant
to
either party inLet
theus
next
election.
[We will
refer to this
as Example
13.
an absorbing (a)
chain.
In the
classification
Find the
transition
matrix. of states, the equiv~lence
(b) Find
thetransient
probability
that
a mansets.
who votes
Democratic
classes were divided
into
and
ergodic
The former,
oncethis yea
abstain
three while
years from
now. once entered, are never
left, are never again
entered;
the latter,
Classify
the states.
again left. If(c)
a state
is the
only element
of an ergodic
set, then it 112
is Repub
of the population
votes Democratic,
(d) I n a given year
called an absorbing
state·.
For such
a state
Sj the entry
mustinbe
What
proportions
do youPii
expect
the1,next pr
the rest abstain.
and hence all other
entries
? in this row of the transition matrix are 0.
election
non-transient
statesisare
absorbing,
is called
antwo fai
A chain, all of24.whose
A sequence
of experiments
performed,
in each
of which
sl indicate
that two
up, chapter.
sz that a head and
are tossed.
absorbing chain.
TheseLet
chains
will occupy
us heads
in thecome
present
come up, and s3 that two tails turn up. [We will refer to this as Examp
3.1.1 THEOREM. In any finite 1'rlarkov chain, no matter where the
(a) Find the transition matrix.
process starts, the probability after n steps thai the process is in an ergodic
(b) If two heads turn up on a given toss, what is the probability
state tends to 1 as
n tends
to infinity.
heads
turning
up three tosses later ?
(c)
Classify
the
states.
PROOF.
If the process once
reaches an ergodic state, then it can
never leave its equivalence class, and hence it will at all future steps be
in an ergodic state. Suppose that it starts in a transient state. Its
equivalence class is not minimal; hence there is a minimal element
below it. This means that it must be possibJe to reach some ergodic set.
Let us suppose that from any transient state it is possible to reach an
ergodic state innot more than nsteps. (Since there are only a finite number of states, n is simply the maximum of the number of steps required
from each state.) Hence there is a positive number p such that the probability of entering an ergodic state in at most n steps is at least p, from
any transient state. Hence the probability of not reaching an ergodic
state in n steps is at most (1 - p), which is less than 1. The probability
of not reaching an ergodic state in kn steps is less than orequal to (1 - P )A;,
and this probabili ty tends to as k increases. Hence the theorem follows.
°
There are numbers b>O, O<c< 1 such that
any transient states Si, s1'
This is a direct consequence of the above proof. It shO"ll's the rate
at which p(n)ij tends to O.
3.1.2
COROLLARY.
p(n)ij:( b ·e n , for
43
FINITE MARKOV CHAINS
44
CH
It is convenient to consider the canonical form of the matrix P
We unite
all the ergodicCHAP.
sets, III
and all the
FINITE
MARKOV
CHAINS
aggregated
version.
sient sets. (Let us say that there are s transient states, an
It is convenient to consider theThe
canonical
formbecomes
of the matrix P in an
form then
ergodic states.)
aggregated version. We unite all the ergodic sets, and all the transient sets. (Let us say that there are s transient states, and r - s
ergodic states.) The form then becomes
44
f-S
S
~
.-"---,
Here again
p= the region 0 consists entirely of 0's. The s x s subma
concerns the process as long as i t stays in transient state
s x (r - s) matrix R concerns the transition from transient t o e
Here again the region 0 consists entirely of O's. The S x s submatrix Q
states, and the (r -s) x (r-s) matrix deals with the process after
concerns the process as long as it stays in transient states, the
reached an ergodic set. From Theorem 3.1.1 we see t h a t t h e p
s x (r - s) matrix R concerns the transition from transient to ergodic
of Q tend to 0. Hence as we raise P t o higher and higher powe
states, and the matrices
(r - s) x (r- s) matrix deals with the process after it has
approach a matrix whose last s columns are all 0. This
reached an ergodic set. From Theorem 3.1.1 we see that the powers
matrix version of Theorem 3.1.1.
of Q tend to O. Hence as we raise P to higher and higher powers,
definition w
Let us now consider a n absorbing chain. By itsthe
matrices approach a matrix whose last s columns are all O. This is the
that S is I(,-,,,(,-,), i.e. a n identity matrix of the appropriate dime
matrix version of Theorem 3.1.l.
Thus its canonical form is
Let us now consider an absorbing chain. By its definition we see
that S is I (r-s)X(r-s), i.e. an identity matrix of the appropriate dimension.
Thus its canonical form is
r-s
s
And byp the
of the powers of P we know that the region I re
= nature
_I_~_)
I . This corresponds
\ R I toQthe fact
s t h a t once a n absorbing state is en
i t cannot be left. From Theorem 3.1.1 we know t h a t the prob
And by the nature of the powers of P we know that the region 1 remains
that such a state is entered, in an absorbing chain, tends to 1.
I. This corresponds to the fact that once an absorbing state is entered,
we may say t h a t with probability 1 the chain will enter a n abs
it cannot be left. From Theorem 3.1.1 we know that the probability
state and stay there, i.e. that i t will be "absorbed."
that such a state is entered, in an absorbing chain, tends to 1. Hence
Let us write some of our examples from Chapter 11, fi 2.2, in t
we may say that with probability 1 the chain will enter an absorbing
canonical form. I n Example 1 the states sl and ss are abso
state and stay there, i.e. that it will be "absorbed."
hence these must be written first. We thus have
Let us write some of our examples from Chapter II, § 2.2, in the new
s 1 S1s5and 85
S2 are
S3 s4
canonical form. In Example 1 the states
absorbing,
1
0
0
0 0
hence these must be written first. We thus have
S1
0 0 0
85
S2 sa 0S4 1
r-s
(_I
0P = 0
81
85
0
82
q
0
P
0
0
0
q
o
O Y O
0
0
q
0
~-----
0
0
p
0
o
p
O P
O q O
0 0 I , 0,
q R,0and
where S3
the regions
P Q have been marked off.
0 P
0 q 0
S4
where the regions I, 0, R, and Q have been marked off.
44
FINITE MARKOV CHAINS
CH
It is convenient to consider the canonical form of the matrix P
We uniteCHAINS
all the ergodic sets, and45all the
aggregated
version. MARKOV
ABSORBING
sient sets. (Let us say that there are s transient states, an
The matrix
for Example
is form
already
canonical form in § 2.2.
The
theninbecomes
ergodic
states.) 10
The first two states are absorbing. Hence R is 4 x 2 and Q is 4 x 4 in
SEC. 2
this example.
Example 9 is not an absorbing chain. It has a single ergodic set,
consisting of the first three states. The matrix appears in canonical
form in § 2.2. If we want to study this process only until it enters the
ergodic set, Here
then again
we may
make 0
theconsists
ergodicentirely
states of
absorbing.
x s subma
the region
0's. The sThe
resulting transition
is
concernsmatrix
the process
as long as i t stays in transient state
s x (r - s) matrix R concerns the transition from transient t o e
0
0
0
0
0
states, and the (r -s) x (r-s) matrix deals with the process after
1 set.
0 From
0
0 Theorem
0
0
3.1.1 we see t h a t t h e p
reached an ergodic
of Q tend to 00. Hence
P0 t o higher and higher powe
1 as we
0
0 raise
0
_ _ _ _I
p= approach a matrix whose last s columns are all 0. This
matrices
0 of0Theorem
0
0 1/2
3.1.1.
matrix version
Let us now
chain. By its definition w
0
1/4
0 a n absorbing
0
1/4 consider
that S is I(,-,,,(,-,), i.e. a n identity matrix of the appropriate dime
1/ 4 0 form is 0 1/2
Thus its canonical
If we do not even care at which state the ergodic set is entered, we may
lump the three ergodic states into a single one, obtaining the much
simpler matrix
And by the nature of the powers of P we know that the region I re
I . This corresponds to the fact t h a t once a n absorbing state is en
i t cannot be left. From Theorem 3.1.1 we know t h a t the prob
that such a state is entered, in an absorbing chain, tends to 1.
we may say t h a t with probability 1 the chain will enter a n abs
state and stay there, i.e. that i t will be "absorbed."
The former matrix
proscn'es
Q and
R, examples
while it modifies
S; the11,
Jatter
Let us
write some
of our
from Chapter
fi 2.2, in t
presen'es onl~'
Q. Thisform.
is in good
agreement
withstates
the interpretation
I n Example
1 the
sl and ss are abso
canonical
gi ven for Q, R,
andthese
S earlier
this
section.first. We thus have
hence
mustin be
written
It should now be clear that absorbing chains serve to answer all
s 1 s5
S2 S3 s4
questiolls of the second type (concerning transient states) raised in
1
0
0
§ ~.5. But absorbing chains are also important in 0
t.he 0study of an
ergodic set. S\1ppose that we wish to ask a0question
happens
1
0about
0 what
0
P 5j
= Thell \\ e lllay wish to "stop" the
as the process goes from SI to
O Y by
O making Sj
process as soon as it reaches s;, which qwe oaccomplish
an absorbing Rtate. Alld since 5j can be 0reached
from
all
0
q o p states of its
equivalence class, thE" resulting chain will be an absorbing Markov
O q O
chain. This trick will be develored in § O
6.1.P
where the regions I , 0,R, and Q have been marked off.
§ 3.2
The fundamental matrix. The following basic theorem is a
direct consequence of the matrix theorem we proved in § 1.11.1, if we
recall that Qk tends to n.
CH
FINITE MARKOV CHAINS
46
3.2.1 THEOREM. For any absorbkng Markov chain, I-&
FINITE
CHAP. III
inverse,
and MARKOV CHAINS
46
3.2.1
THEOREM.
SQ~.
For any absorbing
Markov
(I-Q)-1 =
I + Q + Qchain,
Z + - I-Q
. = has an
inverse, and
k= 0
co
2
3.2.2 DEFINITION.
Markov chain we de$
(l-Q)-l
= 1 +Q+Q2+For...a n= absorbing
QIc.
fundamental matrix to be N = k=O
( I -&)-I.
3.2.2 DEFINITION.
an absorbing
the giving th
W eMarkov
deJine njchain
to be we
the define
function
3.2.3 For
DEFINITION,
fundamental matrix
to be
(l_Q)-l.
number
of N""
times
that the process i s in s f . ( T h i s is defined o
transient state sf.) ukj is dejined as the function that i s 1 if the
3.2.3 DEFINITION. We define Dj to be the function giving the total
is in state sj after k steps, and is 0 otherwise. (See $$ 1.7 and
number of times that the proceS8 i8 in Sj. (This is defined only for
the notation
used in this section.)
Ukj i8 defined as the function that is 1 if the process
transient state Sj.)
now and
giveisa0 probabilistic
interpretation
t o for
N.
is in state Sf We
afterwill
k steps,
otherwise. (See
§§ 1.7 and 1.8
in this
section.) states.
the notation used
the set
of transient
We le
We will now give3.2.4
a probabilistic
interpretation
N. st,We
THEOREM.
{Mi[nj]}= N, to
where
sj Elet
T. T be
the set of transient states.
3.2.4
THEORE;'!.
{JUI[njJ}=N, where St, Sj E T.
<D
PROOF.
It is easily seen that n, =
2: Ukj.
k=O
Hence
{Mi[nj]} =
{Mi L~ Ukj]}
= t~)Mi[U~j]}
= t~ ((1-P(k)if)'O+P(k)w 1)}
=
2:'" {p(k)ii}
k=1J
=
2 Q* since st, sj are transient
e=o
= N by 3.2.1, 3.2.2.
'" Qk since Si,
= 2:
Sj are transient
This completes
the proof.
k=U
= N
by
3.2.1, 3.2.2.
This theorem establishes
the fact t h a t the mean of the total
of times
This completes the
proof.the process is in a given transient state is always fin
that
these means are simply given by N.
This theorem establishes the fact that the mean. of the total number
There is a n interesting alternative proof for this result. To c
of times the process is in a given transient state is always finite, and
Mi[nj], we may add up the original position's contribution, pl
that, these means are simply given by N.
of the steps' contribution. The original position contributes
There is an interesting alternative proof for this result. To compute
only if i =j. I t is convenient to define dij, the constant function
I1I;[nj], we may add up the original position's contribution, plus each
1 if i = j , 0 otherwise. Then we can say t h a t the original
of the steps' contribution. The original position contributes 1 if and
only if i = j. It is convenient to define d;j, the constant fUllction that is
1 if i = j, 0 otherwise. Then we can say that the original position
SEC. 2
CH
FINITE MARKOV CHAINS
46
3.2.1 THEOREM. For any absorbkng Markov chain, I-&
inverse, and
SQ~.
ABSORBING YLL\.RKOV CHAINS
47
(I-Q)-1 ------------~--------------= I+Q+QZ+ - . =
contributes d ij • After one step we move to Sk with probability
Pilc.
k= 0
If the new state
is
absorbing,
it
contributes
nothing
to
our
mean,
butwe de$
3.2.2 DEFINITION. For a n absorbing Markov chain
if it is transient,fundamental
then it contributes
MIc[nj].
Hence
we
have
matrix to be N = ( I -&)-I.
L
i [ 111] = d ij + W ePikMk[
deJinen))
nj to be the function giving th
3.2.3 M
DEFINITION,
number of times that SkE
theT process i s in s f . ( T h i s is defined o
transient state sf.) ukj is dejined as the function that i s 1 if the
Hence
is in state sj after k steps, and is 0 otherwise. (See $$ 1.7 and
{Mi[nj]}used
= (1in_Q)-l
= N.
this section.)
the notation
\Ve will apply
thesenow
results
examples of
the last section.
We will
givetoa the
probabilistic
interpretation
t o N. InWe le
the random walk,
Example
1
of
§
2.2,
the set of transient states.
3.2.4
-:\IJ
I {Mi[nj]}
1
-P = N, where st, sj E T.
THEOREM.
1
(1 -Q) = { -q
\
0
-q
and hence
52
S3
S4
P
p2
p2+q2
p2+q2
1
p2+q2
q
p2+q2
P
p2+q2
q+p 2
p2+q2
[Since p + q = 1, and hence (p + q)2 = 1, we have that 1- 2pq = p2+ q2.]
= starts
Q* since
st, sjmiddle
are transient
We see that, for example, if the process
in S3 (the
state),
e = o of 1/(p2 + q2) times. This
then it will be in the middle state an a\'erage
N minimum
by 3.2.1, of
3.2.2.
quantity is always between 1 and 2. =The
1 is achieved if
p = 0 or 1, theThis
maximum
of 2 the
if p proof.
=
In the former case the process
completes
starts at S3 and goes directly to one of the boundaries, hence it will be
This theorem establishes the fact t h a t the mean of the total
in state 83 only at the beginning. But even in the case p = 1/z we
of times the process is in a given transient state is always fin
expect the process to return on:y once on the average.
that these means are simply given by N.
There
n interesting
alternative
proof
this result. To c
3.2.5 EXAMPLE
laois aAs
an illustration
we give
theforfundamental
we2/may
add
up
the
original
position's
contribution,
pl
,
matrix for theMi[nj],
case p=
i.e.
when
it
is
twice
as
hkely
to
move
to the
3
the steps' contribution. The original position contributes
right as to theofleft.
only if i =j. I t is convenient
to define dij, the constant function
82
S3
84
1 if i = j , 0 otherwise.
Then
we can say t h a t the original
6/ 5
82
2
(/5
i5
N=S3\3/5 9! 5 "6/ 5
S4
1/5
)
7/5
.
F I N I T E MARKOV CHAINS
48
CH
I n the college example, Example 10 of 5 2.2, rememberin
48
CHAP. III
CHAINS
p + q + rFINITE
= 1, welVIARKOV
have
In the college example, Example 10 of § 2.2, remembering that
p+q+r= 1, we have
(I -Q) =
C'
~
T
53
1
53
54
N = (1 _Q)-l
P+1'
r
(p + r)2
1"2
55
(1'+r)3
0
° °0
p+r
-r
p+r
o)
0
-r
1':r
54
85
S6
°
0
0
SENIOR
0
0
JUNIOR
°
SOPHOMORE
p-t- r
r
(1'+r)2
1
1'+r
r2
The
zeros r3
in N indicate thatr no one 1is demoted
in the college.
FRESHMAN
S6
(1'+r)4
(1'+r)3
(1' +spend
r)2 p+r
for example, a junior
cannot
any time as a sophomore o
man in the future. As an illustration we compute N (approx
The zeros in N indicate that no one is demoted in the college. Thus,
for the case we will call Example 10a, where the probabilities of
for example, a junior cannot spend any time as a sophomore or freshout,
repeating, and being promoted are p = .2, q = .1, r = .7, respe
man in the future. As an illustration we compute N (approximately)
for the case we will call Example lOa, where the probabilities of flunking
out, repeating, and being promoted are l' = .2, q= .1, r= .7, respectively.
N=
0
0
.86
1
1.11
0
.52
.67
.86
0
~w.
o ) JUNIOR
.67
.86 1.11
o SOPHOMORE
In the urn example, Example 9 of 5 2.2, we have
C
1.11
FRESHMAN
In the urn example, Example \) of § 2.2, we have
(2,0,0)
(1,1,0).
(1,0,1)
F I N I T E MARKOV CHAINS
48
SEC. 3
CH
I n the college example, Example 10 of 5 2.2, rememberin
49
MARKOV CHAINS
p + q ABSORBING
+ r = 1, we have
If the process reaches (1,1,0) or (1,0,1), then from then on it is expected
to be in that state 4/3 times, and in the other state 2/ 3 times. (The 4/ 3
includes the original position.) From neither of these states can the
state (2,0,0) be reached, since a painted ball always remains painted.
If the process starts in (2,0,0), which is its natural starting position, it
will be in this position only once. It is expected to be in each of the
other two states once, which is the average of 4/ 3and 2/3.
These fimdamental matrices 'will be used throughout this chapter
for illustrations.
§ 3.3 Applications of the fundamental matrix. We will show that a
number of interesting quantities can be expressed in terms of the
fundamental matrix. These results will here be jllustrated in terms of
thc random walk Example la (see § 3.2.5), and all the absorbing chains
will be worked out in the next section.
3.3.1
DEFINITION.
We define' the following new matrices and vectors:
s X s matrix
s x (r-s) matrix
The zeros in N indicate
that no one
is demoted
T = Ng
s component
column
vector in the college.
for example,
spend
any time
T2 = (2N
-1)7-7,0.a junior scannot
component
column
vectoras a sophomore o
man in the future. As an illustration we compute N (approx
3.3.2 THEOREM.
QN =we
NQwill
= Ncall
- I.Example 10a, where the probabilities of
for the case
out,3.2.1,
repeating,
PROOF. From
3.2.2, and being promoted are p = .2, q = .1, r = .7, respe
N2 = N(2Ndg-I)-NsQ
B = NR
N = I +C) +Q2.t
Hence,
which is the original series without 1.
3.3.3
THEOREM.
{Varj[nj]}=N 2 , where Sj, Sf E T.
In the urn example, Example 9 of 5 2.2, we have
PROOF. We recall that Varj[nJJ=M j [n 2j]-M1[njJ2.
3.2.4 we see that
{Mt[n:_
From Theorem
= N sQ ,
hence we need only show that
{Mt[n 2j]} = N( 2N dg-1).
We will assume that these means are finite. A proof of this fact will
be given at the end of this section. To compute these means we again
ask where the process can go in one step, from its starting position Sj.
It can go to 5); with probability PC{. If the new state is absorbing, then
we can never reach Sf again, and :he only possible contribution is from
the initial state, which is d ij . If the new state is transient, we will be
F I N I T E MARKOV CHAINS
50
in sf dtjtimes from the original position, and n , times from the la
MARKOV
CHAP. III
drl is a constant function
and dij =
Nence,FINITE
remembering
that CHAINS
50
in Sj dtj times from the original positio~, and nj times from the later steps.
Hence, remembering that d'j is a constant function and dis = d 2 jj,
{M,[n2j]} =
{2:. Pt"d 1f + 2: Pt"Mk[(nj+dts)2J}
2
8keT
=
{2:
'A:eT
= Q{W[n2jl}
+ 2(QN)dg+I.
Plk(M,,[n21J + 2M,,[ns]' dis) + diS}
Hence 8,tET
2jJ}+2(QN)dg+
1.
= Q{Mt[n{Mt[n2,1)
= ( I -Q)-1(2(QN)dg+I)
= N ( 2 ( N - I ) d g+ I ) = N ( 2 N d g - 1).
Hence
2S]} = (I -Q)-1(2{QN)dg+1)
{M,[n
The
matrix (QN)dgappeared above since the factor daj has
of setting
all elements
off the
main diagonal equal to 0.
= N{2{N
-1)dg+1)
= N{2Ndg-1).
In our Example l a , we have already computed N , and
The matrix (QN)dg appeared above since the factor d'j has the effect
N z as
well.diagonal equal to O.
compute
of setting all elements
off'the
main
In our Example la, we have already computed N, and we now
compute N 2 as well.
N =
C "')
(':' ,D
("'" ",,, "'' )
6/ 5
3/5
9/ 5
1/5
3/5 7/5
C. "'' )
36/ 25
6/ 5
Nsq =
0
N dg =
9/s
9/ 25
81/25
36/ 25
1/25
9/ 25
49/ 25
c .D
i"
0
2Ndg-1 =
0
:
13/ 5
s2
s3
0
S2
"
('"
S3
42/ 25
84
"'l
27/ 25 117/ 25 54/ 25
N2 =' sa, 18/ 25 36/25 18/ 25 .
Thus we see that for any state as initial state the variance is
the
N 2 is14/
quite
9/ 25middle
39/25 state.
63/ 25 We alsoS4note8/ that
25 large com
3°/25
25
N,,;
hence the means are fairly unreliable estimates for th
Thus we see that for any state as initial state the variance is largest for
chain. This will often be the case.
the middle state. We also note that N2 is quite large compared to
Let testimates
be the function
the numb
3.3.4areDEFINITION.
Nsq; hence the means
fairly unreliable
for thisgiving
Markov
(including
original position) in which the process is in
chain. This will often
be thethe
case.
state.
3.3.4 DEFINITION. Let t be the function giving the number of steps
If the position)
process starts
in an
t = 0. If t
in which
theergodic
process state,
is in athen
transient"
(including the original
starts in a transient state, then t gives the total number of ste
state.
to reach an ergodic set. In an absorbing chain this is th
If the process starts in an ergodic state, then t=o. If the process
absorption.
starts in a transient state, then t gives the total number of steps needed
to reach an ergodic set. In 'an absorbing chain this is the time to
absorption.
N{2Ndg-1) =
F I N I T E MARKOV CHAINS
50
SEC. 3
3.3.5
PROOF.
in sf dtjtimes from the original position, and n , times from the la
ABSORBING
MARKOV
CHAINS
51 dij =
is a constant
Nence,
remembering
that drl
- - function and
THEOREM.
{Mi[t)}=T; {Vari[tJ}=TZ, where Si.E T.
It is easily seen that t =
) ' nj.
sjET
Hence
Hence
= Q{W[n2jl}
+ 2(QN)dg+I.
{Mi[t]} = ~',2 Mi[ntf}
.S,-E T
= Nf, = ( I -Q)-1(2(QN)dg+I)
{Mt[n2,1)
since this gives the row sums of N. = N ( 2 ( N - I ) d g+ I ) = N ( 2 N d g - 1).
For the variance
carry (QN)dg
out an appeared
argument above
similarsince
to that
§3.3.3,daj has
Thewe
matrix
theinfactor
but here the firstofstep
always
counts. off the main diagonal equal to 0.
setting
all elements
In our Example l a , we have already computed N , and
{Mi[t2]} = {
1 + )' p,d'l:h[(t+ 1 )2J}'
well.
compute NSk zE TasPik"
Sk E T
.L _
=
{L PikPli;[t2J + 2lHk[t]) + I}
sk ET
= Q{Mi[t 2]} -'- 2Q.,-+t.
Hence
{Mi[t 2 ]} = (1-Q)-1(2QT+g)
= 2NQT+Ng
= 2(N -1)1'+1'
= (2N -J)T.
Thus
{Var,Et]} == {Mi[t2]-M;Ct]2} = (2N-1)T-1'8q.
s2
s3
In' our example,
Thus we see that for any state as initial state the variance is
the middle state. We also note that N 2 is quite large com
N,,; hence the means are fairly unreliable estimates for th
(2N This
-1) will
=
chain.
often be the case.
3.3.4 DEFINITION.Let t be the function giving the numb
(including the original position) in which the process is in
state.
(:::;::)
If the process starts in7Sq
an =ergodic
state, then t = 0. If t
starts in a transient state, then t gives
121/ 25the total number of ste
to reach an ergodic set. In an absorbing chain this is th
absorption.
F I N I T E MARKOV CHAINS
52
C
We see that one expects to reach the boundary most quic
FINITE
CHAP. the
III bound
This
is not MARKOV
surprising,CHAINS
since i t is easier to reach
an outside state than from the middle, and i t is more probable
We see that one expects to reach'the boundary most quickly from
process moves to the right. But we again note that the va
S4. This is not surprising, since it is easier to reach the boundary from.
an outside state sizable.
than from the middle, and it is more probable that the
only that
for measures
in which
the
process moves to We
the have
right.computed
But we means
again note
the variance
is
starts in a given state st. But i t is easy to obtain from this th
sizable.
and variances for an arbitrary initial probability vector.
We have computed means only for measures in which the process
starts in a given state
But it is easyIftonobtain
this the
means vecto
3.3.6Sf. COROLLARY.
is thefrom
initial
probability
and variances for an
arbitrary
initial
probability
absorbing
chain,
and
n' consistsvector.
of the last s components of
gives the initial probabilities for the transient states, then
3.3.6 CoROLLARY. If TT is the initial probability vector for an
52
sq.
absorbing cha,in, and TT' CDnsist.s of the last .~ components of TT, i.e. TT'
gives the initial probabilities for the transient states, then
{M,,[n,]} = TT'N
{Va.r,,[n,]} = TT'N(2Ndg-1)- (TT'N)sq
{M..[t]} This
=TT'Tis an immediate consequence of the fact tha
PROOF.
{Var,,[tJ}
= TT'(2N-I)T-(TT'T)Sq.
function
f, M,[f]=nMf[f],
which follows from the nature of
measure. The right sides contain n' rather than n,since th
PROOF. This is an immediate consequence of the fact that for any
means are 0 if the initial state is absorbing.
function f, M,,[f] = TTM,[f], which follows from the nature of the tree
Our remaining applications will concern the question
measure. The right sides contain TT' rather than TT, since the various
absorbing state is likely to capture the process.
means are 0 if the initial state is absorbing.
Our remaining 3.3.7
a.pplications
will If
concern
theprobability
question that
of which
THEOREM.
bij is the
the process s
absorbing state is likely
to capture
the process.
transient
state st ends
u p in absorbing state sf, then
3.3.7 THEOREM. If bij is the probability that the process starting in
transient state St ends up in absorbing state Sf. then
PROOF.
Starting in si, the process may be captured in sr
{b ij} =steps.
B = NR,
SjE T, of capture
S, E T. on a single step is pij
more
The probability
does not happen, the process may move either to another a
PROOF. Starting in Sl, the process may be captured in Sf in one or
state (in which case i t is impossible to reach sj), or to a transi
more steps. The probability of capture on a single step is Pij. If this
I n process
the latter
case
thereeither
is probability
of being captur
sk. the
does not happen,
may
move
to anotherbkjabsorbing
right state. Hence we have
state (in which case it is impossible to reach si), or to a transient state
Sk.
In the latter case there is. probability bkj of being captured in the
right state. Hence we have
which can be written in matrix form as
blj = Pi! +
2: Plkbtf,
SteT
B = R+&B.
whjch can be written
in
matrix
form
as
Thus
B = ( I - Q ) - l R = NR.
B = R+QB.
Thus
An alternative proof is based on the following observatio
B = (I -Q)-lR = NR.
An alternative proof is based on the following observation: Every
F I N I T E MARKOV CHAINS
52
C
We see that one expects to reach the boundary most quic
ABSORBING
:\L<\.RKOV
CHAINS
This
is not surprising,
since
i t is easier to reach the53bound
an outside state than from the middle, and i t is more probable
time that the process
process moves
is in transient
state Sk,
it has
probability
Pk! of
But
we again
note that
the va
to the right.
going to sf. Hence it is possible to show that
sizable.
We havebijcomputed
meansPic!.
only for measures in which the
= ') Mj[nk]'
"-'
starts in a given s,i:ET
state
st. But i t is easy to obtain from this th
and variances
for an arbitrary initial probability vector.
This gives directly
that
SEC.
3
sq.
3.3.6 COROLLARY.
If n is the initial probability vecto
B = NR.
absorbing chain, and n' consists of the last s components of
In our example
gives the initial probabilities for the transient states, then
R =
1/3
( 0
82 (7/15
0 )
0
B = NR = 83
1/5
S4
\ 1/15
o 2/31
It is worth noting that for each starting state the sum of the two absorption
PROOF.
This is an immediate consequence of the fact tha
probabilities is 1. By Theorem 3.1.1 it will always be true that NRtr-s =
function f, M,[f]=nMf[f], which follows from the nature of
ts. It is also easy
to verify this directly.
measure. The right sides contain n' rather than n,since th
The further to the right we start, the more probable it is, of course,
means are 0 if the initial state is absorbing.
that the process will end up at the right end. It is interesting to see
Our remaining applications will concern the question
that even in the leftmost transient state the probability is somewhat
absorbing state is likely to capture the process.
greater for capture on the right.
3.3.7 THEOREM.If bij is the probability that the process s
~.3.8 COROLLARY. If p" is the a-th column of R, i.e. pa = pia for
transient state st ends u p in absorbing state sf, then
Si in T and fOT fixed a, then S p" gives the probabilities of absorption in
the g1:ven absorbing state Sa, for any transient state as initial state.
This corollary PROOF.
is useful if
we are interested
in a single
absorbing
state. in sr
Starting
in si, the process
may
be captured
more steps.
probability
of whose
capture
on ab*jj
single
If B* The
is the
r x r matrix
entry
givesstep
the is pij
3.3.9 THEOREM.
does
not
happen,
the
process
may
move
either
to
another
a
probability of being absorbed l:n 5;, starting in si,for all states Si and Sj,
state (in which case i t is impossible to reach sj), or to a transi
then
case= there
sk. I n the latter
PB*
B*. is probability bkj of being captur
right state. Hence we have
If Sf E T, then b*i1 = O. Hence the last s columns of B* are O.
Consider Sj absorbing. If Si E T, ,.,hen b*;j=Oij, as in § 3.3.7. If Si is
also absorbing, then b*jj=d ij . Hence we have
PROOF.
which can be written in matrix form as
Thus
B* =
I I ! 0 ) B = R+&B.
\Bi~o-
B = ( I - Q ) - l R = NR.
An alternative proof is based on the following observatio
But R+QB=R+QNR=R+
F I N I T E MARKOV CHAINS
54
Hence
PB* =. BX.
FINITE
MARKOV
CHAINS
We thus see that the r-component column CHAP.
vectorIIIgiving t
bilities
of
absorption
in
an
absorbing
state
sj
is
a fixed vector
PB* =.B*.
Hence
its first r - s components are 0, except the j-th, which is 1. T
We thus see that
thethe
r-component
column
vector
givingthe
theabsorption
probamines
vector. This
method
of finding
pro
bilities of absorption
in an
absorbing
Sj is ain
fixed
vector
is useful
if we
are not state
interested
finding
N. of P, and
its first r - s components
0, except
j-th,verified
which isthat
1. This deterIn ourare
example
it isthe
easily
mines the vector. This method of finding the absorption probabilities
is useful if we are not interested in finding N.
In our example it is easily verified that
54
are fixed vectors of P.
We will now supply the missing step for Theorem 3.3.3.
are fixed vectors of P.
We will now supply the missing step for Theorem 3.3.3.
3.3.10
THEOREM.
Mi[n 2 j] i8 finite for any absorbing chain, and any
PROOF.
Si, 5j E T.
PROOF.
Mi[n 2 j] =
M{C~ UkJ)
= M{~ ,~ UkjUli]
m
m
k-0 1x0
M t [ u k j u l j ] is the probability that the process is in sj both on s
= in st. IfMi[UkjUlj].
on 1, starting
we let m = min(k, I ) , d = Ik - 11, then t
k=O 1-0
probability of being in sj after m steps, and of returning d st
Mi[UkjUl j ] is the Hence
probability
M t [ u kthat
j u l i ] =the
p ( mprocess
) i j p ( d ) jisj . in Sj both on step k and
2.: 2.:
"'
<>0
on I, starting in Sj. If we let m =min(k, I), d = Ik -II, then this is the
probability of being in Sj after m steps, and of returning d steps later.
Hence M i [ Ukjulf] = p(m)jjp(d) ji'
00
00
k~O
1=0
2: 2:
p(m)jjp(d)jj
= bz 2 2 cn where n = max ( l , 1 )
: : : 2: 2.: (b·cm)(b.e
ro
a)
d)
k=O l = U
k=O 1=0
co
= b2
ro
k=O 1=0
<X)
= b2
m
2.: 2.: cn = where
= max
(k, I)which is finite.
b2 2 n (2121
1)cn,
n=O
2: (2n+ 1len, which is finite.
110=0
F I N I T E MARKOV CHAINS
54
Hence
PB* =. BX.
CHAINS column vector giving
55
WeABSORBING
thus see thatlI-LARKOV
the r-component
t
bilities of absorption in an absorbing state sj is a fixed vector
§ 3.4 Examples
its first r - s components are 0, except the j-th, which is 1. T
mines(Example
the vector.1 ;)f
This
method
of finding
EXAMPLE 3.4.1
§ 2.2
continued).
Inthe
theabsorption
random pro
walk we find:is useful if we are not interested in finding N.
In our example itP+Q2
is easily
P verified that
SEC. 4
N =
p2~q2 (
q
q2
1
q
1+2p
2
1 + 2q
are fixed vectors of P.
We will now supply the missing step for Theorem 3.3.3.
1+2P )
4pq
(
TZ = (p2 + q2)2
2
1 +2q
PROOF.
p3
)
pZ
pq+p3
In particular, if p = 1/2 (Example 1b), then
':')
m
m
2
'
;
'
)
C c:
(:) C'I,)
~
k-0 1x0
[ u k j u l1j ] is2the probability
N 2 =that the 2process is in sj both on s
NM t=
on 1, starting in st. If we let m = min(k, I ) , d = Ik - 11, then t
1/ 2 1of being
2 3/ 4
3/21 in sj after m steps,
and of returning d st
probability
Hence M t [ u k j u l i ] = p ( m ) i j p ( d ) j j .
T
(;)
B=
"
\8
1/2 1/ 2 .
1/4 3( 4
And if p= 1 (Example Ie), then
= bz
22
cn
where n = max ( l , 1 )
k=O l = U
N
m
= b2
2 (2121 1)cn, which is finite.
n=O
and the variances are all O.
This last case is easily interpreted if we remember that the process
in this case must move to the right.
F I N I T E MARKOV CHAINS
56
EXAMPLE3.4.2 (Example 10 of. $ 2.2 continued). In th
FINITE
MARKOV
CHAP. III
process
we have,
lettingCHAINS
t =
p+r'
r
56
EXAMPLE 3.4.2 (Example 10 of. § 2.2 continued).
In the college
process we have, letting t = _r_:
p+r
1
N __
1_ ( t
- (p+r)
t2
t3
:, ~ f)
o
qt+:-tz
1
N2 = (p+r)2 ( qt2 + t2_t4
qt3+t3_t6
(
P
q
o
o
qt + t - t2
q
qt2+t2_t4
qt + t - t 2
~)
l_t)
7'=~·1-t2
i_t3
1_t4
1
"7"2
= p(p+r)
(
q(l-t)
q(l_t2)+t-2t2+t3
q(1-t3)+t+t2-4P+t4+t.
)
q( 1- t 4) + t + t 2 + t 3 - 6t 4 + t 5 + t 6 + t 7
B=
( ~=:2 :2)
.
I - tof3 graduating
ta
The probability
from each class depends o
r
I-t4
t4
ratio t = -.
This ratio is the conditional probability tha
p+r
The probabilityis of graduating
fromthan
eachflunked
class depends
only
on he
theleaves h
rather
out, given
that
clilss.ratio
Having
of this
ratio
ratio t = _r_. This
is the successive
conditionalpowers
probability
that
the can
manbe inte
p+r saying t h a t each time he leaves his class he must be promo
is promoted rather
than
flunked
leaves
hislong
present
than
flunked
out,out,
but given
i t doesthat
not he
matter
how
he stays in h
class. Having successive
powers
of this
ratio greatly
can be ifinterpreted
as the po
class. The
formulas
simplify
we eliminate
saying that each having
time he aleaves
class he
be promoted
man his
repeat
themust
class,
that is ifrather
q = O . In
than flunked out, but it does not matter how long he stays ill his present
class. The formulas simplify greatly if we eliminate the possibility of
having a man repeat the class, that is if q = O. In that case,
r
t = P + r = r, and
F I N I T E MARKOV CHAINS
56
EXAMPLE3.4.2 (Example 10 of. $ 2.2 continued). In th
r
ABSORBING
process
we have,MARKOV
letting t =CHAINS
p+r'
SEC. 4
(, ~)
(' ) ('
0
N=
0
1
r3
0
r
r2
N2 =
1 +r+r2
0
0
0
~)
r-r2
1
= pr
1"2
0
0
r3_ r 6 r2- r 4 r-r2
r
l+r
T=
( r-r2
0
r2- r 4
57
1 + 3r+r2
1 +r +r2 + r3
)
.
1 + 3r + 6r 2 + 3r3 + r 4
B is unchanged.
In the numerical Example lOa (cf. § 3.2.5) we have:
1.11
o
.86
.67
1.11
o
o
.86
1.11
.52
.67
.86
.12
o
o
( .31
.37
.12
o
.31
.12
.37
.37
.31
N = (
Nz =
T
=
.12)
(
=
(~:~~\
.43
The probability
of graduating
72
1.13from each class depends o
2.65 )
r
ratio t = 3.17)
-.
This ratio is the
2.22conditional probability tha
p+r
is
rather than flunked out, given that he leaves h
FLUNK
GRADUATE
clilss. Having successive powers of this ratio can be inte
OUT
saying t h a t each time he leaves his class he must be promo
than flunked out, but i t doesSENIOR
not matter how long he stays in h
class. The formulas .60
simplify
greatly if we eliminate the po
JUNIOR
B- a man repeat the class, that is if q = O . I n
having
( :~~
.78)
.53
.4 7
SOPHOMORE
.63
.37
FRESHMAN
Thus a student must reach the junior year before he has a better than
even chance of graduating.
F I N I T E MARKOV CHAINS
58
CH
EXAMPLE
3.4.3 (Example 9 of 5 2.2). In the urn example th
CHAP. III
MARKOV
vectorsFINITE
and matrices
are : CHAINS
58
EXAMPLE 3.4.3 (Example 9 of § 2.2).
vectors and matrices are:
N
~ 4/ 3 ,;,)
(:
2/a 4/3
r=
N,
In the urn example the five
~
(:
')
1/3 -/3
4/9 2/ 3
2ja
4/9
G) r, ~ (;)
Sa
Since the process must leave s4 immediately and cannot return, t
0 variance for the number of times in this state. Of the rem
variances the diagonal elements are smallest-this
is due
Since the process must leave S4 immediately and cannot return, there is
stabilizing effect of having to count the original position.
o variance for theThe
number of times in this state. Of the remaining
B matrix needs special interpretation in this case. Sin
variances the states
diagonal
smallest-this
the proce
sl, elements
sz, and s3are
were
not absorbingis indue
thetooriginal
stabilizing effect of having to count the original position.
"absorption probabilities" must be interpreted as probabilit
The B matrix needs special interpretation in this case. Since the
entering the ergodic set a t the given state. Thus, for example
states 81, 82, and S3 were not absorbing in the original process, the
process starts with both balls unpainted (state s4), then there
"absorption probabilities" must be interpreted as probabilities of
bability
that the first time both balls are painted there will
entering the ergodic
setcolor,
at the given
state.willThus,
example,
if that
the they wi
of each
that they
both for
be red,
and 114
process starts with
both
balls
unpainted
(state
54), then there is probe black. It should be noted that these probabilities are the
bability 1/2 that the first time both balls are painted there will be one
as if we had assigned the two balls colors independently a
of each color, 1/4 that they will both be red, and 1/ 4 that they will both
random.
be black. It should be noted that these probabilities are the same
as if we had assigned
the two' balls
colors We
independently
andresults
at
5 3.5 Extension
of results.
will see that
obtai
random.
5 3.3 can be applied to a wider variety of problems.
§ 3.5 Extension
of results.
We will
see S that
results
in from eve
3.5.1
DEFINITION.
A set
of states
is a nobtained
open set if
§ 3.3 can be applied
variety
i n Stoit aiswider
possible
to go toofaproblems.
state i n S .
3.5.1 DEFINITION.
A set
S of staies
is an open
set ifsets
from: Aevery
slale
It is easy
to think
of examples
of open
set consisting
of a
in S it is possible
to
go
to
a
state
in
S.
state is open (unless the state is absorbing), so is a set of tra
states, so is a proper subset of an ergodic set, etc. The fol
It is easy to think of examples of open sets: A set consisting of a single
theorem characterizes these sets.
state is open (unless the state is absorbing), so is a set of transient
states, so is a proper subset of an ergodic set, etc. The following
theorem characterizes these sets.
58
SEC. 5
F I N I T E MARKOV CHAINS
CH
EXAMPLE
3.4.3 (Example 9 of 5 2.2). In the urn example th
vectorsABSORBING
and matrices MARKOV
are :
CHAINS
59
3.5.2 THEOREM.
is a subset of S.
A set S of states is open if and only if no ergodic set
PROOF. If an ergodic set is contained in S, then there is no escape
from this set once it is entered; hence S is not open.
On the other hand we know that from every state we can reach an
ergodic state. And from an ergodic state we can reach all the elements
of its ergodic set. Hence if there is no ergodic set contained in S,
then for every element of S we can find an ergodic state in 8 which can
be reached from the given state. Hence S is open.
3.5.3 THEOREM. If S i8 an open set of states, and all the states in S
are made ab80rbing states, then the resulting Markov chain is absorbing,
and its transient states are the elements of S.
PROOF. Since S is open, from every state of it we can reach a state
in 8-which must be an absorbing state. Hence the chain is absorbing.
Sinceeach
the process
leave
immediately
and cannot
t
And since from
elementmust
of S we
cans4reach
an absorbing
state, return,
the
fortransient
the number
in process.
this state. Of the rem
elements of0S variance
must all be
statesofintimes
the new
variances the diagonal elements are smallest-this
is due
3.5.4 THEOREM.
S of
be having
an open
of sthestates.
Q be the
stabilizing Let
effect
to set
count
originalLetposition.
s x s submatrix
P corresponding
to these
states. Letinpathis
be the
8- Sin
The of
B matrix
needs special
interpretation
case.
componentstates
column
vector
components
pia, where
the Sj
are theproce
sl, sz,
andwith
s3 were
not absorbing
in the
original
elements of
S and Sa E S.
Lee :he process
start
Sj.
Then: as probabilit
"absorption
probabilities"
must
beininterpreted
entering the ergodic set a t the given state. Thus, for example
(1) The ij-component of N = (1 _Q)-l is the mean number of times
process starts with both balls unpainted (state s4), then there
the process is in Sf beJore leaving S.
bability
that the first time both balls are painted there will
(2) The
N2=N(2Nag-I)-Nsq
i8 and
the variance
of
of ij-components
each color, oj
that
they will both be red,
114 that they wi
thebe
same
function.
black. It should be noted that these probabilities are the
(3) The
of T = 111 ~the
is the
mean
number
of steps
needed
asi-th
if component
we had assigned
two
balls
colors
independently
a
to random.
leave S.
(4) The i-th component of Tz=(2N -I)T-TsQ is the variance of the
3.5 Extension of results. We will see that results obtai
same5function.
5 3.3
be applied
variety ofthat
problems.
(5) The
i-thcan
component
of Nto
po, aiswider
the probability
the proceS8 goe8
to So, when it leaves S.
3.5.1 DEFINITION.A set S of states is a n open set if from eve
PROOF. The
parts of
this
i n Svarious
it is possible
to go
to theorem
a state i n are
S . a direct consequence
of the corresponding results in § 3.3, due to Theorem 3.5.3.
It is easy to think of examples of open sets : A set consisting of a
As an application
of this
theorem
consider
the following
state is open
(unless
the state
is absorbing),
so isproblem.
a set of tra
Let 81 and states,
S/c be any two states in a regular Markov chain.
Assume
so is a proper subset of an ergodic set, etc.
The fol
that the process
is started
at a third
theorem
characterizes
thesestate.
sets. What is the probability
of reaching 8/c before 811 This probability may be found from
3.5.4(5) by choosing S to be the set of all states in the ergodic set except
Sj and Sic.
F I N I T E MARKOV C H A I N S
60
60
C
3.5.5 EXAMPLE.Consider the random walk Example 6 of
II. The
transition
matrixCHAINS
is
FINITE
MARKOV
CHAP. III
0 lExample
'14
/4
'!4
3.5.5 EXAMPLE. Consider the random walk
6 of '14
Chapter
0
II. The transition matrix is
'13
0
l/3
l/3
"C ')
1/4
S2
1/3
1/3
1/4
0
1!4
l/3
9'3
113
0
1/ 3
00
0
l/3
I13
I13
'14
0
0 1/ 3 1/3 1/'14
3 o
' ,1 4
l/4
0
0
1/3
84
l/S
l/S
and since from any state we can move to any other state in tw
P = 8a
the Markov
chain
is regular.
any proper subset of the
0
1/4 1/4
1/4 1/4 Hence
85
open. Let S consist of the last three states.
and since from any state we can move to any other state in two steps,
the Markov chain is regular. Hence any proper subset of the states is
open. Let S consist of the last three states.
"C ,~.)
1/ 3
Q = s"
1/3
1/3
S5
1/4
1/4
C· "')
12/ 9
N =
N2 =
15/ 9
24/ 9
8/9
9/9
9/9
12/ 9
("""
27°/ 81
216/ 81
T =
C')
47/9
324/ 81
36°/ 81
27°/81
T2 =
3°/9
"'' )
66/ 81
36/ 81
C"")
1114/ 81
1062/ 81
C)
NPI = 2/ 9 .
The N matrix tells us the mean number of times that the.pro
3/ 9 before it goes to one of the first tw
each of the last three .states,
We see that the numbers are small if the process starts in the la
The N matrix tells us the mean number of times that the 'process is in
But this is intuitively clear, since in this case it has a 112 pro
each of the last of
three
states, before
goesstep.
to oneFor
of the
states.
"escaping"
on theitfirst
the first
sametwo
reason,
the mean
We see that theofnumbers
are
small
if
the
process
starts
in
the
last
state.
times that it is in the last state is small, no matter
where the
But this is intuitively
sincefrom
in this
it has
1/2former
probability'
starts. clear,
However,
Nzcase
we see
thatathe
numbers ha
of "escaping" ongreater
the first
step.
For
the
same
reason,
the
mean
number
variances than the latter.
of times that it is in the last state is small, no matter where the process
starts. However, from N 2 we see that the former numbers have much
greater variances than the latter.
60
SEC. 5
F I N I T E MARKOV C H A I N S
C
3.5.5 EXAMPLE.Consider the random walk Example 6 of
II. The
transition :.'IJ:ARKOV
matrix is CHAINS
ABSORBIKG
61
'14
l / 4 from
'!4 which
'14 has no
From T we see that it takes lO<lgest to 0
escape
S4,
0 steps to
' 1number
3
0
connection to S. Indeed, the differencesl / 3in mean
of
l/3
escape can be accounted for by the number
0 lof
/ 3 connections
113
0the three
9'3
states have with outside states. Note 0that while the means differ
0 l / 3 I13 I13
considerably, the variances are roughly the same.
'14 for
0 state 81,
'14
'14
l
/4
Finally, the vector NPI gives us the "exit probabilities"
i.e. the probabilities
(depending
on
starting
state)
of
going
to
81 state
when in tw
and since from any state we can move to any other
the process leaves
S; or, chain
statedisotherwise,
probability
of hitting
the Markov
regular. the
Hence
any proper
subset Sl
of the
before hittingopen.
82.
These
to depend
Let S probabilities
consist of theseem
last three
states.very simply on
the number of steps necessary to reach Sl from the starting state
(going through S).
3.5.6 THEOREM. Let fi be the function giving the number of times that
the process remains in the non-absorbing state Si once the state is entered
(including the entering step). Then
(a)
(b)
And the conditional probability of the process going to Sj, given that it
leaves St, is
(c)
PROOF. The set whose only element is Si is an open set. We apply
Theorem 3.5.4 to this set. In this case N is a 1 x 1 matrix, and hence
identical with T; its only component is 1/(1- Pill. Hence (a) is a
consequence of either (1) or (3) of Theorem 3.5.4. Similarly, N 2 =T2,
and (b) is a consequence of either (2) or (4) of Theorem 3.5.4. We
obtain (c) from 3.5.4(5) by choosing the vector Pi whose only component is Pij. Since Si is not absorbing, Pii < 1, hence our quantities
are well defined.
One type of concept th"t we ha ve not investigated as yet is illustrated
matrix the
tellsprocess
us the mean
numbera ofgiven
timestransient
that the.pro
The
by the question
of N
whether
ever enters
the last
three states,
before up
it goes
to one of the
first tw
state. This each
and of
related
questiOY1S
are taken
in Theorems
3.5.7,
We seeFor
that
the numbers
if the
process
starts inofthe la
3.5.8, and 3.5.9.
these
theorems are
we small
will let
nj be
the number
Butprocess
this is isintuitively
clear,
since
in be
thisthe
case
it has
a 112 pro
times that the
in transient
state
Sj, m
total
number
"escaping"
on ever
the first
the mean
of transientofstates
it will
be step.
in, andForhijthebesame
the reason,
probability
of times
the last state
no matter
where the
that the process
will that
ever itgoistoin transient
stateisSf,small,
starting
in transient
starts. the
However,
from Nz we see that the former numbers ha
state Sj (not counting
initial state).
greater variances than the latter.
3.5.7 THEOREM.
FINITE MARKOV CHAINS
62
FINITE MARKOV CHAINS
62
----------------------
C
CHAP. III
PROOF.
{Mi[DJ]} = {di;} + {hjJMj[nj]}
or
{niJ} = I + {h[Jnfj}
or
Hence
II = (N -1)Ndg- 1 •
3.5.8 THEOREM. {Prt[n,-d1j=k]}=
of going to a given tr
E-H This theorem determines the probabilityilk=O
{
state-Hdg]
exactly
k times.
theorem is if
ank immediate
consequ
2 (I -Ndg-1)k-l
H.H dg k-l[1
= (N
-I)Ndg-The
> 0
the following consideration : To go to a given state k times one
This theorem determines
theonce,
probability
of must
goingreturn
to a given
there a t least
then one
k - 1 transient
times, and one m
state exactly kreturn
times. again.
The theorem is an immediate consequence of
the following consideration: To go to a given state k times one must go
there at least once, then one must return k-l times, and one must not
return again.
3.5.9
THEOREM.
PROOF.
The me&n number of transient states occupied is e
sum of the
probabilities
being in the various states.
JL the
= {M;[mJ}
= [H
+(1 -Hdg)Jgof =ever
NNdg-lg.
process starts in st, the probability of ever being in sf is htj
PROOF. The and
me~n
is 1number
if i =j. of transient states occupied is equal to
the sum of the probabilities
ever being
into
theExample
various 1,
states.
If the
If we applyof
Theorem
3.6.7
we obtain
process starts in St, the probability of ever being in sJ is hif if i =I j,
and is 1 if i = j.
If we apply Theorem 3.5.7 to Example 1, we obtain
pq
1-pq
H=
q
I-pq
p2
I-P7
P
2pq
q2
We see, for example,
that
q if q = 0, then all entries on and below t
l-pq
diagonal are 0.
This means that if the process is sure to mov
right, then it can never re-enter the starting state, nor can i t
We see, for example,
that
q = of
0, the
thenstarting
all entries
on and below the main
state to
theifleft
state.
diagonal are O. This means that if the process is sure to move to the
right, then it can never re-enter the starting state, nor can it enter a
state to the left of the starting state.
1 (1 +P +P3)
= -2
f-L
I-pq
2-pq
1 + q2+ q3
.
FINITE MARKOV CHAINS
62
SEC. 5
ABSORBING }l,gRKOV CHAINS
If q = 0 the vector f.L =
(D'
C
63
which is obvious in this case since it moves
directly to the right boundary, passing through the intermediate states
onlyonce.
3.5.10 THEOREM. The mean and variance of the number of changes of
state in an absorbing chain can be calculated by setting Pit = 0 for all
transient states, and dividing each row by its row-sum. The i-th
component of the new 7 gives the mean number of changes of state for the
of the
functionofisgoing
giventobya the
original process.
The variance
This theorem
determines
thesame
probability
given tr
new 72. state exactly k times. The theorem is an immediate consequ
the following consideration : To go to a given state k times one
PROOF.
Assume that the Markov chain is started in a non-absorbing
there a t least once, then one must return k - 1 times, and one m
state. \Ve form a new process in which the n-th outcome funct"lon is
return again.
defined as follows: If the original chain is absorbed at state Sk before
making n changes of state, th~n in = Sk;. If not, in is the state to which
the process moved on the n-th change of state. The new process is
clearly a Markov chain. The transition probabilities are the same as
PROOF.
me&n number of transient states occupied is e
P for s, absorbing.
ForThe
s, non-2bsorbing
the sum of the probabilities of ever being in the various states.
PH = starts
Pr,[f1=iJ
<)
process
in st, the
probability of ever being in sf is htj
and is 1 if i =j.
If we apply Theorem 3.6.7 to Example 1, we obtain
From this new transition matrix we can obtain the mean and variance
of the time to absorption for the process ii, £2, . . .. This time represents the number of changes of state in the original chain started in
state St.
We can also find the mean number of times that the process does not
change its state while it is among the transient states. This is found
by taking the mean number of times to reach the absorbing states and
subtracting the mean number of changes of state.
q = 0, then
all entries
and below t
Wetosee,
for example,
that if3.5.10
If we want
illustrate
Theorem
by the
college on
example,
areset
0. PitThis
therenormalize:
process is sure to mov
Example 10 diagonal
of § 2.2, we
= 0, means
i = 3,4,that
5, 6,ifand
right, then it can never re-enter the starting state, nor can i t
0
0
state to the1 left 0of the starting0 state.
1
0
F=
°
0
0
0
0
.22
.78
0
0
0
0
.2')
0
.78
0
0
0
.22
0
0
.78
0
0
2':>
0
0
0
.78
0
~
64
F I N I T E MARKOV CHAINS
64
FINITE MARKOV CHAINS
N =
.
C'
O.
0
1
0
•61
.78
.47
.61
1.78
7
=
72
=
CHAP. III
~)
.78
(1.00\
C
C)
.17
.68
\ 2.38) these results with Example 3.4.2, we note t
By comparing
2.85 of steps t o absorption
1.51 is somewhat higher than th
mean number
number of changes of state (but not by much, since repetiti
By comparing
these
results
3.4.2,
that
the
state
is rare),
andwith
thatExample
the variance
of we
the note
former
is considerabl
mean number of
steps
is somewhat higher than the mean
than
t hto
a t absorption
of the latter.
number of changes
of state
(but not use
by much,
since repetition
of a for ab
Another
interesting
of conditional
probabilities
state is rare), and
that
the
variance
of
the
former
is
considerably
higher
chains is the following. Assume t h a t for a n absorbing chain w
than that of theinlatter.
a non-absorbing state and compute all probabilities relativ
Another interesting
uset hof
for absorbing
hypothesis
a t conditional
the process probabilities
ends up in a given
absorbing state
chains is the following.
Assume
that
for
an
absorbing
chain
we start
Then we obtain a new absorbing chain with a single
absorbing
in a non-absorbing state and compute all probabilities relative to the
The non-absorbing states will be as before, except that we ha
hypothesis thattransition
the process
ends up in a given absorbing state, say SI.
probabilities.
We compute these as follows. Let p
Then we obtain a new absorbing chain with a single absorbing state Sl.
statement "the original process is absorbed in state sl." Th
The non-absorbing
will be as state,
before,the
except
that weprobabilities
have new for t
is a states
non-absorbing
transition
transition probabilities.
\-Ve
compute
these
as
follows.
Let
p be the
process are
statement "the original process is absorbed in state SI." Then if Si
is a non-absorbing state, the transition probabilities for the new
process are
Pri[P I[1 = Sj] . Pri[rl = Sj]
l'ri[p]
.
This formula applies for j= 1 if we interpret bll = 1. The s
form for I' may be obtained as follows. The matrix R is a
}
Let D ob lbe
matrix with d
This formulavector
applieswith
for j R= =
I if we interpret
The standard
l = a1. diagonal
form for P may be obtained as follows. The matrix 11 is a column
entries bll, for sj non-absorbing. Then
Let Do be a diagon.al matrix with diagonal
0 = D-loQDo.
. h R- = {'Pill
vector WIt
~r
entries bjl, for Sj non-absorbing.
Then
From this we see that
$a
From this we see that
= D-10&nDo
64
SEC. 5
F I N I T E MARKOV CHAINS
C
65
ABSORBING MARKOV CHAINS
and
N = D-1 o[I +Q+Q2+ ... ]Do
=
D-1 0 NDo.
B = g and T may be obtained from R.
EXAMPLE. Consider Example la, § 3.2.5. Let us consider the
process obtained by assuming that the original chain is absorbed in
state S1. Then the new ma,trix Q is
By comparing these results with Example 3.4.2, we note t
mean number of steps
is somewhat higher than th
S3absorption
82 t o
S4
number of changes of 2state
(but
not
by much, since repetiti
;
()
13
state is0rare), and that the
variance of the former is considerabl
than
t15/
h a3t of othe latter.
0
3/15 0 "
Q =
()
lis 0 2/s
Another interesting use of conditional probabilities for ab
()
15h
0 lisAssume
o t h0a t for0 a n 1/15
chains 0is the
following.
absorbing chain w
in a non-absorbing state
and
compute all probabilities relativ
2j;
hypothesis t h a t the process ends up in a given absorbing state
Then we obtain a new0 absorbing chain with a single absorbing
The non-absorbing states
will be as before, except that we ha
1
transition probabilities. We compute these as follows. Let p
80 that
statement "the original process is absorbed in state sl." Th
is a non-absorbing
the S4transition probabilities for t
S2
S3
81 state,
process are
. C'
O)C
OW"
(,:.
,~,)
"('
IV =
0
2/7
0
p = S2
5~7
0
Ss
\.J
7/ 9
S4
0
0
0
C'
0
°)
,~,)
6/5
O)C 3!5 ''7/fM
'5) e"
0
() \
This formula applies for j= 1 if we interpret bll = 1. The s
15/ 3
3/15
o be
3/ 5 9! 5
0 for
form
I' may
obtained as follows.
The matrix R is a
c
}
0
vector with 15
R =,1/ 5
0
0
9/5 'f")
18/35
entries bll, for sj non-absorbing.
C
2! 5
7! 5
7! 5 9/5 7/s
From this we see that
Then
0 = D-loQDo.
C")
$a
T=
0
'f:.)
Let D o be a diagonal matrix with d
ISis .
23/ 5
= D-10&nDo
Exercises for Chapter III
66
FINITE MARKOV CHAINS
CHAP. III
For § 3.1
for Chapter
1. P u t Exercises
the following
matricesIIIin the canonical form for absorbing
For § 3.1
1. Put the following matrices in the canonical form for absorbing chains.
51
51
(a)
P = 82
53
81
(b)
p=
82
52
Sa
l/S
,~,)
e
1
1/3
1/2
1/6
81
52
53
0
0
84
C D
0
0
1/4 t o
8S
1/4a n absorbing chain with a single ab
2. Apply Theorem
3.1.1
state.
0
0
84
3. Apply the result of the previous exercise t o an ergodic chain i
2. Apply Theorem
3.1.1has
to been
an absorbing
chain with
single absorbing
(Seea Chapter
11, Exercise 21.)
one state
made absorbing.
state.
4. I n Example 8 of $ 2.2 make state R into a n absorbing state.
3. Apply the result
of the previous
exercisetotothe
anresulting
ergodic chain
in which
3.1.1, applied
absorbing
chain, say ab
does Theorem
one state has been
made absorbing.
Chapter
II,is,
Exercise
(That
what do21.)
we learn about the
weather
in the Land (See
of Oz?
4. In Examplechain?)
8 of § 2.2 make state R into an absorbing state. What
does Theorem 3.1.1, applied to the resulting absorbing chain, say about the
weather in the Land of Oz? (That is, what do we Jearn about the original
For 5 3.2
chain?)
5. Compute the fundamental matrix for the absorbing chain with tr
matrix.
For § 3.2
s1 S2 S3
5. Compute the fundamental matrix for the absorbing chain with transition
matrix.
6. Compute the fundamental matrix for Example 11 of Chapter
c = 0 , and d # 0 .
7. Make Example 9 of $ 2.2 into an absorbing chain by making a
6. Compute the
fundamental
matrix for Example
of Chapter matrix
II whenand inter
Find the 11
fundamental
ergodic
states absorbing.
c=O, and d=l=O. entries of the first row of this matrix.
7. Make Example
of § 2.2
an fundamental
absorbing chain
by making
all for
of the
N is given
an absorbin
8. 9Show
t h ainto
t if the
matrix
ergodic states absorbing.
Find and
the Q
fundamental
= I - N-1. matrix and interpret the
then N-1 exists
entries of the first row of this matrix.
9. Prove t h a t NQ = N - I.
8. Show that if the fundamental matrix N is given for an absorbing chain,
Check
the results of Exercise 9, above, in Example 9 of 3 2.2.
then N-1 exists and10.
Q=I
-N-1.
9. Prove that NQ=N -I.
10. Check the results of Exercise 9, above, in Example 9 of § 2.2.
SEC. 5
Exercises for Chapter III
ABSORBING :MARKOV CHAINS
For § 3.1
67
1. P u t the following
in the canonical form for absorbing
Formatrices
§ 3.3
11. If an absorbing chain has only one absorbing state, what can be said
about the matrix B? In Example 8 of § 2.2 make R an absorbing state,
compute Nand B, and verify your statement.
12. Change Example 7 of § 2.2 into an absorbing chain by assuming that
the process is stopped if a 0 or 9 is reached. Construct the new transition
matrix, in canonical form.
13. In the example of Exercise 12, above, compute N, N2, B, 'T, 'T2.
14. In Example 8 of § 2.2 make N into an absorbing state. Compute the
fundamental matrix for the resulting Markov chain. Find N 2 , B, 'T, 'T2'
Interpret the results in terms of the original chain.
15. Compute N for the tank duel (Exercise 2 of Chapter II). From this
find the mean length of the duel and the probability of each possible ending.
16. Carry out the computations of Exercise 15, above, for the moa.ified
tank duel (Exercise
3 ofTheorem
Chapter 3.1.1
II). tWhich
duel is more
to
2. Apply
o a n absorbing
chainfavorable
with a single
ab
tank A?
state.
17. In Example lOa (of § 3.2.5) find the probabilities of graduation by the
3. Apply the result of the previous exercise t o an ergodic chain i
method resulting from Theorem 3.3.9, that is, by finding a certain fixed
one state has been made absorbing. (See Chapter 11, Exercise 21.)
column vector for the transition matrix.
4. I n Example 8 of $ 2.2 make state R into a n absorbing state.
18. The chain
Example3.1.1,
la (cf.applied
§ 3.2.5)toisthe
started
by means
of a random
resulting
absorbing
chain, say ab
doesofTheorem
device which make all five states equally likely as starting states. Find the
weather in the Land of Oz? (That is, what do we learn about the
means and variances of the number of times in the various transient states,
chain?)
and of the number
of steps to absorption.
For 5 3.2
For § 3.5
5. Compute the fundamental matrix for the absorbing chain with tr
9 of § 2.2, assume that initially both balls are unpainted.
19. In Example
matrix.
Find the mean number of draws before the first time that both balls are
S2 both
S3 balls are red?
s1 that
painted. When this occurs, what is the probability
20. It is snowing in the Land of Oz today. Find the mean number of
changes of weather that will occur before the next rainy day. Find the probability that there is at least one nice day before a rainy day.
21. For Example 1 of § 2.2 with p=1/2, assume that it is known that the
process is absorbed in state 81. Find the transition matrix for the new
6. Compute
themean
fundamental
matrix for Example 11 of Chapter
Find the
time to absorption.
conditional process.
c = 0 , and d # 0 .
22. Compute the following quantities for the tank duel (see Exercise 2 of
7. Make Example 9 of $ 2.2 into an absorbing chain by making a
Chapter II).
ergodic states absorbing. Find the fundamental matrix and inter
(a) The mean
andofvariance
theofnumber
of rounds for which all three
entries
the firstofrow
this matrix.
tanks remain active.
N be
is given
anBabsorbin
8. Show
t h aatt ifsome
the fundamental
that
stage A and Cmatrix
will still
active,for
but
is
(b) The probability
N-1 exists and Q = I - N-1.
then
no longer
active.
(e) The probability
that
9. Prove
t hat
a t some
NQ =stage
N - I.A and B will still be active, but C is
no longer 10.
active.
Check the results of Exercise 9, above, in Example 9 of 3 2.2.
(d) The probability that A and B will be eliminated on the same round.
(e) P, il, T, assuming that C wins the duel.
(f) P, N, T, assuming that no tank survives.
23. I n the tank duel (Exercise 2 of Chapter 11)let tank A have prob
B probability
315, and tank C anCHAP.
unspecified
proba
of hitting,
tank
III
FINITE
lvL<\RKOV
CHAIN8
(with p < 3 1 5 ) .
68
23. In the tank duel
(Exercise
of Chaptermatrix.
II) let tank A have probability
(a) Set
up the2transition
3/ 4 of hitting, tank B probability 3/ 5 , and tank C an unspecified probability p
(b) Find the probability that tank C is the survivor.
(c) I n the answer obtained in (b), let p tend to 0.
Interpret
(a) Set up the transition
matrix.your result.
(with p < 3/ 5 ).
(b) Find the probability that tank C is the survivor.
(c) In the answer obtained in (b), let p tend
O. entire chapter
Fortothe
Interpret your result.
24. Seven boys are playing with a ball.
The first For
boy the
always
throws
it to the second boy.
entire
chapter
The second boy is equally likely to throw it to the third or the sev
Theplaying
third boy
keeps
the ball if he gets it.
24. Seven boys are
with
a ball.
The throws
fourt,h boy
throws
The first boy always
it to always
the second
boy.i t to the sixth.
The
fifth boy
is equally
o throw
fourth, sixth, or
The second boy is
equally
likely
to throwlikely
it to tthe
third it
or to
thethe
seventh.
The third boy keeps
the ball if he gets it.
boy.
The s ~throws
x t hboy it
always
The fourth boy always
to thethrows
sixth. it to the fourth.
The seventh
equally
likely
throwsixth,
it to the
first or fourth bo
The fifth boy is equally
likelyboy
to is
throw
it to
the to
fourth,
or seventh
boy.
(a) Set up the transition matrix P.
The sixth boy always throws it to the fourth.
(b) Classify the states.
The seventh boy is equally likely to throw it to the first or fourth boy.
(c) P u t P into canonical form.
(d) Give an
interpretation
for the chain ending up in one of the ergo
(a) Set up the transition
matrix
P.
(e) The ball is given to the fifth boy. Find the mean and varianc
(b) Classify the states.
number
of times that the seventh boy has the ball, and find th
(c) Put P into canonical
form.
and variance
of theending
time to
set. sets.
for the chain
upreach
in onean
of ergodic
the ergodic
(d) Give an interpretation
(e) The ball is given
to thean
fifth
boy. Find
the mean
of the
25. Given
absorbing
Markov
chain,and
we variance
play a game
as follows:
number of times that the seventh boy has the ball, and find the mean
We
start
in
a
specified
state,
and
carry
the
chain
out
till
it
reaches
an
and variance of the time to reach an ergodic set.
ing state. If we reach s,, we receive a payment of c,. Form the
Markov
we playa
game
as of
follows:
25. Given an absorbing
vector y whose
i-thchain,
component
is the
mean
the payment if we sta
\Ve start in a specified
state, and
(a) Prove
thatcarry
Py=the
y. chain out till it reaches an absorbSa, we receive a payment of Ca.
the column of y is c,.
ing state. If we reach
(b) Prove
that for absorbing state sa theForm
a-th component
vector y whose i-th component is t.he mean of the payment if we start in Si.
(c) Prove that these two conditions determine y.
(HIST: Cons
limit of Pny.)
(a) Prove that Py=y.
(d)absorbing
Let y, be
theSavector
giving
the probabilities
(b) Prove that for
state
the a-th
component
of y is Ca. of absorption
y can be
expressedy. in (HI~T:
terms ofConsider
the y,. the
(c) Prove that theseShow
two that
conditions
(letermine
limit of pn y .)
(d) Let Ya be the vector giving the probabilities of absorption in Sa.
Show that y can be expressed in terms of the Ya.
23. I n the tank duel (Exercise 2 of Chapter 11)let tank A have prob
of hitting, tank B probability 315, and tank C an unspecified proba
(with p < 3 1 5 ) .
(a) Set up the transition matrix.
(b) Find the probability that tank C is the survivor.
(c) I n the answer obtained in (b), let p tend to 0.
Interpret
your result. IV
CHAPTER
For the entire chapter
24. Seven boys are playing with a ball.
REGULAR
MARKOV
CHAINS
The
first boy always
throws it to the
second boy.
The second boy is equally likely to throw it to the third or the sev
The third boy keeps the ball if he gets it.
The fourt,h boy always throws i t to the sixth.
The fifth boy is equally likely t o throw it to the fourth, sixth, or
boy.
§ 4.1 Basic The
theorems.
this section
shall
s ~ x t hboyInalways
throwswe
it to
thestudy
fourth.the behavior of
a regular Markov
recall
thatlikely
a regular
Markov
chain
Thechain.
seventh We
boy is
equally
to throw
it to the
first is
or one
fourth bo
that has no transient sets, and has a single ergodic set with onfy one
(a) Set up the transition matrix P.
cyclic class. (b) Classify the states.
(c) P u t P into canonical form.
4.1.1 DEFINITION.
transition for
matrix
for ending
a regular
Markov
(d) Give anThe
interpretation
the chain
up in one
of the ergo
regular
transition
matrix.
chain is called
(e)aThe
ball is
given to the
fifth boy. Find the mean and varianc
number of times that the seventh boy has the ball, and find th
4.1.2 THEOREM.
transition
regular
if and set.
only if for
and Avariance
of thematrix
time tois reach
an ergodic
some N, P N 25.
has Given
no zero
entries.
an absorbing Markov chain, we play a game as follows:
We start
in a specified
state,
and carry
the chain
out till itifreaches
It was shown
in Chapter
II that
a Markov
chain
was regular
and an
ing state.to be
If in
weany
reach
s,, after
we receive
payment
Form
only if it is possible
state
some anumber
N of
of c,.
steps,
no the
y whose i-th component is the mean of the payment if we sta
matter whatvector
the starting
state. That is, if and only if pN has no zero
(a)N.
Prove that Py= y.
entries for some
(b) Prove that for absorbing state sa the a-th component of y is c,.
(c) Prove
determine
y. (HIST:
4.1.3 THEOREM.
Letthat
P bethese
an rtwo
x r conditions
transition matrix
having
no zero Cons
limit of Pny.)
entries. Let(d)
€ be the smallest entry of P.
Let
x
be
any
r-component
Let y, be the vector giving the probabilities of absorption
maximum
Moterms
and ofminimum
column vector, having
Show that
y can becomponent
expressed in
the y,. component mo, and let M 1 and ml be the ma.ximum and minimum components for the vector Px. Then lFI 1 ~ M o, ml;" mo, and
M1-ml ~ (1-2€)(Mo-mo).
PROOF. Let x' be the vector obtained from x by replacing all components, except one mo component, by Mo. Then x ~ x'. Each
component of Px' is of the form
a·mo+(l-a)·Mo = Mo-a(Mo-mo)
where a;" E. Thus each such component is ~ M 0 - €(M 0 - mo).
since x ~ x', we have
Ml ~ .ll1 o-€(Mo-mo).
(\9
But
(1)
If we apply this result to the vector - x we obt,ain
FINITE MARKOV CHAINS
CHAP. IV
70
If we apply this result to the vector - x we obtain
Adding (1) and (2) we have
-mi ~ -mo- .. (-mo+M o}.
(2)
Adding (1) a.nd (2) we have
This
theorem
gives us a .. simple
proof of the following funda
M1-ml
~ Mo-mo-2
(~~fo-mo)
theorem for =regular
chains.
(1- 2Markov
.. )(M 0 - mol·
This theorem gives
a simple proof
following
fundamental
4.1.4 usTHEOREM.
If P of
i s the
a regular
transition
matriz then
theorem for regular Markov chains.
(i) T h e powers Pn approach a probability matrix A .
u of A itransition
s the samematrix
probability
4.1.4 THEOREM. (ii)
If Each
P is ar oregular
then vecto? a = { a l , az, . .
tha.t is A = ( a .
(i) The powers pn approach a probability matrix A.
(iii) T h e components of a are positive.
(ii) Each row of A is the same probability veetot a = {aI, a2, ... , an},
We shall first assume that P has no zeros. Let c
that is A PROOF.
= ga.
minimum of
entry.
Let p j be a column vector with a 1 in t)
(iii) The components
a are positive.
component and 0 otherwise. Let M, and 7nn be t,he maximu
PROOF. \Ve shall first assume that P has no zeros. Let .. be the
minimum components of the vector P p j . Since P n p j = P . P
minimum entry. Let Pi be a column vector with a 1 in the j-th
we ha,ve, from Theorem 4.1.3, that
2 X 28 M 3 8 .
component and 0 otherwise. Let 111n and mn be the maximum and
m l < 7722 < m s < . . . and
minimum components of the vector pnpj. Since pnpi=p·pn-Ipi,
we have, from Theorem 4.1.3, that MI~M2-;;,M3~ .. ·and
ml,;:;m2,;:;m3';:; ... and
for n 2 1 . If we let, d, = 1W,- nz,, this tells us that'
Mn-mn';:; (1-2 .. )(Mn - I -m,,-I)
for n ~ I.
If we let dn = M n - m n , this tells us that
Thus as n tends t o infirtit,y d , goes to 0,111, and m, approach a co
(1-2 .. )n·dPnpj
tends.. )n.
to a vector with all componen
limit, dand
n ,;:; therefore
o = (1-2
same. Let ai be this common value. I t is clear that, for
Thus as n tends to infinity d n goes to 0, jVI" and mn approach a common
and ilill < 1, we hav
m , < aj < M,. I n particular, since O <
limit, and therefore Pnpi tends to a vector with all components the
0 < ai < I . Now P n p j is the j-th column of Pn. Thus the j - t h c
same. Let at be this common value. It is clear that, for all n,
of Pn tends to a vector with all components the same value ni. T
m n ,;:;aj,;:;ltl n . In particular, since O<ml and ~~11< 1, we have that
Pn tends to a nlatrix A with all rows the same vector a = { a l , aa: .
0< aj < 1. Now pnpj is the j-th column of pn. Thus the j-th column
Since the row-sums of Pn are always 1 , the same must be true
of pn tends to a vector with all components the same value aj. That is,
limit. This completes the proof for t,he case where the matrix
pn tends to a matrix A with all rows the same vector a = {aI, a2, ... , ar}.
positive entries.
Since the row-sums of pn are alway.s 1, the same must be true of the
Consider next the case that P is only assumed to be regular.
limit. This completes the proof for the case where the matrix has all
be such t h a t P" has no zero entries. Let E' be the smallest en
positive entries.
PN. Applying the first part of the proof t o the matrix PN, we h
Consider next the case that P is only assumed to be regular. Let N
be such that plY has no zero entries. Let E' be the smallest entry of
plY. Applying the first part of the proof to the matrix PN, we have
Therefore the sequence d,, w h ~ c his norl-increasing, has a subse
(3)
Therefore the sequence d n , which is non·increasing, has a subsequence
SEC. 1
If we apply this result to the vector - x we obt,ain
REGULAR MARKOV CHAINS
71
tending to O. Adding
Thus d",(1)tends
to zero
and the rest of the proof is the
and (2)
we have
same as in the proof for all positive entries.
4.1.5 COROLLARY. Let P be a regular transition matrix. Let
aj = lim p("')jj. Then there are constants band r with 0< r< 1 BUCh that
This theorem gives us a simple proof of the following funda
theorem for regular Markov chains.
with le(lI)jil ::;;; br".
4.1.4 THEOREM.
If P i s a regular transition matriz then
PROOF. We know that le(n)jjl ::;;; dll • Let N be such that pN has no
(i) the
T h esmallest
powers Pn
approach
matrix
A.
zero entries. Let" be
entry
of PN. a probability
Choose r= (12E)1/N
(ii)
Each
r
o
u
of
A
i
s
the
same
probability
vecto?
a
= { a l , az, . .
N
ll
and b=1/(l-2,,)=r- • If n=kN, then from (3), dn ::;;;r • If
a.
tha.t is A = (then
n=kN+nl> where O::;;;nl::;;;N,
since d", is non-increasing,
(iii)
T
h
e
components
of a obtained
are positive.
d.,::;;; ""-""::;;; rn. r- N = brn. The bound here
for e(n)/j is useful
for proving theorems,
is very
estimate
for Let c
PROOF. but
Weit shall
firstconservative
assume thatasPanhas
no zeros.
p(n)/J. Let p j be a column vector with a 1 in t)
the rate of convergence
minimum of
entry.
component
otherwise.
andand
7nn Abeand
t,he 0:maximu
Let M,
4.1.6 THEOREM.
If P and
is a0 regular
transition
matrix
minimum
components
of
the
vector
P
p
j
.
Since
P n p j= P . P
are as given in Theorem 4.1.4, then
we ha,ve, from Theorem 4.1.3, that
2 X 28 M 3 8 .
(a) For any
vector 1T, 1T'P" approaches the vector 0: as
m l <probability
7722 < m s < . . . and
as n tends to infinity.
(b) The vector ex is the unique probability vector such that o:P = 0:.
PA =AP=A.
for n 2 1 . If we let, d, = 1W,- nz,, this tells us that'
PROOF. If 1T is a probability vector, then 1Tg= 1; hence 1TA =1Tgo: = ex.
But 1T' pn approaches 7T' A. Hence it approaches ex. This proves
Thus as n tends t o infirtit,y d , goes to 0,111, and m, approach a co
part (a).
Pnpj
tends to a approaches
vector withA,allbutcomponen
limit, of
and
P therefore
approach A,
p.,+l=pn.p
Since the powers
ai AP=A.
be this common
I t proving
is clear(c).
that, for
same.
it also approaches
AP;Let
hence
Similarlyvalue.
PA =A,
I n particular,
since
andnow
ilillshow
< 1, we hav
m this
, < aj
< M,. equation
Anyone row of
matrix
states that
exP O= <0:. We
0 < ai Let
< I . {3 Now
P nprobability
p j is the j-thvector
columnsuch
of Pn.
Thus
the j - t h c
that 0: is unique.
be any
that {3P
= {3.
to a But
vector
with
same0: =value
of Pn tends ex.
By (a), {3. P'" approaches
since
f3Pall
= components
f3, f3P" = f3. the
Hence
f3. ni. T
tends(b).
to a nlatrix A with all rows the same vector a = { a l , aa: .
Thus we have Pn
proved
row-sums
are always
1 , the
same matrix
must be true
The matrix Since
A andthe
vector
0: will of
be Pn
referred
to as the
limiting
limit. forThis
the proof
for t,heby
case
and limiting vector
the completes
Markov chain
determined
P. where the matrix
positive
Theorem 4.1.6
showsentries.
that for a regular transition matrix there is a
Consider
next the
case that
is only assumed
regular.
row vector 0: which
remains
"fixed"
whenP multiplied
by P.to be
Any
such that
t h a t ex'P"P =has
zero entries. toLet
be the smallest en
other vector 0:'be such
ex' no
is proportional
theE' probability
Applying
the first
part
of the
o the matrix
PN, we h
vector 0:. ThePN.
following
theorem
shows
that
anyproof
fixedt column
vector
for P is proportional to f
(c)
4.1.7 THEOREM. If P is a regular transition matrix and p ={r/} is a
Therefore
column vector
such thatthe sequence d,, w h ~ c his norl-increasing, has a subse
Pp = p
then p = c . g jor some constant c.
72
Since P p = p, P e p = P p = p and in general Pnp = p.
PROOF.
But this states that
all I\'
compone
CHAP.
?liAHKOV
also A pFIXITE
= p . Thus
r,=ap. CHAINS
have the same value. That is p = c[ for some constant c.
PROOF.
Since Pp=p. P2p=Pp=p and in general pnp=p Hence
In T;Chapter
I1 we
saw
thatthat
if the
is st,arted
also Ap = p. Thus
= ap.
But
this
states
all process
components
of p in each
states with
probabilities
T , then
have the same value.
That
is p = ct for given
some by
constant
c. the probabilities for
each of the states after n steps are given by n P n .
For large n T
In Chapter II we saw that if the process is started in each of the
4.1.6 states that x P n is approximately a. Since a depends on
states with probabilities given by 7T, then the probabilities for being in
and not on 7 , this may be described by saying tha,t, for a
each of the states after n steps are given by n pn. For large n Theorem
Markov chain, the long range predictions are independent of th
4.1.6 states that npn is approximately a. Since a depends only on P
vector. Let us illustrate this in terms of Example 8 of Cha
and not on 7T, this may be described by saying that, for a regular
The transition matrix for this example is
Markov chain, the long range predictions are independent of the initial
N 8 of
S Chapter II.
vector. Let us illustrate this in terms of Example
The transition matrix for this example is
a
R
N
S
:(:;: 1:4 :;:).
To find the vector a = ( a l , a2, a s ) , we must find a probability
a. 1That
is, we must satisfy the following set o
such that aSP =14
/4 liz
tions :
To find the vector a = (aj, az, as), we must find a probability vector
1 =
al+ az+ as
such that ct.P = ct..
tions:
That is, we must satisfy the following set of equaal+
al
az+
a3
)!2al+l/zaz+J/4aS
The unique
solution
az = 1/4aJ
+ 1/4ato
3 these equations is
a3 = J/ 4a) + l/ZaZ+ I/ZaS.
The unique solution
to matrix
these equations
A is thenis
The limit
CZ/5, lis, z/s).
ct. =
The limit ma.trix A is then
.4
.2
.4)
(
Corollary A4.1.5 states that this hmit is reached geometrically
\': :: ::
being a very fast kind of convergence,
we would expect that, e
.
moderately large values of n , Pn should be approximated by A
Corollary 4.1.5 states that this limit is reached geometrically. This
matrix P5 is
being a very fast kind of convergence, 1ve would
a expect
N that.seven for
moderately large values of n, pn should be approximated by A. The
R /.4004 .2002 .3904\
matrix p5 is
R
N
S
Reo,
.2002
ps = N .4004
. 1992
.4004
S
.2002
.4004
.3994
3904) .
SEC. 2
Since P p = p, P e p = P p = p and in general Pnp = p.
PROOF.
But this states that all compone
also AREGULAR
p = p . Thus
r,=ap. CHAINS
yL\RKOV
73
have the same value. That is p = c[ for some constant c.
Each row of p5 gives the probability of each kind of weather five days
Chapter
I1 weFor
sawexample,
that if the
st,arted
after a particularInkind
of day.
the process
first rowis gi\'es
thein each
states
given
byafter
T , then
the probabilities
probabilities for
eachwith
kindprobabilities
of weather five
days
a rainy
day. The for
n steps
arethat
given
by n Pweather
n . For large
nT
fact that the each
rows of
arethesostates
nearlyafter
equal
means
today's
in
that x P n tois have
approximately
a. Since
a depends
on
the Land of Oz4.1.6
maystates
be considered
very little effect
on our
preand
notfrom
on 7now,
, this may be described by saying tha,t, for a
dictions for five
days
§ 4.2
Markov chain, the long range predictions are independent of th
Law of
large numbers
regularthis
Markoy
chains.
As we have
vector.
Let us for
illustrate
in terms
of Example
8 of Cha
seen in § 4.1, for
regular Markov
there
is a limiting
probability
Thea transition
matrixchain
for this
example
is
aj of being in state S1 independent of the starting state.
In this section
S time that the
we shall prove that aj also represents the fractionNof the
process can be expected to be in state Sj for a large number of steps.
This result will also be independent of the starting state.
To state the above result precisely, we shall need to introduce some
new functions. Let u(n)j be a function with domain the tree Un and
with value I To
if the
wasa =to( astate
otherwise.
'\\-e
findn-th
the step
vector
l , a2, a5js )and
, we 0must
find a probability
a
n
such
that a P = a. That is, we must satisfy the following set o
U(k)j'
Then y(n 1j is again a function with domain
tions :
=
al+
az+ asthe initial
the tree Un and value the number of1 times
(not counting
define y(n)j =
.2
k~l
position) that the process is in state 5j during the first n steps. The
function y(n)j = y(n)j!n gives the fraction of times in the first n steps that
the process moves to state 8J.
4.2.1 THEOREM (The Law of Large Numbers). Consider a regular
The unique solution to these equations is
Markov chain with limiting vector a = (aI, az, ... , arlo For any
initial vector 7T,
(a)
The limit matrix A is then
and for any t > 0
(b)
as n tends to infinity,
PROOF. According to Theorem 1.8.10, to prove this theorem it is
sufficient to prove
that Mn[(y(n'j
as hmit
tendsis to
infinity.
To
4.1.5 states
that 0this
reached
geometrically
Corollary
prove this it is being
sufficient
to
prove
that,
for
every
i,
Mi[(v(n\
aj)2J~0.
a very fast kind of convergence, we would expect that, e
aj)2J . . . .
n
moderately large values of n , Pn should be approximated by A
matrix P5 is Thl; [ (
aj
a
N
s
k~ (U(k)j/n) - fJ
R /.4004
=
.2002
~2M;[C~(U(k)j-aj))1
Let mk,I=Mi[(u(k)j-Uj)(u(l)j-aj)l
Then we must prove that
i i mk,l""'"
1~1
k~l
.3904\
0
(1)
as n tends to infinity. Multiplying out the expression for m k S 1we
FINITE MARKOV CHAINS
CHAP. IV
74
as n tends to infinity. Multiplying out the expression for mk.1 'we have
Let m = min ( k , 1 ) and d = lk - El. Then
mk,1 =
Mi[U(k)ju(l)j] - ajMi[u(kl i ] - 0iM/[U(I)j]
+ 02 j .
Let m=min (k, I) and d= Ik-ll. Then
Using Corollary 4.1.5, we have
Using Corollary 4,1.5, we have
where je(n)ij/< brn with O < r < 1. Hence for a suitably chosen con
mk,lc,=
aj(e(m)ji
+ e(d);} - elk);} - eU)ii) + e(m)ije(d)jj
where le(n);il ~ br n with 0< r< 1.
c,
c(rm+ r d chosen
+ r k + r 2constant
).
lmk,rl
Hence
for <
a suitably
Each value of m , d, k , and 1 occurs < 2 n times in the sum in (1). M
c(rm+rd+rk+rl).
(2)
2 ) , we ~
have
using ( Imk,!;
Each value of m, d, k, and I occurs ,-;; 2n times in the sum in (1).
using (2), we have
Hence,
4c inequality
2n
8etends to 0 as n tends to inf
The right side of this
n 2 "1r to
n(lr)'
hence, also the left~side,
as was
be proved.
Let us apply this theorem to the Land of Oz example. We fou
The right side of this inequality tends to 0 as n tends to infinity;
$ 4.1.5 that a = (2j5, ' I 5 , 2 i 5 ) Thus we can now say that for a
hence, also the left side, as was to be proved.
number of days we can expect about ' I s of the days to be rain
Let us apply this theorem to the Land of Oz example. We found in
of the days to be nice, and 215 of the days to be snowy.
§ 4.1.5 that a=(2/s, 115, 2 tS ). Thus we can now say that for a large
Consider the special case of an independent trials process. S
number of days we can expect about 2/5 of the days to be rainy, 1/5
process is a Markov chain with transition mat,rix having all row
of the days to be nice, and 2/5 of the clays to be snowy.
same vector a and with initial probability vector chosen to be a.
Consider the special case of an independent trials process. Such a
law of large numbers for a n independent trials process is thus a s
process is a Markov chain with transition matrix having all rows the
case of the theorem just proved. The proof for this case is very
same vector a and with initial probability vector chosen to be a. The
simpler. I n fact, in this case P l l , [ ~ ( n ) ~ ] = afor
j a,ll n; hence
law of large numbers for an independent trials process is thus a special
M , [ V ( ~=
) ~aj.
] Also mkVl= 0 for k # 1 and m k , k= 02, a constant fo
case of the theorem just proved. The proof for this case is very much
Hence
simpler. In fact, in this case Malu(n)j]=Oj for all n; hE:nce also
Ma[v(n)jJ=aj.
Also mk,Z=O for k#l and mk,k=a 2 , a constant for all k.
Hence
This tends to 0 as n tends to infinity.
Another special case of interestn is a general Narkov chain p
which is started by an initial probability vector n = a. In this cas
This tends to 0 as n tends to infinity.
. [ ~ ( n )=
~ ]M , [ V ( ~ =
) ~aj] for all n . Hence
Another special case of interest is a general Markov chain process
which is started by an initial probability vector 1T = a. In this case also
Ma[u(n)j] = Ma[v(n)j] =aj for all n. Hence
However, i t is not possible, in this case, to give a simple expressi
Ma[(v(n)}as
- OJ)2]
= Vara[v(n)j].
this variance
a function
of n , as was possible in the independen
However, it is not. possible, in this case, to gi\'e a simple expression for
this variance as a function of n, as was possible in the independent case.
as n tends to infinity.
SEC. 3
Multiplying out the expression for m k S 1we
REGULAR :\LJ...RKOV CHAINS
75
We shall consider
this ( variance
§ 4.6,
Let m = min
k , 1 ) and din= lk
- El. where
Then we shall give an
asymptotic expression for it.
§ 4.3 The fundamentai matri.x for regular chains. In Chapter III
Using
4.1.5, chain,
we have
we found that,
forCorollary
an absorbing
the matrix (l-Q)-l played a
fundamental role. (Q was the matrix obtained by truncating the
transition matrix to include only the non-absorbing states.) We shall
Hence for
for regular
a suitably
chosen con
where
je(n)ij/< brn with
O < r < 1. matrix
see that there
is a corresponding
fundamental
chains.
c,
4.3.1 THEOREM. Let P be the transition matrix for a regular Markov
< c(rm
+ r d +Zr=k +(1r 2 )(P
. - A»-l
chain. Let A be the limiting lmk,rl
matrix.
Then
exists and Each value of m , d, k , and 1 occurs < 2 n times in the sum in (1). M
.,
using ( 2 ) , we have
Z = 1+
(pn-A).
.I
n~l
PROOF.
We shall prove that (p-A)n=pn-A. Since pn-A--+O,
our theoremThe
willright
thenside
follow
frominequality
the matrix
theorem
proved
in to inf
of this
tends
to 0 as
n tends
§ 1.11.1. We
have
A2=tata=ta=A,
hence
Ak=A,
and
hence, also the left side, as was to be proved.
We fou
n theorem to the Land of Oz example.
Let us apply this
l)n-tp'An-t
(p-A)n
Thus we can now say that for a
$ 4.1.5
that a== (2j5, ' I 5 ,(_
2i5)
number of days i-a
we can expect about ' I s of the days to be rain
of the days to be nice, and 215 of the days to be snowy.
Consider the special case of an independent trials process. S
process is a Markov chain with transition mat,rix having all row
pn-A.
same vector a and
with initial probability vector chosen to be a.
4.3.2 DEFINITION.
P befor
a aregular
transition
matrix.
law of large Let
numbers
n independent
trials
process The
is thus a s
case of the theorem
just proved.
The proofmatrix
for thisforcase
is called
the fundamental
theis very
matrix Z=(l-(P-A»-l
simpler.
I n fact,
j a,ll n; hence
J11 arkov chain
determined
by P.in this case P l l , [ ~ ( n ) ~ ] = afor
M , that
[ V ( ~the
=
) ~aj.
] AlsoZ mkVl
= 0 for k # 1 and m k , k= 02, a constant fo
We shall see
matrix
is the
basic quantity used to compute
Hence
most of the interesting descriptive quantities for the behavior of a
.I
regular Markov chain. We shall first establish certain important
properties of the matrix which will be useful in later work.
4.3.3 THE
ORE'>!,
fundamental
tends
to infinity.matrix jar a regular
This
tendsLet
to 0Z asben the
limitingis vectoT
a, andNarkov
limitingchain p
Markov chain
with transition
Another
special matrix
case of P,
interest
a general
matrix A. which
Thenis started by an initial probability vector n = a. In this cas
~ ( n )=
~ ]M , [ V ( ~ =
) ~aj] for all n . Hence
(a)
PZ =. [ZP
(b)
Zt= t
(e)
aZ However,
=a
i t is not possible, in this case, to give a simple expressi
(d) J-Z this
= A-PZ.
variance as a function of n , as was possible in the independen
PROOF.
Part (a) follows from the infinite series representation for Z
and the fact that P commutes with each term in this infinite series.
CH
FINITE MARKOV CHAINS
7G
P a r t (b) states that Z has row-sums 1. This again follows fro
MARKOV CHAINS
IV
infiniteFINITE
series representation
for Z since the CHAP.
first matrix
I has
sums 1 and each of the matrices Pn-A have row-sums 0. Pa
Part (b) states that Z has row-sums 1. This again follows from the
follows from tshe
representation
for rowZ, since a1 =
infinite series representation
for infinite
Z sinceseries
the first
matrix I has
76
m
1
sums I and each
of the
- A have
row-sums
a(PnA )matrices
= O . To pn
prove
(d) we
multiplyO. Z Part
= I + (c) (Pn- A
n- 1
follows from the infinite series representation for Z, since (XI = a and
-
00
L (pn-A) by
I Pprove
obtaining
To
(d) we multiply Z=I+
a(pn-A)=O.
( I - P ) Z = ( 1 - Pn-l
)+(P-A)
= I-A.
1- P obtaining
(I-P)+(P-A)
We(I-P)Z
shall use =the
Land of Oz example as our standard example
=
I-A.
applications of the Z matrix. For this example P and A are
\Ve shall use the Land of Oz example as our standard example of the
applications of the Z matrix. For this example P and A are
R
N
S
T'
1/4
'I,)
1/2
N
R
1/5
C
S
'f')
R
A
liz 0
2/5 1/5 2/5 N.
find the inverse of the matrix
TO find the matrix Z we must
2I
S 1/4 1/4 l/Z
S
/5 1/5 2/5
P = N
.9
-.05
To find the matrix Z we must filld the inverse of the matrix
1.2
-.l
(9
-.05
I-P+A
- 1
Doing this we obtain
1.2
.15
-.05
-:)
.15
-.05
.15
-.I
.9
.9
Doing this we obtain
R
N
S
86
3
-1:)
R
Z
63
N.
,,(
6
While"the
fundamental
matrix Z has several properties in eo
-14 matrix,
86 seeS from this example that it do
3
with a transition
we
necessarily have non-negative entries.
While the fundamental matrix Z has several properties in common
An example where the Z matrix turns out to be a very simple
with a transition matrix, we see from this example that it does not
is the case of an independent trials process. In this case P = A
necessarily have non-negative entries.
Z = ( I - ( P - A ) )-"I.
Thus for an independent trials .proce
An example where the Z matrix turns out to be a very simple matrix
Z
matrix
is
the
identity
matrix.
is the case of an independent trials process. In this case P = A so that
be the number of times t h a t the process is in state s
Let
Z=(l-(P-A»-l=l. Thus for an independcnt trials 'process the
n
stages,
i.e. the initial position plus n - l stages.
first
Z matrix is the identity matrix.
Let y(n)j be the number of times that the process is in state s1 in the
first n stages, i.e. the initial position plus n - l stages.
SEC. 3
CH
FINITE MARKOV CHAINS
7G
P a r t (b) states that Z has row-sums 1. This again follows fro
infiniteREGULAR
series representation
for Z since the first matrix
~1ARKOV CHAINS
77 I has
sums 1 and each of the matrices Pn-A have row-sums 0. Pa
THEOREM.
For any
regular series
Markov
chain, and any
initial
follows from
tshe infinite
representation
for Z,
since a1 =
4.3.4
vector "IT,
m
a(PnA ) = O . To prove (d) =we17Z-a.
multiply Z = I +
{M,,[y(n)f]}-ncc-+17(Z-A)
1 (Pn- A
n- 1
For any i,
I - P obtaining
PROOF.
11-1
n-1
}=o
..-0
=
I-A.
( I - P ) Z =2:( 1 - P ) + ( P - A )
2: Mt[U(k)j]
p(lC)/j.
Thus
»-1
We shall use the Land
of Oz example as our standard example
(Pk-A)For
-+this
Z-A.
example P and A are
applications of the Zk=O
matrix.
{Mi[y(n),)-naj} =
2:
Therefore
77{M.[y(n)j1- na,} -+ 77(Z - A) = 77Z - a.
An immediate consequence of this theorem is the following:
4.3.5
COROLLARY.
For any two initial distributions 77 and 77'
{Mw[y(r.)j] - M~·[y(1l)jJ} -+ (7T - 17')Z.
TO find the matrix Z we must find the inverse of the matrix
If we choose a particular starting state, say i, then Theorem 4.3,4
.9
-.05
.15
states that
fill[Y(")j] - naj ->- (Zij -aj).
-.I
-.l 1.2
Thus wc see that for large n the mean number
times in .9
state Sj,
.15of -.05
starting at state SI, differs from lWj by approximately Zij - aj We recall
that by Theorem
mean of the fraction of times in state Sj
Doing4.2.1
this the
we obtain
approaches aj indepe:'ldent of the starting state. Thus the entries of
(Z - A) give us an interesting quantity for regular chains for which the
initial state does have an influence. We can compare two starting
states, since by Corollary 4.3.5
MI[y(n)j] - M k[y(1l)J] ->- ZiJ -
Zkj·
Another interesting corollary to Theorem 4.3.4 is the following.
While the fundamental matrix Z has several properties in eo
4.3.6 COROLLARY.
Let c =matrix,
ZJj.
with a transition
weThen
see from this example that it do
necessarily
have non-negative-rentries.
(Mj[y(n)j]-M,,[y(n)j])
c-1
An example
where the Z matrix turns out to be a very simple
j
is the case of an independent trials process. In this case P = A
as n approaches infinity, independent of 77.
Z = ( I - ( P - A ) )-"I.
Thus for an independent trials .proce
PROOF. ByZCoronary
matrix is4.3.5
the identity matrix.
be the number of times t h a t the process is in state s
LetM;[y'(n)j]
- M o [y(n)5] -+ Zjj - (77Z),.
first n stages, i.e. the initial position plus n - l stages.
Therefore the sum approaches
2:
2:
2: ZjJ - Z~ = c - 1.
17
This corollary has the following interpretation. For a
IV possible
FINITE
MARKOV
CHAINS
Mj[y(")j]
M,[jF(*)j],
Hence
j[jF(*)jJ gives theCHAP.
largest
number of times in sj. The corollary states that the deviation
This corollary
the following
any 77,
this has
maximum,
summed interpretation.
over all states, For
approach
a limit w
Mj[y(n)j] ~ M~[j(n) j].
Hence
l'fIj[y(n) j] gives the largest possible mean
independent of the choice of n.
78
number of times in Sj. The corollary states that the deviations from
I n this section
wewhich
shall study
5 4.4 First
passage
times. approach
this maximum, summed
over
all states,
a limit
is the le
We shall s
si
to
a
state
sj
for
the
first
time.
time
t
o
go
from
a
state
independent of the choice of 77.
the mean of this time is easily obtained from the fundamental
First passage times. In this section we shall study the length of
4.4.1 Si to
DEFINITION.
Forfirst
a regular
the first p
time to go from a state
a state sf for the
time. Markov
We shallchain,
see that
is
a
function
whose
value
is
the
number
of
steps
time
fk
the mean of this time is easily obtained from the fundamental matrix. before e
§ 4.4
sk,for thejirst time after the initial position.
4.4.1 DEFINITION. For a regular Markov chain, the first passage
4.4.2 whose
THEOREM.
any i, Mofl [steps
f k ]isjinite.
time fk is a function
value is For
the number
before entering
Sk for the first time
after the
initial position.
PROOF.
Assume
first that i# k. Form a new Markov ch
The resulting Marko
is a n absorbing Markov chain with a single absorbing state, sk
PROOF. Assume
f k. siForm
a the
newgiven
Markov
meanfirst
timethati
to go from
to s j in
chainchain
is the by
same as th
making state Sktime
into before
an absorbing
state. in The
Markov
absorption
the resulting
new chain.
Thechain
mean time
is an absorbing Marko\· ehain with a single absorbing state, Sk. The
absorption is finite by Theorem 3.2.4.
4.4.2
state.
makingFor
state
THEOREM.
anys ki, into
Mi[fka]nisabsorbing
finite.
mean time to go from 5i to Sj in the given chain is the same as the mean
If i = k , in
then
time before absorption
the new chain. The mean time before
pikMk[fi]
absorption is finite by Theorem 3.2.4.Mt[fi] = pii +
2
k#i
Ifi=k, then
which is finite by the first part of the proof.
lUt[f;] = Pi; +
PikMk[f;J
4.4.3 DEFINITION.
k¢i The mean first passage matrix, denoted
matrix
entries
mi, = M,[fJT
which is finite by is
thethefirst
partwith
of the
proof.
I t then follows from 1.8.9 that, for a n initial vector n, th
4.4.3 DEFINITION. The mean first passage matrix, denoted by M,
first passage times are the components of the vector nM.
is the matrix with entries m l ; = Mj[fj).
Theanmatrix
satisjies11,thethe
equation
It then follows4.4.4
from THEOREM.
1.8.9 that, for
initialM vector
mean
first passage times are the components of the vector 771tI.
.L
T
4.4.4
THEOREM.
The matrix
M satisfies
the by
equation
We calculate
Mi[fj]
taking the mean of the con
PROOF.
means, given the outcome of the first experiment.
This gives
(1)
PROOF. We calculate Mi[fj ] by taking the mean of the conditiona,l
means, given the outcome of the first experiment. This gives
2: Pik(Mk[f + 1) + Pi}
= L PikMk[f + 1
L PiklUklfj] - PijMj[fj] + 1..
l\i;[fj] =
j ]
k¥j
j]
k-:tj
k
That is,
This proves the theorem.
mlj =
)J/kmkj - ptjm}} + 1.
This matrix
is denoted by 2 in later works by the authors.
k
.L
This proves the theorem.
t
This matrix is denoted by M in later works by the authors.
This corollary has the following interpretation. For a
Mj[y(")j]
M,[jF(*)j],MARKOV
Hence CHAINS
j[jF(*)jJ gives the largest 79possible
SEC. 4
REGULAR
number of times in sj. The corollary states that the deviation
4.4.5 THEOREM.
Let a = {ai,
az, ...over
, a r} be
limiting
probability
this maximum,
summed
all the
states,
approach
a limit w
Then mii =ofIjai.
vector for P.independent
the choice of n.
PROOl?
Multiplying
equation
(1)times.
above by
a we
have we shall study the le
I n this
section
5 4.4 First
passage
si to a state sj for the first time. We shall s
time t o go
from
a state-Mdg)+aE
alJ,1
= aP(M
the mean of this
time
is easily obtained from the fundamental
= a(~i}J -Mctg)+aE.
4.4.1
DEFINITION.
For a regular Markov chain, the first p
Therefore
time fk is a function whose value is the number of steps before e
sk,for thejirst time after the initial position.
This states that aim/j = 1 for e,"ery i, or mji = l,'(z".
4.4.2 THEOREM. For any i, M l [ f k ]isjinite.
4.4.6 THEOREM. Eq'<wtion (1) of Theorem 4.4.4 has a unique
PROOF. Assume first that i# k. Form a new Markov ch
soi'ution.
making state s k into a n absorbing state. The resulting Marko
PROOF. Let
and M' beMarkov
two solutions
for a(1).
Then
from the
is aJ1
n absorbing
chain with
single
absorbing
state, sk
proof of Theorem
have
aMsidgto
=aM'ag=7).
Hence
mean4.4.5
time we
to go
from
s j in the given
chainMdg=M'dg.
is the same as th
This gives time before absorption in the new chain. The mean time
- M' by
= P(JJJ
- M').
absorption M
is finite
Theorem
3.2.4.
But this means If
that
column of M - jW is a fixed column vector for
i =each
k , then
P. Hence by Theorem 4.1.7 each Mt[fi]
column=ispii
a+
constant
vector. Since
pikMk[fi]
2
M -.1."[, has 0'8 on the diagonal, these vectors must
all be 0 vectors.
k#i
Hence M =.M'.
which is finite by the first part of the proof.
DEFINITION.
The mean
firstM passage
denoted
The
mean first pa.ssage
matrix
is given matrix,
by
is the matrix
with
entries
mi,
= M,[fJT
(2)
M = (1 - Z +EZdg)D
I t then follows from 1.8.9 that, for a n initial vector n, th
where Dis the
diagonal
matrix
with
diagonal
elements
d/j=
1
la/.
first passage times are the components of the vector nM.
PROOF.
By Theorems
4.4.4 andThe
4.4.6
we need
only show
that M
4.4.4 THEOREM.
matrix
M satisjies
the equation
4.4.7
4.4.3
THEOREM.
as defined by (2) satisfies equation (1) above.
Let
We
Mi[fj] by taking the mean of the con
PROOF.
M calculate
= (1- Z+EZag)D.
means, given the outcome of the first experiment. This gives
Then
M-D = (-Z+EZdg)D
and
P(M-D) = (-PZ+EZdg)D
J1 +(-1 +Z-PZ)D.
By Theorem 4.3.3(d) this is
P(M-D) = M-AD
=
.1.1{ -E.
proves the theorem.
By (2), D=]rlThis
ctg . Hence cvI =.P(M -Mag)+E.
,
This matrix is denoted by 2 in later works by the authors.
4.4.8 THEORE)!.
Lei P be the transition matrix for an indE.pendent
trials process. Then llJ = {lPij}.
FINITE MARKOV CHAINS
80
From Theorem 4.4.7 and the fact that Z = I for
PROOF.
an independ
IV
FINITE
MARKOV
pendent
trials process,
weCHAINS
have M = ED. ForCHAP.
process the limit matrix A = P. Hence pi! = ar and l/ai = l / p
PROOF. From Theorem 4.4.7 and the fact that Z = 1 for an indeM = ED = { l / p t l ) .
pendent trials process,
we illustrate
have M =ED.
For an independent
We now
the calculation
of the mean trials
first passag
process the limit
matrix
A
= P.
Hence
Pij =Gj and l/aj = l/ptj.
to be ( 2 1 5 ,
for the Land of Oz example. We have found. a Thus
M=ED={ljpjj}.
Hence the matrix D is
We now illustrate the calculation of the mean first passage matrix
for the Land of Oz example. We have found. a to be (2/ 5, 1/ 5, 2/5).
Hence the matrix D is
80
o
D
5
The Z matrix was found in $ 4.3. From this, usihg Theor
we obtain M by M =o(I- Z + E Z d g ) D . Carrying out this ca
we obtain
The Z matrix was found in § 4.3. From this, usihg Theorem 4.4.7
R outN thisS calculation
we obtain M by M = (1 - Z + E Zdg)D. Carrying
we obtain
R
RC
N
4
S
f 3
'O! )
8/ s 5 8/ 3 .
Thus, for example, if i t is raining in the Land of Oz today
5/z day is 4. The mean number
number of days
before4 a nice
S 1°/3
1'rf = N
before another rainy day is 5 1 2 ; before a snowy day
Thus, for example,
if it isnext
raining
inathe
Land which
of Oz today
thethe
mean
We shall
prove
theorem
connects
diagonal
number of days before a nice day is 4. The mean number of days
of Z with the mean time to reach sl for the initial probabili
before another rainy day is 5/ 2; before a snowy day 1°/3.
n = a . We have seen previously that the mean number of
We shall nextstate
prove
which
connects
thecase.
diagonal
elements
sj aistheorem
particularly
simple
in this
This
choice of init
of Z with the mean
to reach
Sj for the initial probability vector
is of time
special
importance
for the following reason. Assum
7T = a.
vVe have
seen previously
thathas
the
mean
number
of times
in of ste
regular
Markov chain
gone
through
a large
number
state Sj is particularly
simple
in
this
case.
This
choice
of
initial
vector
it is observed. Then Theorem 4.1.6 suggests that a natural c
is of special importance
for the
following
Assume that
a later
the new initial
vector
n is a. reason.
The probabilities
for any
regular Markov chain has gone through a large number of steps before
then also given by a. In this case we say that the process is
it is observed. Then Theorem 4.1. 6 suggests that a natural choice for
i n equilibrium.
the new initial vector 7T is a. The probabilities for any later time are
then also given by4.4.9
a. InTHEOREM.
this case weFor
sayathat
the Markov
process chain
is observed
regular
in equilibrium.
4.4.9
THEOREM.
Por aMultiplying
regular Markov
chain
(2) by a we have
PROOF.
aJ.lf =
PROOF.
{M.[fj ]} = 7JZdgD = {Zjj/aj}.
Multiplying (2) by cx we have
cxJ.l1 = ct.(I - Z +EZag)D
=
(a-a+7J Z dg)D
aM = 'l}ZdgD.
FINITE MARKOV CHAINS
80
From Theorem 4.4.7 and the fact that Z = I for
PROOF.
REGULAR
MARKOV
CHAINS
M = ED. For an independ
pendent
trials process,
we have
81
process the limit matrix A = P. Hence pi! = ar and l/ai = l / p
4.4.10 THEOREM.
Let
Then MaT=cf
M = ED =
{ l /c=
ptl).
We now illustrate the calculation of the mean first passag
PROOF.
for the Land of Oz example. We have found. a to be ( 2 1 5 ,
T
= (I-Z+EZdg)Da
Hence MaT
the matrix
D is
SEC. 4
2Zi/.
= (1- Z + EZdg)g
= g('7Zd~) = eg.
In § 4.3 we compared the mean number of times y(n), in a state Sj
under the assumption of two different starting distributions. We can
From this, usihg Theor
The
Z matrix for
wasthe
found
in $ 4.3.
make the same
comparison
function
ff.
we obtain M by M = (I- Z + E Z d g ) D . Carrying out this ca
4.4.11 THEOREM.
For any two initial probability vectors 7T and TT'
we obtain
R
N
S
PROOF.
{M.[ffJ-M,,{f,]} = TTM-TT'M
= (TT-TT')(I-Z+EZdg)D.
= (TT-TT')(I-Z)D.
for example,
i t is raining in the Land of Oz today
In the Land ofThus,
Oz example
we see ifthat
number of days before a nice day is 4. The mean number
. snowy day
MN[fs]-MR[fsJ
= day
8/ S _10/
- 2/ 3 a
before
another rainy
is 5 1S2 ;=before
We
shall
next
prove
a
theorem
which
Thus the mean time to the first snowy day is shorterconnects
starting the
withdiagonal
a
of
Z
with
the
mean
time
to
reach
sl
for
the
initial
probabili
nice day than it is starting with a rainy day.
n = a . We have seen previously that the mean number of
We will conclude
this section by showing that the Markov chain is
sj is particularly simple in this case. This choice of init
state
completelY determined by the numbers mij, for i # j. We shall use
is of special importance for the following reason. Assum
these numbers as the non-zero entries of the matrix.M =M -D. This
regular
Markoventries,
chain has
gonesuffices
throughtoa determine
large number
matrix has n(n-l) non-zero
which
the of ste
4.1.6
suggests
that
a
natural c
it
is
observed.
Then
Theorem
chain. When we give the chain in terms of P, we specify n Z entries;
the
new
initial
vector
n
is
a. The probabilities for any later
but these satisfy n relations, since P must have row-sums 1. But
then way
also given
by a. In
this
case we
say that
thewhile
process is
there is no natural
of specifying
just
n(n-l)
entries
of P,
i
n
equilibrium.
fJ. is a natural way of giving this minimum information.
4.4.12
4.4.9 For
THEOREM.
For.Markov
a regular
Markov chain
any regular
chain
THEOREM.
(a) The matrix Iv! has an inverse
Multiplying (2) by a we have
PROOF.
(b) a = (e-I)
(M-lg)T
(c) P = 1+ (D-E)JVJ-l.
PROOF.
From equation (1) in § 4.4.4 we have
!Yf+D = pM +E;
hence
(P-I)M = D-E.
(3)
FINITE MARKOV CHAINS
82
C
If &' has no inverse, then there is a non-zero column vector y su
HenceMARKOV
from ( 3 ) CHAINS
CHAP. IV
M y = 0.FIXITE
82
E ) ~= ( Pcolumn
- I ) Mvector
~= o y such that
If 111 has no inverse, then there( isDa- non-zero
Dy = E y
My=O. Hence from (3)
(D-E)y = (P-I)i1,1y y= =0 D-lEy = D-llvy = (vy)aT,
whereDy
1 ==vyEy
is a number.
y
t.' ~.
~'
And since y # 0 , 1 # 0.
= D-IEy = D-l('1Y = (7)y)a T ,
where l=7)Y is a number.
And since y;fO, l;f 0.
But, clearly,a7'A ?= c ~(lll)y
> 0 and
,
we have a contradiction. Therefore
an inverse.
MaTFormula
= (lll)ICity(c)=isO. then a n immediate consequence
To prove ( b )we make use of $ 4.4.9 and the fact that DaT= [.
But, clearly, iVIa7' > 0, and we have a contradiction. Therefore, 111 has
an inverse. Formula (c) is then an immediate consequence of (3).
To prove (b) we make use of § 4.4.9 and the fad that Da T
=r
(J1 + D)r:t7' = cg
l1,1a7' = (c-l)g
r:t7' find
= (c-l)M-lf
We can now
a from formula ( b ) ,and the condition tha
This determines
D ,(c-l)(l'Ci!-lg)7'.
and then formula (c) will yield P. Thus th
r:t =
is determined by A?.
We can now find a from formula (b), and the condition that ~ = 1.
n the
previous sec
Variance
of the
This determines D,$4.3
and then
formula
(c) first
will passage
yield P. time.
ThusI the
chain
is determined byfound
.M. that the Z matrix enabled us to find the mean first passa
from si to sj. I n this section we shall show that tJhe Z mat
§ 4.5 Variance
of the first
passage
time. Inofthe
section
we
provides
us with
the variance
theprevious
first passage
time.
found that the Z matrix
enabled
us
to
find
the
mean
first
passage
time
We recall that fj is the function whose value'gives the nu
from Si to 5j. In
thisrequired
section to
wereach
shall sjshow
thatfirst
t,hetime
Z matrix
alsoinitial ste
steps
for the
aftmerthe
provides us withhave
the variance
of
the
first
passage
time.
found Rf,[fj]. Hence to find Vari[ff] i t is only necessary
'rVe recall that
fj is the
whose
the number
of We de
Mi[f2?]
and function
use the fact
thatvalue'gives
Vari[fj] = Mi[Pj]
- Mi[fj]2.
steps required toW'reach
Sj for the first time after the initial step.
We
the matrix
W = {Mi[fZj]}.
have found l\i,~fiJ. Hence to find Vari[f)] it is only necessary to find
THEOREM.
T h e matrix TV scctisjles
equation
fact that
Vari[fjJ=lHi[f2j]-l\-IMjJ2.
We the
denote
by
Mi[f2j] and use the4.5.1
TV the matrix TV = {Mi[f2j]}.
4.5.1
THEOREM.
The matrix
W conditional
satisfies the 6quation
Taking
means we have
FROOF.
TV = P[W - Wdg] -
PROOF.
2P[Z-EZdg}D + E.
(1 )
Taking conditional means we have
2: PikMk[(f + 1)2] + PiJ
= 2: PikThI ,,[f2 + 2'2 PikMAf + l.
lVI/[f2j] =
j
k,c}
j]
k:;r,j
j]
k ':1':-j
Or
TV = P[TV - TVdg} + 2P[M -Mdg) + E.
(2)
FINITE MARKOV CHAINS
82
If &' has no inverse, then there is a non-zero column vector y su
83
from ( 3 ) CHAlKS
M y =REGULAR
0. Hence l\gRKOV
SEC. 5
( D -J·J-Mdg=(-Z+EZdg)D.
E ) ~= ( P - I ) M ~= o
From Theorem 4.4.7 we have
this in (2) we have (1).
Dy = E y
4.5.2
C
THEOREM.
Putting
=i ] D-lEy
= D-llvy
= (vy)aT,
The values for Myi [f2
are given
by
where 1 = vy
1 ). since y # 0 , 1 # 0.
Wdgis =a number.
D( 2Z dg D - And
(3)
PROOF.
Multiplying equation (1) through by a and using the fact
that aP=a, we have
But,= clearly,
A dg]
? c ~- >2a[Z0 and
,
we
a contradiction.
Therefore
EZhave
aW
a[W- W
dg]D +
7];
(4)
an inverse. Formula (c) is then a n immediate consequence
or, since aZ = a,To
and
aD =( baE
= '/,make use of $ 4.4.9 and the fact that DaT= [.
prove
)we
This gives
Cf.in'U
= - 1 + 2zii /ai
or
We can now find a from formula ( b ) ,and the condition tha
This determines D , and then formula (c) will yield P. Thus th
Written in matrix form, this is (3).
is determined by A?.
Theu.nique solution to (1) is
$4.3 Variance of the first passage time. I n the previous sec
found
that the Z matrix enabled us to find the mean first passa
W
= },J(2ZdgD-1)+2(ZM-E(ZM)dg).
from si to sj. I n this section we shall show that tJhe Z mat
PROOF.
The provides
uniqueness
proofthe
is variance
the sameof "s
given for
the
us with
thethat
first passage
time.
matrix JI in Theorem
It isfjthen
onlyfunction
a matterwhose
of verifying
that the nu
We 4.4.6.
recall that
is the
value'gives
the expression given
W s"tisfies
(l). sj
We
of this.
stepsfor
required
to reach
foromit
the the
firstdetails
time aftmer
the initial ste
From the matrix
it is an
easy matter
to find Vari[ff]
the {Var;[fjJ}.
Wenecessary
haveWfound
Rf,[fj].
Hence to
i t is only
denote by M2={VarilfjJ}.
;JI fact
-111Vari[fj]
Mi[f2?] and Then
use the
= Mi[Pj] - Mi[fj]2. We de
z = Wthat
sQ '
Let us find these
variances
Land of Oz example. \Ve have
W' the
matrix for
W =the
{Mi[fZj]}.
previously found Jl, D, and Z for this example, so that to find W, the
4.5.1
T hise matrix TV scctisjles the equation
only new matrix we
needTHEOREM.
is Z],I. This
4.5.3
THEOREM.
176 1 / 3
FROOF.
Z ·/'11 =
303
259 2
/3)
Taking
conditional means we have
(
1/75
203'
363
203
/3
303
176 1 / 3
259 2
.
From the formula TV = M(2ZdgD -1) + 2(ZM - E(ZM)dg) we find
W=
(:~:
54'~
::
28
:~:),
71/ 6
F I N I T E MARKOV CHAINS
84
and subtracting Maqfrom this we obtain
84
FINITE MARKOV CHAINS
CHAP.
IV
and subtracting 1Il.q from this we obtain
67/12
12
62/9)
( 56/ 9the variance
We observe
in this example depe
J.V z = that
12 56/ Vari[fj]
.
9
67~12
little on the choice of the starting state si. The first passage t
12
9
regular chain are62/quite
similar to absorption times for an
They
bothinhave
which are
in gen
We observeMarkov
that thechain.
variance
Vari[fil
this variances
example depends
very
compared
their means.
little on the choice
of theto
starting
state Si. The first passage times for a
and for M ztimes
are very
simplified fo
Thequite
formulas
fortoW absorption
regular chain are
similar
for much
an absorbing
independent
process.
Markov chain.of an
They
both havetrials
variances
which are in general large
compared to their
means.
4.5.4
THEOREM.For an independent trials process
The formulas for Wand for jl-1 2 are very much simplified for the case
= ((l/ptj)(d/p<j-l)j
of an independent trials process.If' = E D ( " - I )
and
4.5.4 THEORE)'1. For an independent
trials
process
M Z = E(D219)
= { ( I / I ~ ? ) ~I/~A,}.
IV = ED(2D-I)
= ((l/pi})(2/PIJ-l»)
We recall that
for an independent trials process
PROOF.
and
identity matrix and ill = ED. Thus, using Theorem 4.5.3,
Mz = E(D2_D) = {(l/Pij)2-I/piJ ).
PROOF.
We recall that for an independent trials process Z is the
identity matrix and j'.1 = ED. Thus, using Theorem 4.5.3,
W = ED("2D-I)+2(ED-ED)
= ED(2D-J).
From this we obtain
lvl2 = W -M. Q
The alternative
expressions for W and M z given in the sta
= 2ED2_ED-(ED)sQ
the theorem follow from the fact that pij =l)lj for all i.
= E(D2_D).
8 4.6 Limiting covariance. Let f and g be two functions
The alternative
Wand
.M 2 given
in =
the
statement
of L
Let f(si)
f i and
g(si)= 9%.
theexpressions
states of a for
regular
chain.
the theorem follow
from
the
fact
that
Pij = pjj for all i.
g(nJbe the values of these functions on the n-th step. We are
in finding
§ 4.6 Limiting
covariance. Let f and g be two functions defined on
the states of a regular chain. Let f(s;) = ii and g(sd = gi. Let f(n) and
gin) be the values ofthe~e functions on the n-th step.
We are interested
in finding
I t can be shown that this limit exists and is independent of n.
" f(kJ, 1':;1
" g(k) ] •
}~~ 1 Cov~ [ k~1
n
It can be shown that this limit exists and is independen~ of 1T.
4.6.1
THEOREM.
.2
lim - 1 Cov" [ "
H-+OO
n
k-l
f(k),
1.2~
1 g(k)
"
]
.
.2 iiciJgj
i, j = 1
F I N I T E MARKOV CHAINS
84
and subtracting Maqfrom this we obtain
SEC. G
REGULAR MARKOV CHAINS
85
where
eli
= atZij + aJZj; - aid ij - aial'
PROOF.
We shall assume
the result
thatvariance
the limitVari[fj]
is independent
of
We observe
that the
in this example
depe
and prove the theorem
case of
7r =
a. starting state si. The first passage t
little onfor
thethe
choice
the
regular chain are quite similar to absorption times for an
Markov chain. They both have variances which are in gen
compared to their means.
and
The formulas for W and for M z are very much simplified fo
of an independent trials process.
7r
4.5.4
Hence,
THEOREM.For an independent trials process
If' = E D ( " - I )
= ((l/ptj)(d/p<j-l)j
and
M Z = E(D2- 19)= { ( I / I ~ ? ) ~I/~A,}.
We recall that for an independent trials process
identity matrix and ill = ED. Thus, using Theorem 4.5.3,
PROOF.
"
n
r
L LThe(Pro[u(k)i
= 1/\
= 1Jlig'j - Pra[u(k)j = l]j;ajg'j
alternative expressions for W and M z given in the sta
U(l)j
<,,1=1 ;,j=1
the theorem follow from the fact that pij =l)lj for all i.
n
l'
8 4.6 Limiting covariance. Let f and g be two functions
'r
(1) = 9%. L
L Lthe(Pro[u(k)i=
= lJfi9'i
Let f(si)= and g(si)
states of 1a ;\u(l)j
regular
chain.- aiad/gi)'
fi
... k,l= 1 1,)=1
Now
g(nJbe the values of these functions on the n-th step.
in finding
if
Ie < I
if
Ie > I
We are
(2)
k = l.
I t can be shown that this limit exists and is independent of n.
f(k),
2:'.' g(I)J'1
1=1
r
+ L (aid;! - aiaf)!igJ'
i,j= t
F I N I T E MARKOV CHAINS
86
Collecting terms with the same d = 11- kl, we have
FINITE MARKOV CHAINS
86
CHAP. IV
Collecting terms with the same d= Il-kl, we have
-1 COva [ n £(.1:),
n
2:
2:n] =
k=l
1=1
g(l)
r
m
Since Z =
2+ (2:P - (a,dfj-afaj)!,gj.
converges, it is Cesaro-summable (see §
d=O
A)d
i,;=l
<X)
Since Z =
that is
L: (P-A)d
converges, it is Cesaro-summable
,-d (see § 1.10),
2 -( P- A)&.
a-1
Z = Iim
d~O
n+ m d = ~
that is
Hence
L: n-d
- n (P-A)d.
,,-1
Z = lim
n _ ct:I
d-O
Hence
"i,1
Z-I = lim
n-d (Pd-A),
"'-+- co d=1 n
,,-1
Then from ( 3 )
Zfj - d'j = lim
L:
n-+co d=l
r:.
Then from (3)
i
lim .!. COV.[
f(.I:) ,
n
k-l
n .... oo
i g(l)]
l-l
r
L: [adlgj(zIJ =-d2jj ) +fi(atzij
aJ!tgj(zj,+ (a,dlj
+ W I Idlj)
. -a&
- - aja'J)!,gj]
aiaj)gj.
i,;-l
i,j=l
r
=
completes the proof.
L: This
!,(a,z,j+ajzJj-atd'J-ataJ)gj.
I f f and g are the same function, then the above theorem gives u
This completes the4.6.2
proof.COROLLARY.
If f and g are the same function, then the above theorem gives us
1
lirn - ~ a r . [ W]
~
=
ftci~.
4.6.2 COROLLARY.
n - t m n.
i,3=1
i,;=l
2
We shall
a slight
1 need
[n
] extension
r of this last result. Suppose
"li.:n",simply
Var"a function
k~ f(k) =of i,"f:./,Cjj/J.
is not
the state, but f = 1 with probability
state si and 0 with probability 1- f*. We may think of f as
We shall need a slight extension of this last result. Suppose that f
mined as follows: We carry out the Markov chain, and if the pro
is not simply a infunction
the step
state,
1 with coin
probability
It onfi for hea
st on a of
given
webut
flipf=
a biased
(probability
state Sf and 0 with probability 1-k
We may think of f as determined as follows: We carry out the Markov chain, and if the process is
in 5, on a given step we flip a biased coin (probability /1 for heads) to
n
F I N I T E MARKOV CHAINS
86
Collecting terms with the same d = 11- kl, we have
REGULAR MARKOV CHAINS
SEC. 6
87
determine whether f is 1 or O. Again, f(lI) is the value of the function
on the n-th step. Then the above argument for the limiting variance
applies except that in (1) a slight change must be made when k=l.
Here, the termf2 j should be simply fl> and we have a correction term
,
L: aifl(l-f,)·
;-1
Thus we have proved
m
Since Z =
( P - A ) d converges, it is Cesaro-summable (see §
4.6.3 THEOREM. If fd=O
is a function that takes on the value 1 with
probability f.that
in s;,
is and is 0 otherwise, then
a - 1 ,-d
-ft}·
( P - A)&.
Z = Iim r
lim -1 Var"
f(le) =
f,C/lft
+
n + m d =aifj(l~
n-POO
n
.1:-1
i, j=1
i=l
Hence
We can also extend the result to two such functions. If these take
on their values independently of each other, then the proof of 4.6.1
applies exactly. This proves
2
[n2:
]
2:2
2:r
4.6.4 THEOREM. If f and g are functions that take on the value 1
in St with probabilities
Then from (ft3 ) and gj, respectively, independently of each
other, and if the functions are 0 otherwise, then
lim -1 Cov"
n~ 00 n
[n2:
f(k) ,
k=l
2:
n
gIl)
1=1
]
=
2:r f,CtJgj.
i,j=l
One application of the covariance is to obtain correlation coefficients.
Let f and g be as in Theorem 4.6.1. Then
=
fi(atzij + W I I . -a& - aiaj)gj.
4.6.5 DEFINITION. i , j = l
2
cov,,[i
This completes the proof.
f(k), .Ig(Z)]
I f f and g are the same function,
k-1then the
1=1 above theorem gives u
4.6.2 COROLLARY.
2
1
lirn - ~ a r . [ W]
~
=
ftci~.
n - t m n.
i,3=1
Dividing numerator and denominator of the right side by n, and using
shall need a slight extension of this last result. Suppose
4.6.1 and 4.6.2, We
we have
is not simply a function of the state, but f = 1 with probability
4.6.6 THEOREM.
state si and 0 with probability 1- f*. We may think of f as
mined as follows: We carry out the Markov
f,ctjgJ chain, and if the pro
step we flip a biased
in st on af(k)given
i.j=l coin (probability fi for hea
lim
,
n -t> !Xl
corr,,[ i
,\:=1
i gO)]
1=1
i
Another important application of Theorem 4.6.1 is t h
Let A FINITE
and R be
any two CHAINS
sets of states. Let yCHAP.
@ ) and
~IV y (
MARKOV
tively the number of times in set A in the first n steps and
Another important
application
Theorem
4.6.1 isFor
thethese
following:
of times
in set B inofthe
first n steps.
functions
Let A and B be following
any two sets
of
states.
Let
yin) A and yin) B bc respectheorem.
88
tively the number of times in set A in the first n steps and the number
of times in set B in the first n steps. For these functions we have the
following theorem.
4.6.7
THEOREM.
.
1
2:
hmPROOF.
- Cov,,[y(n)
Cjf' is 1 on the states o
Let A,f beyin)
a B]
function which
n ..... 00 n
PROOF.
A
all other states. Let g be asf ininfunction
which is 1 on state
B
otherwise.
Let f be a function which is 1 on the states of A and 0 on
all other states.
otherwise.
Sj
Let g be a function which is 1 on states of Band 0
n
Hence the
theorem
follows from
4.6.1.
" f(k)
and yin) B
g(l).
Fromk~lthis theorem we see that
1=1
yin) A =
2:
2:
Hence the theorem follows from 4.6.1.
From this theorem we see that
4.6.8
COROLLARY.
2: CiJ
sjin A
lim Corr,,[y(n)A, y(n)B]
si in B
Taking A and B in 4.6.7
t o beSi~B
setsCliwith a single ele
Si~AClj'
Sj in A
sJ in B
covariance
for the numbe
that cij represents the limiting
states i and j in the first n steps. The values of cii give
Taking A andvariances
Bin 4.6.7
sets with
a single
element,
we are
see oft
for to
thebenumber
of times
in state
st. We
that Clj represents
the
covariance
the them
number
only
in limiting
these variances,
we for
denote
by of
thetimes
vectorin,I3.
n ..... '"
J
states i and j in the first n steps. The values of Cii give the limiting
the in
number
of times
in sgoften
and interested
sf is variances for thecorrelation
number offor
times
state St.
\Ve are
dTcC13'
only in these variances, we denote them by the vector {3. The limiting
correlation is 1.
. f or t h CHum
b erindependent
. 8ttrials
'IS • ; _
C/j
corre1atlOn
0 f" tllnes m
and
Sj
Hi-=j,
For an
process
cij_=. aidll
aiaj,the
and
Cit· Cjf
formulas found above simplify. v For
example, if i f j ,
correlation is 1. correlation is
For an independent trials process CIj = aid!j - ala" and hence all the
formulas found above simplify. For example, if i '" j, the limiting
correlation is
J
The diagonal
entries of C , i.e. the
limiting variances, have
- a l Cl f
ajaj
important use. Let
yaf(l-a;)aj(l-al)
= ,I3- = { b(l-at)(l-aj)'
f )= {cfj). Then ,6 is a vector wh
limiting va.riances for the number of times in each s
The diagonal entries of 0, i.e. the limiting variances, have the following
important use. Let {3 = {bj} = {Ci'}' Then {3 is a vector which gi'Q'es the
limiting variances for the number of times in each state. These
Another important application of Theorem 4.6.1 is t h
Let
A and R be
any twoCHAINS
sets of states. Let y @ )89
and
~ y(
REGULAR
MARKOV
tively the number of times in set A in the first n steps and
variances appear in
following
important
theorem
(called
Central
of the
times
in set B
in the first
n steps.
Forthe
these
functions
Limits Theorem for
Markov
Chains).
following theorem.
4.6.9 THEOREM. For an ergodic chain, let y<n)j be the number of
times in state Sj in the first n steps and let a = {aj} and f3 = {b1} respectively be the fixed vector and the vector of limiting variances. Then
if bJ'I=O, for any numbers r< s,
SEC. 6
PROOF.
Let f be a function which is 1 on the states o
all other states. Let g be a function which is 1 on state
otherwise.
as n-+(jJ, for any
choice of starting state k.
The proof of this theorem is beyond the scope of this book and appears
only in the more advanced books on probability theory. However, for
a discussion of this theorem in the case of independent trials processes
Hence the theorem follows from 4.6.1.
see FMS Chapter 3. It is not possible to evaluate the integral in this
From this theorem we see that
theorem exactly, but for illustmtive purposes we mention that the
value for r = - 1. and.5 = 1 is approximately .681, for r= - 2 and s= 2 it
is approximately. 954, and for r = - 3 and s = 3 it is .997.
EXAM:FLE.
Let us consider the Land of Oz example.
example we have found,
For this
and
R B in
N 4.6.7
S t o be sets with a single ele
Taking A and
limiting covariance for the numbe
that cij represents
I 86 the
3
n steps. The values of cii give
states i and j in the first-14\)
6
63
. in state st. We are oft
J!
75
(
variances
z= for the number of 6times
only in theseI\ variances,
we
denote
3
So} them by the vector ,I3.
-14
correlation for the number of times in sg and sf is -
This is all the information we need to compute the matrix C = {elj}.
dTcC13'
Carrying out this computation
correlation iswe
1. find
For an independent trials process cij = aidll - aiaj, and
It
N
S
formulas found above simplify. For example, if i f j ,
!
134
J
. ...
-18
R
correlation is
C=
>!",(
18
36
-18
"6)
N.
-116
-18
134
S
The diagonal
entries
of limiting
C , i.e. theva;'iance
limiting /3={134/
variances,
The diagonal entries
of C give
us the
375 ,have
Then
,6 is aLimit
vector wh
important
use.of the
36/ 375 , 134/ 375 } for being
in each
Central
= { b f )=Thus
{cfj). the
Letstates.
,I3
limiting
va.riances
Theorem would say,
for example,
thatfor the number of times in each s
y<n)N-n/5
V (36/s 75 )n
FINITE MARKOV CHAINS
90
C
would for large n have approximately a normal distribution.
FI~ITE
MARKOV
CHAP.
IVwhic
this we may
estimate
that theCHAINS
number of days in 375
days
be nice would be unlikely (probability about .046) to deviate
would for large
n have
a normal distribution. From
by more
thanapproximately
12.
this we may estimate
that
the
number
of
days only
in 375
would
Assume that we are interested
indays
bad which
weather
or only
be nice would
be unlikely (probability about .046) to deviate from 75
weather.
Then we would want t o consider the number of ti
by more than
12. is in the set A1 = (R,S} and the number of times i t is in
process
Assume that we are interested only in bad weather or only in good
A2= {N). Let ti, be the limiting covariance for the number
weather. Then
would
to consider
of times
in setweA*
and want
A,. Then
from the
4.6.7number
we know
that the
we can
process is in&
the{tii}
set by
Al =
{R,
S}
and
the
number
of
times
it
is
in
the set
simply adding elements of C. For example,
A2 = {N}. Let Gtj be the limiting covariance for the number of times
in set Ai and A j • Then from 4.6.7 we know that we can obtain
(j = {Cti} by simply adding elements of C. For example,
90
C12
= CRN+CSN = _18/375-18/375 =
Al
= _12/125'
_36/ 375
A2
125 ).
_12/
It is easily verified that the
row-sums
of C must be 0. Si
For a 2 x 2 ma
symmetric, the column-sums must
12/ 125also be 0.
tells us t h a t the entries must all have the same absolute value
It is easily
that the
C must
n. we
Since
C is
weverified
would expect
6 torow-sums
have the of
special
formbethat
found.
symmetric, the column-sums must also be O. For a 2 x 2 matrix this
5 4.7
Comparison
twothe
examples.
I n thisvalue.
section we
shall c
tells us that the
entries
must all of
have
same absolute
Thus
the basic
quantities
for form
two regular
we would expect
C to have
the special
that we Markov
found. chains with th
limiting vector.
One will be an independent trials example
§ 4.7 Comparison
examples. example.
In this section
we shall
compare
oGher willof
betwo
a dependent
The two
examples
are the
the basic quantities
for
two
regular
Markov
chains
with
the
same
walk Example 3 of Chapter 11 with ~ = l (denoted
/ ~
by Exam
limiting vector.
One will7 be
an independent
trials for
example
and3athe
. The
transition matrix
Example
is
and Example
other will be a dependent example. The two examples are the random
walk Example 3 of Chapter II with p = 1/2 (denoted by Example 3a)
and Example 7. The transition matrix for Example 3a is
S1
Sa
S4
0
1
0
0
1/ 2
0
82
"C
82
liz
55
0)
0
o
.
0
liz
P = 8a
0 1/2
0
0 for 1/Example
84
0
2
The transition
matrix
7 is
1~2
0
0
85
0
The transition matrix for Example 7 is
Sl
s3
84
.2
.4
.2
.2
.2
.4
.2
.2
.2
.2
.4
S2
S5
"C )
82
.1
P = 83
.1
84
.1
S5
.1
.4
.4
.2
.2
.1
.1
.1
.1
.
FINITE MARKOV CHAINS
90
C
would for large n have approximately a normal distribution.
this REGULAR
we may estimate
that the
number of days in 375 91
days whic
MARKOV
CHAINS
be nice would be unlikely (probability about .046) to deviate
The limiting vector
forthan
each12.
of these chains is a= (.1, .2, .4, .2, .1).
by more
Thus by the Law Assume
of Largethat
Numbers
can expect
ill in
each
about or only
we areweinterested
only
badcase
weather
.1 of the steps toweather.
be to 51, about
82, etc.want t o consider the number of ti
Then .2
wetowould
For a regularprocess
Markov
chain
theA1
fundamental
is given
by i t is in
is in
the set
= (R,S} and matrix
the number
of times
Z = (1 - P + A )-1.
If
this
is
computed
for
Example
3a
we
0
btain
A2= {N). Let ti, be the limiting covariance for the number
in set A*
and 82
A,. Then
from
4.6.7
we know that we can
Ss
81
S4
85
& {tii}by simply
adding
elements
of
C.
For example,
-.04 '0')
SEC. 7
.12
-.14
-
.33
.88
.02
.86
Z = S3
.16
.72
.16
S4
-
.17
-.14
. 12
.86
-")
S5
-
.12
-.04
.32
-.04
.8&
82
"
(
.,,~
-.04
- .17
- .02
.
.33
For an independent
is the
matrix.
Hence
It is trials
easilyprocess,
verified Zthat
theidentity
row-sums
of C must
be 0. Si
for Example 7, symmetric,
Z = I.
the column-sums must also be 0. For a 2 x 2 ma
The first information
obtain
frommust
Z relates
thesame
number
of value
tells us t hwe
a t the
entries
all havetothe
absolute
times in a statewe
in would
the first
n steps.
Let y(n)j
be the form
number
expect
6 to have
the special
thatof
wetimes
found.
in state Sf in the first n steps (counting the initial state). Then by
5 4.7 Comparison of two examples. I n this section we shall c
Theorem 4.3.4
the basic quantities for two regular Markov chains with th
M,[y(n)j] - naj ---+ Zij - aj.
limiting vector. One will be an independent trials example
For the independent
case
limit is replaced
equality.
In the are the
oGher will bethis
a dependent
example.by The
two examples
dependent case,walk
Zlj gives
us
a
comparison
of
Mt[y(n)jJ
for
fixed
Example 3 of Chapter 11 with ~ = l (denoted
/ ~8j and
by Exam
different starting
Sj.
For
example,
in Example
3a, Example
ZI1 > Z21 >3a is
7 . The
transition
matrix for
andstates
Example
Z31 > Z51 > Z41.
Thus for large n
M 1[y(1l) d > M 2[y(n) 1J > M3 [y(n) 1] > l't1 5[y(n) 1] > M 4[y(n) 1].
The fact that the process may be expected to be more often in 81
starting from S5 than from S4 may be seen also from the fact that to
reach· Sl from either of these states it is necessary to go through S3.
From Ss t!1e first step is to S3 while from 54 it is either to S3 or to Ss.
We can also find from Z the limiting variance for y(n)j/vn. This
is given by
The. transition matrix for Example 7 is
f = {aj(2z jj -l-aj)}.
In Example 3a this gives
f3 = (.066, .104, .016, .104, .066).
For the independent trials process f3 = {aj( 1- aj)}.
Thus for Example 7,
f3 = (.09, .16, .24, .16, .09).
The variance in the independent case is in each case larger than the
corresponding variance in the dependent case. The variance for S3
is much larger. This means that we can make more accnrate predictions about y(n)3 in Example 3a than in Example 7. For example, the
92
FINITE MARKOV CHAINS
CHA
Central Limit Theorem tells us that in 1000 steps the number of o
FI::\ITE
lItIARKOV
CHAINS
CHAP. IVnot de
rences of SQ
in Example
3a will,
with probabilityz.95,
from 400 by more than 22/1000. ,016 = 8. I n Example 7 we c
Central Limit Theorem tells us that in 1000 steps the number of occurwith the same probability, only say that the number of occurrenc
rences of S3 in Example 3a will, with probability~ .95, not deviate
sa would deviate from 400 by less than 22/-4=
31.
from 400 by more than 2 Y lOOO·.016=8. In Example 7 we could,
Let
us
next
compute
the
covariance
matrix
and
some
correlat
with the same probability, only say that the number of occurrences of
For
Example
3a
we
have
S3 would deviate from 400 by less than 2
lOOO· .24~31.
92
Y
Let us next compute the covariance matrix and some correlations.
For Example 3a we have
(
C =
_03)
.042
- .016
- .058
.066
.042
.104
.008
- .096
-.058
-.016
.008
.016
. 008
- .016
.
.008
.104s l and.042
For the-.058
limiting-.096
correlation
between
each of the five state
have (rounded):
(1.00,
- .49, .70, - .066
.52).
-.034
.058 .51,
-.016
.042
For Example 7 the covariance matrix is
For the limiting correlation bet\veen 81 a.nd each of the five states we
have (rounded): (1.00, .51, - .49, - .70, - .52).
For Example 7 the covariance matrix is
C=
(-::
- .04
_0)
- .02
-.04
- .02
.16
-.08
-.04
-.02
-.08
.24
-.08
-.04
.
-.02correlations
. 16 (1.00,
- .02 - .IT, - .28, - .17, -.04 -.08
The limiting
with sl are:
I t is to-.01
be expected
that
often
the
limiting
-.02 -.04 -.02
.00 correlations between
different states will be negative, since-generally-the
more ofte
The limiting
correlations
withstate,
S1 are: (1.00, - .17, - .28, - .17, - .11).
process
enters one
the less often it will be in the other state.
It is to be
expected that often the limiting correlations between two
the independent process all the correlations between pairs of diff
different states will be negative, since-generally-the more often the
states are negative, though quite small. But for Example 3a
process enters one state, the less often it 'will be in the other state. For
correlations are fairly large, and the correlation between s l and
the independent process all the correlations between pairs of different
positive.
states are negative, though quite small. But for Example 3a the
We next consider the function f j which gives the number of
correlations are fairly large, and the correlation between S1 alld 8'2 is
taken to reach sj for the first time. The values of Mi[fj] are give
positive.
the matrix M = (I- Z + EZd,)D. For Example 3a this is
We next consider the function fj which gives the number of steps
taken to reach Sj for the first time. The values of M;[fj] are given by
the matrix M = (1- Z +EZdg)D. For Example 3a this is
S1
"C
S2
S3
S4
4.5
1
4.5
S5
5.5
5
1.5
5
10.5 )
10
9
3.5
2.5
3.5
9
S4
10.5
5
1.5
5
5.5
85
10
4.5
1
4.5
10
82
;,VI = sa
.
92
FINITE MARKOV CHAINS
CHA
Central Limit Theorem tells us that in 1000 steps the number of o
REGULAR
MARKOV CHAINS
93 not de
rences of
SQ in Example 3a will, with probabilityz.95,
8. M
I n reduces
Exampleto7 we c
from 400 bytrials
moreprocess
than 22/1000.
,016 =for
For an independent
the formula
with for
the Example
same probability,
only say that the number of occurrenc
M=ED. Thus
7
sa would deviate from 400 by less than 22/-4=
31.
S1
S2 the
S3 covariance
84
S5
Let us next compute
matrix and some correlat
For Example 3a we have
5 2.5 5
82
10 5 2.5 5 10
.
10 5 2.5 5 10
M = S3
S4
10 5 2.5 5 10
Ss
10 5 2.5 5 10
SEC. 7
1)
"C
In the independent case the mean time to reach S1 is independent of
the starting For
state.
This is not
true for between
the dependent
fact,
the limiting
correlation
s l and case.
each ofInthe
five state
the mean time
to reach
S2 is .70,
only-about
haverequired
(rounded):
(1.00, 81
.51,from
- .49,
.52). half that
required for any
starting
We matrix
observeis that the mean
For other
Example
7 the state.
covariance
time to return to a state, Mt[fl ], is the same in the two examples.
This is because these means depend only on Ct.
The variances Varj[f1] are given by
M2 = M(2ZdgD-1)+2(ZM-E(ZM)dg)-Msq.
For Example 3a this is
S2
S1
S3
84
15/ 4
20
20
20
20
20
S5
The limiting correlations with sl are: (1.00, - .IT, - .28, - .17, I t is to be expected
/ 4 limiting correlations between
12 3the
12 3/4that 0often
different
will
be
negative,
since-generally-the
more ofte
"S2 ( states
66
1
1 /4 )
13
13
1/
53 /4
4
6666
process enters one state,
the
less
often
it
will
be
in
the
other
state.
Ma = S3
66
66
.
12 3 /4 1/ 4 12 3 / 4
the independent process all the correlations between pairs of diff
13though
1/ 4 quite
13 small.
66 1/ 4
53 1/4 But for Example 3a
states54are negative,
0 and
55
66 fairly
66
correlations
are
correlation
between s l and
12 3 /4large,
12 3 /the
4
positive.
For the independent
trials case
formula
for .M
2 reduces to
We next consider
the the
function
f j which
gives
the number of
Thus
for Example
we have
J.1Iz=E(DLD).
sj for the 7first
time. The values of Mi[fj] are give
taken to
reach
the matrix M = (I
+ EZd,)D.
81- ZS2
Sa
54 For
55 Example 3a this is
51
S2
JJ.f 2 = 53
54
55
CO
90
90
90
90
20
20
20
20
20
15/ 4
15/ 4
15/ 4
15/ 4
9)
90
90
90
90
.
As in the case of the means, in the independent case, Example 7, the
variances do not depend on the starting state. Unlike the case of the
means this is almost true for the variances in the dependent case
Example 3a. We note finally that, as in the case of the varian
y(n)j, the variances
fj are inCHAINS
each case greater for
theIV
indep
FINITE for
l\IARKOV
CHAP.
case than for the dependent.
Example 3a. \Ve note finally that, as in the case of the variances for
jj 4.8 The
y(n)" the variances
for fJgeneral
are in two-state
each casecase.
greater
independent
I nfor
thisthe
section
we give for
reference
the basic quantities for Example 11 of Chapter I
case than for
the dependent.
recall that this was the general Markov chain with two states
§ 4.8 The
general matrix
two-state
case.
In this
section
transition
was
written
in the
form we give for future
reference the basic quantities for Example 11 of Chapter II. '\Ve
recall that this was the general Markov chain with two states. The
transition matrix was written in the form
94
I-C
We assume that
p=O <( c < l and O < d < 1 but c and d are not b
This will give us the general
regular two-state Markov chain.
d
The limiting
is 1 but c and d are not both 1.
We assume that
0< c"; 1 vector
and 0<a d,.;
This will give us the general regular two-state Markov chain.
The limiting vector a is
. de) + A)-' is
The fundamental matrix Z = (I- P
a = (c+d' c+d .
The fundamental matrix Z = (I - P + A )-1 is
1
z=-
C
d+_
c+d
'-':d).
(
The meanc+d
first passaged matrix itl is
c+-d-c+d
c+d
The mean first passage matrix 11-1 is
M=
and the variance matrix for the first passage time is
and the variance matrix for the first passage time is
(
C(2-d~-d)
.112 =
I-d for the number of times in state sj is g
The limiting variance
d2
The limiting variance for the number of times in state Sf is given by
{3 = (Cd(2-C-d),
(e + d)3
Cd('2-e-d l ).
(c +d)3
Example 3a. We note finally that, as in the case of the varian
y(n)j, the
variancesMARKOV
for fj are CHAINS
in each case greater for the
REGULAR
95 indep
case than for the dependent.
Compare this variance for the dependent case with the independent
4.8 limiting
The general
two-state
case. I n thisprocess
sectionwould
we give for
case having the jjsame
vector.
This independent
reference
have transition
mc.t.rix the basic quantities for Example 11 of Chapter I
recall that this was the general Markov chain with two states
transition matrix was written in the form
C!d
p= (
SEC. 8
':d).
d
c+d
c+d
We assume that O < c < l and O < d < 1 but c and d are not b
and the limiting variance for the number of times in Sj would be
This will give us the general regular two-state Markov chain.
The limiting vector a is ed )
(C+d)2 .
Thus the limiting variance for the number oHimes in S1 will be greater
in the dependent case if and only if
The fundamental matrix Z = (I- P A)-' is
2-c-d > c+d.
+
That is, if the sum of the diagonal elements is greater than the sum of
the off-diagonal elements. Or in other words, if the probabilities for
remaining in a state have a sum greater than the probabilities for a
change of state.
The covariance matrix is
The mean first passage matrix itl is
c= cd(2-c-d)
(C+d)3
-1
1
( 1 -1).
Thus eii > 0 if i =j, but elj < 0 if i 0# j. The limiting correlations are
+ 1 and - 1 in the two cases, respectively.
Exercises
for Chapter
and the variance
matrix
for theIV
first passage time is
FOT § 4.1
1. Find the lilniting r!1atrix A for Example 13. (See Exercise 23,
Chapter II.)
2. Find the limitin"b matrix A for Example 14. (See Exercise 24,
Chapter 11.)
3. Show that the four-state chain in Example 12 is regular. Find the
The limiting
variance
forvector
the number
of times
infor
state
fixed vector a. vVhat
is the relation
of this
to the fixed
vector
the sj is g
two-state chain which determiDed the four-state chain?
4. Show that if a is the fixed probability vector for a chain with transition
matrix P, then it is also a fixed vector for the chain with transition matrix
pn.
5. Prove that if a transition matrix has column sums 1, then the fixed
vector has equal components.
96
6. Given a probability vector a with positive components, dete
regular transition
matrix
which will
have this as its fixed
vector.
.FINITE
MARKOV
CHAINS
CHAP.
IV
6. Given a probability vector a with positive
determine a
4.2
For $components,
regular transition matrix which will have this as its fixed vector.
7. For Example 14 find the mean and variance for the number of
state s1 in the first n steps.
For § 4.2
8. Consider the Markov chaiil with transition matrix
7. For Example 14 find the mean and variance for the number of timcs in
state 81 in the first n steps.
8. Consider the Markov chain with transition matrix
81
$2
(0 1)
Start the process in s2, and compute the mean of v(*)l for n = l , 2, 3
P = 81with al.
Compare these results
82
1/2 1/ .
2
Start the process in 82, and compute the meanFor
of V(n)1
$ 4.3 for n = 1, 2, 3, 4, 5, 6.
Compare these results with al.
9. Find the fundamental matrix for Example 1 1 when c = and d
10. Find the limit For
of the
§ 4.3difference between t8he mean number
days in the Land of Oz in the first n days, starting with a rainy d
9. Find the
fundamental
for Example II when c=l/z andd= 1/ 4.
starting
with a matrix
nice day.
10. Find the11.limit
the fundamental
difference between
number
nice 8
Findofthe
matrix the
for mean
the chain
in of
Exercise
days in the Interpret
Land of Oz
in the first n days, starting with a rainy day and
211 -221.
starting with a12.
nice
day.the fundamental matrix for Example 14.
Find
(Use the r
11. Find Exercise
the fundamental
2 above.) matrix for the chain in Exercise 8 above.
Interpret Zn-Z21'
13. Find the fundamental matrix for Example 13. (Use the r
12. Find Exercise
the fundamental
1 above.) matrix for Example 14. (Use the result of
Exercise 2 above.)
13. Find the fundamental matrix for Example
13. (Use the result of
For $ 4.4
Exercise 1 above.)
14. Find the mean first passage matrix for Example 14. (Use th
of Exercise 12 above.)
4.4
15. Find the mean For
first§passage
matrix for Example 13. (Use th
14. Find of
theExercise
mean first
passage matrix for Example 14. (Use the result
13 above).
of Exercise 1216.
above.)
Verify Theorems 4.4.9 and 4.4.10 for Example 13.
first
passage
for Example
13. (Use
the all
result
15. Find the17.mean
Prove
that
for a n matrix
independent
trials process,
it1 has
rows th
of Exercise 13 above).
18. Given that the mean first passage matrix of a chain has the fo
16. Verify Theorems 4.4.9 and 4.4.10 for Example 13.
17. Prove that for an independent trials process, )vI has all rows the same.
18. Given that the mean first passage matrix of a chain has the form
determine the transition matrix
19. Give two different transition matrices which have the same
mental matrix, and hence show t h a t the fundamental matrix does no
determine the
transition
matrix.matrix.
mine
the transition
19. Give two
matricescolumn
which sums
haveifthe
20. different
Prove t h transition
a t P has constant
andsame
only funda·
if
has c
mental matrix,
and hence show that the funda.mental ma.trix does not deter·
row-sums.
mine the transition matrix.
20. Prove that, P has constant column sums if and only if Ii? has constant
row·sums.
SEC. 8
6. Given a probability vector a with positive components, dete
regular
transition matrix which will have this as its fixed vector.
REGULAR MARKOV CHAINS
97
For $ 4.2
For § 4.5
7.
For
Example
14
find
the
mean
and variance for the number of
21. Find 1'>12 for Example 14.
state s1 in the first n steps.
22. Using the result of Exercise 15 above, find },f 2 for Example 13.
8. Consider the Markov chaiil with transition matrix
23. Find .J1 2 for Example 11 ,vith c=1/2 and d=l/4.
24. A die is rolled a number of times. Find the mean and variance for
the number of rolls between occurrences of two 6's.
25. Find the mean and variance of the rust passage times in Exercise 8
above.
Start the process in s2, and compute the mean of v(*)l for n = l , 2, 3
Compare these results with al.
For § 4.6
26. Find the limiting covariance matrix for Exam
For $pie
4.311 with c = lIz and
d=lh·
9. Find the fundamental matrix for Example 1 1 when c = and d
27. Find the limiting covariance matrix for Example 13. Interpret the
10. Find the limit of the difference between t8he mean number
diagonal entries.
days in the Land of Oz in the first n days, starting with a rainy d
28. On a nicestarting
day a man
withinathe
niceLand
day.of Oz takes his umbrella with proba·
bility liz, on a rainy day with probability 1 and on a snowy day with probability
11.
Find
the
fundamental
matrix
thehechain
in Exercise
8
3/4. Find the limiting
variance for the number
of daysfor
that
will take
his
Interpret 211 -221.
umbrelia.
12. Findchain
the fundamental
the r
29. For an absorbing
let ilj be the matrix
numberforofExample
times in 14.
state(Use
Sf
Exercise
2
above.)
before absorption. Using the method of proof for Theorem 4.6.1, show that
13. Find the fundamental matrix for Example 13. (Use the r
.M.,[nt·nj]
= nkjnjl+nUnjj-dtjnkU
Exercise
1 above.)
where N = {njf} is the fundamental matrix. Find Covk[ni,nj] and Var,,(ntJ.
For $ 4.4
14. Find the mean
For first
§ 4.8passage matrix for Example 14. (Use th
of Exercise 12 above.)
30. Find the limiting variance of the number of times in a state when c= d.
15.with
Findc 1theInterpret
mean first
passage
matrix
for Example 13. (Use th
How does this vary
your
formula
as ~.
of Exercise 13 above).
31. Find the limiting vector and the mean rust passage matrix for the
Verify
for Example
13. as
case where c = 2d.16.How
do Theorems
these vary4.4.9
withand
c? 4.4.10
Interpret
your results
17. Prove that for a n independent trials process, it1 has all rows th
c-+O.
18. Given that the mean first passage matrix of a chain has the fo
For the entire chapter
32. Considcr the following transition matrix for a iyfarkov chain.
t/2
determine the
matrix
P =transition
\3~4
1 transition matrices which have the same
19. Give two different
mental
matrix, and hence show t h a t the fundamental matrix does no
(a) Is the chain
regular?
mine
(b) Find «, A,
andthe
Z. transition matrix.
(c) Find M and
20.N Prove
t h a t P has constant column sums if and only if
2.
has c
(d) Find therow-sums.
covariance matrix.
(e) Use absorbing. chain methods to find the mean time to go from 83 to
Sl.
Check your answer against part (e).
33. Let PI and P2 be two different transition matrices for a three. state
Markov chain. By a random device we select one of these matrices
out the resulting
is selected with probability
p.)
(Say PI
CHAP. IV
FINITEchain.
MARKOV
CHAINS
(a) Is this process a Markov chain 1
Markov chain. Bya
random
we seJect one
thesein
matrices
carry
(b) Show
thatdevice
the probability
ofofbeing
a givenand
state
tends
out the resulting chain.
with probability
l is selected
and (Say
showPhow
these probabilities
mayp.)be obtained from
of the
two matrices.
(a) Is this process vectors
a Markov
chain?
thatofinbeing
Exercise
33 westate
use the
random
device b
(b) Show that 34.
the Suppose
probability
in a given
tends
to a limit,
step,
to these
decideprobabilities
which rnatrixmay
to apply
on that step.
and show
how
be obtained
from the fixed
98
vectors of the two matrices.
(a) Is this process a Markov chain ?
34. Suppose that
in Exercise
33limiting
we use probabilities
the random for
device
(h) Show
that the
beingbefore
in theeach
various
matrix tonot
apply
that,asstep.
step, to decide whichnormally
the on
same
those obtained in Exercise 33(b).
(a) Is this process a Markov chain 1
(b) Show that the limiting probabilities for being in the various states are
normally not the same as those obtained in Exercise 33(b).
Markov chain. By a random device we select one of these matrices
out the resulting chain. (Say PI is selected with probability p.)
(a) Is this process a Markov chain 1
(b) Show that the probability of being in a given state tends
and show how these probabilities may be obtained from
vectors of the two matrices.
34. Suppose
that in Exercise
CHAPTER
V 33 we use the random device b
step, to decide which rnatrix to apply on that step.
(a) Is this process a Markov chain ?
(h) Show that the limiting probabilities for being in the various
ERGODIC
MARKOV
CHAINS
normally not
the same as those
obtained in Exercise 33(b).
§ 5.1 Fundamental matrix. 'Ve will now generalize the results
obtained in the last chapter. There they were proved for regular
chains, and now we will extend· them to an arbitrary chain consisting
of a single ergodic set, i.e. to an ergodic chain. \Ve know that such a
chain must be either regular or cyclic. A cyclic chain consists of d
cyclic classes, and a regular chain may be thought of as the special
case where d = 1. The results to be obtained will be generalizations of
the previous results in the sense that if we set d = 1 in them, we obtain
a result from the previous chapter. As a matter of fact, in most of
the results d will not ,"ppear explicitly, so that the result ofthe previous
chapter will be shown to hold for all ergodic chains.
An ergodic chain is characterized by the fact that it consists of a
single ergodic class, that is, it is possible to go from every state to
every other state. Howcver, if d> 1, then slich transition is possible
only for special n-values. Thus no power of P is positive, and different
powers will have zeros in different positions, these zeros changing
cyclically for the powers. Hence pn cannot converge. This is the
most important difference between cyclic and regular chains.
But while the powers fail to converge, we have the following weaker
result.
5.1.1 THEOREM. For any ergodic chain the sequence oj powers pn
is E1tler-s1tmmable to a limiting matrix A, and this limiting matrix is
of the form A = get, l{.·ith a a posiliDe probab·iliiy Dector.
PROOF.
Consider the matrix (kI+(l-k)P), for some k, O<k<l.
This matrix is again a transition matrix. Since it has positive entries
in all places where P is positive, the new matrix also represents an
ergodic chain. And since the diagonal entries are positive, it is
possible to return to a state in one step, and hence d = 1. Thus the
new chain is regular.
99
100
From $4.1.4 we know t h a t ( k I + (1 - k)P)n tends to a matrix
with aFINITE
probability
vector CHAINS
a > 0. Thus
MARKOV
CHAP. V
= lirn (tends
k I + ( lto
- k )aPmatrix
)n
From § 4.1.4 we know that (kI +A(I-k)p)n
A =;a,
with a probability vector a> O. Thus n--r m
A
= k)p)n
lim
lim (kI +A(1n-m
2 (1
i = ~
But this states precisely that the sequence P n is Euler-summa
(1)
(see $ 1.10). Indeed, i t is Euler-summable for every value of k
But this states precisely
that the sequence
A
5.1.2 THEOREM.
If P ipn
s aisn Euler-summable
ergodic transitiontomatrix,
and
(see § 1.10). Indeed,
it
is
Euler-summable
for
every
value
of
k.
a are as i n Theorem 5.1.1, then
5.1.2 THEOREM. (a)
Ij For
P isany
anprobability
ergnrhc transition
and xPn
A and
vector n, matrix,
the sequence
i s Euler-su
a are as in Theorem 5.1.1,
to a. then
(b) T h e vector a is the unique $xed probability vector of P .
(a) For any probability
vector
77, the sequence 77pn i.s Euler-summable
(c) P A =
AP=
A,
to a.
(b) The vectorPROOF.
a is the If
unique
fixed probability
vector
oj P.t h a t the Euler sun
(1) by x we
obtain
we multiply
(c) PA=AP=A,
sequence nPn is xA = n[a= a, which proves (a).
Since a was obtained from the limiting matrix of ( k I + (1 PROOF. If we multiply (1) by 71 we obtain that the Euler sum of the
is the unique fixed probability vector of this regular transition
sequence lTpn isBut
lTA this
=71fa=a,
(a).same fixed vectors as P,since
matrixwhich
must proves
have the
Since a was obtained from the limiting matrix of (kJ + (l-k)P), it
n(kI
+ regu:ar
(1- k)P)transition
= n
is the unique fixed probability vector of
this
matrix.
that the same fixed vectors as P, since
But this matriximplies
must have
x ( l - k ) P = n(l-k)
1T(kJ
k)P) = 17
and since
k +1(1,
implies that
n P = n.
+
IT(l-k)P = 7T(l-k)
and since k# 1,This proves (b). P a r t (c) follows from the fact that P[= [
transition matrix, and
h a71.
t aP = 1.
lTP t =
thus
that from
a andthe
A fact
havethat
nearly
the for
same
propertie
This proves (b). We
Part
(0) see
follows
Pf=;
any
ergodic
they did for regular chains ; only, (a) had to be w
transition matrix,
and case
that as
aP=a.
to summability in place of convergence.
We will now show tha
We thus see chains
that a have
and A
have nearly matrix
the same
properties
the like th
a fundamental
which
behavesin just
ergodic case as they
did
for
regular
chains;
only,
(a)
had
to
be
weakened
mental matrix of regular chains.
to summability in place of convergence. We will now show that ergodic
5.1.3 THEOREM.
chains have a fundamental
matrix which
justtransition
like the mutris,
funda- then th
If P ibehaves
s a n ergodic
= (I- (P- A))-1 exists, and
mental matrix of matrix
regularZchains.
P
THEOREM. (a)
Jj PP Zis=
anZergodic
transition matrix, then the inverse
matrix Z=(I-(P-A))-l
and
(b) ZE =exists,
!$
(c) aZ = a
(a) PZ = ZP (d) ( I - P ) Z = I - A .
(b) Zt = t
(c) aZ = a
(d) (J-P)Z = I-A.
5_1.3
SEC. 1
From $4.1.4 we know t h a t ( k I + (1 - k)P)n tends to a matrix
with a probability
a > 0.CHAINS
Thus
EEGODIC vector
MARKOV
101
( k I + ( l - k ) P ) to
n A by § 5.1.1,
PROOF. Since the sequence Apn=is lirn
Euler-summable
n--r m
and since (p-A)n=pn-A by § 5.1.2(c), the sequence (P-A)" is
Euler-summable to o. Hence
the lim
inverse Z exists (see § l.Il).
A =
n-m
i = ~
Furthermore, the series
2 (1
'" (Pi-A)
But this states 1+
precisely
that the sequence P n is Euler-summa
( ')\
(see $ 1.10). Indeed,
i=l i t is Euler-summable for every value of k
2:
~I
is Euler-summable
Z. Then (b) and (c) follow from the fact that
5.1.2 toTHEOREM.
If P i s a n ergodic transition matrix, and
1~=~, a1=a, and multiplying Pi-A by either~ on the right or by ex
a are as i n Theorem 5.1.1, then
OIl the left yields O.
Result (d) is obtained by multiplying (2) by
(a) For any probability vector n, the sequence xPn i s Euler-su
I-P.
to a.
While for the theorems so far, Euler-summability of pn sufficed, we
(b) T h e vector a is the unique $xed probability vector of P .
results.
will need the following
(c) P Astronger
=AP=A
,
5.1.4
THEOREM.
If P is an ergodic tran.s':tion matrix,
PROOF. If we multiply (1) by x we obtain t h a t the Euler sun
(a) The sequence pn is Cesaro-summable to A.
sequence nPn is xA = n[a= a, which proves (a).
'" (.Pi-A.)
obtainedisfrom
the limiting to
matrix
of ( k I (1 Since1 a+ was
(b) The series
Cesaro-summable
Z.
+
.L
PROOF.
i==l fixed probability vector of this regular transition
is the unique
But
this
matrix
have
the
same fixed
If n = lcd, then in must
n steps
after
starting
at SI vectors
we mustasbeP,
insince
a
state in the cyclic class of St. Andn(kI
if k +
is (1sufficiently
we may be
k)P) = large,
n
in any stateimplies
in the that
class. Hence pd may be thought of as the transition matrix of a Markov chain with d separate ergodic sets, each of
x ( l - k ) P = n(l-k)
which is non-cyclic.
pkd tends to a limiting matrix Ao,
and since kTherefore,
1,
whose ij-entry is 0 if Si and Sf are not in the same cyclic class, and
n P = n.
otherwise the ij entry is gotten by taking the components of a belonging
(b). P a r t (c)
to the cyclicThis
classproves
and renormalizing
them.
follows from the fact that P[= [
1.
matrix,
t h a t aP
If 0 ~ k d,transition
then pka+1
tends and
to PlAo
as k=tends
to infinity. Hence
the sequence pn has these d convergent subsequences, and hence
We thus see that a and A have nearly the same propertie
(see § 1.10) pn is Cesaro-summable to the average of the limits. But
ergodic case as they did for regular chains ; only, (a) had to be w
two different summation methods cannot give different answers, hence
to summability
in place
A must be this
average; that
is, of convergence. We will now show tha
+
chains have a fundamental matrix which behaves just like th
mental matrix of regulard-Ichains.
A = (lid)
.L PIAo,
(3)
5.1.3 THEOREM. If 1=0
P i s a n ergodic transition mutris, then th
matrix Z = (Ito
- (P
A))-1
exists,
and Pis- Cesaro-summable
A. - It
is then
an and
immediate consequence
that since P'-A is Cesaro-sum mabIe to 0, (h) must hold.
(a) P Z = Z P
Let us restate the summability result (b) as a limit.
(b) ZE = !$
1L
n-i
(c) aZ = a
5.1.5 COROLLARY.
lim
(d) ( I - 1P +
) Zn-+oo
= I i=l
- A . n (Pi-A) = Z.
.L --
Since we have now succeeded in generalizing many basic properties
of Z to ergodic chains, and since d did i10t appear explicitly, we may
102
FINITE MARKOV CHAINS
CH
now assert that many of the results of Chapter I V hold for all erg
chains. I FINITE
n particular
this applies
t o all results concerning
the m
MARKOV
CHAINS
CHAP. V
first passage time matrix M and the variance of first passage
now assert that
many
the is
results
Chapter
holdand
for 4.5.
all ergodic
H zof
, that
to allofresults
in IV
$5 4.4
We also hav
matrix
particular
this
applies
to
all
results
concerning
the and
meancovaria
chains. In the
results of § 4.6 concerning limiting variances
first passagesince
timeinmatrix
jr] and the variance of first passage time
the proof
of § 4.6.1 we needed only the surnmability (5.1
matrix llh', the
thatinfinite
is to all
results
4.5. We also
have
series
for in
Z , §§
not4.4itsand
convergence.
And
thusallall the
the results formulas
of § 4.6 of
concerning
limiting
variances
and
covariances,
$5 4.4, 4.5, and 4.6 may be applied to any ergodic cha
§ 4.6.1making
we needed
only comment
the summability
(5.1.5)We
of now
since in the proof
I t isofworth
a special
on $ 4.4.12.
the infinite that
series fl.
for determines
Z, not its convergence.
thu~ all the basic
the transition And
matrix
of any ergodic chai
§§ 4.4, 4.5,
and formula
4.6 may be
any) R
ergodic
formulas of means
P=applied
I + ( Dto- E
-~
Thus,
. chain.
in particula
of the
It is worth
making awhether
special (or
comment
§ 4.4.12.
We This
now isknow
determines
not) theon
chain
is cyclic.
quite surpr
that .if? determines
the transition
matrix ofto any
chain
and i t wouId
be highly desirable
find ergodic
necessary
and by
sufficient
means of the
formula
P=] +
in matrix
particular,
ditions
(i) that
be(D-E)jfI-l.
the mean firstThus,
passage
of a n .if?
ergodic c
determines whether
not)
chain isacyclic.
is quite
and (ii) (or
that
i t the
represent
regular This
rather
thansurprising,
a cyclic chain.
and it wouldwould
be highly
desirable
to findtonecessary
andthan
sufficient
con- P fro
like these
conditions
be simpler
computing
J!
be
the
mean
first
passage
matrix
of
an
ergodic
chain,
ditions (i) that
and then checking P.
and (ii) that itWhat
represent
rather
thanhave
a cyclic
chain.
One
resultsa regular
on regular
chains
we not
generalized
so
would like these
conditions
to besuch
simpler
than
computing
P fromestimate
Jf
$
The most
important
results
are:
The geometric
and then checking
P. of Large Numbers 4.2.1, and the results in §$ 4.3.4the Law
What results
wethese
not we
generalized
so tofar?
n ) , . regular
To be chains
able to have
discuss
shall have
find some
on Y (on
The most important
such
results
p@)tlThe
- a$.geometric estimate § 4.1.5,
of a n upper
bound
on are:
the Law of Large
Numbers
4.2.1,
and thebound
results
§§ 4.3.4-4.3.6
I t is clear
that §the
geometric
of in
$ 4.1.5
cannot apply to
on y(n)j. Todifference
be able toindiscuss
thesechain,
we shall
have
to find
sort be 0
the cyclic
since
p(n),?
will some
frequently
of an upper hence
boundthe
on p(n)jj-aj.
difference-in absolute value-is ar infinitely often.
It is clearever
thati tthe
4.1.5 cannot
applyofto§ this
5.1.4, that
cangeometric
be shown,bound
using of
the§ ideas
of the proof
difference inadd
theupcyclic
chain, since
p(n)ij will frequently be 0, and
d consecutive
terms,
t h a t is, form
hence thc difference-in absolute value-is aj infinitely often. However it can be shown, using the ideas of the proof of § 5.1.4, that if we
add up d consecutive terms, that is, form
1=0
102
then this sum d··l
is bounded geometrically. This suffices to prov
(p(n+l)jj - aj),
Law of Large Numbers,
if in 5 4.2.1 we take sums d terms a t a
1=0
This method also allows us to prove analogues of $9 4.3.4-4.3.6
then this sum
bounded
This suffices to prove the
we iswill
not takegeometrically.
these up.
Law of Large Numbers, if in § 4.2.1 we take sums d terms at a time.
This method also
us to of
prove
of §§simplest
4.3.4-4.3.6,
but cyclic
$5.2allows
Examples
cyclicanalogues
chains. The
possible
we will not take
these up.
is obtained
from the two-state Example 11 by choosing c = d = I.
will call this Example I la. The transition matrix is
§ 5.2 Examples of cyclic chains. The simplest possible cyclic chain
is obtained from the two-state Example 11 by choosing c""d= 1. We
"l'ilI call this Example lla. The transition matrix is
2
(~ ~)-
From the proof of $ 5.1.1 we know that A may be obtained a
p =
From the proof of § 5.1.1 we know that A may be obtained as the
102
FINITE MARKOV CHAINS
CH
now assert that many of the results of Chapter I V hold for all erg
chains. I n particular this applies t o all results concerning the m
ERGODIC :MARKOV CHAINS
SEC. 2
103
first passage time matrix M and the variance of first passage
H z+, that
is to all results
We also hav
matrix
limiting matrix
of (liz)]
(1/2)P=(1/2)E.
B"G.t in
this$5is4.4
itsand
own4.5.
limiting
the results of § 4.6 concerning limiting variances and covaria
matrix. Hence
since in the proof of § 4.6.1 we needed only the surnmability (5.1
d = 2,
= (lh)E,
theAinfinite
series for Z , not its convergence.
And thus all the
formulas of $5 4.4, 4.5, and 4.6 may be applied to any ergodic cha
( 3/ 2making a special comment on $ 4.4.12. We now
IZ
t is=worth
_1/ 2
that fl. \determines
the transition matrix of any ergodic chai
Thus,
. in particula
means of the formula P= I + ( D - E ) R - ~
determines whether (orj11
not)
This is quite surpr
2 =the chain is cyclic.
\0 to find necessary and sufficient
and i t wouId be highly desirable
ditions (i) that
be the mean first passage matrix of a n ergodic c
It is very easy
to (ii)
findthat
M directly,
and ato regular
see thatrather
M 2 must
all chain.
and
i t represent
thanhave
a cyclic
components O. Similarly, the limiting variances are o.
would like these conditions to be simpler than computing P fro
As a less trivial
example
we take
and then
checking
P. up the random walk Example 2
by
Example
2a).
Its transition
matrix
is generalized so
for p = 1/2 (denoted
What results on regular
chains have
we not
The most important such results are: The geometric estimate $
0
0 4.2.1, and the results in §$ 4.3.4the Law of Large Numbers
0 to\'2
0 these
0
1/2 able
. To be
discuss
we shall have to find some
on Y ( n ) ,82
a$. o
of Pa n=upper
bound
on p@)tl
1/2
Sa
0
.
0 -1/2
I t is 84
clear that
the
geometric
bound
of $ 4.1.5 cannot apply to
0 l/Z
0
0
1/2
difference in the cyclic chain, since p(n),? will frequently be 0
0
0absolute
S5
0
1
hence the
difference-in
value-is ar infinitely often.
§ 5.1.4, that
i t can be shown,state,
using the
the process
ideas of can
the proof
Starting from ever
an even-numbered
be in of
evenadd only
up d in
consecutive
terms, t hof
a t steps,
is, form
numbered states
an even number
and in an oddnum bered state in an odd number of steps; hence the even and odd
states form two cyclic classes. Computing the other quantities we
1=0
find:
This suffices to prov
then thisa sum
is bounded
geometrically.
1/4,1/4,1/4,
l/S)
= (1(8.
Law of Large Numbers, if in 5 4.2.1 we take sums d terms a t a
18 -2
This method also allows
us to-14
prove analogues of $9 4.3.4-4.3.6
we will not take these
22 up. 2 -10 -7
-1
14
2 -1
.
2
Z = 1/ 16
$5.2 Examples of cyclic chains. The simplest possible cyclic
2
-7 the
-10two-state
22
is obtained from
Example
11 by choosing c = d = I.
-14 I la.
-2 The18transition matrix is
will call this -9
Example
(O 0°).
0)
"C
01
-)
( 2:
2:
Sl
81
82
S2
S3 S4
85
(; 1:;
,~
9
4
3
8
9
4
8
16~
15
From the proof of $ 5.1.1 we know that A may be obtained a
5 4 5
M = 83
15 8 3 4,
S4
85
\16
160
112
0 8
FINITE MARKOV 112
CHAINS
24 8
48
48
160 CHAP. V
"(2
104
l6)
0
815248 40
8
40
152
816048 481608
24
112
816040 481528
.0
112
S2
112
24
M2 = S3
152
40
48 8 some
24 ll2
I t is interesting 160
t o examine
of the entries of ilf and
S4
From either 55
end state
112 state only by passing t
0 any
160 we48can8 go to
the neighboring state. Hence the first row of 31 is, with one exc
It is interesting to examine some of the entries of ill and of M 2.
1 greater than the second row, and similarly for the fifth and
From either end state we can go to any state only by passing through
rows. The one exception is stepping into the neighboring stat
the neighboring state. Hence the first ro,v of M is, with one exception,
The third row is the average of the second and fourth, plus 1
1 greater than the second row, and similarly for the fifth and fourth
for stepping into a neighboring state.
rows. The one exception is stepping into the neighboring Some
of the eq
I n Hz i t is worth noting the equal entries. state itself.
The third row
is the
of the second
fourth,
plus
except
are due
to average
the symmetry
of the and
process.
But
this1, does
not a
for steppingforinto
a neighboring
example,
for thestate.
third column being constant. The seco
In M 2 it is worth noting the equal entries. Some of the equalities Th
fourth entries are the same in this column by symmetry.
are due to three
the symmetry
the process.
Butofthis
notinaccount,
are also 8,ofbecause
from one
thedoes
states
the first cy
for example, for the third column being constant. The second and
we must enter the second cyclic set, and then the variance is
fourth entries are the same in this column by symmetry. The other
two 0 entries are due to the fact that from an end state we althree are also
8, neighbor
because from
of the states in the first cyclic set
to its
in oneone
step.
we must enterItthe
second
cyclic
set,
then
is 8. The
is also interesting to and
think
of the
the variance
middle column
in Jif
two 0 entries
due from
to themaking
fact that
an endand
stateasking
we alwa.ys
go me
as are
arising
sa from
absorbing,
for the
to its neighbor
in
one
step.
variance of the number of steps needed for absorption. The r
It is also interesting to think of the middJe column in M and J"f1 2
process behaves in all essential features like 5 3.4.1 (with p =
as arising from making 83 absorbing, and asking for the mean and
hence the numbers 3, 4, and 8 are the samc. as the entries of 7
variance of the number of steps needed for absorption. The resulting
there obtained.
process behaves
all conclude
essential by
features
like § the
3.4.1covariance
(with p= matrix.
liz), and
computing
We in
shall
hence the numbers 3, 4, and 8 are the same as the entries of 7 and 72
there obtained.
\Ve shall conclude by computing the covariance matrix.
~
8
-2
-8
12
0
-12
0
1
- ,')
4
0
-)-;
-8
G
'/" ( - :
0
12
From this we-8
obtain the limiting
variances
-8 -2
-5
P = ( 7 / 3 2 , 3/81 !'8,8 3/8, i / 3 d .
From this we obtain the limiting variances
The fact that ~ 2 =3 c 4 3 = 0 means that the limiting corr
between sz f3and
and sq l/S,
and31s,
s3, 7/32
are ).0. On the other band the
= SQ,
(7/32,3/8,
tion between sl and sz is 8 1 4 8 4 c . 8 7 . The reason for this i
The fact that C23 = C43 = 0 means that the limiting correlations
obvious from the transition matrix.
between 82 and 53, and 84 and 83, are O. On the other hand.the correlation between 81 and 82 is 8/Y84:::::.87. The reason for this is fairly
obvious from the transition matrix.
~
112
0
ERGODIC MARKOV CHAINS
112 24
SEC. 3
8
48
160
8
48
160
105
40 a 152
§ 5.3 Reverse l\Iarkov chains. \Ve saw152
in §402.1 8that
Markov
process observed in reverse order would be
process
112 with
160a Markov
48 8 24
transition probabilities given by
0 112
160 48 8
Prn[fn-l=Sj]Prw[fn=s¥,,-l=Sj]
()
Pif
n
=
--=----'-''=~
__
-+---'-'
I t is interesting t o examine some of the entries of ilf and
Pr,,[fn = St]
From either end state we can go to any state only by passing t
where fn is the
outcome state.
function.
It was
observed
that,
31 is,
withifone exc
then-th
neighboring
Hence
the first
row ofalso
the forward process
is athan
Markov
chain, row,
the reverse
processfor
willthe
befifth
a
1 greater
the second
and similarly
and
Markov chainrows.
only ifThe
PI" ,,[fn
sf] does not
depend on
n. the
This
will be stat
one= exception
is stepping
into
neighboring
the case if the
process
in equilibrium.
caseplus 1
The third
rowisisstarted
the average
of the secondIn
andthis
fourth,
Pr a[fn = 5t] = atfor
forstepping
all n, and
Pij(n)
becomeS state.
into
a neighboring
I n Hz i t is worth noting the equal entries. Some of the eq
are due to the symmetry of the process. But this does not a
for example, for the third column being constant. The seco
5.3.1 DEFIXITIO~.
Let P
thesame
transition
ergodic
Th
fourth entries
arebethe
in thismatrix
columnforbyansymmetry.
~il arkol' chain.
be the
probabnity
vector
for states
P. Then
thefirst cy
three Let
area also
8, fixed
because
from one
of the
in the
renrse ::Harkov
chainenter
for Pis
a l;J arkov
chain
with
we must
the second
cyclic
set,
andtransition
then the matrix
variance is
given by
two 0 entries are due
to
the
fact
that
from
an end state we alr
.,
to itspneighbor
step.
= {p'e~ in
= jone
(iJpjt
~ = DPT D-l.
JJ
l at )to think of the middle column in Jif
It is also interesting
as arising
from making
sa absorbing,
for the me
To justify the
above definition
we must
show that and
P is asking
a transition
variance
of
the
number
of
steps
needed
for
absorption.
matrix. By Theorem 5.1.2 the a/s are all positive, so that Pi! is The r
process behaves in all essential features like 5 3.4.1 (with p =
defined and non-negative.
hence the numbers 3, 4, and 8 are the samc. as the entries of 7
Ft=DPTD-1t=DPTaT=D(aP)T=DaT=r
Hence P is a tranthere obtained.
sition matrix. We shall conclude by computing the covariance matrix.
5.3.2 DEFINITION. A Markov chain i8 reversible if P= P.
5.3.3
THEOREM.
A Markov chain is reversible if and only if
D-lP ,is a symmetric matrix.
PROOF.
F=DPTD-l. Hence P=P if and'only if
D-lP = pT D-l = (D-IP)T.
That is, if and only if D-lP is a symmetric matrix.
From this we obtain the limiting variances
A reversible Markov cl;tain in equilibrium win appear the same
= ( 7 / 3 2 , 3/81 !'8, 3/8, i / 3 d .
looked at backwards as forwards. P An
alternative way to describe
reversibility is the
follo\ving.
is 0reversible
if, in equilibrium,
The
fact thatA ~process
2=
3c43 =
means that
the limiting corr
for any andbetween
Sf the probability
of SIsqfollowed
the the
same
as the
sz and SQ, and
and s3, by
are Sf0. is On
other
band the
probability oftion
Sf followed
Si, Sf
betweenbyslSt·andThat
sz isis,8 if
1 4for
8 4 cevery
. 8 7 . n,The
reason for this i
s,
obvious
transition
Pra[fn
= 5ifrom
/\fn+lthe
= Sf]
= Pralfnmatrix.
= Sf /In+l = 5i].
This last equation will be true if a/po = a;pJi or if PH = ajpjI/at.
is, if Pit = Pif for every i, j.
That
It is obvious that any periodic chain with period greater
MARKOV
CHAP.
V whi
cannot be FIXITE
reversible.
I n factCHAINS
for such a chain, any
state
be reached on the next step could not have been the result of t
It is obvious
periodic
with 1period
2
Thusany
only
chains chain
of period
and 2greater
can be than
reversible.
step. that
cannot be reversible. In fact for such a chain, any state which can
clear that if such a chain is reversible i t will have the same perio
be reached on the next step could not have been the result of the last
As a n example of a reversible chain of period 1, we can consi
step. Thus only chains of period 1 and 2 can be reversible. It is
Land of Oz example. I n this case we find:
106
clear that if such a chain is reversible it will have the same period.
As an example of a reversible chain of period 1, we can consider the
Land of Oz example. In this case we find:
c
0
Z/5
1/4 1/4 l/Z
R
N
S
R
:,
1/ 10
O)C 'I,)
0
D-1P =
1/ 4
1/5 o
:
C
'/
1/2
0
1/2
'1/10I,") ,
=N
which is a symmetric matrix. Hence by Theorem 5.3.3 the c
of a1/5reversible chain with period 2 is
reversible. An
1/ 10
S example
I,' 10
by Example 2a. In this case the matrix D-IP is
0
which is a symmetric matrix. Hence by Theorem 5.3.3 the chain IS
reversible. An example of a reversible chain 'with period 2 is given
by Example 2a. In this case the matrix D-lP is
liS
D-'P ~
1,/ S
0
0
0
l/S
0
lis
0
1/8
C
:
0)
0
o
.
1i
0 matrix.
0
! 8
This is again a symmetric
J~8
Given an ergodic cl~ain,
we
0
0 now
J ! 8 ask for the relation betwe
chain and the associated reverse chain. We shall find the r
This is again
a symmetric
matrix.
between
the fixed
vectors and the fundanlental matrices and
Given anfrom
ergodic
chain,
now askwhich
for the
relation
betweenWe
thisshall
these, any we
quantities
depend
on them.
chain and A
the
associated
reverse
chain.
We
shall
find
the
relation
, 2 , M , etc. for the reverse chain by A, 2,i@, etc.
between the fixed vectors and the fundamental matrices and hence,
5.3.4
THEOREM.
for Pdenote
and is th
from these, any
quantities
whichThe,fixed
depend probabiliiy
on them. vector
\Ve shall
A, Z, M, etc. for the reverse chain by A, Z, M, etc.
5.3.4
PROOF.
Then
THEOREM.
Let
The .fixed probability veeto)' for P and j> is the same.
aP = a.
aj' = aDPT D-l
= 1)PT D-l
= (Pf)T j)-l
=
1)D-l
= a.
It is obvious that any periodic chain with period greater
ERGODIC
¥..ARKOV
cannot be
reversible.
I n factCHAINS
for such a chain, any state
107 whi
be reached on the next step could not have been the result of t
2 chains
= DZTD-l.
5.3.5 THEOREM.
of period 1 and 2 can be reversible.
step. Thus only
clear
that
if
such
a
chain is reversible i t will have the same perio
PROOF.
2 = (I-P+A)-l.
As a n example of a reversible chain of period 1, we can consi
From the form
ofofA Oz
it example.
is clear that
A = DAT
From Theorem
Land
I n this
case D-l.
we find:
5.3.4, A=A. Thus
SEC. 3
2 = (I -DPTD-l+DATD-l)-l
= D(I-PT+AT)-lD-l
= DZTD-l.
5.3.6 THEOREM. Any quantity whose value depends only on Zdg
and A is the same for the reverse process as for the forward process.
PROOF.
By Theorem 5.3.5, tag= Zctg, and by § 5.3.4, A =A,
An example of the application of the above theorem is the mean
and variancewhich
of the isfirst
passage time
to stateHence
St, if we
in Sl 5.3.3
or if the c
by start
Theorem
a symmetric
matrix.
we have as initial
vectorAn
cx. example
Similarly,
limiting variance
forperiod
the 2 is
of the
a reversible
chain with
reversible.
number of times
in a state2a.depends
only
Zdgmatrix
and A.D-IP
Hence
In this
caseonthe
is these
by Example
quantities are the same for the forward and reverse processes. Additional examples are provided by the following theorem.
5.3.7
PROOF.
THEOREM.
eii = aizi}
+ ajZji - aid!} - alaj
= ai(ajZji!a,) + aj(aiZti!aj) - ajd tj - aiaj
= ajZji + a jZjj - a/dtj - a,at
This=isCij·
again a symmetric matrix.
Given an ergodic cl~ain,we now ask for the relation betwe
Thus all results
only on reverse
the covariance
are find
the the r
chain that
and depend
the associated
chain. matrix
We shall
same for the between
reverse process.
the fixed vectors and the fundanlental matrices and
from these, any quantities which depend on them. We shall
5.3.8 THEOREM.
£1-M
A , 2 , M , etc.
for =the(ZD)-(ZD)T.
reverse chain by A, 2,i@, etc.
PROOF.
£1-M
5.3.4= (l-t+EZdg)D-(I-Z+EZag)D
THEOREM. The,fixed probabiliiy vector for P and
= (Z-t)D.
The theorem then follows from § 5,3.5.
5.S.9
THEOREM.
lV- W = (ZD-(ZD)T)(2ZdgD-31)
+ 2(Z2D- (Z2D)T).
PROOF.
w-W = (£1-M)(2Z dg D-I)+2(Z£1-ZM)-2E(t£1-ZM)dg. (1)
M-M = ZD-(ZD)T.
(2)
is th
108
Since Z.@= (2- g2+ EZd,)D, and Z M = (2- Z2+ EZdg)D, we h
FINITE MARKOV CHAINS
CHAP. V
ZB-ZM
=
(2-z)D+(z~-Z~D
+
(ZD)T - ( Z D ) (Z2D)-we
(Z2D)*.
Since Zk=(Z_Z2+EZdg)D, and=ZM=(Z_Z2+EZdg)D,
have
Zk- this
ZM is
= the
(Z - difference
Z)D + (ZL
Since
of Z2)D
a matrix and its transpose, i t
(3)
= (ZD)T_(ZD)+(Z2D)-(Z2D)T.
diagonal entries,
hence
( Z B -and
ZN).l)dg
= 0.
Since this is the difference of a matrix
its transpose,
it has 0
diagonal entries,
hence
We obtain our theorem by combining ( I ) , (2), (3), and (4).
(Zk- Z21·f)dg = O.
(4)
We shall now illustrate the application of the above theorem
We obtain process
our theorem
(1), (2),Such
(3), and
(4). is the random
whichbyis combining
not reversible.
a process
Example 3a. Here
We shall now illustrate the application of the above theorems for a
process which is not reversible. Such a process is the random walk
Example 3a. Here
A
81
82
81
0
82
0
('1'
P = 83
l/z
0
and a = (.1,84.2, .4, .2, .I).
Ss
and a= (.1, .2, .4, .2, .1).
84
85
0
liz
0
0
1/2
,1)
liz
From 0this we find
0
0
From this we find
81
82
81
0
S2
P = S3
83
G·
S3
84
0
0
0
1
Ss
'D
0 1/4
1
0 are obvious.
0
Most of the entries in this matrix
For example,
0
0
8s
process is ever in state s l i t must h a r e come from sz, hence l
The fixed vector for j? is again a = (. 1, .2, .4. .2, . I ) .
Most of the entries
in this
are2obvious.
For example,
I n Chapter
I V matrix
we found
for this example
to be if the
process is ever in state 81 it must have come from 82, hence P12 = 1.
Thefixedvectorforpisagaina=(.l, .2,.4, .2, .1).
In Chapter IV we found Z for this example to be
84
.33
.86
.12
-.04
- .14
S3
-.02
. 16
.72
.16
84
- .17
- .14
.12
.86
-')
Ss
- .12
-.04
.32
-.04
.88
81
81
52
Z =
1/ 4
(88
82
83
84
-.04
.32
8S
- .17
-:~:
.
Since Z.@= (2- g2+ EZd,)D, and Z M = (2- Z2+ EZdg)D, we h
ZB-ZM
= ( 2CHAINS
- z ) D + ( z ~ - Z ~ D 109
ERGODIC
MARKOV
= (ZD)T - ( Z D ) (Z2D)- (Z2D)*.
From this we find t=DZTD-l to be
Since this81 is the82difference
of S4
a matrix
and its transpose, i t
S3
S5
diagonal entries, hence
.66 -.08 -.34 -.12\
( Z B - ZN).l)dg= 0.
SEC. 3
+
"(88
S2
t
-
.02
.32
.86
-.14
-.02
00,/
our theorem
by
( I ) , (2), (3), and (4).
.06
.72combining
.06
=We
S3 obtain.08
S4
- .02
.86 -.02of the above theorem
We shall
now-.14
illustrate.32
the application
A
process
not .reversible.
a process is the random
-.08
.66Such .88
S5
-which
.12 is-.34
Example 3a. Here
We found M to be:
S1
S1
82
M = Sa
S4
(:5
S5
82
S3
S4
S5
4.5
1
5
1.5
4.5
5
10 )
3.5
3.5
9
10.5
5
2.5
1.5
10
4.5
1
10.5
5
4.5
.
.
5.5/
10
From this we obtain
=11,1
(ZD(ZD)T):
.2,+.4,
.2, .I).
and a =M(.1,
From this we find
81
52
S3
S4
S5
(
1
2
6
5
1
5
4
2.5
4
1~)
5
1
5
9
.10
6
2
1
10
51
52
S3
84
(66~~~.5
28
2.5
28
83.5
S5
1611
33
1
33
166
S1
S2
Sf = S3
54
S5
8
.
We found W to be
84
55
Most of the entries in this matrix are obvious. For example,
33 1s l i t 33
51 is ever in state
must h a r e come from sz, hence l
process
176.5
83.5for38
) .4. .2, . I ) .
The fixed
vector
j? is2.5
again38a =166
(. 1, .2,
82
I
n
Chapter
I
V
we
found
2
for
this
example
to be
.
25 6.5 25 147
W = Sa
From this we obtain
W = TV + (ZD- (ZD)T)(2ZdgD-3J) + 2(Z2D-
(Z2D)T) :
/166
M
147
W=
\130
38
1
38
29
6.5
29
16)
38
147
1
166
I
147
38
166
49
4
4,
49
147
130
.
FINITE MARKOV CHAINS
110
Hence
llO
FINITE MARKOV
66 CHAINS
0 0
66
Hence
("
a
2
=0
13
066 1313
0
13 66
13 66
'14
13
6)
CHAP. V
66
13 066 1313 660 13 66
66 1313 660 . 0 66
lW 2 =
66 13 1/4
13 a66
13 variances
0
T h e zeros a n d t h e66equal
r e easily deduced from p.
66
66
13
0
0
66
The zeros and the equal variances are easily deduced from P.
Exercises for Chapter V
For $ 5.1
Exercises
for
Chapter
1. Compute the limiting matrix AV and the fundamentai matrix
ergodic chain with transition matrix
For § 5.1
I. Compute the limiting matrix A and the fundamental matrix for the
ergodic chain with transition matrix
2 . Compute M and M z for the chain in Exercise 1 above.
3. Find the covariance matrix 6'for the chain in Exercise 1 above
4.
I n Example 2 let p=2l3. Find the fixed probability vector
2. Compute M and M 2 for the chain in Exercise 1 above.
fundamental matrix.
3. Find the covariance matrix C for the chain in Exercise 1 above.
5. For the example of Exercise 4 above find the mean first passag
4. In Example
let presults
= 2h. byFind
the fixed
probability
P from
g. vector and the
Check 2your
obtaining
fundamental matrix.
6. Given t h a t for a n ergodic chain
5. For the example of Exercise 4 above find the mean first passage matrix.
Check your results by obtaining P from M.
6. Given that for an ergodic chain
(~
M =
:
:)\,
show that the chain is cyclic.
144
7. Prove t h a t the matrix
show that the chain is cyclic.
7. Prove that the matrix
is not the first passage matrix of a n ergodic chain.
8. Let P be the transition matrix for an ergodic chain. Let t'be t
P with diagonal entries replaced by 0's and the rows renormalized
is not the firstsum
passage
matrix
of an
chain;
1. Show
that
theergodic
resulting
chain is again ergodic ; and if n =
8. Let P be the transition matrix for an ergodic chain. Let P be the matrix
P with diagonal entries replaced by O's and the rows renormalized to have
sum 1. Show that the resulting chain is again ergodic"; and if a = {aj} is the
110
FINITE MARKOV CHAINS
Hence
ERGODIC ~fARKOV66CHAINS
0 0
SEC. 3
111
13 66
66
fixed vector for the original chain, then a={aJ(l-pjj)}
proportional
to
66 13 0 is13
the fixed vector for the new chain. What is the interpretation of the
13
661
'14
a 2 =
13 original
components of the new fixed vector
in terms66
of the
chain
13 for
66 the Land
13 0exercise
9. Carry out the procedure indicated in the66previous
of Oz example.
0 66
66 13 0
T h e zeros a n d t h eFor
equal
variances a r e easily deduced from p.
§ 5.3
10. Find the reverse transition matrix for the chain in Exercise 1 above.
Compute the fundamental matrix for this reverse chain from the fundamental
matrix for the original chain.
Exercises for Chapter V
n. For which values of p is the chain in Exercise 1 reversible 1
For $ 5.1
12. Find the reverse transition matrix for Example 2 withp=2/J. Compute the fundamental
matrix the
for the
reverse
chain Aand
your resultmatrix
1. Compute
limiting
matrix
andcompare
the fundamentai
with the result ergodic
of Exercise
above.
chain4with
transition matrix
13. Compute M fo!' the example of the last exercise directly from the
fundamental matrix there found'. Compute lit from M (see Exercise 5)
using Theorem 5.3.8, and compare your answers.
14. Prove that every independent process is reversible.
15. Prove that every two,state ergodic chain is reversible.
an ergodic
chain
a symmetric
transition
matrix
16. Prove that2 .ifCompute
M and
M zhas
for the
chain in Exercise
1 above.
(i.e., Pij = Pit), then
the chain
is reversible.
3. Find
the covariance
matrix 6'for the chain in Exercise 1 above
17. Show for an
chain that
4. ergodic
I n Example
2 let p=2l3. Find the fixed probability vector
fundamental
matrix.
(a) If the chain
is reversible,
then PijPjkPkl = PjiPkjPik.
(b) If the transition
has ali of
positive
entries,
then
the
above
4 above
find
the
meancondifirst passag
5. Formatrix
the example
Exercise
tion assures
reversibility.
[HINT:
ShowP that
from for
g.fixed i the row
Check
your results by
obtaining
vector'\ = 6.
{JiijiPii}
of P.chain
Hence this vector must be
Givenist haafixed
t for avector
n ergodic
proportional to a. J
For the entire chapter
18. The general (finite) random walk is defined as follows. The states are
num bered So, 51, ' .. , Sn. If the process is in St, then it moves to 51-1 with
probability ql, it stays in 5i with probability Tj, and moves to 8Hl with
show that the chain is cyclic.
probability Pi. (Where Pi +qi +Tj = 1, qo = 0, pn = 0.)
7. Prove t h a t the matrix
(a) Under what conditions is a random walk ergodic?
(b) From the equation aP = a prove, by mc.thematical induction, that
aHlqH1 =
alpi·
(0) Prove that an ergodic random walk is reversible.
(d) Find a formula for the fixed vector a.
is not the first passage matrix of a n ergodic chain.
8. Let P be the transition matrix for an ergodic chain. Let t'be t
P with diagonal entries replaced by 0's and the rows renormalized
sum 1. Show that the resulting chain is again ergodic ; and if n =
CHAPTER VI
ER RESULTS
FURTHER RESULTS
$6.1 Application of absorbing chain theory to ergodic chain
have seen that the 2 matrix enables us to find the mean and v
of the first passage time to state si. Assume now that we are int
§ 6.1 Application
absorbing
chain of
theory
to ergodic
chains.
in more of
detailed
behavior
the process
in going
to sj.We For ex
have seen that the Z matrix enables us to find the mean and vatiance
we might ask for the mean number of times that it will be in
of the first passage
time states
to statebefore
Sj.
Assume
no,"v
thatthe
we first
are interested
s, for
time. The an
the other
reaching
in more detailed
behavior
of
the
process
in
going
to
sf. For example,
this and other similar questions is furnished by applying the ab
we might askMarkov
for the chain
mean theory.
number of
thatwe
it change
will be in
of by
Totimes
do this
oureach
process
the other states
before
reaching
Sj for the first time.
The
a.nswer
to
s, into an absorbing state. The resulting process will be an ab
this and otherprocess
similar with
questions
is furnished
bystate.
applying
thebehavior
absorbing
a single
absorbing
The
of this
Markov chainbefore
theory.
To do this
we change
our process
makingof the
absorption
is exactly
the same
as the by
behavior
Sj into an absorbing
resulting
be a.nHencc
absorbing
processstate.
before The
hitting
sj for process
the firstwill
time.
we can tr
process with all
a single
absorbing
state.
The
behavior
of
this
processMarkov
of the information we have about an absorbing
is exactly the
sameour
as original
the behavior
original it p
before absorption
into information
about
chain.of the
In particular
Sj for the first time.
Hence
we
can
translate
process beforeushitting
with an alternative way to find the mean and variance of
all of the information
we from
have s.~
about
absorbing
chain
to sj,an
these
being theMarkov
mean and
varianc
passage time
into information
original chain.
In particular
provides
timeabout
beforeour
absorption
in the new
process. it
Since
any prope
us with an alternative
way set
to find
mean
of the
the results
first
of an ergodic
is anthe
open
set,and
we variance
can apply
of
passage time obtain
from Si the
to Si,
these
being
the
mean
and
variance
of
the for
behavior of our process before it hits a subset
time before absorption
in the new p"'ocess. Since any proper subset
time.
of an ergodic setLet
is an
open
set, we
apply
the with
results
§ 3.5 of
to Oz e
us illustrate
the can
above
ideas
theofLand
obtain the behavior
of
our
process
before
it
hits
a
subset
for
the
first
Assume that we arc interested in the behavior of the proces
time.
the first rainy day. We make state 1IC absorbing and have
Let us illustrate
the Markov
above ideas
Land matrix
of Oz example.
:
absorbing
chain with
with the
transition
Assume that we are interested in the behavior of the process before
the first rainy day. vVe make state It absorbing and have the new
absorbing Ma.rkov chain with transition matrix:
R
N
R
0
P=N
0
S
C'
1/4 1/4
112
S
,;,)
1j 2
SEC. 1
FURTHER RESULTS
113
The basic results for this absorbing chain are obtp.ined from the
fundamental matrix N = (/ _Q)-l where Q is the matrix obtained by
considering only non-absorbing states. For example, let nj be the
num ber of times the process is in state s1 before being absorbed. Then
the values ofM;[nj] are given by the matrix N, ER
in thisRESULTS
case
$6.1 Application of absorbing chain theory to ergodic chain
matrix
us number
to find the
mean and v
seen that
thea 2
For example,have
calculated
from
nice
day, enables
the mean
of nice
of the first passage time to state si. Assume now that we are int
days before the next rainy day is 4/ 3. We can find Vari[u;] from the
in more detailed behavior of the process in going to sj. For ex
matrix
we might ask for the mean number of times that it will be in
the other states before reaching s, for the first time. The an
this and other similar
is furnished by applying the ab
is" questions
S
Markov chain theory. To do this we change our process by
s, into an absorbing state. The resulting process will be an ab
process with Sa single
absorbing
state. The behavior of this
"/3 4°/9
.
before absorption is exactly the same as the behavior of the
Let t be the ftmction
which hitting
gives the
number
of steps
beforewe can tr
process before
sj total
for the
first time.
Hencc
absorption. Then, from Theorem 3.3.5, we have that the column
all of the information we have about an absorbing Markov
veotor 7=[lUi (tJ}
is information
given by 7=Nf
In the
example
we In
areparticular
coninto
about our
original
chain.
it p
sidering, this is
us with an alternative way to find the mean and variance of
passage time from s.~to sj, these being the mean and varianc
time before absorption in the new process. Since any prope
of an ergodic set is an open set, we can apply the results of
The function t obtain
represents
in the original
time
to reach
a subset for
the behavior
of our process
process the
before
it hits
state R for the time.
first time. Thus the mean first passage time to R,
starting in state N, Let
is 8/3,usand,
startingthe
in state
S it,ideas
is lOp.withThese
illustrate
above
the values
Land of Oz e
agree with thoseAssume
found inthat
the we
matrix
1.f calculated
frombehavior
the Z matrix
arc interested
in the
of the proces
1IC the
absorbing
the first
day. 3.3.5,
We we
make
state
in § 4.4. Similarly,
fromrainy
Theorem
have
that
Var![t] and
is have
given by the column
vector
72 = (2N
-1)7
- Tsq.
Calculating
this,
: we
absorbing
Markov
chain
with
transition
matrix
N(:/!} 4)
ha,ve
72 = (56/ 9').
\ 52/9/
The vector 72 gives us the variance of the time before absorption. In
terms of the original process this is the variance of the first passage
time to state R. Again the above values check with those found from
the matrix .1.l1z obtained in § 4.5.
By successively making each state absorbing we could find all the
non-diagonal elements of jYf and Jfz for an ergodic chain. However,
FINITE MARKOV CHAINS
114
CH
the use of the 2 matrix is much more natural and convenien
FINITE
CHAINS
CHAP.
VI th
would normally
useMARKOV
the absorbing
methods only to
obtain
detailed information not available by the Z matrix methods.
the use of the Z matrix is much more natural and convenient. We W
As an example of a cyclic chain we consider Example 2a.
would normally use the absorbing methods only to obtain the more
states sl and sz absorbing. We then have
detailed information not available by the Z matrix methods.
As an example of a cyclic chain we consider Example 2a. We make
states 81 and S2 absorbing. We then have
114
2
YI =
3
84
S5
S
(222
4
4
The entries of N and N z give the mean and variance of the num
times that the process is in each state before reaching sz. (Th
sl can only be reached through sz.) The vectors T and T Z g
The entries of Nand N 2 give the mean and variance of the number of
mean and variance of the steps needed to reach sz, hence of t
times that the process is in each state before reaching S2. (The state
passage times. We ean verify that the components of T and T
SI can only be reached through S2.)
The vectors T and T2 give the
with the corresponding entries (in the second column) of M an
mean and variance of the steps needed to reach S2, hence of the first
$ 5.2.
passage times. We can verify that the components of T and T2 agree
As a second application of absorbing theory to ergodic
with the corresponding entries (in the second column) of}J and M 2 ·in
consider the following problem. Assume that we have an
§ 5.2.
Markov chain with r states and that the process is observed on
As a second
application
ergodic chains.
A new
i t is in
a subset Sofofabsorbing
the statestheory
havingtos elements.
consider the following problem. Assume that we have an ergodic
chain is obtained: A single step in the new process correspond
Markov chain with r states and that the process is observed only when
old process t o the transition (not necessarily in one step) from
it is in 11 subset S of the states having s elements. A new Markov T
in S to another state in S. Let sj and st be two states of S.
chain is obtained: A single step in the new process corresponds in the
transition probability will be found by finding the probability t
old process original
to the transition
(not necessarily
onethe
step)
process starting
in s, hits Sinfor
firstfrom
timeaastate
t state s
in S to another state in S. Let Sf and Sj be two states of S. The new
is the probability that it goes to sj in one step, plus the pro
transition probability will be found by finding the probability that the
that i t goes to a state in g and from this state enters S for t
original process starting in Si hits S for the first time at state Sf. This
time a t state sj. Using the results of Chapter 111 we can eas
is the probability that it goes to Sf in one step, plus t,he probability
these transition probabilities. To do this we relabel the s
that it goes to a state in S and from this state enters S for the first
that those in S come first. We then write the transition m
time at state Sf. Using the results of Chapter III we can easily find
in the form
these transition probabilities. To do this we relabel the states so
s transition
5
that those in S come first. We then write the
matrix P
in the form
S
P=
S
FINITE MARKOV CHAINS
114
CH
the use of the 2 matrix is much more natural and convenien
FVRTHER
would normally
use theRESULTS
absorbing methods only to 115
obtain th
detailed information not available-----------------by the Z matrix methods.
The new process
will
be an ofs-state
:vfarkov
cha,in
with Example
transition2a. W
As an
example
a cyclic
chain we
consider
matrix whichstates
we denote
byabsorbing.
P. We shall
now have
find this ma,trix.
We then
sl and sz
Assume a starting state in S. Then the probability of going to each
of the states in S on the first step is given by the matrix T. To take
more than one step, it must enter a state of S, with probabilities given
by U. Then from a given state of S it enters S for the first time at
SEC. 1
state 81 with probabilities given by (I-Q)-lR (see Theorem 3.3.7).
Putting all of this information together, we have that
J5 = T+ U(l-Q}-lR.
It is easily seen that P again ;-epresents an ergoclic chain.
6.1.1. THEoREM. Lei 0: = (aI, a2, ... , as, asH, ... ,a,.) be the fixed
probability vector for P. Then cq = (aI, a2, ... ,as), normalized to
have S7lrn 1, is the fixed probability vector for J5.
The entries of N and N z give the mean and variance of the num
Since
an that
ergodic
hasisa in
unique
fixed
sz. (Th
times
thechain
process
each probability
state beforevector
reaching
point, it is sufficient
to prove
that althrough
is a fixed
for P. T Let
sl can only
be reached
sz.) vector
The vectors
and T Z g
az = (asOl, ... mean
', arlo and
Then
we can of
write
= (aI,needed
,"2). Since
a is sz,
a fixed
hence of t
variance
the asteps
to reach
vector for P we
have
passage times. We ean verify that the components of T and T
PROOF.
with the corresponding entries (in the second column) of M an
$ 5.2.
<X2 = al U + '"zQ.
As a second
application of absorbing theory to ergodic
consider
the
problem.
Assume that we have an
From this last equation we following
have a2(I-Q)=a
1 U or
Markov chain with r states and that the process is observed on
a2 = S
a 1of
U(ithe
_Q)-l.
i t is in a subset
states having s elements. A new
chain
is
obtained:
A
single
Putting this result in the first equation westep
havein the new process correspond
old process t o the transition (not necessarily in one step) from
(il = a1T+o:jU(I -Q)-lR
in S to another
state in S. Let sj and st be two states of S. T
transition
probability
which states that <Xl is a fixed vectorwill
forbeP.found by finding the probability t
original
process
starting
in s, let
hitsusS consider
for the first
time a 6.
t state s
As an example
of the
above
procedure
Example
the probability
that itwalk
goes eX2.mple
to sj in is
one step, plus the pro
The transitionismatrix
for this random
that i t goes to a state in g and from this state enters S for t
82
54
53
s5 of Chapter 111 we can eas
time a t state sj.SI Using
the results
0
1/4
51
these transition
probabilities.
To do this we relabel the s
1 j 41
that those
in S11/3come
0We 0then write the transition m
1/ 3
1/3 first.
52
in the form
1--·-and
P = S3
S4
S5
0
1/ 3
1/ 3
1/ 3
0
0
0
1/3 1/3
l'
1/4
1j4
1/4
0
1/ 4
s
5
i3
The fixed vector for this Markov chain is a= (4/ 38 • 9/3s, 12/ 38 • 9/3s, 4/38).
Assume now that the process is observed only when it is in 81 and 82.
FINITE MARKOV CHAINS
116
CHA
Then we find the new transition matrix as follows. From the
MARKOV
CHAINS
CHAP. VI
cussion of FINITE
this example
in 3.5.5
we have
116
Then we find the new transition matrix as follows.
cussion of this example in 3.5.5 we have
(
(l-Q)-l =
21/9
12/9
15/ 9
24/ 9
9/ 9
9/ 9
From the dis-
Thus the new transition probabilities are
Thus the new transition probabilities are
P = T + U(I -Q)-lR
1/4) +
= C;3
C4
1/ 3
4
'1,)('
12/9
'
0)
0
I/ S
C
1/
1/4
C'
8/ 9
0
9/9 9/ 9 12/ 9
1/ 4
24/ 9
15/ 9
5i n )
D
vector for F" is oi= (4113, 9/13) which is simply the firs
hThe17~fixed
27 . of a normalized to have sum 1.
= 1°1z7
components
For a cyclic example, let us consider the random walk Exampl
The fixed vector for Pis a= (4/ 13, 9/d which is simply the first two
We observe the process in S
components of a normalized to have sum 1.
For a cyelieexample, let ns consider the random walk Example 2a.
We observe the process in S {SI' 82, S3}.
0
1
0
liz
0
0
0
p=
0
0
l/Z
0
0
1/2
0
liz
0
0
l/Z
0
0
0
1
"/ 2
0
0
1
P=
C'
0
l'
2
~
I
,~,)+( :
o/
\ 1/2
0
IX
~)c:
,; )
=
(1,'8, "/4,1/4, 1/4, lfg).
-1~2r"G 0 "~2).
0
; oi = (l/5, 2 / 5 , 2 / 5 ) .
a = (1/5,2/
5,2/ 5).
2
,
P Pdiffers( :only
' slightly from the first three states of P.
The pr
1 .., 2
can leave S only through
s3, and must return there. Hence on
is clianged. We note that, in accordance with 5 6.11.1, oi consi
P differs only
slightly from the first three states of P. The process
the first three components of a normalized.
P33
can leave S only through S3, and must return there. Hence only
is changed. We note that, in accordance with § 6.1.1, ii consists of
the first three components of a normalized.
FINITE MARKOV CHAINS
116
SEC. 2
CHA
Then we find the new transition matrix as follows. From the
cussion of this
example inRESULTS
3.5.5 we have
FURTHER
H7
§ 6.2 Application of ergodic cha.in theory to absorbing Markov
chains. In the preceding section We saw that absorbing Markov chain
theory could furnish us with new information about ergodic chains.
We shall now show that certain results of absorbing chain theory can
be obtained by using the theory of ergodic chains.
Thus the new transition probabilities are
We will need the following generalization of § 5.1.2(b).
6.2.1 THEOREM. Every .i1>larkov chain ledh a single ergodic sd has a
unique probability vector fixed point. ThislJector has positive components for the ergodic states, and zero for the transient states.
PROOF.
Let us write the transition matrix in canonical form.
The fixed vector for F" is oi= (4113, 9/13) which is simply the firs
The matrix Scomponents
is the transition
matrix ofto
thehave
ergodic
sum set;
1. hence it has
of a normalized
a limiting yector
al
>
o.
Let
a
=
(aI,
az),
where
a2
is
vector with
For a cyclic example, let us consider thearandom
walk sExampl
components We
all observe
O. Thenthe
weprocess
see that
a
is
a
pro
babiIity
vector
fixed
in S
point of P. Conversely, let us suppose that /3= (/3I, /3z) is a probability
vector fixed point. Then {3zQ={3z. Hence ,B 2Qn={32, and (3z=
lim /32Qn = O. Thus f31 is a probability vector fixed point of S; hence
by § 5.1.2(b) we have /3 1=al, and f3=a.
Assume now that we have an absorbing Markov chain with r states,
r - 8 of which are absorbing, and 8 non-absorbing. As usual we shall
label the absorbing states so that they come first. The transition
mat.rix then has the form:
\Ve now change this process into a new process as follows.
17
Let
= {Pl, P2, ... ,Pr} be the initial probability vector for the given
2 / 5 , 2 /it
5 ) . is
; oi = (l/5,state
process. \Vhenever this process reaches an absorbing
started over again with the same initial vector 17. The resulting
process is a new Markov chain with transition matrix given by
P differs only slightly from the first three states of P. The pr
...through
, pr-s I s3,
PI,S Pz,
Pr-Hl,
. . , Prreturn there. Hence on
can leave
only
and .must
We
note
that,
in
accordance
with 5 6.11.1, oi consi
is clianged.
Pl, PZ, ... , pr-s \ Pr-s+l, ... ~ , Pr
P' the (first three components of a normalized.
~~, ... ,Pr-8 I pc-HI,· .. , Pr
R
\
Q
)
FINITE MAILKOV CHAINS
118
CHA
The matrix P' is obtained by making all rows corresponding to ab
ing stat,es FINITE
the same
vector CHAINS
n. Let nl= ( P I ,. .CHAP.
. ,p,,) VI and
MARKOV
( P ~ - ~.+. .~, p, r ) Then P' may be written in the form :
The matrix P' is obtained by making all rows corresponding to absorbing states the same vector 7T. Let 7Tl = (PI, ... ,Pr-.) and 7T2 =
(Pr-s+l, ... ,Pr)' Then P' may be written in the form:
118
TBEOREM. The matrix P' represents a Markov chain w
single ergodic set.
6.2.2
Let Imatrix
be the set
states fora which
n has
positive
6.2.2 THEOREM.
pi of
represents
Markov
chain
with compon
a
PROOF. The
Let d beset.the set of all states to which the process can go starting
single ergodic
It is clear that from s,, a = 1, 2 , . . . , r -s, we can go only to s
PROOF. Let I be the set of states for which 7T has positive components.
in J. However, from any state we can go to some s,, since the ori
Let J be the set of all states to which the process can go starting in I.
chain was absorbing. ... , r-s,
3 are
transient.
all states
in go
It is clear that from Sa, a= 1,2, Hence
we can
'only
to states
Since from any state we can go to an s,, and from this to all s
in J. However, from any state we can go to some Sa, since the original
in I, and hence in J, we see that J is an ergodic set. Hence the
states inset
j are
transient.
chain was absorbing.
Hence
all ergodic
process has the
single
J, and
at least one s, E J.
Since from any state we can go to an Sa, and from this to all states
Letina J,bewe
theseefixed
vector
P'. Write
a in the
in I, and hence
thatprobability
J is an ergodic
set.for Hence
the new
process hasathe
single
ergodica1set
=( a
~UZ)
, where
= ( J,
a l ,and
az, .at
. .least
, a,.-s) one
andSaazE =J.(a,-,+I, a,-,+2, . . .
Then, since aP' = a , we have the two equations
Let a be the fixed probability vector for P'. Write a in the form
a= «((1, ((2) where al = (aI, a2, ... , ar-B) and Ct2 = (ar-HI, a T- s+2, ... , aT)'
Then, since CtP' = Ct, we have the two equations
cqgT-s7Tl + CtzR = Ctl
(1)
Let a=(2) a -cQgr-s11'2 + Ct2Q = a2·
1€
The result ci=(cl, ciz) will still be a fixed vector; remembering
1
81Sr-s
1, ourthat
equations
By §6.2.1
we =know
al>O, become
hence algr-s>O. l ..et a=-/;-a.
By $6.2.1 we know that al>O, hence a&,-,>O.
alsr-.
The result cl = {aI, a2} will still be a fixed vector; remembering that
algr-s = 1, our equations become
7Tl +a 2 R = 0:1
From equation (2') we have
(1 ')
+ cl2Q = a2.
(2')
11'2
From equation (2') we have
This inverse exists because the original chain was absorbing.
Theorem 3.3.5 we
cl2 see
= 11'2(1
that_Q)-l.
c2 gives us the mean number of tim
each of the states before absorption, for the given initial proba
This inversevector
existsn.because the original chain was absorbing. From
Theorem 3.3.5
we see
gives
us the in
mean
number
of times in
Using
the that
resulta2just
obtained
equation
(1') we have
each of the states before absorption, for the given initial probability
vector 7T.
Using the result just obtained in equation (I') we have
FINITE MAILKOV CHAINS
118
CHA
The matrix P' is obtained by making all rows corresponding to ab
ing
stat,es the
same vector
n. Let nl= ( P I ,. . . ,p,,)
I<'URTHER
RESULTS
SEC. 2
119 and
( P ~ - ~.+. .~, p, r ) Then P' may be written in the form :
The vector 71'1 gi.ves the probability in the original process of being
absorbed at each absorbing state on t.he initial step, and Tf2(I - Q)-l R
gives this probability for being absorbed in each absorbing state if the
initial step is to a non-absorbing state. Hence al gives the probabilities for absorption in each of the given states, for the initi8J
6.2.2 TBEOREM. The matrix P' represents a Markov chain w
probability vector
17.
single
ergodic set.
We thus see that the single vector a furnishes us with both absorption
probabilities and
the mean
ofof
times
infor
a transient
sta,te
beforecompon
Let Inumber
be the set
states
which n has
positive
PROOF.
absorption. Let
This
economical
the method
d bemethod
the set isof more
all states
to whichthan
the process
can goof
starting
Chapter III,Itifisweclear
are that
interested
from in
s,, aa given
= 1, 2 ,initial
. . . , rprobability
-s, we canvector.
go only to s
It must beinremembered,
furnishes
the the ori
J. However,however,
from any that
state Chapter
we can goIII
to some
s,, since
solution for chain
any initial
vector.
was absorbing.
Hence all states in 3 are transient.
Let us carry
out from
this procedure
for can
the go
random
walk
andExample
from this1.to all s
Since
any state we
to an s,,
The transition
is
and hence
in J, we see that J is an ergodic set. Hence the
in I,matrix
process has the single
set 54J, and at least one s, E J.
81 ergodic
85 82 Sa
/~
0 0 0 vector for P'. Write a in the
Sl fixed probability
Let a be the
a = ( a ~UZ)
, where
a1 = ( a1l , az,
0 .0. . , a,.-s) and az = (a,-,+I, a,-,+2, . . .
S5
Then, since
= a , we0have
p=aP'
0 the
p two equations
52 M q
~
0 q 0
S3 \. 0
\0 P 0 q
S4
D
Let 17=(0,0,0,1,0). Then the new Markov chain obtained by
Let a=By $6.2.1
the above procedure
is we know that al>O, hence a&,-,>O.
a -1€
The result ci=(cl, ciz)
will82still
S3 be
S4 a fixed vector; remembering
Sl 55
81Sr-s = 1, our 81equations become
85
P' = S2
83
GlHD
From equation (2') we have
S4
This is the SaGlE' as Example 3 of § 2.2, with the states reordered.
This inverse
exists because the original chain was absorbing.
The fixed vector
is
Theorem 3.3.5 we see that c2 gives us the mean number of tim
1., before
a
each of the
states
'1,
+ 2 (2
q absorption,
,p 2, q, 1, P ) .for the given initial proba
~'P"
q
vector n.
Thus
Using the result just obtained in equation (1') we have
1
a = ~+
2 (q2, p2, q, 1, pl.
p q
From this we see that, in the original process, the absorption probabilities are q2/(p2+q2) for state 81 and p2/(p2+q2) for state S5. The
FINITE MARKQV CHAIKS
120
mean number of times in each of the states sz, S Z , S* are q
FINITE
MARKOV
CHAINS
l/(p2+q21,
and p/(p2+
q2) respectively.
CHAP.
VI
This is in
agreement
results found in $ 3.4.1.
mean number of
each
the0, states
S2, S3, S4 are
+ q2),is cy
If times
we letin n=
( 0 ,of
0 , 0,
1), then
the resulting chain
Ij(p2+q2), andpJ(p2+
q 2) same
respectively.
is in
d = 2. The
is true if This
n = (0,
0, agreement
c, 0 , d ) , d =with
l - c.theM7e w
results found out
in §this
3.4.1.
example.
If we let 17 = (0, 0, 0, 0, 1), then the resulting chain is cyclic with
d = 2. The same is true jf 17 = (0, 0, C, 0, d), d = 1- c. We will work
out this example.
120
:: (~ ~ : ~ ~)
P' = S2
q
0
0
pO.
Sa
g O p not to find cc first. I n sol
I n calculating
cO
it O
is simplest
0 condition
P 0 q GI
0 + t i 2 = 1 is very helpful.
equation aP'S4=a, the
In calculating a it is simplest not to find a first. In solving the
equation uP' = a, the conclition al + a2 = 1 is very helpful.
1
a = ~+2 (q2+p2qc_pq 2d,p2+pq 2d_p2qc,
P
q
If we let p = q =
we obta,in
q+p 2c-pqd,1-qc-pd,p+q 2d-pqc).
If we let p=q= liz, we obtain
The first two components furnish the probabilities of abso
s l , s g for the chain starting with x. As is to be expected, the
is, the more likely i t is that the process is absorbed in sl.
The first twothree
components
furnish
the the
probabilities
of absorption
components
furnish
mean number
of times inina sta
Sl, S5 for the absorption.
chain startingI twith
1T.
As is to
expected,
sa larger
this isc 1, n
is interesting
to be
note
that forthe
is, the more what
likely citis.is that the process is absorbed in Sl. The last
three components
the application
mean number
in a sta.te
before
,4n furnish
interesting
of of
thetimes
last result
can be
made to
absorption. chain
It is interesting
to P
note
S3 this is 1, no matter
theory. Let
be that
the for
transition
matrix for an ergod
vvlmt cis.
with fixed probability vector a. Let us make one of the state
An interesting
of state.
the lastThen
resultevery
can be
made
ergodicreach
into application
an absorbing
time
thistoprocess
chain theory.willLet
P ibe
the transition
matrix for an
ergodic Then
chain by t
t again
with t'he probabilities
x = {pljf
start
with fixed probability
vector
a.
Let
us
make
one
of
the
states,
say
result the fixed vector for this new process will give51,us th
into an absorbing
state.
Then
everystate
timebefore
this process
reachesBut
81 we
in each
absorption.
the new
number
of times
probabilities
= [Plj}.
the above me
will start it again
is justwith
thethe
original
process,'iT and
timeThen
beforebyabsorption
result the fixed
vector
for this new
IJrocess
give
the mean a
bet'ween
occurrences
of state
sl. ""ill
Hence
by usre-normalizing
number of times
each stat.e 1before
absorption.
the number
new process
we will
obt'ain theBut
mean
of times
firstincomponent
is just the original
process,occurrences
and time before
absorption
time
state between
of state
sl for themeans
original
ergod
between occurrences
of state
51.
Hence
by re-normalizing
a theorem
to have :
Since sl was
arbitrary,
this gives
us the following
first component 1 we will obtain the mean number of times in each
THEOREM.
state between 6.2.3
occurrences
of stateLet81 afor
the $xed
original
ergodic vector
chain.for a
be the
probability
Since 81 was arbitrary, this gives us the following theorem:
6.2.3
THEOREM.
Lei a be the fixed probability vector for an ergodic
FINITE MARKQV CHAIKS
120
mean number of times in each of the states sz, S Z , S* are q
FURTHER
RESULTS
121
l/(p2+q21,
and p/(p2+
q2) respectively. This is in agreement
results found in $ 3.4.1.
chain. Then theIfmean
number
of times in stale Sj between occurrences
we let
n= ( 0 , 0 , 0, 0, 1), then the resulting chain is cy
of state Si is ajlai.
d = 2. The same is true if n = (0, 0, c, 0 , d ) , d = l - c. M7e w
outtransition
this example.
Note that if the
matrix for the chain has column sums 1,
SEC. 2
then the fixed vector has all components equal. This means, by this
theorem, that the mean number of times in each of the other states,
between occurrences of a given state, is the same.
6.2.4 COROLLARY. Let Ii be the vector obtained from a by deleting
component l; let p be the l-th row of P with component l deleted; let Q
be the matrix obtained from P by deleting row l and column l; and let
N=(I-Q)-l, -:-=Ng. Then
I n calculating c i t is simplest not to find cc first. I n sol
equation aP' =a,
1 the condition GI+ t i 2 = 1 is very helpful.
pN,
(a)
-al = l+pT
(b)
-Ii =
ai
1
PROOF. In (a) the left side is the mean number of times in each
If we let p = q =
we obta,in
of the other states between occurrences of S/. The right side is the
same quantity computed from absorbing chain theory. In (b) we
have mil computed from regular and absorbing chain theory, respect- .
ively.
The first two components furnish the probabilities of abso
s l , s g for
chain
starting with
x. As is to be expected, the
If Zthe
is the
fundamental
matrix
of an ergodic chain,
is,
the
more
likely
i
t
is
that
the
process matrix
is absorbed
and A is its limiting matrix, and N is the fundamental
of thein sl.
three components furnish the mean number of times in a sta
absorbing chain obtained by making SI absorbing, and we construct N*
absorption. I t is interesting to note that for sa this is 1, n
f)'om N by inserting an l-th row and l-Ih column of all zeros, then
what c is.
Z = A+(I-A)N*(I-A).
(3)made to
,4n interesting
application of the last result can be
chain theory. Let P be the transition matrix for an ergod
PROOF.
\Vithout loss of generality Vie may choose i= 1. Then,
with fixed probability vector a. Let us make one of the state
using the notation of § 6.2.4, P and A are of the form
into an absorbing state. Then every time this process reach
will start i t again with t'he probabilities x = {pljf Then by t
for this new process will give us th
P=result the fixed vectorA=
~-Q~
\ Q in each state before absorption. But the new
number
of times
is just the original process, and time before absorption me
and hence,
bet'ween occurrences of state sl. Hence by re-normalizing a
first component 1 we will obt'ain the1 mean number of times
1_ A =
-_a__ _ _-_a_-_),
state between occurrences of state sl for the original ergod
-al~
1-~a
Since sl was arbitrary, this gives us the following
theorem :
6.2.5
THEOREM.
(-~-r-p)
(_1
and
6.2.3 THEOREM. Let a be the $xed probability vector for a
N*
(
0 ' 0 \
i
\
ON)"
122
FINITE MARKOV CNAIKS
Then
~FINITE
122
MARKOV CHAINS
CHAI'.
VI
Then
\01
/ 0 I -IPN )
alp7
-pN+ p 7 ~
(I-P)X* =
(I-P)N" (I-A' - a ~ [ I-&
aWT
- pN +
Making use of 5 6.2.4,
(I-P)N* (I-A) = (
( I - P )1N *git( I - A ) = I - A
-al~
[
I
P
+
A
]
[
A
+
(
I
-A)N*(I-A)]
= A+(I-P)N*(I-A
Making use of § 6.2.4,
= A +(I-A)
(I -P)N* (I-A) = I-A
= I.
[I-P+A][A+(I-A)N*(I-A)) = A+(I-P)N*(I-A)
Hence
= A+(I-A)
A + ( I - A ) N * ( I - A ) = ( I - P + A ) - 1 = Z.
= I.
I t is interesting to note that in (3) we may use any N* ob
Hence
making any one state
A+(I-A)N*(I-A)
= absorbing.
(I-P+A)-l = Z.
4PTa).
6.2.5
weN*
let obtained
N = {n(')ig)
If may
in use
It is interesting to noteCOROLLARY.
that in (3) we
any
by and N
= t(0, = 0 if i = l , then
and n(Otr
=
making anyone state
absorbing.
6.2.6
If in § 6.2.5 we let N = {n(l)tj} and Ng = {t(l)d,
= l(l)i = 0 1/ i = l, then
COROLLARY.
and n(l)tj = n(l)j;
Zij
= aj + n(/)ij- 2: akn(llkj-aj[U)j +aj
2:
akt(l)k.
(a)
k *1
k*I
These quantities
are obtained directly
from $ 6.2.5, mak
formulas
may
he used t o derive(b)many i
5 4.4.7
for (b). These
mij
= (l/a'j)(n(l)jj
- n(l)lj
+ dij ) + t(lli
- t(l)j.
results. A few of these are given below. The number n
These quantities are obtained directly from § 6.2.5, making use of
number of times the process is in s f , starting in si, before r
§ 4.4.7 for (b). These formulas may be used to derive many interesting
for the first time.
results. A few of these are given below. The number n(lllj is the
number of times the process is in Sj, starting in 5i, before reaching Sl
for the first time.
6.2.7
COROLLARY.
n(l)ii
%
n(l)jj
n(l)ii
n(i)11
ai.
at
(c) (a) m/j+mjl = m
l + +mlf
- - (I-h(l)ij),
for i',j
7'= l. = =
(b)imjl
aj
n(l)jj
PROOF.
(b) mjl+mlj = - _ .
aj if i, j # l ,
From 3n(l)jj
6.2.6(b),
(c) - - = - .
n(j) II
aj
PROOF.
al
From § 6.2.6(b), ifi,j7'=l,
+ mjl = mij + t(l)j = (lfaj)(n(l)jj - n(l)ij + d ij ) + (nil
I-h(lltj = 1- n(l)jj-d ij
n(l)jj-n(l)ii+dij
mij
n(l)jj
Hence (a) follows.
n(l)}}
122
FINITE MARKOV CNAIKS
Then
SEC. 3
FURTHER RESULTS
123
Hence (0) follows.
1II11+m11
n(l)jj/U1
mlj + mjl = n(J)Il!al'
4
alp7
-pN+ p 7 ~
(I-P)N" (I-A' -
Hence (c) follows.
- a ~ [ I-&
If in the Land
of Ozuseexample
we make R absorbing, we obtain
Making
of 5 6.2.4,
(see § 6.1),
(I-P)N* (I-A) = I - A
[I-P+A][A+(I-A)N*(I-A)]
= A+(I-P)N*(I-A
= A +(I-A)
= I.
Hence
A+(I-A)N*(I-A)
= (I-P+A)-1
= Z.
I t is interesting to note that in (3) we may use any N* ob
making any one state absorbing.
COROLLARY. If in 6.2.5 we let N = {n(')ig)and N
= t(0, = 0 if i = l , then
and n(Otr=
, I
- -/5
86
(
3
-14)
6These
63 quantities
6 = are
Z obtained directly from $ 6.2.5, mak
5
4.4.7 for (b). These formulas may he used t o derive many i
-14
86 of these are given below. The number n
results.3 A few
f , starting
before r
of times
thefrom
process
in sfrom
as we saw in § number
4.3. Using
results
§ 4.4isand
§ 6.1, inwesi,can
for the
firstand
time.
illustrate Corollaries
6.2.6
6.2.7.
mss = (l/a·s)(nss-nss) +tN-ts
= (5/z)(8/ 3 _4/s)+8/ 3 _10/3 = 8/3
nSS
.
1°/3+ 1°/3 = (8/s)/(2/5)'
n(l)ii
n(l)ii ai.
(c) - = (b) mjl +mlf =
%
n(i)11
§ 6.3 Combining states. Assume that we are givenat an r-state
Markm' chain with
transition
P and
vector 'IT. Let
PROOF.
Frommatrix
3 6.2.6(b),
if i, jinitial
#l,
A={Al' A 2 , .•• , At} be a partition of the set of states. We form a
new process as follows. The outcome of the j-th experiment in the
new process is the set Ak that contains the outcome of the j-th step
in the original chain. We define branch probabilities as follows: At
the zero level we assign
mSR
+ mRS = -
as
or
(1)
At the first level we assign
Pr~[fl E A11fo E AiJ-
FINITE MARICOV CHAINS
124
I n general, a t the n-th level we assign branch probabilities,
:FINITE MARKOV CHAINS
124
CHAP. VI
In general, at the The
n-thabove
level we
assign branch
probabilities,
procedure
could be
used to reduce a process with
large
numberE of
number o
Pr,,[f
As stat,es
1\ ... to
,\f1a Eprocess
Aj I\fo with
E Ai]. a smaller (2)
n E Atifn-1
We call this process a lumped process. It is also often the
The above procedure could be used to reduce a process with a very
applications that we are only interested in questions which r
large number of stat,es to a process with a smaller number of states.
this coarser analysis of the possibilities. Thus i t is importan
We call this process a lumped process. It is also often the case in
able to determine whether the new process can be treated by
applications that we are only interested in questions which relate to
chain methods.
this coarser analysis of the possibilities. Thus it is important to be
W e shall
saytreated
that a by
Markov
chain i s lu
DEFINITION.
able to determine 6.3.B
whether
the new process
can be
Markov
chain methods. with respect to a partition A={A1, A z , . . . , AT) if for every
vector 71 the lumped process de$ned by ( 1 ) and ( 2 ) i s a Marko
6.3.1 DEFINITION.
shall sayprobabilities
that a Markov
chain
is lumpable
and theWe
transition
do not
depend
on the choice of 7
with respect to a partition A={A), A 2 , ••• ,Ar} if for every starting
We process
shall seedefined
in thebynext
section
that,
a t leastcha.in
for regular
vector 7T the lu.mped
(1) and
(2) is
a Markov
the condition
t h do
a t not
the depend
transition
and the transition
probabilities
on theprobabilities
choice of 7T. do not depen
follows from the requirement that every starting vector give a
We shall see chain.
in the next section that, at least for regular chains,
the condition that the transition probabilities do not depend on 7T
Then pi^, represents the probability of
Let pi,,=
pix.every
follows from the requirement
that
starting vector give a Markov
2
chain.
Let PiA, =
state st into set Aj in one step for the original Markov ch
2:from
Then
represents the probability of moving
6.3.2 THEOREM.
A necessary and suficient condition for a
Pik·
PiAj
Sk E Aj
respect Markov
to a partition
from state Si into chain
set Aj to
in be
onelumpable
step for with
the original
chain. A={A1, Az,
is that for every pair of sets At and Ai, pkn, have the same
6.3.2 THEOREM. A necessary and sufficient condition for a lVarkov
form the transitio
every st in At. These common values
chain to be lumpable wl:th respect to a partition A = {AI, Az, ... , As}
for
the
lumped
cha,in.
is that for every pair of sets Ai and Aj , PkA; have the same value for
For thevalues
chain{pjj}
to beform
lumpable
i t is clearly
necessary
These common
the transition
malTix
every Sk in A j • PROOF.
for the lumped cha,in.
Pr,[fi Ai 1 fo E A ]
For the
chain
to befor
lumpable
is clearly
that Call this
be the
same
every nitfor
which necessary
i t is defined.
I n particular
this
must
be
the
same
for x having a
value $if. Pr~[fl
E Ailfo E Aj]
k-th component, for state sk in At. Hence P E A , = h [ f l E Af]
be the same for every 7T for which it is defined. Call this common
Thus
is necessary. To pr
every sk inthis
At.must
value Pij. In particular
be the
the condition
same for given
7T having a 1 in its
sufficient,
we
must
show
that
if
the
condition
is satisfied the pro
k-th component, for state Sk in Ai. Hence P1c4/ = Prk[fl E Aj] = Pii for
( 2 ) depends only on A, and At. The probability ( 2 ) may be
every SA; in Ai. inThus
the condition given is necessary. To prove it is
the form
sufficient, we must show that if the condition is satisfied the probability
Pr,,[fl 6 At1
(2) depends only on As and At. The probability (2) may be written
where
x'
is
a
vector
with
non-zero
components only on the sta
in the form
I t depends on Pr,,·[f
n and1 on
the first n outcomes. However, if
EAt]
for all sk in A,, then i t is clear also that Pr,,[fl E At] =
where 7T' is a vector
with non-zero
components
the Ai.
states of As.
onlyonly
on A,onand
probability
in (2) depends
It depends on" and on the first n outcomes. Howe"l-er, ifPrk[fl EAt] =
pst for all Sk in As, then it is clear also that Pr",[fl E AIJ = pst. Thus t.he
probability in (2) depends only on As and At.
PROOF.
FINITE MARICOV CHAINS
124
SEC. 3
6.3.3
I n general, FURTHER
a t the n-th RESULTS
level we assign branch probabilities,
125
~~~~------------~~
Let us consider the Land of Oz example.
Recall
The above procedure could be used to reduce a process with
that P is given by
large number of stat,es
to
a process with a smaller number o
R
N
S
EXAMPLE.
We call this processI a lumped process. It is also often the
R we:' 2 are
1/4only interested in questions which r
applications that
this coarser
analysis
of
possibilities.
Thus i t is importan
0the 1/
P=N
2 .
able to determine whether the new process can be treated by
1"
S
chain methods.
.' 1/4 1/2
'I,)
Assume now 6.3.B
that ,ve
are interested
and "bad"
W e only
shall in
say "good."
that a Markov
chain i s lu
DEFINITION.
weather. This suggests
lumping
R.and S. A={A1,
We note
with respect
to a partition
A zthat
, . . . ,the
AT)probaif for every
bility of moving vector
from 71
either
of theseprocess
states de$ned
to N is by
the( 1same.
the lumped
) and ( 2Hence
) i s a Marko
if we choose forand
ourthe
partition
A
=
({N},
{R,S})
=
(G,
.8),
the
condition
transition probabilities do not depend on
the choice of 7
for lump ability is satisfied. The new transition matrix is
We shall see in the next section that, a t least for regular
G transition
B
the condition t h a t the
probabilities do not depen
follows from the requirement
that
every
starting vector give a
G
P=
chain.
B 1/4
Let pi,,=
pix. Then pi^, represents the probability of
(0
2
Note that the condition for lumpabihty is not satisfied for the
intoPSA
set, Aj
in one step
for the
original Markov ch
from{N,S})
state st
partition A=({R]
since
=PNR=1/
2 andpsA
, =PSR=1/4.
6.3.2weTHEOREM.
A necessary
and suficient
condition
Assume now that
have a :vlarkov
chain which
is lumpable
with for a
chain
to
be
lumpable
with
respect
to
a
partition
A={A1,
respect to a partition A={Al, .... As}. We assume thaJ, the original Az,
is that
pairchain
of sets
and Ai, pkn,
have
chain had r states
and for
the every
lumped
hasAt8 states.
Let U
be the
the same
form the
transitio
everyi-th
st in
At.is the
These
common values
8 x r matrix whose
row
probability
vector having
equal
for theinlumped
components for states
Ai andcha,in.
0 for the remaining states. Let V be
the r x s matrix PROOF.
with theFor
j-ththe
column
with 1i t'sisinclearly
the comchain atovector
be lumpable
necessary
ponents corresponding to states in Aj and O's otherwise. Then the
Pr,[fi Ai 1 fo E A ]
lumped transition matrix is given by
be the same for every n for which i t is defined. Call this
P = UPV.
value $if. I n particular this must be the same for x having a
In the Land ofk-th
Oz example
this
is state sk in At. Hence P E A , = h [ f l E Af]
component, for
is necessary. To pr
every sk inUAt. Thus thePcondition given
V
sufficient, we must show that if the condition is satisfied the pro
1/4
( 2 ) depends only (/2
on A, and At. The probability ( 2 ) may be
in the form
pO
1/2 0 liz 1
0 1 2/ i
2
Pr,,[fl 6 At1
\ 1/4 1/ 1/ 2 0
where x' is a vector with non-zero components only on the sta
I t dependsUon n and onPV
the first n outcomes. However, if
for all sk in A,, then i t is clear also that Pr,,[fl E At] =
o depends
'
only on A, and Ai.
probability in (2)
C
C~2
'1,)("
°\
4
0
~)
'
:
'
)
C C~4 3~J
1J 1°{4
3/4
Note that the rows of P V corresponding to the elements i
be VI
true in ge
CHAP.
MARKOVareCfL\IXS
same set FINITE
of the partition
the same. This will
for a chain which satisfies the condition for lumpabilit,y. The m
Note that the rows of PI' corresponding to the elements in
U then simply removes this duplication of rows. Thethechoice of
same set of the partition are the same. This will be true in general
by no means unique. I n fact, all that is needed is that the i-t
for a chain which satisfies the condition for lumpability. The matrix
should be a probability vector with non-zero components onl
U then simply removes this duplication of rows. The choice of U is
states in Ai. We have chosen, for convenience, the vector with
by no means unique. In fiLct, all that is needed is that the i-th row
components for these states. Also i t is convenient for proo
should be a proba,bility vector with non-zero components only for
assume that the states are numbered so that those in A1 come
states in Ai. \Ve have chosen, for convenience, the vector with equal
those in A z come next, etc. I n all proofs we shall unttersta,nd tha
components for these states. Also it is convenient for proofs to
had been done.
assume that the states are numbered so that those in A) come first,
The following result will be useful in deriving formulas for lu
those in A2 come
next, etc. In all proofs we shall understand that this
chains.
126
had been done.
The following6.3.4
result,THEOREM.
will be useful
formulas
for lumped
I f Pin i deriving
s the transition
matrix
of a chain lum
chains.
w i f h respect to the partition A, and i f the matrices U nnd,V are d
-as above-with vespect to this partition, then
6.3.4 THEOREM. If P is the transition matrix of a chain lumpable
with respect to the partition A, and If theVmalriccs
are defined
U P V = UP and'V
V.
-as above-with resljecl to this partition, thin
PROOF. The matrix V U has the form
VUPV = PV.
PROOF.
(3)
The matrix VU has the form
Wl I
VU = (
0
'
0
--~--iw:--o-
--1-----
)
,
where W 1 , Wz, and W 3 are probability matrices. Condition ( 3 )
0 fixed
I Wavectors of V U . But since the
.that the columns o
of P V are
s lumpable, the probability of moving from a state, of At to the
where W), Wis
2, and W3 are probability matrices.
(3) states
the same for all st,ates in At. 'lenceCondition
the components
of a colu
.that the columns of PV are fixed vectors of VU. But since the chain
P V corresponding to Aj are all the same. Therefore they f
IS lumpable, the probability of moving from a state of Ai to the set Aj
fixed vector for Hrj. This proves ( 3 ) .
is the same for all states in Ai. hence the components of a column of
PV corresponding to Ai are all the same. Therefore they form a
6.3.5 THEOREM.If Y ,A, U , and T' are as i n Theorem 6.3.4
fixed vector for Wj. This proves (3).
corrdition (3) i s eq~iimlentto lumprrbility.
6.3.5 THEOREM.
If We
P, A,have
[T, and
V are
as that
in Theorem
6.3.4, then
PROOF.
already
seen
(3) is implied
by lumpa
condition (3) is
equivalent let
to lumpability.
Conversely,
us suppose that ( 3 ) holds. Then the columns
are fixed vectors for V U . But each HIj is the transition matrix
We have already seen that (3) is implied by lumpability.
ergodic chain, hence its only fixed column vectors are of the fo
Conversely, let us suppose that (3) holds. Then the columns of PV
Hence all the components of a colun~nof P V corresponding to o
are fixed vectors
for Vbe
U. theBut
each That
Wj is is,
thethe
transition
of an
chain ismatrix
lumpable
by $ 6 . 3 .
Aj must
same.
ergodic chain, hence its only fixed column vectors are of the form
Hence all the components of a column of P V corresponding to one set
Aj must be the same. That is, the chain is ]nmp"ble by § 6.3.2.
PROOF.
cr
Note that the rows of P V corresponding to the elements i
FURTHER
true in ge
same set of the
partitionRESULTS
are the same. This will be 127
The m
for
a
chain
which
satisfies
the
condition
for
lumpabilit,y.
Note that from (3)
U then simply removes this duplication of rows. The choice of
f>2 = UPVUPV
by no means unique. I n fact, all that is needed is that the i-t
VP 2 vector
V
should be a probability
with non-zero components onl
and in general
states in Ai. We have chosen, for convenience, the vector with
= UPnV.
components for p71
these
states. Also i t is convenient for proo
assume
that
the
states
are numbered
so process.
that those in A1 come
This last fact could also be verified directly
from the
those in A z come next, etc. I n all proofs we shall unttersta,nd tha
Assume now
is an absorbing chain. "We shall restrict our
had that
been Pdone.
discussion to the
whereresult
we lump
statesin of
the same
kind. for lu
Thecase
following
will only
be useful
deriving
formulas
That is, anychains.
subset of our partition will contain only absorbing states
or only non-absorbing states. We recall that the standard form for
an absorbing chain
6.3.4is THEOREM. I f P i s the transition matrix of a chain lum
w i f h respect to the partition A, and i f the matrices U nnd,V are d
-as above-with
p= vespect to this partition, then
RIQ
SEC. 3
(_1I~).
VUPV = PV.
We shall write U in the form
PROOF. The matrix V U has the form
(
V=
U 1 ." 0
\\
-o-I~t
where entries of VI refer to absorbing states and entries of U 2 to nonabsorbing states. Similarly we write V in the form
where W 1 , Wz, and W 3 are probability matrices. Condition ( 3 )
.that the columns of P V are fixed vectors of V U . But since the
s lumpable, the probability of moving from a state, of At to the
is the
same for
st,ates inforAt.lumpability,
'lence the components
of a colu
Then, if we
consider
the all
condition
VU PV = PV,
V corresponding
Aj are all
same. Therefore
we obtain inPterms
of the aboveto matrices
the the
equivalent
set of con-they f
ditions:
fixed vector for Hrj. This proves ( 3 ) .
rr
Ii
(4a)
1
6.3.5 THEOREM.If Y ,A, U , and T' are as i n Theorem 6.3.4
(4b)
corrdition (3) i s eq~iimlentto lumprrbility.
(4c)
PROOF. We have already seen that (3) is implied by lumpa
Since U 1 V 1 =Conversely,
I, the first let
conciiti01,
is automatically
satisfied.
Then the columns
us suppose
that ( 3 ) holds.
The standard
form vectors
for the for
transition
matrix
P HIj
is obtained
from
each
is the transition
matrix
are fixed
V U . But
ergodic chain, hence its only fixed column vectors are of the fo
f> =Hence
UPV
all the components of a colun~nof P V corresponding to o
Aj must be the same. That is, the chain is lumpable by $ 6 . 3 .
)(_11_0)(~I~_).
°\
f> = (~I_O
\
0
I U2
R
I Q
V2
128
Multiplying this out we obtain
FINITE ..YLO\RKOV CHAINS
CHAP. VI
Multiplying this out we obtain
11
0
)
Hence we have
(
p = [T2RVI I U2QV-;fi .= U z R B l
& = U2QVz.
Hence we have
R (4c)
= UZRV
1
From condition
we obtain
Q = U ZQV 2•
From condition (4c) we obtain
Q2 = U
2 U 2QV 2
More generally
we2QV
have
= U 2Q2V 2 •
More generally we have
From the infinite series representation for the fundamental
N we have Qn = U zQnV 2 •
ig = for
I + Qthe
+ @fundamental
+ .. .
From the infinite series representation
matrix
N we have
= hl2PV2+ UzQV2+ . . .
R = I +Q+(P+ ...
= Uz(Z+Q+Q2+ . . . )Vz
10
= ZU2NVz.
U zIV 2 + U 2QV
+
[72(1
+Q+Q2+ ... )V2
From this we
obtain
B = U 2NV 2 .
4 = UzNVz[
From this we obtain
f
= U2NV2~
= UzN~
f
= UZT
f
i = U2Nl
4 = Uzr
and
B = BR = U ZNV ZU 2 RV 1
L'2NRVl
Hence all threefJof=the
quantities X, T ,and B are easily obtained
fJ = the
U 2 BV
j •
lumped chain from
corresponding
quantities for the origina
An important consequence of our result i'=U ~ is
T the fol
Hence all three of the quantities S, T, and B are easily obtained for the
Let Ai be any non-absorbing set, and s k be a state in Af. W
lumped chain from the corresponding quantities for the original chain.
to be a probability vector with li in
choose the i-th row of
An important consequence of our result f = U 2T is the following.
component. But this means that t,= tk for all s k in A,. Henc
Let Ai be any non-absorbing set, and Sk be a state in Ai. \Ye can
a chain is lumpable, the mean time to absorption must be th
choose the i-th row of U 2 to be a probability vector with 1 in the Sk
for all starting states s k in the same set At
component. But -4s
this
that
= Ik for all Sk in Ai.
Hence
when walk e
anmeans
example
of tithe
above, let us consider
the random
rJ2
a chain is lumpable, the mean time to absorption must be the same
for all starting states Sk: in the same set Ai
As an example of the above, let us consider the random walk example
Multiplying this out we obtain
FURTHER RESULTS
SEC, 3
129
with transition matrix
Hence we have
S5
o
0
I~
~ fi =~ U z R B l
& = U2QVz.
p = 52
-;1-;-O-I~!2~;From condition (4c) we obtain
53
0
0
11/2
0
1/2
54
0
1/2
0
liz
0
\Ve take the partition A=({Sl, S5}, {52, 54}, {S3}), For this partition
the condition for
lumpabiIity
Notice that this would not
More
generally is
wesatisfied,
have
have been the case if we have unequal probabilities for moving to the
right or left,
From the original
chain
foundseries representation for the fundamental
From
the we
infinite
N we have
S2
84
83
r ';')
ig =
1 I+Q+@+ . . .
2= hl2PV2+ UzQV2+ . . .
l'
= Uz(Z+Q+Q2+ . . . )Vz
3/z
\\ 2
10 = U2NVz.
52
N = S3
S4
From this we obtain
T=
(:13)
51
52
B= 53
54
4 = UzNVz[
i = U2Nl
S5 4 = Uzr
:;:)
C
1/4
3/4
The corresponding quantities for the lumped process are
Hence all three of the 0quantities
0
0 X, T ,and B are 0easily obtained
lumped chain
from
the
corresponding
for the origina
o quantities
1 0
0\ /
0
0
'I, An
0 0important consequence of our result i'=
U ~ is
T the fol
P=
0
0 1 '2 0
O·
0
0
liz and
0
0
any non-absorbing
s k be a state in Af.
W
Let Ai be
'/
'/0 0 1/2 set,
0
0 a probability
1/2
0 vector
to
be
with
li
in
the
i-th
row
of
o 0 choose
0
0component. But this
1/2means
liz t,=
0 that
o tk for all s k in A,. Henc
chain
Al aAz
..13 is lumpable, the mean time to absorption must be th
for all starting states s k in the same set At
Al
1 0-4s an example of the above, let us consider the random walk e
0
= A2
~
C'
I
\
A3
\~"
1~2
,;.)
0
rJ2
)(
\0
~)
CHAP. VI
FINITE ~IARKOV CHAINS
130
0\
'~')C :::)(: ~)
1
C~Z
N
0
2
1/2
I
Az As
Az
A3
G :)
C~2
0
=
B= C~2
0
f
'~')(:)
A2 C\
A3
4)
·c :;:)Gl
l~Z) 1/2
1/4
j 4
Az now that we have an ergodic chain which satisfies
Assume
dition for
As lumpability for a partition A. The resulting chain
ergodic. Let 2 be the limiting matrix for the lumped chain
}.. ssume now that
we have
we know
thatan ergodic chain which satisfies the con-
C)·
dition for lumpability for a partition A. The resulting chain will be
ergodic. Let A be the limiting matrix for the lumped chain. Then
we know that
n
A = lim UP~+UP2V+
+UPnV
n
I n particular, this sta,tes t h a i the components of G are obtain
simply adding components in a given set. Similarly f
Aa =byUAV.
infinite series representation for the fundamentnl matrix 2 we
In particular, this RtR,t.es that thc components of a are obtained from
a by simply adding components in a givcn set, Similarly from the
infinite series repref'entatioll for the fundamental matrix Z we have
There is in general no simple relation between 31 and &. H
the mean timeZ to
from a state in At to the set Aj, in the
= go
UZV.
process, is the same for all states in A,. To see this we need on
There is in general
no simple
between
fif and M.
absorbing.
We know
that theHowever,
mean time to ab
the statcs
of A, rehtion
the mean time to
go from
state
At to the
set Achosen
j , in the original
is the
same8. for
allinstarting
states
from a given set
process, is the same for all states in Ai, To >'ce this we need only make
the state" of Ai absorbing. We know that the mean time to absorption
is the same for all starting states chosen from a given set. If, in
FURTHER RESULTS
SEC. 3
131
addition, Aj happens to consist of a single state, then '"LIt may be
found from M.
\Ve can also compute the covariance matrix of the lumped process.
As a matter of fact we know (see § 4.6.7) that the covariances are
easily obtainable from C even if the original process is not lump able
with respect to the partition, that is, if the lumped process is not a
Markov chain. In any case
CO =
2:
Okl·
sk in Ai
S l in Ai
Let us carry out these computations for the Land of Oz example.
For A we have
A=
Ole: 'I,) (0 ;)
1/5
1
c;z 0 liz
2!.
10
1/5 2/ 5 \ 1
l! 5 2/5 0
Cs 4~5)
Assume now that we have an ergodic chain which satisfies
1/5 4 i 5
dition for lumpability for a partition A. The resulting chain
86/ 75
the limiting
matrix for the lumped chain
ergodic. Let 2 be
3/75
/ 0know that
we
z= I
\ 1/2
0
,;J(
\
6/ 75
--
63/ 75
3/ 75
-"1")("
6/ 75
I
86/ 75
0
CJ,s 12/75)
3
3/ 75
72/ 75
- 12/ 125)
c=
I n particular, this sta,tes t h a i the components of G are obtain
12/ 125
(
12/ 125
_12/125
a by simply adding components in a given set.
Similarly f
R N for
S the fundamentnl matrix 2 we
infinite series representation
R
C'
4
1')
Sja 5 >0S/3 .
M =N
There is in general no simple relation between 31 and &. H
10/3from4 a state
5/2 in At to the set Aj, in the
the mean time Sto go
process, is the same for all states in A,. To see this we need on
From the fundamental matri" Z we find,
the statcs of A, absorbing. We know that the mean time to ab
N B states chosen from a given set
is the same for all starting
N
kf=
B
15
\4 5~J
13~
Note that the mean time to reach N from either R or S is 4. H
N in the lumped
is a single
element set. CHAP.
This common
v
VI
:FINITEprocess
l\IARKOV
CHAINS
-----------------is the mean time in the lumped chain to go from B to N. Simila
Note thatthe
thevalue
mean5 time
to reach N
fromMeither
or S is that
4. Here
is obtainable
from
. WeRobserve
the mean t
N in the lumped
process
is B
a single
element set.
This the
common
to go from
N to
is considerably
less than
mean yalue
time to go f
is the meanN time
in the
chain
from
B to process.
N. Similarly,
to either
of lumped
the states
in Btoingothe
original
the value 5 is obtainable from M. We observe that the mean time
lumpability.
to go from N §to6.4B isWeak
considerably
less than
the mean
time
to go to
from
I n practice
if one
wanted
apply Xar
for which
the states have been combi
of the ideas
statestoin aB process
in the originaJ
process.
N to either chain
with respect to a partition A={.41, Az, . . . ,
it is most natura
§ 6.4 Weak
lumpability.
In practice
if onebe
wanted
to apply
Markov
require
that the resulting
process
a MarBov
chain
no matter w
chain ideaschoice
to a process
the states
have been
combined,
is madeforforwhich
the starting
vector.
However,
there are s
with respect
to a partition
A = {AI,considerations
A2, ... , An}, when
it is most
natural to
interesting
theoretical
we require
only tha
require thatleast
the one
resulting
process
be lead
a Markov
chain nochain.
matterWhen
what this is
starting
vector
to a Markov
choice is made
forshall
the say
starting
vector.
are with
some
lumpable
respect to
case we
that the
processHowever,
is weakly there
interesting partition
theoretical
considerations
when
we
require
only
that
A. We shall investigate the consequences of atthis wea
least one starting
vector
a Markov
this is the to reg
assumption
in lead
this tosection.
Wechain.
restrictWhen
the discussion
case we shall
say that
theresults
processofisthis
weakly
lumpable
with respect
to the
chains.
The
section
are based
in part
on results
"VeBurke
shall and
investigate
the
consequences
of
this
weaker
partition A.C. K.
M. Rosenblatt.?
assumption in
section.
We restrict
rcgul&ror not
Forthis
a given
starting
vector nthe
, to discussion
determine towhether
chains. The
results
this section
aremust
based
in part
on results of
ofthe form
process
is aof
Markov
chain we
examine
probabilities
C. K. Burke and M. Rosenblatt. t
For a given starting vector 77", to determine ,.,-hether or not the
process is aFor
Markov
chain
we must examine probabilities of the form
a given
?r the process will be a Markov chain if these probabili
do not depend upon the outcomes before the n-th.
(1)
We must find conditions under which the knowledge of the outco
For a givenbefore
77" the process will be a "Markov chain if these probabilities
the last one does not affect the probability (1). Let us see h
do not depend
the outcomes
before
n-th. the information in ( I ) , we kn
suchupon
knowledge
could affect
it. theGiven
We mustthat
find after
conditions
under
which
the
knowledge
outcomes
n steps the underlying chain is inofathe
state
in A,, but we
before the last
one does
not affect
(1).however,
Let. us assign
see how
not know
in which
statethe
i t probability
is. We can,
probabili
such knowledge
could
affect
it. Given
information
(1),aswefollows:
know For
for it#s
being
in each
state the
of A,.
We do in
this
that after nprobability
steps the vector
underlying
is by
in pj
a state
in As, but vector
we do formed
8, wechain
denote
the probability
not know in
which all
state
it is. Wecorresponding
can, however,toassign
making
components
statesprobabilities
not in A3 equal t
for it,s being
each
state ofcomponents
As. We do this as follows: For any
andinthe
remaining
proportional to those of p. We s
probabilitysay
vector
fJ,
we
denote
by fJj
vector
formed 0byin Aj we
to the
Aj. probability
(If ,R has all
components
that ,i3j is p restricted
making allnot
components
to
states
not
in
Ai equal to 0
define pi.)corresponding
Consider now the information given in (1). The f
and the remaining
proportionalastochanging
those of our
fJ. initial
We shall
Ai may be interpreted
vector to
that f o E components
say that fJiLearning
is /8 restn:cted
to
Ai.
(If
fJ
has
all
components
0
in
Ai
we do this vec
then that f l E Ai may be interpreted as changing
not define to
fJj.)(nipif.
Consider
t,he information
in (1).
fad
We now
continue
this process given
until we
have The
taken
into acco
that fo E Aiallmay
beinformation
interpretedgiven
as changing
initial
to 77"i.assignm
We are
led vector
to a certain
of the
in (1). our
Learning then
that fl E Aj for
maythe
bestates
interpreted
changing
vector
of probabilities
in A,. asFrom
thesethis
probabilities
we
to (n/P)i. vVe continue this process until we have taken into account
7 C. K. Burke
and M.
function of
a Markov chain," An
all of the information
given
in Rosenbiatt,
(1). We "A
areMarkovian
led to a certain
assignment
of Mathematical Slalistics, 29: 1 1 12-1 122, 1958.
of probabilities for the states in As. From these probabilities we can
t C. K. Burke and. M. Rosenblatt, ·'A Markovian function of a Markov chain,·' Annals
oj Mathematical Statistics, 29: 1112-1122, 1958.
SEC. 4
Note that the mean time to reach N from either R or S is 4. H
N in the lumped
process RESULTS
is a single element set. This 133
common v
FURTHER
is the mean time in the lumped chain to go from B to N.
Simila
easily compute
probability
of a transition
At on
the next
5 is obtainable
from M .to We
the the
value
observe
thatstep.
the mean t
But note thattothis
probability
be quite different
for different
go from
N to Bmay
is considerably
less than
the meankinds
time to go f
of information.
example,
our information
plaee
high probaN to For
either
of the states
in B in the may
original
process.
bility for being in a state from which it is certain that we move to At.
§ 6.4 ofWeak
lumpability.
I n practice
if one wanted
apply Xar
A different history
the process
may place
low probability
on to
this
chain
ideas to agive
process
for which
the states
have
been combi
state. These
considerations
us a clue
as to when
we could
expect
with
respect
to a partition
. . . are
, suggested.
it is most natura
that the past
could
be ignored.
Two A={.41,
differentAz,
cases
resulting
process begained
a MarBov
no matter w
First would require
be the that
case the
where
the information
fromchain
thc past
made :For
for the
starting
vector.
However,
there are s
would not dochoice
us anyis good.
example,
assume
that the
probability
considerations
for moving interesting
to the set A/theoretical
from a state
in As is the when
same we
for require
all statesonly tha
least
one starting
vector lead
a Markov
in As. Then
clearly
the probabilities
for tobeing
in eachchain.
state When
of As this is
weakly lumpable
with respect to
case we
say thatfor
thethe
process
would not affect
ourshall
predictions
next is
outcome
in the lumped
partition
We shall
the ability
consequences
of this wea
process. This
is the A.
condition
we investigate
found for lump
in § 6.3.2.
assumption
in this section.
We restrictAssume
the discussion
A second condition
is suggested
by the following:
that no to reg
results ofis, this
sectionend
are up
based
on results
matter whatchains.
the pastThe
information
we always
withinthepart
same
C. probabilities
K. Burke and
Rosenblatt.?
assignment of
forM.being
in each of the states in As. Then
givennostarting
vector
n , predictions.
to determine V/e
whether
again th" pastFor
cana have
influence
on our
shall or not
is aalso
Markov
see that thisprocess
case can
arise. chain we must examine probabilities of the form
We have indicated above that the information given in (1) can be
represented by a probability vector restricted to As. This vector is
obtained from
initial
vector
1T bywill
a sequence
of transformations,
For the
a given
?r the
process
be a Markov
chain if these probabili
each time taking
account
bit ofbefore
information.
do notinto
depend
uponone
themore
outcomes
the n-th. That is,
we form the sequence
We must find conditions under which the knowledge of the outco
before the last one
1T! not affect the probability (1). Let us see h
07 1 does
such knowledge1T2
could (affect
it. Given the information in ( I ) , we kn
1T IP)}
}
the(07ZP)k
underlying chain is in a state in (2)
A,, but we
that after n steps
1T3
not know in which state i t is. We can, however, assign probabili
for it#s being in11meach
of A,. We do this as follows: For
= ( 1Tstate
m-I P l s
probability vector 8, we denote by pj the probability vector formed
vVe denote making
by Y s the
of vectors
obtainedtobystates
considering
aJl equal t
all totality
components
corresponding
not in A3
Ai, Aremaining
ending in A".
finite sequences
j , . . . . As,components
and the
proportional to those of p. We s
is plumped
restricted
to Aj.
(Ifarkov
,R haschain
all components
0 in Aj we
say that ,i3j'l'he
6.4.1 THEORE)!.
chain
is a 111
for the initial
not define pi.) Consider now the information given in (1). The f
vector 1T if and only if for every sand t the probability PrtJ{f1 E At] is
that f o E Ai may be interpreted as changing our initial vector to
the same for every f3 in Y s. This common valve is the transition
Learning then that f l E Ai may be interpreted as changing this vec
probability for moving from set As to set At in the lumped process.
to (nipif. We continue this process until we have taken into acco
PHOOF.
The
(1) given
can be
represented
in tothe
form assignm
We are led
a certain
all of probability
the information
in (1).
Pr~[fl E At] of
forprobabilities
a suitable f3 in
1',.theTo
do this
the first
n outcomes
for
states
in we
A,. useFrom
these
probabilities we
for the constmction (2). By hypothesis this probability depends only
7 C. K. Burke and M. Rosenbiatt, "A Markovian function of a Markov chain," An
on.s and t as
required. Slalistics,
Hence the
lumped process is a Markov chain.
of Mathematical
29: 1 1 12-1 122, 1958.
Conversely, assume that the lumped chain is a Markov chain for initial
vector 7T. Let j3 be any vector in Ys . Then j3 is obtained from a
possible sequence, say of length n, Ai, A j , . . . , As. Let these be the
given outcomes used to compute a probability of the form (1).
CHAINS
CHAP. VI must
probabilityFINITE
is PrB[fMARKOV
E At] and
by the Markov property
depend upon the outcomes before At. Hence it has the same v
given outcomes
used,Bto
a probability of the form (1). This
in compute
Y,.
for every
134
probability is Prp[f1 E At] and by the Markov property must not
EXAMPLE.
Consider
a Markov
chain
ma
depend upon 6.4.2
the outcomes
before
At. Hence
it has
the with
sametransition
value
for every fJ in Y s .
6.4.2
EXAMPLE.
Consider a :Markov chain with transition matrix
ftq
~
A2
:~:)
Let A = ({sl), (sZ,s3)). Consider any vector of the form (1 - 3
P
:: ( :;: i :;:
2a). Any such vector multiplied by P will again be of this f
Also any such vector restricted to A l or A2 will be such a ve
Let A ({S1}, {52, 53}). Consider any vector of the form (1- 3a, a,
Hence for any such starting vector the set U1 will contain the s
2a). Any element
such vector
multiplied by P will again be of this form.
(1, 0, 0) and Yz the single element (0, 113, 2 1 3 ) . Thus
Also any such
vectorof restricted
to Al or trivially
A2 will for
be any
such such
a vector.
condition
5 6.4.1 is satisfied
starting ve
Hence for any such starting vector the set Y 1 will contain the single
On the other hand assume that our starting vector is n = (0,
element (1,0,0) and Y2 the single element (0,1/ 3 ,2/ 3 ). Thus the
Let n l = ( n P ) 2 =(0, 1, 0) and 7 r 2 = ( n l P ) 2 = ( 0 , 'I6, 5/6). Then nl
condition of
§ 6.4.1 is satisfied trivially for any such starting vector. Hence
772 are in Yz and Pr,,[& E All = 0 while Pr,,[fl E .A1] = 35/48.
On the other hand assume that our starting vector is 7T = (0, 0, I).
choice of starting vector does not lead to a Markov chain.
Let 1Tl=(7TP)2=(O, 1,0) and "Z=(7T1P)2=(0, 1/ 6 , 51s). Then 7Tl and
We see that i t is possible for certain starting vectors to lea
7T2 are in Y 2 and Prrr.[f l EO A 1 ]=0 while Pr"2[f1 E A 1]=35j4S. Hence this
Markov chains while others do not. We shall now prove that if t
choice of starting vector does not lead to a Markov chain.
;s any starting vector which gives a Markov chain, then the
We see that
possible for certain starting vectors to lead to
vectorit aisdoes.
Markov chains while others do not.
We shall now prove that if there
6.4.3
THEOREM.
Assume
that a chain,
regular then
Narkov
chain is w
is any starting
vector
which gives
a }Larkov
the fixed
vector a does.lumpable with respect to A=(A1, A2, . . . , -4s: Then the sta
vector a will give a Markov chain for the lumped process. The
6.4.3 THEOREM.
Assume that
sition probabilities
willaberegular N arkov chain is weakly
lumpable with resped to A={Al' A 2 , . . • ,As}. Then the starling
vector a will give a Markov chain for the lumped process. The transition probabilities will be
A n y other starting vector ,which yields a Markov chai?~for the lw
same transition
probabilities.
process will give
pJj the
= Prai[fl
E Aj].
Sincewhich
the chain
weakly chain
lumpable
there
must be
Any other starting
yields is
a Markov
for the
lumped
PROOF.vector
which leads
to a Markov chain. Let its trans
starting
vector
process will
give the
same ntransition
probabilities.
matrix be ($ir).
For this vector n
Since the chain is weakly lumpable there must be some
starting vector 7T which leads to a :\larkov chain. Let its transition
matrix be {Pi;}.
for all For
setsthis
for vector
which 7Tthis probability is defined. But this ma
PROOF.
written as
P,p- a [ f ~ E &Ifl E Ai Afo E Ah].
for all sets for which this probability is defined. But this may be
written as
given outcomes used to compute a probability of the form (1).
135 must
probability isFURTHER
PrB[f E At]RESULTS
and by the Markov property
depend upon the outcomes before At. Hence it has the same v
Letting n tend to infinity we have
for every ,B in Y,.
Pra[fz E Allfl E All\fc E A k ] = pjj.
6.4.2 EXAMPLE. Consider a Markov chain with transition ma
We have proved that the probability of the forn (1), with a as
starting vector, does not depend upon the past beyond the last outcome for the case n = 1. The general case is similar. Therefore, for a
as a starting vector, the lumped process is a Markov chain. In the
course of the proof we showed that Pij for a starting vector '1T is the
same as for a, hence it will be the same for any starting vector which
yields a Markov chain.
Let A = ({sl), (sZ,s3)). Consider any vector of the form (1 - 3
such vector
multiplied
byweak
P will
again be we
of this f
2a). Any
By the previous
theorem,
if Vie are
testing for
lumpability
l or A2vector
will be
Also
vector
restricted
may assume
thatany
thesuch
process
is started
withto
theA initial
a. such
In a ve
U1 will contain the s
for any
suchPstarting
vector in
thethe
setform
this case theHence
transit,ion
matrix
can be written
element (1, 0, 0) and Yz the single element (0, 113, 2 1 3 ) . Thus
UPV trivially for any such starting ve
condition of 5 6.4.1Pis=satisfied
the other
where V is On
as before
but hand
U is aassume
matrix that
with our
i-th starting
row a i • vector
When is
we n = (0,
Let n l there
= ( n Pis
) 2a=(0,
1, 0)
andof 7freedom
r 2 = ( n l P )in
2 =the
( 0 , 'I6,
5/6).of Then
nl
have lumpability
great
deal
choice
U
arewe
in chose
Yz anda Pr,,[&
E All = 0 while
.A1]have
= 35/48.
and in that772
case
more convenient
U. Pr,,[fl
\Ve do Enot
this Hence
of starting vector does not lead to a Markov chain.
freedom forchoice
weak lumpability.
seeconditions
that i t isforpossible
for can
certain
starting
vectors
We considerWe
now
which we
expect
to have
weak to lea
others
do not.chain
Wewhen
shalllumped
now prove
lumpability.Markov
If thechains
chain while
is to be
a Markov
thenthat if t
;s any P2
starting
gives aitMarkov
we can compute
in twovector
ways. which
Computing
directlychain,
from then
the the
does.
underlying vector
chain awe
have P2= UpzV. By squaring P we have
UPVUPV. Hence it must be true that
6.4.3 THEOREM. Assume that a regular Narkov chain is w
lumpable with
respect = toUPPV.
A=(A1, A2, . . . , -4s: Then the sta
CPVUPV
vector a will give a Markov chain for the lumped process. The
One sufficient
condition
for this
sition
probabilities
wilisl be
SEC. 4
VUPV = PV.
(3)
This is the condition
for starting
lumpability
in terms
of ourchai?~
new for
U. the lw
A n y other
vectorexpressed
,which yields
a Markov
It is necessary
and
sufficient
for
lumpability,
and
hence
sufficient
for
process will give the same transition probabilities.
weak lumpability. A second condition which would be sufficient for
the above is PROOF. Since the chain is weakly lumpable there must be
which =leads
starting vector nUPVU
UP.to a Markov chain. Let(4)its trans
For this vector n
matrix be ($ir).
This condition states the rows of UP are fixed vectors for VU. The
matrix V U is now of the form
for all sets for which this probability is defined. But this ma
WI
0
0 \
written as
( ----~--- \
P,pa
[
f
~
E &Ifl E Ai Afo E Ah].
vu
Wz
i,
\ -----1-)
W3
.
0
\
0
0
Ii
o!
where W fis a transition matrix having all rows equal to a.'.
that the
i-th row
of U P CHAINS
is a fixed vector for
IrU VI
means
FINITE
MARKOV
CUAI.'.
vector, restricted to A,, is a fixed vector for W,. But this m
where Wj is a the
transition
matrixofhaving
all rows
to a f . To say
components
this vector
mustequal
be proportional
to a,. H
that the i-th TOW
of
UP
is
a
fixed
vector
for
Vl}
means
that this
have
vectoI', restrict,ed to Aj , is a fixed vector for (a"),
Wi. But
this means that
= a?.
136
the components of this vectot must be proportional to ctf. Hence we
This means that if we st,art with a , the set Yi, obtained by con
have
(i?), consists, for
each= i a, iof
. a single element, namely(5)a$. Co
(alp)j
if each such set has only a single element, then (5) is sat,i
This means thathence
if we also
start(4).
with To
a, the
fi'Yobtained
byone
construction
sayset
that
i has only
element for ea
(2), consists, for each i, of a single element, namely a i . Conversely,
say t h a t when the last outcome was Ai the knowledge of
if each such set
has onlydoes
a single
element, the
thenassignment
(5) is satisfied and
outcomes
not influence
of the probab
hence also (4). being
To say
that
Y
i has only one element for each i is to
in each of the sta.tes of Ai. Hence we have found t
say that when necessary
the last outcome
was Aiforthethe
knowledge
of previous
and sufficient
past beyond
the last ou
outcomes does provide
not influence
the
assignment
o£
the
probabilities
for lump
no new information, and is sufficient for weak
being in each of Example
the states6.4.2
of Ai.
Hence
we
have
found
that
(4)
is that
Recall
is a case where ( 4 ) is satisfied.
necessary and that
sufficient
forhad
theonly
past
the last outcome to
each Pi
onebeyond
element.
provide no new information,
and is sufficient
for weak
lumpability.
We can sumn~arize
our findings
as follows:
We stated in
I':xamvle 6.4.2 is a case where (4) is satisfied. RecaJ1 that we found
duction that there are two obvious ways t,o make the inf
that each Yi had
only oneinelement.
contained
the outcon~esbefore the last one useless. One
We can summarize our findings as follows: We stated in the introrequire that even if we know the exact state of the origina
duction that there are two obvious ways to make the information
our predictions would be unchanged. This is condition
contained in the outcomf>S before the last one useless. One way is to
other is to require t,hat we get no information a t all from
require that eyen
if wethe
know
exactThis
state
the original
is of
condition
(4). process
Each leads
except
lastthe
step.
our predictions would be unchanged.
This
is
condition
(3). The
lumpability. We have thus proved :
other is to require that we get no information at all from the past
6.4.4. This
THEOREM.
Either (4),
condition
or condition
except the last step.
is condition
Each (3)
leads
to weak ( 4 ) is
weak
lumpabilily.
lumpability. \Vefor
have
tlms
proved:
There
is ancondition
interesting
(3) and (4) in
6.4.4. THEOREM.
E£ther
(:3) connection
or cand'ilionbetween
(4) is s1Lfficient
the process and its associated reverse process (see 5.3).
Jor weak lumpability.
regular (3)
chain
(3) ifofa d o
6.4.5 THEOREM.
There is an interesting
connection Abetween
and satisjies
(4) in terms
chain reverse
satisjks process
(4).
the process and itsreverse
associated
(see § 5.3).
PROOF.
Assume
that satisfies
a process(3)satisfies
) . Then
THEOREM.
A regular
chain
iJ and( 3only
iJ the
reverse chain satis.fies (4).
V U P V = PV
6.4.5
PROm'.
Assume that a process sati;;fies (3).
Then
Let Po be the transition matrix for the reverse process,
DPToD-1. VCPV
Hence = PV.
TUDPToD-1V
= DPToD--'V
Let Po be the transition matrix for
the reverse process,
then P =
DPToD-l. Hence
VUDPTOD-IV
or, transposing,
SEC. 4
and
where W fis a transition matrix having all rows equal to a.'.
that the i-th
row of RESULTS
U P is a fixed vector for IrU137
means
FURTHER
- ' - vector
---- W,. But this m
for
vector, restricted to A,, is a-fixed
the components of this vector must be proportional to a,. H
VTD-1P oD[}TVTD-l = VTD-IP O•
have
(a"),
= a?.
\Ve observe that VT D-1 = [;--IU. Furthermore, VUD is a symmeans
that if we st,art
a , the set Yi,
obtained
metric matrix This
so that
VUD=DUTVT
orwith
DUTVTD-l=
vU.
Usingby con
for each
i , of a single element, namely a$. Co
these two facts,(i?),
ourconsists,
last equation
becomes
if each such set has only a single element, then (5) is sat,i
D-1UPOVU
= D-lUP
hence also
(4). To say
that Y Oi •has only one element for ea
say t h a t when the last outcome was Ai the knowledge of
Multiplying on the left by [; gives condition (4) for Po. The proof
outcomes does not influence the assignment of the probab
of the converse is similar.
being in each of the sta.tes of Ai. Hence we have found t
necessary
the past
beyond
last ou
If a and
givensufficient
pj·oce.ss isfor
weakly
lurnpable
with the
respect
6.4.6 THEOREM.
provide
no
new
information,
and
is
sufficient
for
weak
lump
to a partition A, then so is the reverse procu;s.
Example 6.4.2 is a case where ( 4 ) is satisfied. Recall that
PROOF.
We that
musteach
prove
probabilities
of the form
Pithat
had all
only
one element.
We can sumn~arizeour findings as follows: We stated in
Pra[fl
e Adf2
/\f3 are
E A,,!\
/\f" EAt]
duction
thate Aj
there
two...
obvious
ways t,o make the inf
the
the last one
useless.
depend only 011contained
A, and A jin
We outcon~es
can write before
this probability
in the
form One
.
require that even if we know the exact state of the origina
Pra[f l E Ai !\f2 Eour
Aj 1\£3
E An l\ ...
/\{n be
E Atl
predictions
would
unchanged. This is condition
other
is /\to ...
require
t,hat we get no information a t all from
Pr a [f2 E Aj /\f3
e A"
!\f" EAt]
except the last step. This is condition (4). Each leads
Pralfn E At
II ... !\f3 E A,,]f
/\£1 Eproved
Ad Prolfl
e Ai l\f2 E Aj]
2 E Ajthus
We have
:
lumpability.
Pl'a[fn eAt \ ... /f3 e A"ffz E Aj) Pr.[f2 E Aj]
6.4.4. THEOREM. Either condition (3) or condition ( 4 ) is
By hypothesis the
process is a :'.'larkov chain, .so that the
for forward
weak lumpabilily.
first term in the numerator does not clepend em At. Hence this whole
There is an interesting connection between (3) and (4) in
expression is simply
5.3).
the process and its associated reverse process (see
Prarfl E Ai /I,[z E Aj]
6.4.5 THEOREM. A regular chain satisjies (3) if a d o
--Pr.[f2 E Ai]
reverse chain satisjks ( 4 ) .
which depends only
on Ai Assume
and Aj . that a process satisfies ( 3 ) . Then
PROOF.
6.4.7 THEOREM.
when lumped.
A reversible Tegular ..M
chain
V a.rkov
UPV =
PV is reversible
Let Po be the transition matrix for the reverse process,
PROOF.
By reversibility,
DPToD-1. Hence
TUDPToD-1V = DPToD--'V
P=DPTD-l
and
p = l.lPV.
Hence
P = UDPTD-IV.
F I N I T E MARKOV CHAINS
138
We have seen that VTD-l= b - ' U .
Thus we CHAINS
have
D-1V =FINITE
U T ~ - 1 . MARKOV
138
We have seen that VTD-l=D-lU.
Thus we have
CH
Nence U D = b V T .
CHAP. VI
Hence UD=DVT.
Also
D-l V = UT jj-l.
f> = DVT pTUT D-l,
and hence
This means that the lumped process is reversible.
6.4.8 THEOREM. FOT a reversible regular Markov chain
lumpability
This means that the
lumped implies
process Eump,ability.
is reversible.
P be the transition
matrix for
a regular
PROOF.
6.4.8 THEOREM.
For Let
a reversible
regular Markov
chain,
weakreversible
Then, if11Imp,ability.
the chain is weakly lumpable,
lumpability implies
PROOF. Let P be the transition matrix
reversible
U Pfor
P Va regular
= UPV
U P V chain.
Then, if the chain is weakly lumpable,
or
UPPV = UPVUPV
Since U .- d VTD-1, we have
U P(I - VU)PV
B V T=Do.- ~ P ( IV-U ) P V = 0,
Since U = DVT D-l, we have
or, multiplying through by d-1 and using the fact that for a rev
chain jjVT
D I PD-IP(I
= PTD-1,
we have= 0,
- VU)PV
VTPTD-I(I
V U )for
P Va =reversible
0.
or, multiplying through by f)-I and using
the fact-that
chain D-lP= pT D-l, we have
Let W = D-1- D-1VU. Then W = D-1- U r b - l U . We sha
is semi-definite.
that W
VTPTD-l(I
- VU)PVThat
= O.is, for any vector ,B, BTWP
negative. It is sufficient to prove that
Let W=D-l_D-lVU. Then W=D-LUTD-IU. ""Ve shall show
nkb2k
> & f3, f3TWf3 is nonthat W is semi-definite. That is, for any
vector
I; in A,
kin Ai
negative. It is sufficient to prove that
1
where & is the
b2 i-th diagonal entry of b, or equivalently
L ak k ~ rl/ ( :2 a kbk)2
k in At'
where rl t
2 at&b2r > 2 (~kCZtbk)~.
k in Ai
iJ,
t i n Ai or equivalently
k In A i
is the i-th diagonal entry of
L
But since the coefficients
akdt are ) non-negative
and have sum
k 2.
a k d;b 2k ~
is a standard
inequality
of (a;;d;b
probability
theory. It can be pro
k [n Ai
kin Ai
considering a function f which takes on the value bk with pro
But since the coefficients
akd i areinequality
non-negative
and have
sum I, this
expresses
that.
a&. Then Ae
is a standard inequality of probability theory. It can be proved by
$ 1.8.5, this simply asserts that the variance of f is non-negati
considering a function
which
takes on the W
value
probability
Sincef Wt
is semi-definite,
= X Tbl<X with
for some
matrix X. Th
akd!. Then \.;,," inequality expresses that. M[f2] ~ (M[f])2; and, by
§ 1.8.5, this simply asserts that the variance of f is non-negative.
Since WI is semi-definite, W =XTX for some matrix X. Thus
.L
VT PTXTXPV = 0
or
(XPV)T(XPV) = o.
F I N I T E MARKOV CHAINS
138
SEC. 4
We have seen that VTD-l= b - ' U .
Thus we
have
D-1V = U T ~FURTHER
- 1 .
RESULTS
CH
Nence U D = b V T .
----- ----
139
This can be true only if
XPV = O.
Hence
or
= 0,process is reversible.
This means thatXTXPV
the lumped
THEOREM.
FOT a = reversible
regular Markov chain
D~l(I - VU)PT'
0,
lumpability implies Eump,ability.
6.4.8
or
Hence
(1 P
- be
VU)PV
= O.
Let
the transition
matrix for a regular reversible
PROOF.
Then, if the chain is weakly lumpable,
PV = VUPV.
U P P V and
= U sufficient
P V U P V conditions
Sote that while we have given necessary
for lumpability with respect to a partition A, we have not given
necessary and' sufficient conditions for weak lumpability. We have
given two different
conditions
Since U sufficient
.- d VTD-1,
we have (3) and (4). It might be
hoped that for weak lumpaoility one of the two conditions would have
B VtoT get
D - ~anPexample
( IV-U ) P Vwhere
= 0, neither
to be satisfied. It is, however, easy
is satisfied asor,
follows:
If we through
take a Markov
and the
findfact
a method
multiplying
by d-1 chain
and using
that for a rev
of combining chain
states DtoI give
a ),![arkov
P = PTD-1,
wechain,
have we can then ask whether
the new chain can be combined. If so, the result can be considered a
VTPTD-I(I
- Vour
U ) counterexample,
P V = 0.
combining of states in the original chain.
To get
We sha
we t?"ke a chain
can combine
states
condition
Then
W =by
D-1U r b - l (3)
U . and
LetforWwhich
= D-1-weD-1VU.
then combinethat
states
new chain That
by condition
(4); vector
the result
,B, BTWP
is, for any
W in
is the
semi-definite.
considered asnegative.
a lumpingItof
the original
chain that
will obviously be a
is sufficient
to prove
Markov chain, but it will satisfy neither (4) nor (3). Consider a
nkb2k > &
Markov chain with transition matrix
1
kin Ai
I; in A,
where &Al
is the
of b, or equivalently
( i-th
1/4 tdiagonal
1/16 3/16entry
l~)
2
2
o I 1/12 at&b2r
1/12 i 5/6
>
(~kCZtbk)~.
t i n Ai
o 1/12 1/12 I~ k In A i
P = A2
But since the coefficients akdt are non-negative and have sum
A3
7/ 8
1/32 of3/32
i 0
is a standard
inequality
probability
theory. It can be pro
considering
a
function
f
which
takes
on
the value (3)
bk with
For the partition A= ({Sl}, {S2' S3}, {S4}) the strong condition
is pro
inequality
that. matrix
a&. we Then
obtainAe
a lumped
chainexpresses
with transition
satisi'led. Hence
$ 1.8.5, this simply asserts that the variance of f is non-negati
Al Az W
A3= X T X for some matrix X. Th
Since Wt is semi-definite,
1( 4 1/ 2\
Al
1/ fi
P = Az
(~.
A3
7(8
5~6r
But this is Example 6.4.2, which satisfies (4), Hence we can lump it
by ({At), {A2' A3})' The result is a lumping of the original chain by
140
FINITE MARKOV CHAINS
CH
A=({sl), {sz, sg, sq)). It is easily checked that neither (3) nor
FINITE
MARKOV
CHAINS
satisfied in
the original
process
for this partition. CHAP. VI
We conclude with some remarks about a lumped process whe
A = ({81}, {S2, condition
83, 54}). It
easily
checked that
neither
(3) nor
is
forisweak
lurnpability
is not
satisfied.
We (4)
assume
tha
satisfied in the originalThen
process
for process
this partition.
if
the
is
started
in
equilibrium
regular.
We conclude with some remarks about a lumped process when the
condition for weak lumpability is not
satisfied.
assume
$tj =
Br,[f%+~EWe
&Ifn
E At] that Pis
regular. Then if the process is started in equilibrium
)
s
is the same for every n. Hence the matrix p ~ s = { & ~may
Also
interpreted
one-step
transition
Pii as
= aPra[fn+l
E Ajlfn
E Ad matrix.
140
is the same for every n. Hence the matrix P = {Pii} may still be
interpreted as a one-step transition matrix. Also
is the same for all n. The vector &=(a^i)will be the unique
vector for p. I t s components may be obtained from a by s
adding the ccmponents corresponding t o each set. Similarly w
for all two-step
n. The vector a = {at} will be the unique fixed
is the same define
transition probabilities by
vector for P. Its components may be obtained from IX by simply
E At].we may
$(2)ir to
= Pr,[fn+z
adding the components corresponding
each set.E Similarly
define two-step transition probabilities by
The two-step transition matrix will then be p(2)=($(Z)tj). I t w
p2=Ep(2).
longer be
f.Pltrue
jj = that
Pra[fn+z
Ajlfn EAt].
We can also define the mean first passage matrix J?l for the l
The two-step transition matrix will then be F(2) = {p(2){j}. It will no
process. It cannot be obtained by our Markov chain formula
longer be true that P2 = P(2).
obtain & i t is necessary first to find m t , ~ ,the
, mean time to g
We can also define the mean first passage matrix M for the lumped
state i to set A$ in the original process. We can do this by m
process. It cannot be obtained by our Markov chain formulas. To
allnecessary
of the elements
Aj mi,A
absorbing
and find the mean time to a
obtain M it is
first to of
find
j , the mean time to go from
tion. (A slight modification is necessary if i is in Aj.) From
state i to set Aj in the original process. We can do this by making
we obtain the mean time to go from Ai to A?, by
all of the elements of Aj absorbing and find the mean time to absorpmil = if i is ina*kmk,A,
tion. (A slight modification is necessary
Ad From these
we obtain the mean time to go from At to A j k, InbyA ,
where a*k is the k-th component of at.
mij =
a*kmk,A,.
5 6.5 Expandingkin aAj arkov chain. I n the last two sectio
showed
that under ofcertain
conditions a Markov chain wou
where a*k is the
k-th component
ai .
lumping states together, be reduced to a smaller chain whic
§ 6.5 Expanding
a f¥larkov
chain.about
In the
the original
last twochain.
sections
interesting
information
By we
this proc
showed that under certain conditions a Markov chain would, by
obtained a more manageable chain a t the sacrifice of obtaini
lumping states together, be reduced to a snmller chain which gave
precise information. In this section we shall show that i t is p
interesting information
about
original chain.
process
to go in the
otherthe
direction.
That is,By
to this
obtain
from we
a Markov
obtained a more manageable chain at the sacrifice of obtaining less
a larger chain which gives more detailed information about the p
precise information.
In this section
we shall
that it is possible
being considered.
We shall
baseshow
the presentation
on results ob
to go in the other direction. That is, to obtain from a Markov chain
by S. Hudson in his senior thesis a t Dartmouth College.
a larger chain which
givesnow
morea detailed
about sl,
thesz,
process
. . . , s,. W
Consider
Markov information
chain with states
being considered.
We
shall
base
the
presentation
on
results
obtained
a new Markov chain, called the expanded process, as follows. A
by S. Hudson in his senior thesis at Dartmouth College.
Consider now a Markov chain with states 81, S2, . . . , Sr- \Ve form
a new Markov chain, called the expanded process, as follows. A state
2
2:
FINITE MARKOV CHAINS
140
CH
A=({sl), {sz, sg, sq)). It is easily checked that neither (3) nor
satisfied in the original process for this partition.
SEC. 5
FURTHER
RESULTS
141
We conclude
with some
remarks about a lumped process
whe
condition for weak lurnpability is not satisfied. We assume tha
is a pair of states (Si' Sj) in the original chain, for which Pij>O. We
Then if the process is started in equilibrium
denote these regular.
states by S(!j). Assume now that in the original chain the
transition from SI to Sf and from Sf to
on Etwo
successive
steps.
$tj Sk=occurs
Br,[f%+~
&Ifn
E At]
We shall interpret this as a single step in the expanded process from
the matrix transition
p ~ s = { & ~may
)
s
thethe
same
for S(j}().
every -With
n. Hence
the state S (tj)is to
state
this convention,
Also
interpreted
as ainone-step
transition
matrix.
from state 8(/1)
to state S(kl)
the expanded
process
is possible only if
j = k. Transition probabilities are given by
Or
P(ii)(}l)
n. The vector &=(a^i)will be the unique
is the same
for =allPjl,
for j
k.may be obtained from a by s
vector forP(ij)(kl)
p. I=t s 0components
adding the ccmponents corresponding t o each set. Similarly w
define two-step transition probabilities by
'*
6.5.1 EXAMPLE. Consider the$(2)ir
Land= of
Oz example.
E The
At]. states
Pr,[fn+z
E
for the expanded process are RR, R!". RS, NR, NS, SR, SN, SS. Note
It w
Thea two-step
transition
then be
p(2)=($(Z)tj).
that NN is not
state, since
PUN = matrix
0 in thewill
original
process.
The
p2=
p(2).
longer
be
true
that
transition matrix for the expanded process is
We can also define the mean first passage matrix J?l for the l
RR It
RNcannot
RS NIt
NS SR by
SN ourSSMarkov chain formula
be obtained
process.
obtain1/&
i t is necessary
first
to
, mean time to g
.I ~
RR
0
0
0 find
0 m t , ~ ,the
2 1/ 4 1 J"
i
to
set
A$
in
the
original
process.
We
can do this by m
state
1/2
RN
0
0
0
1/2 0 0
all of the elements of Aj absorbing and find the mean time to a
RS
0
0 1/4
0
0modification
1/4
is necessary
if i is in Aj.) From
tion. 0(A slight
NR
0
0
1/4
1/4
0
0
1/2
we obtain the mean time to go from Ai to A?, by
NS
0
SR
1/2 l' 4 1/4
0
J
0
0
0
)'\r
1~2
0 =1/4 1/4 a*kmk,A,
mil
0k In A0,
0
2
1/2 0 of0at.
SN
where 0a*k is0 the 0k-thlizcomponent
1/
l~J
0
4 I n the last two sectio
i 4
arkov
chain.
showed
that
under
certain
conditions
Markov
chain wou
Let us first see how the classification of states for a the
expanded
lumping
states
together,
be
reduced
to
a
smaller
chain
whic
process compares with the original chain. We note that p(ll) (ij)(kl) =
interesting information about
the if
original
chain.chain
By is
this proc
p(n-l)jkPkl > 0 if and only if p(n-l)jk > O.
Henee
the original
obtained
a moreprocess
manageable
t the sacrifice
ergodic, so will
the expanded
be, andchain
if the aoriginal
chain isofofobtaini
precise
information.
In
this
section
we
shall
that i t is p
period d, then the expanded chain will also be of period d. show
A state
to
go
in
the
other
direction.
That
is,
to
obtain
from
a Markov
S(ij) in the expanded process is absorbing only if i =j and only if state
a larger chain which gives more detailed information about the p
Sj is absorbing in the original chain.
being
shall
base thechain.
presentation
on be
results ob
Assume that
the considered.
original chainWe
is an
absorbing
Let S(ij)
Hudson
in his
senior thesis
a t Dartmouth
by S.
a non-absorbing
state
in the
expanded
process.
Since the College.
original
Consider
now
a Markov
chain with
states
sl, sz, . . . , s,. W
chain was absorbing,
there
must
be an absorbing
state
Sk ,mch that it
A
a
new
Markov
chain,
called
the
expanded
process,
as follows.
is possible to go from Sf to Sk. Thus it is possible to go from 8(lj)
to
SS
0
0
0
5 6.50 Expanding
a
l'
in the expanded process. Thus the expanded process is also
absorbing.
It is interesting to observe that from the expanded process we can
S(kk)
FINITE MARKOV CHAINS
142
CK
go back to the original process by lumping states. For this w
A = { AMARKOV
1 , A t , . . .CHAINS
, A,) of the states inCHAP,
the extended
the partition
FINITE
VI
with Ai the set of all sta,tes of the form s ( k ~ . Then the condit
go back to the
original
process
states,
For
p ( by
k ~lumping
should
~,
not
depend
onthis
k. we
Butform
this is tr
lumping
is that
the partitionthe
A =Markov
{Al, Az,property
. , . , Ar} for
of the
in the
extended
chain, pro
thestates
original
chain.
The lumped
with Ai the then
set ofthe
all same
statesasofthe
theoriginal
form 8(.1:1).
for parti
chain.Then
I n the
our condition
example, the
lumping is that P(kl)A, should not depend on k. But this is true by
= original
{(BE,SIR,chain.
NR), (SN,
(RS,process
NS, SS)}.is
the Markov property for Athe
TheRN),
lumped
then the same We
as the
original
chain,
our quantities
example, the
is
next
compare
the In
basic
for partition
our expanded
p
W
with
the
corresponding
quantities
for
the
original
chain.
A = {(RR, SR, NR), (SN, RN), (RS, NS, SS)}.
treat only the regular case. The okher cases may be treated sim
\Ve next compare the basic quantities for our expanded process
{ac}
be the fixed
a regular
6.5.2 THEOREM.
with the corresponding
quantitiesLetfora =
the
original
chain.vector
"Veforshall
Letmay
8 = {be
a (treated
i j ) )be the
fixed vector f
with transition
matrix
treat only the regular
case. The
otherP.cases
similarly.
expanded chain. Then
6.5.2 THEOREM. Let a = {at} be the fixed vector for a regular chain
acrr) = aipu.
with transition matrix P. Let a={a(/j)} be the fixed vector for the
Then
expanded chain.
PROOF.
It is obvious that arpgj is positive. Also,
142
at/f)
PROOF.
=
alplj.
It is obvious that a/Plj is positi\·e. Also,
Hence we need only prove that 8 = { ~ , ~ p iisj } a fixed vector f
= for a/pi}
=
aj =process.
1.
That is,
transitiona(jj}
matrix
the expanded
2:
2:
2:
i,j
(iiI
j
Hence we need only prove that a= {a'IPif} is a fixed vector for the
transition matrix for the expanded process. That is,
2:
a(lf)p(ij) (kl)
=
a(kl).
(ij)
But
2:
(ij)
a(lj)p(/f) (kl)
2: a,p1lPlld
== 2:
=
jk
i, j
ajpj1djk
)
In our example, this
= gives
akPlc1 for the fixed vector
=
a(kl)'
In ou'r example,
this gives
for the
Kote that
the result
wefixed
havevector
proved is intuitively obvious, sin
represents
that
a large number of steps the p
a =the(.2,probability
.1, .1, .1, .1,
.1, after
.1, .2).
will be in state si and then move to state s,. The probability th
Note that the
willresult
occurwe
is have
clearlyproved
arpij. is intuitively obviou~, since a(jj)
represents the probability that after a large number of steps the process
T H Emove
OREM
.state
funda.menta1
matrix for that
the expanded
c
will be in state 6.5.3
s, and then
toThe
Sj.
The probability
this
will occur is clearly ajp!!.
6.5.3
THEORE:'L
The fundamental matrix for the expanded chain is
2 = {Z(tj)(kl)} = {d(ij)(kl) + (Zjk-ak)Pkl}.
FINITE MARKOV CHAINS
142
SEC. 5
PROOF.
CK
go back to the original process by lumping states. For this w
the partition A = { A 1 , A t , . . . , A,) of the states in the extended
set of all sta,tes
of the form s ( k ~ . Then 143
the condit
with Ai theFURTHER
RESULTS
~,
not depend on k. But this is tr
lumping is that p ( k ~ should
the Markov property for the original chain. The lumped pro
'"
then the same as the original
chain. I n our example, the parti
(p(n) (4j)(k1) - a(kl»)
Z(/i) (kl)
= dO}) (.1:1) + L
A = {(BE,SIR, NR), (SN, RN), (RS, NS, SS)}.
n~l
00
2:
- akPkt) for our expanded p
== dlljl(k!)
+ the(p(n-l)jkPkl
We next
compare
basic quantities
,,-1 quantities for the original chain. W
with the corresponding
The okher cases may be treated sim
treat only the regular case. (p(n)jk-ak)
00
= d(jj)(kl) + Pk1 2:
n-O
a = {ac}be the fixed vector for a regular
THEOREM.Let
Let 8 = { a ( i j ) )be the fixed vector f
= dW)(kl) matrix
+ PkZ(ZjkP.
- Ok)'
with transition
expanded chain. Then
For our example
= aipu.
RR
RN
RS
NR
NSacrr) sa
SN
S8
6.5.2
RR
RN
RS
NR
Z=
NS
SR
SN
S8
PROOF.
obvious
arpgj
is positive.
Also,
.187
- .OSO that
-.080
-.147
- .1<11 -.293
.181It is
/1.373
-.160
.920 -.080
.320
.320 -.080 -.080 -.160
-.293 - .147
.853 -.080 -.080
.187
.187
.373
Hence
prove
that- .147
8 = { -.147
~ , ~ p iis-.293
j } a fixed vector f
-.080
.373 we
.187 need
.187only.920
That
is,
transition matrix for the expanded
process.
.920
.373
.187
.187
- .293 -.147 -.147 -.080
.373
.187
.187 - .080 -.080
- .160 -.080 -.080
.320
.853 -.147 -.293
.320 - .080
- .293 - .147 -.147 -.080 - .080
.920 -.160
.187
.187
1.373
We next consider the mean first passage times for the expanded
process.
6.5.4
THEOREM.
m(lj)(.~I)
1
(Zjk-Zl/;')
OkP};!
ak
= - - - "-'----'
example,
this gives
fixed vector
From In
theour
matrix
expression
for for
J1 the
in terms
of the fundamental matrix we have
PROOF.
.
1
1n(/j) (kl)
= (d(ll)«!)
- Z{ii)(kl)
+ Z(k!)
(kl») -is
- intuitively obvious, sin
Kote
that
the result
we have
proved
a(};I)
represents the probability that after a large number of steps the p
[d(j1J (k1) - Pkl(Zjk - ak)
si and then move to state s,. The probability th
will be in= state
will occur is clearly arpij.
6.5.3
T H E O R E MThe
. funda.menta1 matrix for the expanded c
= [1-Pk/(2Jk- ZIli:)]_1O"Pkl
= _1__ (Zjk-Zlk).
OkPki
ak
Again, as was to be expected, m(I!)(k/) does not depend on i.
For our example, we obtain
FINITE MARKOV CHAINS
144
C!HAP
F I N I T E MARKOV CHAINS
144
CHAP. VI
For our example, we obtain
RR
RN
RS
NR NS
SR
SN
SS
RR
5
10 2 /3 8 1 / 3\
7 1 /3
6 2 / 3 10 10 10
-RN
10
6
6
[)l; 3
J ,,3
7 2/3 10
9 1 /3
RS
10 10
8 1! 3 10 2/3 10
G2/s
7 1/3 5
'
NR
5
10 2 / 3 8 1 .3
7 1 /3
6 2/3 10 10 10
NS
10 10
5
71/3
G2/3
8 1/3 10 2 / 3 10
SR
5
71/3
10 2,3 8 1 /3
6 2 / 3 10 10 10
SN
72/a
10
6
6
7 2/3
9 1/ 3
9 1 /3 10
I n comparison, the mean first passage matrix for the original cha
2
1
1
2
SS
6 /3
10 IO
7 /3 5
8 / 3 10 /3 10
~2!
if =
R
N
S
In comparison, the mean first passage R
matrix
original chain is
/ 2 . 5 for4 the3.3\
R
R
N
S
4
~~)
C'
2.7 5
M = N
Consider next the reverse transition
;.~ are
. matrix for the expanded
3.3 4
S probabilities
cess. The transition
Consider next the reverse transition matrix for the expanded pro_ cess. The transition probabilities are
Hence
Hence
(1 (kl)P(kl) (ii)
alii)
and
P(ij)(kl)
= 0
if i =I
and
a(ki)p (tt) (i})
P(ii) (ki\
alii)
(liPi}
=
(1kPki
(Ii
Hence the reverse process
for the expanded process is simpl
=
Pik.
reverse process for the original chain expanded.
One application
theexpanded
expanded process
process isis the
following:
It
Hence the reverse
process forofthe
simply
the
P
for
a
chain
is
not
known
happens
that
the
transitiorl
matrix
reverse process for the original chain expanded.
must be estimated from data. If a large number n of outcom
One application of the expanded process is the following: It often
happens that the transition matri x P for a chain is not known and
must be estimated from data. If a large number n of outcomes for
I
If
C!HAP
F I N I T E MARKOV CHAINS
144
For our example, we obtain
FURTHEg RESULTS
SEC. 5
145
the process are known, then an obvious estimate is
yin),}
Pii = yin),
where y(n)j} is the number of transitions from state Si to state Sf, and
is the number of times the process is in state St. To study the
properties of this estimate it is necessary to study the properties of
y(n)i} and yin)! for a Markov chain.
In particular, the limiting variances and eovariances for these quantities are important. 'Ve know
that we can obtain the limiting eovariances for y(n)j, y(n)j from the
fundamental matrix for the basic chain. But how do we obtain the
limiting covariances for y(n)(!j) , y(n)(k/) 1 'Ve simply observe that these
comparison,for
thethe
mean
first of
passage
forofthe
original cha
are the limitingI neovada.nees
number
times matrix
in a pair
states
for the expanded chain. Hence we can express
these
covarianees
in
R N S
terms of Z and Q. for the expanded process. We ca·n then use Theorems
R / 2 . 5 4 3.3\
6.5.2, 6.5.3 to express the limiting covariances in terms 'of quantities
relating to the original chain. Carrying out. this computation gives:
yin>;
6.5.5 THEORE;\1.
ai'e given by
C(lt) (kl)
The limiting Cl)varianc(;s for the expanded chains
Consider
next the reverse transition matrix for the expanded
I
= aipijplcZZjk
-;- a.'cpkZpijZZi
+ akPkZd(ij)(kl)
cess.
The transition
probabilities
are - 3atPijaIcPk/.
For our example these covariances are
.300
.001 - .012
.001
.074 - .033 - .041
Hence
- .012and
- .033
c=
.001
.041
.001 - .065 - .01:2 - .065 - .157
.007
.001 - .026 - .065
.061
.001 - .033
.027
.001 - .012
.001
.074 - .026 - .033
.007 - .065
- .065
.007 - .033 -.OZ(i
.074
.001
-.012
.001
. O·)~
_I
- .033
.001
.061 - .033 - .012
-.065
.026
.001
.007
.041 - .033
.074
.001
- .157
.06:3 - .012 - .06:')
.001 - .01:2
.OOJ
.309
.041
.001
Hence the reverse process for the expanded process is simpl
Exercises
for original
Chapter VI
reverse process
for the
chain expanded.
One application of the expanded process is the following: I t
For S5.1
happens that the transitiorl matrix P for a chain is not known
1. Consider must
Example
2 with p = 1/2.
process
is observed
If athe
large
number
n of outcom
be estimated
fromAssume
data. that
only when it is in the set {52. 53, 54}. Find t.he resulting transition matrix.
Find }v! for the new process. What do the entries of }r[ mean in t.erms of the
original chain~
F I N I T E MARMBV CHAINS
146
,
CH
I
2. The following table gives the probability of a team ending u
certain pFINITE
~ s i t ~ i onext
n
year, given
what its posttionCHAP.
is thisVIyear. Fo
MARKOV
CHAE\"S
position in title second division, calculate the mean number of years t
the first
2. The following
tabledivision.
gives the probability of a team ending up in a
1 1 is this
2 year.
3 4 For
5 each
6 7 8 1
certain pOflition next year, given what its position
position in tile second division, calculate the mean number of years to reach
the first division.
;
2
1
3 4 5 6 7 81
146
(
!
PST 1.3
FIRST DIVISlO N
~ 2ND 1.1
.3
.3
.2
.2
.2
.1
.1
.1
.05
.1
',4TH. 0
.1
.2
0
.2
.1
0
.1
.1
.2
.2
0
0
0
0
.1
.1
.1
I
I
.2
.1
.2 .2
.1
f
.1
.05 .2 .3
f5TH
0
.1 .1 .2 .3 .2 .1
6TH 0
0
SECOND DIVISION
3. Consider
a chain
with
a0 single
set, .3
which has transient
17TH.
0
.3 .3
0
0 .1ergodic
Let fj be the time r
transient
Let si be a LSTH
.1 .2 state.
.4 .3
I 0 state0 and0sf a0n ergodic
to reach sf, g the first ergodic state reached and t the time required t
ergodic
Then
3. Consider a the
chain
with set.
a single
ergodic set, which has transient states.
Let 81 be a transient state and an ergodic state. Let fj be the time required
Mdfjlreached
=
to reach s1, g the first ergodic state
and t the time required to reach
si ergodic
the ergodic set. Then
Use this result t o find, for Example 4 with ~ = 2 / ~the
, mean time
Mj[fj] =
') thePr,[g
= sk][MAf, ]+M/[t]g = Sk]J.
first time starting in state ss.
state sl811: for
e;gOdlc
4. It is raining in the Land of Oz. Find the mean number of da
Use this result to find, for Example 4 with p = 2/ 3 , the mean time to reach
each kind of weather has occurred a t least once.
state 81 for the first time starting in state 53.
5. Prove t h a t when a process is observed only when in a subs
4. It is rainingergodic
in the chain,
Land of
Find the
mean
of days chain.
until
theOz.
resulting
process
is number
also a n ergodic
each kind of weather has occurred at least once.
$ 6.2
5. Prove that when a process is observed onlyFor
when
in a subset of an
ergodic chain, the resulting
process
is
also
an
ergodic
chain.
6. Find the mean number of rainy days between nice days in the
3RD
I
s,
02.
For § 6.2
7. For the Markov chain in Exercise 2 of Chapter I I , assume that w
Find
fixed
for of
the resultin
6. Find the mean
of rainy
between
nicethe
days
in vector
the Land
duelnumber
ends a new
duel days
is started.
Oz.
Use this t o determine the absorption probabilities and the mean nu
timeschain
in each
state for2 the
originalII,chain.
7. For the Markov
in Exercise
of Chapter
assume that when the
duel ends a new duel8.is Consider
started. the
Find
the with
fixed transition
vector for the
resulting chain.
chain
matrix
Use this to determine the absorption probabilities and the mean number of
s1
sz
s3
times in each state for the original chain.
8. Consider the chain with transition matrix
51
52
S3
=:: (1~4 ~ 3~4)'
Find theP fundamental matrix Z by making state sl into an absorbi
calculating N , and using t h e result of Theorem 6.2.5. Find mlz
83
0 1 0
using Corollary
6.2.7.
Find the fundamental
matrix6.2.5
Z bydeduce
making
51 into an absorbing state,
9. From
thestate
identity
calculating N, and using the result of Theorem
by
I - A = (6.2.5.
I - P ) N *Find
( I - Am12+m;n
).
using Corollary 6.2.7.
9. From 6.2.5 deduce the identity
I -A = (I - P)N*(1 -A).
F I N I T E MARMBV CHAINS
146
CH
2. The following table gives the probability of a team ending u
RESULTS
certain p ~ s i FURTHER
t ~ i onext
n
year,
given what its posttion is this147
year. Fo
position in title second division, calculate the mean number of years t
the first division. For § 6.3
11
2
3 4 5 6 7 8 1
10. For Example 2 with p = 1/2, which of the following partitions produces
SEC. 5
a Markov chain when lumped!
(a)
A = ({SI' 53, S5},
{S2,54})'
(b) B = ({51, S5}, {Sz, S4}, {S3})'
Which produce Markov chains if p4=-lj2?
11. Show that Example 3 (a) is lump able with respect to the partition
A = ({51, 55}, {S2' S4}, {53}). Find the fundamental matrix for the lumped
chain from the fundamental matrix for the original chain (see § 4.7).
12. Find a three-cell partition which makes Example 61umpable.
13. Let P be the
of an aindependent
trials
a"d
a be
3. transition
Consider amatrix
chain with
single ergodic
set,chain
which
has
transient
Let fj be
the time r
si be a transient
sf a n ergodic state.
Let O<a<l.
any number with
Show state
that and
P'=aP+(l-a)]
islumpable
with
respect to any to
partition.
reach sf, g the first ergodic state reached and t the time required t
the ergodic
14. Show that
Exampleset.12 Then
is lumpabJe with respect to the partition
A = ({S1SI, S2SI}, {SlS2, szsz}). How is the resulting transition matrix related
Mdfjl = the four-state chain 2
to the two-state chain which determined
si ergodic
15. Prove that for a lump able ergodic chain, a=a V.
Use this result t o find, for Example 4 with ~ = 2 / ~the
, mean time
16. Give an example of a Markov chain which is not itself an independent
sl
for
the
first
time
starting
in
state
ss.
state
trials chain, but which can be lumped to an independent trials chain. Check
4. It is raining
the Land of Oz. Find the mean number of da
your answer by computing
Z andinUZV.
each kind of weather has occurred a t least once.
5. Prove t h a t when
process is observed only when in a subs
For §a6.4
ergodic chain, the resulting process is also a n ergodic chain.
17. Show that the Markov chain with transition matrix
For $ 6.2
Sl (1/2 1/4
1/4)
6. Find the mean number of rainy days between nice days in the
S2 1/ 2 1/2 0
02.
7. For the Markov
Exercise 2 of Chapter I I , assume that w
s3 \ 0 chain
1/4 in3/4
duel ends a new duel is started. Find the fixed vector for the resultin
is weakly lumpable,
butt o not
lumpable,
respect
to A = ({51},and
{S2,the
53})'mean nu
Use this
determine
the with
absorption
probabilities
Find the transition
the for
lumped
process. chain.
Show that the reverse
timesm«trix
in eachforstate
the original
process is lumpable.
8. Consider the chain with transition matrix
18. Show that the .Markov chain with transition matrix
s1 sz s3
:: (
~0/4 o
l/S 5/ s
P = 83
Find the84fundamental
matrix Z by making state sl into an absorbi
1/16
calculating N , and using t h e result of Theorem 6.2.5. Find mlz
85
lIs
using Corollary
6.2.7. o 1/ 3
9.
From
6.2.5
deduce
the
is weakly lumpable with respect to
([51, 82,
S3},identity
{S4, S5}).
I-A=(I-P)N*(I-A).
19. For the Land of Oz example, let A1={R, N} and A2={S}, Compute
Pr.[f2 E A11fl E A l ] and Pr a [f2 E AII!'l E Al Afo E AI]. Use the result to show
that the chain is not weakly lumpable wit,h respect to A = (AI, A 2).
148
FINITE MARKOV CHAINS
20. Prove that if P is lumpable with respect to a given partitio
has column
sumsMARKOV
1, then PTCHAINS
is weakly lumpable CHAP.
with respect
to
FINITE
VI
partition.
20. Prove that if P is lurnpable with respect to a given partition, and if P
148
has column sums 1, then pT is weakly lumpable with respeet to t.he same
partition.
Represent this as a
21. A coin is tossed a sequence of times.
Form
the expanded process. Find M and H
Markov chain. For
§ 6.5
expanded chain. Interpret the diagonal entries in terms of t
21. A eoin is tossed
chain. a sequence of times. Represent this as a two-state
.Markov chain. Form
the Pexpanded
proeess. matrix
Find},f
)1 2 for the
22. Let
be the transition
of anand
independent
trials ch
expanded chain. formulas
Interpret
diagonal
in termschain.
of the Check
original
the lat
forthe
Z and
M ofentries
the expanded
chain.
Exercise 21.
tramition
matrix
of an independent
ehain. Find
22. Let P be the23.
Let P be
the transition
matrix of trials
an independent
trials ch
formulas for Z and
Jv[ of the
chain.
Check the
latterexpanded
against chai
a formula
for expanded
the limiting
covariances
of the
Exercise 21.
Write the formula in the form akaL(. . . ).I Use this formula t
the transition
limiting covariances
in Exercise
21. trials chain. Find
23. Let P be the
matrix of an
independent
a formula for the limiting eovarianees of the expanded chain. [HINT:
Write the formula in the form akal( ... ).] Use this formula to compute
the limiting covariances in Exercise 21.
,
,
FINITE MARKOV CHAINS
148
20. Prove that if P is lumpable with respect to a given partitio
has column sums 1, then PT is weakly lumpable with respect to
partition.
21. A coin is tossed a sequence of times. Represent this as a
Form the expanded
process. Find M and H
CHAPTER
VII
Markov chain.
expanded chain. Interpret the diagonal entries in terms of t
chain.
22. Let P be the transition matrix of an independent trials ch
formulas for Z andOF
M ofMARKOV
the expandedCHAINS
chain. Check the lat
APPLICATIONS
Exercise 21.
23. Let P be the transition matrix of an independent trials ch
a formula for the limiting covariances of the expanded chai
Write the formula in the form akaL(. . . ).I Use this formula t
the limiting covariances in Exercise 21.
§ 7.1 Random walks. vVe will consider four simple, related random
walks. The first three are v;alks on a line, with states 0, I, ... , n:
o
2
i
i-I
i+l
I
n-l
n
FIGURE 7·1
In each of the first three types of random walks we have probability p
of moving to the right (from i toi + 1) and probability q of moving to
the left (from i to i-I), for states i = 1, 2, ... , n - 1. The three types
differ in their behavior at the "boundaries," 0 and n.
AAR'V: A random walk having both 0 and n as absorbing states.
APRW: A random w2.lk having 0 as an
absorLing state, while n is "partially reflecting." That is, at n it moves back to
n-1 with probability q and stays at n
3
with probability p.
PPRW: A random walk partially reflecting at both boundaries. That is it is
like APlnV at n, and at 0 it moves to 1
with probability p and stays at 0 with
probability q.
i+1
p
The fourth random walk will move
on a circle, with states numbered 1, 2,
FIGURE 7·2
... , n, as in Figure 7·2.
CR\V: The process moves on the circle, taking one step clockwise
with probability p, and counterclockwise with probability q.
14\l
FINITE MARMOV CHAINS
150
GI
We wid illustrate each of these random walks for n= 5. Si
FINITE
CHAINS
VII
behavior
is oftenMARKOV
quite different
forp = 112 thanCHAP.
for any
other p-va
will carry out the illustration for both p = l / z and p = 2 1 3 . Then
We will illustrate each of these random walks for n = 5. Since the
solve each random walk. It will be convenient to let T stand fo
behavior is often quite different for p = 1/2 than for any other p-value, we
us first for
consider
absorbing
will carry out the Let
illustration
both pAARW.
= 1/ 2and pThis
= 2/ is
we will chain w
3. anThen
n.
absorbing
solve each random
walk. states
It will0 and
be convenient
to let r stand for p/q.
AARW for n = 5 , p="z
Let us first consider AARW. This is an absorbing chain with two
absorbing states 0 and n.
AARW for n=5, p= 1/2
150
p=
o
5
1
234
o
1
0
0
0
0
0
5
o
1
0
0
0
0
1
o 1/2 0
0
2
1/2
0
3
o 1/2 0 1/2
o 0 1/2 0
4
0
1/2
0
B=
o
5
.6
.4
1(.8 .2)
2
B =
AARW for n =35, p.4= 213.6
4
.2
.8
2
AARW for n=5, p=2/ a
p=
o
5
1
o
1
o
5
o
1
o o
o o
1
l/a
2
3
4
o
o o
o o
o 2/3
3
4
o o
o o
o 0
2/ 3 0
o 2/ 3
1/3
0
5
FINITE MARMOV CHAINS
150
GI
We wid illustrate each of these random walks for n= 5. Si
APPLICATIONS
MARKOV
151 p-va
behavior
is often quiteOF
different
forp CHAINS
= 112 than for any other
will carry out the illustration for both p = l / z and p = 2 1 3 . Then
42 36
solve each random walk. It will be convenient to let T stand fo
21 63 54 36
Let usNfirst
AARW. This is an absorbing chain w
31
= 1/consider
:
absorbing states 0 and n.27 63 42
AARW for n = 5 , p="z 9 21 45
SEC. 1
24)
C
5
0
T
= 1/31
C) n
174
16)
1
24
2
28
3
30
4
B = 1/31
141
78
The matrix 1- Q has the form
I-Q =
( -qI
-p
0
1
-p
~ -q
0
1
-q
-D
when n= 5. In general it has entries 811 (i,j = 1,2, ... , n -1) which
are 0 except that 8jj= 1, 8j-1.1= -p, 81+1,j= -q. We note from the
numerical examples that the entries of N decrease
0 on
5 both sides of the
diagonal. N will have the form .
n1
i
1
= (p-q)(rn-l)
. {(ri-l)(rn-l-l)
B =
(r'-I)(rn-i-ri-i)
ifj ~ i
ifj ~ i
(1)
except where p = 1/2, here the solution simplifies to
AARW for n 2= {j(n-i)
5, p = 213 ifj ~ i
nil = ii' i(n-j) ifj ~ i.
Let us verify the solution for P'j61/ z, by computing N(I -Q).
i,j-th entry is
1
(p-q)(rn-l)
[±(rk-l)(rn-i-I)8kJ+
k=l
(2)
Its
nil (ri-l)(rn-i-rk-i)8kJ]'
k=i+1
We recall that 8wl: 0 only for k = j - 1, j, j + 1. Suppose that j < i.
Then all terms in the second sum are O. The first sum simplifies to
rn-i-l
(p _ q)(rn _ I) [(rf-l-l)( - p) + (r l - 1)( I) + (r1+ L 1)( - q)]
=
(p~nq~j(::~ 1) [rl( -~ + l-rq) + (p-l +q)] = O.
The answer for j > i is also 0. Hence all off-diagonal ent
CHAP.
VIIentry
CHAINS terms for the
Let us FINITE
computeMARKOV
the three non-zero
i,i-th
152
The answer for j >i is also O. Hence all off-diagonal entries are O.
Let us compute the three non-zero terms for the i,i-th entry.
1
[(rl-I- l)(rn-t-l)(-p)+(ri-l)(rn-i-l)(l)
(p - q)(rn - 1)
+ (rL l)(r n - i - r)( - q)]
1
lrn(-Z':+l-q)+rn-t(p-l+q)
(p - q) (rn - 1) L
r
rn(p-q)+(q-p)
&)-I
as required.
Thus N = (I-1.
The solution for p=l12 may be verified similarly.
(p-q)(rn-l)
We
obtain i t by a limit process from the general solution: We
Thus N = (l-Q)-l
as-required.
as q(r
I), and we let q - + l / e ,r-1.
The solution for
= 1/next
be verified
similarly. We can also
2 may
Letp us
compute
7.
obtain it by a limit process from the general solution: We write p-q
as q(r-l), and we let q-;.1/2, r-;.l.
Let us next compute T.
tt =
n-l
2: nti = (p-q)(rn-l)
--.,.---~
j=l
[i
(ri - I )(r"-t - 1) +.
)=1
ni,l (ri - 1 )(rn- i
J=~+l
(n-i-
1)rn-i-
l + i simply,
(n - i)l" - nr
Orn -more
(p-q)(rn-l)
Or more simply,
(3)
An interesting question to consider is finding the maxim
= l 1 2 weifsee
middle,
of ti. For
ii = pi(n-i)
p =that
1/2. ti increases till the (4)
decreases symmetrically. If n is even, the mid-point i =
An interesting question to consider is finding the maximum value
ti= (n,'2)2. For the general case we can write down the ra
of /·i. For p = 1/2 we see thctt Ii increases till the middle, and then T
terms, and find the value of i for which this ratio is one.
decreases symmetrically.
If n is even, the mid-point i = n'2 yields
vield an integer
in general, but the i nearest will give a
Similarly,
ti = (n/2)2. For the general case we can write down the ratio of two
terms, and find the value of i for which this ratio is one. This will not
yield an integer in general, but the i nearest will give a maximum.
The answer for j > i is also 0. Hence all off-diagonal ent
LetAPPLICATIONS
us compute the OF
three
non-zeroCHAINS
terms for the i,i-th entry
MARKOV
SEC, 1
'.Ve find the i"pproximate solution
i max = logr((r-l)n).
If p > q, that is r> 1, then the resulting t max is of the order of magnitude
of (n - imax)/(p - q). Thus we see that if p # 1/2 the absorption time is
of a lower order of magnitude for large n than for p = 1/ 2. For example,
if p
and hence r = 2, i max = logzn, and t max is about 3(n -log2n)
for large n. For thc ycry small n= 5 of our example i max = 2.3, and
t
2 is the ma~'imum point, but the value t max obtained is too large
for so small an n.
Finally, weThus
shall
compute
It is sufficient to find bin, since
N=
(I- & ) - I B.
as required.
bto = I - bin. IVe
that rinforis p=l12
0 except
for Tn-l,n
= p.
Thenote
solution
may
be verified
similarly. We
n-]
obtain
i t by a limit process from the general solution: We
narkn
as q(r I), and we let q - + l / e ,r-1.
(p-q)(rn-l)
Let us next compute 7.
bin =
L
k~l
Hence
if P =f. 1
(5)
if P = 1;2.
(6)
Similarly.
For p =
the solution is very intuitive. The probability of ending
up at the right-hand bounda,ry is proportional to the original distance
from the left-h;lIld boundary. But the solution for p =f.
has some
i - 1.1)rn-isurprising features. Let us study the case p>
that (isn -r>
vVe
find that for a giyen starting positioni the probability of ending up at
n is not negligible for any n. Indeed, if we keep ~. fixed and let n
tend to infinity, bin approaches the limit 1- r- t . Hi is fairly large
Or more
simply,
this probability
will be
close to 1 no matter how large n is! Even
fori = 1 we have a probability (r - 1 of ending up at n. The absorption time in this case is about nip, surprisingly smalL For 1) = 2/3
this pro bability is
This means that if p = 2! 3we may put the righthand boundary n as far out as we wish, start the process at i = 1,
and still have a better than enoll chance of ending up at n r,.ther
than at O.
An interesting question to consider is finding the maxim
This particular
is often referred to as "ganlbler's
For p = l\valk
1 2 we see that ti increases till the middle,
of ti. randorn
ruin:" v\' e may
think of
tv;o men playing
repeatedly i =
If na certain
is even,game
the mid-point
decreases
symmetrically.
in which player
A
hels
p
of
winning.
Let
i
dollars
be the
his ra
ti= (n,'2)2. For the general case we can write down
original capital,
n - iand
thefind
fortune
of hisofopponent,
terms,
the value
i for whichand
thisassume
ratio isthat
one. 1 T
dolh" ~s bet vield
each an
tinl(~.
T'hf~n . .4 1S fortune ca.rries out the random
integer
in general, but the i nearest will give a
v;:alk AAR\V with the given p. Absorption at n means that A ends
up with all the money, 'while absorption at 0 means that he is ruined.
\Ve see that for a fair game (p = 1
the probability of ruin is equal to
the fraction of the two fortunes held by the opponent. But th
CHAINS if playerCHAP.
vnan adva
of theFINITE
solutionMARKOV
changes drastically
A has
the game. I n this case he has a good chance of winning out ev
the fraction of the two fortunes held by the opponent. But the nature
opponent has a much greater capital. For example, if he ha
of the solution changes drastically if player A has an advantage in
that is r = 2 (meaning the odds are 2 : 1 in his favor), then
the game. In this case he has a good chance of winning out even if his
better than even chance of ruining a rich opponent even if he
opponent has a much greater capital. For example, if he has p = 2/3,
1 dollar to start with!
that is r= 2 (meaning the odds are 2: 1 in his favor). then he has a
We will briefly mention two applications of this result.
better than even chance of ruining a rich opponent even if he has only
all, gambling houses can exist due to it. They fix the odds
1 dollar to start with!
r > I. Then by making sure that their original capital i
We will briefly mention two applications of this result. First of
enough (measured in terms of the size of one wager), they w
all, gambling houses can exist due to it. They fix the odds so that
probability I - r - f , very near to 1, of staying in business n
r> 1. Then by making sure that their original capital i is large
how much is bet a t their gambling tables. We also see
enough (measured in terms of the size of one wager), they will have
absorption
time is enormous, which is the reason that gamblin
probability l - r - i , very near to 1, of staying in business no matter
have not yet acquired all the money in the world. For r n
how much is bet at their gambling tables. We also see that the
(3) is roughly equal to (4). We may estimate, very conser
absorption time is enormous, which is the reason that gambling houses
t h a t t h e gambling house can cover 10,000 bets, while the
have not yet acquired all the money in the world. For r near to 1,
can provide 1,000,000 bets. Then i(n- i) is about 1010, whic
(3) is roughly equal to (4). We may estimate, very conservatively,
put the absorption time into thousands of years. This leav
that the gambling house can cover 10,000 bets, while the gamblers
opportunity for t h e raising of new gamblers.
can provide 1,000,000 bets. Then i(n-i) is about 1010; which would
A second application is t o a simple model for the p~incipleo
put the absorption time into thousands of years. This leaves ample
selection in the theory of evolution. Suppose that on a n
opportunity for the raising of new gamblers.
island the population of some species is fixed a t n by the s
A second application is to a simple model for the principle of natural
food. Let us suppose that a mutant is born with a slight
selection in the theory of evolution. Suppose that on an isolated
chance of survival than the regular member of the species.
island the population of some species is fixed at n by the supply of
model of the struggle for survival is given by assuming tha
food. Let us suppose that a mutant is born with a slightly better
generation the mutants gain one place with probability p > 1
chance of survival than the regular member of the species. A simple
one place with probability q. We then know that the muta
model of the struggle for survival is given by assuming that in each
probability of more than ( r - I ) / r of taking over the island. I
generation the mutants gElin one place with probability p> 1/ 2 or lose
this probability is .04; if p= .6, the probability is 113. Henc
one place with probability q. We then know that the mutants have
that relatively minor advantages can result in the surviva
probability of more than (r-1
of taking over the island. Ifp=.51,
mutants.
this probability is .04; if p= .6, the probability is 1/3, Hence we see
While this simple model serves to illustrate how mutants m
that relatively minor advantages can result in the survival of the
over a large species, the estimate for the absorption time is un
,
mutants.
Even for p=.6 and n as small as 100 we obtain nip or ab
While this simple model serves to illustrate how mutants may take
generations before the mutants take over. This brings
over a large species, the estimate for the absorption time is unrealistic.
unrealistic nature of the assumption that only one place is ch
Even for p=.6 and n as small as 100 we obtain
or about 167
each generation. For a realistic time for absorption we need
generations before the mutants take over. This brings out the
sophisticated model.
unrealistic nature of the assumption that only one place is changed in
each generation. We
For will
a realistic
time that
for absorption
we need
a more
now show
the solutions
for the
other three
sophisticated model.
walks may be obtained from AARW by various tricks. Le
illustrate APRW for n = 5.
We will now show that the solutions for the other three randor:1
walks may be obtained from AAR\V by varions tricks. Let us first
illustrate APR\V for n = 5.
154
I
I
the fraction of the two fortunes held by the opponent. But th
of
the solution changes
drastically
if player A has an
SEC. 1
APPLICATIONS
OF MARKOV
CHAINS
155adva
the game. I n this case he has a good chance of winning out ev
APRW for opponent
n=5, p=1/2
has a much greater capital. For example, if he ha
that is r = 2 0(meaning
1
2the odds
3
4are 52 : 1 in his favor), then
better than even chance of ruining a rich opponent even if he
0
0
1 with!
0 to start
0
0
0
1 dollar
We will briefly mention two applications of this result.
1
1/houses
1/2 exist
0
0
2 0 can
0 it. They fix the odds
all, gambling
due
to
1/2 0 sure
r > I. 2Then 0by making
that
0their original capital i
1/2 0
p= (measured in terms of the size of one wager), they w
enough
3
0
0 l/ Z 0 1/2 0
probability I - r - f , very near to 1, of staying in business n
4
0
0
0
0 1/ 2
how much is bet a t theirl/Z
gambling
tables. We also see
5
0
0
1/2
0
0 which 1/2
absorption time is enormous,
is the reason that gamblin
have
all the money in the world. For r n
2 acquired
1 not yet
2 2
(3)2is roughly
4 equal
4 4 to (4). We may estimate, very conser
t h a t t h e gambling house can cover 10,000 bets, while the
6 6
N can
= 3 provide 41,000,000
bets. Then i(n- i) is about 1010, whic
4
6 8time into thousands of years. This leav
4
put the absorption
opportunity
5
4for6t h e8 raising of new gamblers.
A second application is t o a simple model for the p~incipleo
APRW for selection
n=5, p= in
2/ 3 the theory of evolution. Suppose that on a n
4,
0
1
2 some
5 is fixed a t n by the s
3
island the population
of
species
food. 0Let us1 suppose
that
a
mutant
0
0
0
0
0 is born with a slight
chance of survival than the regular member of the species.
model 1of thel/Sstruggle
0 2/for
0
0 is0 given by assuming tha
3 survival
generation the mutants gain one place with probability p > 1
0
0
0
2
l/s 0 2/q.3 We
then know that the muta
one place with probability
P=3
0
0 than
lis ( r0- I ) 2/
probability of0 more
/ r3of taking
over the island. I
this probability
is 113. Henc
4
0 is0 .04;0 if p=
l/S .6,0 the2/ probability
3
that relatively minor advantages can result in the surviva
1/
2/3
0
5
0
0
3
0
mutants.
24
6 simple
12
While this
model serves to illustrate how mutants m
over
absorption time is un
,
2 a large
3 9 species,
18 36 the72estimate for the 138
Even for p=.6 and
n as small T as= 100 159
we obtain
nip or ab
.
N = 3
3 9 21 42 84
generations before the mutants take over. This brings
90 assumption that168
4
3 9nature
21 45
unrealistic
of the
only one place is ch
5 generation.
3 9 21 45
each
For 93
a realistic time for171absorption we need
sophisticated
model.
The N matrix
for APRW
may be obtained from AARW through
T~GD
G D
l( 4) ( 93)
the following observations. Consider i >j. If the process
We will now show that the solutions for the other three
walks may be obtained from AARW by various tricks. Le
illustrate APRW for n = 5.
G
e
k
8
®
n
j
FIGURE 7·3
FINITE MARKOV CHAINS
156
starts in i, it must eventually enter j. Hence hcj =nij/n
FINITE
CHAINS
(ThisMARKOV
is conspicuous
in N above.) CHAP,
If k <jVII
, then hk
nil = njj.
gotten as the probability of ending up a t j in AARW with n =j
starts in i, it must eventually enter j. Hence h ij = nij/njj = 1, and
n k j = hkjnjj= b k j n j j , with bki from AARW.
Thus it suffices Lo
n/j=nJ}.
(ThisLet
is conspicuous
in N the
above.)
k<j, then hk ; may be
us first compute
case pIf
= 1 / 2 . We note that & (exce
gotten as the probability of ending up at j in AAH;W with n = j. Hence
last row and column) is a submatrix of Q for larger n. A
nkj = hkjnjj = bkjnjj, with b kj from AARW.
Thus it suffices to find njj.
larger Q these rows and colurnns are filled out with 0's. He
Let us first compute the ca,se p = I! 2. We note that Q (except for the
independent of n. But as we let n+co in AARW, the
last row and column) is a submatrix of Q for larger n. And in the
between i t and APRW disappears. Hence njj for APRTV is
larger Q these rows and columns are filled out with O's. Hence nii is
as n+co of njj for AARW. Thus
independent of n. But as we let n-+oo in AARW, the difference
nfj = 2j Hence njji ffor
i APRW
j 1 is the limit
between it and APRW disappears.
as n-+oo of njj for AARW. Thus
156
niJ =
2j
~
if i ? j
The same argument is applicable
if r <'I,.
1 . Thus
} fo, p.
~. 2j = 2i if i ~ j
nij =
(7)
J
The same argument is applicable if r < 1.
Thus
ri - 1
nij = - -
p-q
ifi?j}
forpof l !2.
(8)
ri - ri-i
- 1 alsori obtain
- r j - i nij for i # j from AT for AARW by lettin
Werican
nij = --_._- = - - -
if i .,,; j
r i - IForp-q
r > 1 the p-q
above argument breaks down, since no matter h
n is in AAIZW, the probability of ending up a t n does no
We can also negligible.
obtain nij forBut
iofjhere
fromweN make
for AAR\V
n-+oo.
use ofby
theletting
fact that
if in A
For r> 1 the renumber
above argument
breaks
down,
since
no
matter
how
large with
state i as 12-i, we have the original process
n is in AAH,W,interchanged.
the probability
ending
at nthedoes
not formulas
become also
Weofthen
findupthat
above
negligible. But
r >here
l . we make use of the fact that if in AARW we
renumber state i For
as nT- we
i, obtain
we have the original process with p and q
interchanged. We then find that the above formulas also hold for
r> 1.
For T we obtain
(2n-i+ 1)£
(0)
there
is only one absorbing
state, a.ii absorption pro
t, =Since
_1_.
[:-"Tl-rn-i+l_
i ] ifp #- 1/ 2.
are I.p-q,
r-l
I n both AARW and APRW there is a simple relation be
Since there isand
only
oneI nabsorbing
state,
all absorption
prohabilities
nji.
fact i t may
be verified
from our
formulas for the
are 1.
tities that
In both AAR\V and APR\V there is a simple relation between ntj
and njt. In fact it may be verified from our formulas for these quanWe shall see that there is a simple probabilistic proof of this
tities that
Assume that j i i. Let d = i - j. Then any path which a
process to go from i to j in n steps must take d more steps t
We shall see that there is a simple probabilistic proof of this fact..
Assume that j < i. Let d = i - j. Then any path which allows the
process to go from i to j in n steps must take d more steps to the left
FINITE MARKOV CHAINS
156
starts in i, it must eventually enter j. Hence hcj =nij/n
APPLICATIONS OF MARKOV CHAINS
157
nil = njj. (This is conspicuous in N above.) If k < j , then hk
gotten as
the probability
of ending
j in AARW
withatn =j
than to the right.
Hence
the probability
that up
thea tprocess
starting
k j = hkjnjj
b k j n j j , with bki from AARW. Thus it suffices Lo
i will follow nsuch
a path= is
Let us first computep"qk+ct
the case p = 1 / 2 . We note that & (exce
last row and column) is a submatrix of Q for larger n. A
where 2k + d larger
= n. On the other hand each such path, looked at backQ these rows and colurnns are filled out with 0's. He
wards, may be
considered of
a path
to i. For the process to follow
n. fromj
independent
But as we let n+co in AARW, the
this path it between
must make
d
more
steps
to the right than to the left.
i t and APRW disappears.
Hence njj for APRTV is
Hence the probability that, starting at j, the process will follow this
as n+co of njj for AARW. Thus
path is
nfj = pk+ctq".
2j
ifi j 1
SEC. 1
There are the same number of paths from i to j in n steps as there
are from j to i, but the ratio of the probability for each path is
p"qk+d is qd
,_ if r < 1 .
The same argument
applicable
- - = - = rJ~.
p/ctdq"
pd
Hence, p(n)ij = ri-ip(n)j;.
2:
But then
2:
'"
nij =
Thus
ro
p(n);j =
rj- i
n=O
p(n)j;
= ri-inji.
n=O
Note that this We
also can
shows
when
p= 1/2, then r= 1, nij=nji.
alsothat
obtain
nij for i # j from AT for AARW by lettin
We now turnFor
to the
regular
chains,
starting
with down,
PPRW.
r > 1 the above argument
breaks
since no matter h
is in p=l/2
AAIZW, the probability of ending up a t n does no
PPRW fornn=5,
negligible. But
2 make
4 of5 the fact that if in A
0 here we
3 use
renumber state i as 12-i, we have the original process with
liz liz 0 0 0 0
0
interchanged. We then find that the above formulas also
liz
0
1/2
0
0
0
For T 2we obtain
0
1/2
p=
0
1/2
0
0
0
r>l.
1
3
0
0
1/2
0
1/ 2
4
0
0
0
1/2
0
5
0
0
0
0
1/2
1
10
6
4
10
Since there is only one absorbing state, a.ii absorption pro
I. (l/s, lis, 1/6, lis, 1/ 6, 1/ 6)
area=
I n both AARW
there
0
1and 2APRW
5 is a simple relation be
3
4
and nji. I n fact i t may be12verified from our formulas for the
0
6
2
6
20 30
tities that
18
28
18 there
8
6is a 6simple
14 probabilistic
24
We shall2 see that
proof of this
M
j
i
i.
Assume
that
Let
d
=
i
j.
24 14
6
6
8 18Then any path which a
3
process to go from i to j in n steps must take d more steps t
4
28
18
10
4
6
10
5
30
20
12
6
2
6
P P R W for n = 5, = 213
FINITE MARKOV CHAINS
0
1
2
158
CHAP. VII
3
4
5
PPRW for n=5, p=2/s
p=
2
1
0
4
3
5
0
1/3 2/3
0
0
0
0
1
1/ 3
0
2/3
0
0
0
2
0
1/3
0
2/3
0
0
3
0
0
lla
0
2/ 3
0
4
0
0
0
lis
0
2/ 3
5
0
0
0
0
l/a
2/ 3
ex = (1/63,2/ 63 , 4/a a , 8/ 63 , 16/ 63 , 32/63)
0
M=
1
2
3
4
123
1
93
3/ 2 15/ 4 51/ 8
63/2 9/ 4 39/ 8
2
138
45
63/ 4
21/ 8
87/I6
3
159
66
21
63/ S
45/16
0
63
5
147/16
116
63/16
168 75
30from9 PPRVV
We may obtain
APRW
by making 0 absorbing.
12
78
33
if we 5know171
a for PPRW, we may3 obtain Ikf from N of AP
Corollary 6.2.6,
We may obtain APRW from PPRW by
1 making 0 absorbing. Thus,
i # j .by
if we know ex for PPRW, we may
from
APRW
mi! =obtain
- ( n j jM
- nit)
+ tt N- t jof for
a!
Corollary 6.2.6,
1
If p = 112,1 we have column sums 1, and hence a=T
mlj = - (njj-njj)+t,-tj for i of- j.
n+1'.
aj
obtain
I
If p= 1/2, we have column sums 1, and hence ex=--Tj. Thus we
n+l
(2n-i+l)i-(2n-j+l)j
i >
obtain
4,
mij
{~2:~i+l)i-(2n-j+1)j ~ :~1
( 11)=rat.
= For p# 112 the above example suggests
p = liz· that ni+l
thisj(j+l)-i(i+l)
yields a fixed vector, as can
i < be
jJ seen from writing our equa
For p #- 1/2 the above example suggests that aHl = mj. Indeed,
this yields a fixed vector, as can be seen from writing our equations as
or
pan-l + pa" = an
or
p
rai+l + qrat- = at
1
1
a,,-l = - an·
r
SEC. 1
P P R W for n = 5, = 213
APPLICATIONS OF MARKOV
0
1 CHAINS
2
3
4
5
159
Thus
Q,
r-1
- - - ri
rn +1 -1 .
=
Then
i = j
I
-.
p-q
[rn-1+1- r,,"':;+l
. .]
-(t-J)
r-1
1
[..
rJ - ri ]
p_q' (J-t - (r-1)r i +1
I
>j
i < j
Let us compare mOn with mno. If p=lh, moC'=mno=n(n+I),
which is of the order n Z for large n. But if p i= liz, we observe a completely different behavior. Say r> 1, then
rnOn = _1_. [n- rn-1 ]
(r-l)r"
p-q
is roughly _1_. n for large n.
Hence it takes a surprisingly short
p-q
time to go "all the way" in the favored direction.
1
[r1>+1- r
]
1
+
m 0 = APRW
pq' --r=-l-n
We may obtain
from PPRVV by making 0 absorbing.
if we know a for PPRW, we may obtain Ikf from N of AP
which is roughly
(p-q;'(r-l)
Corollary
6.2.6, ·r" for large n. Since r> 1, this increases
71
exponentially in n. For very
large
it nit)
will tttake
tremendously
i # j.
mi! =
- ( n j j- t j a for
a! wrong direction. This type of
long time to go "all the way" in the
1
behavior is If
typical
of random
will 1,
seeand
another
of
p = 112,
we havewalks.
columnWe
sums
henceexample
a=n
+
1'.
this in the Ehrenfest model.
obtain
Finally we consider CRW.
CRW for n= 5, p= 1/2
( 2 n - 2i + l ) 3i - ( 2 4n - j +5l ) j
i >
liz
T
o
o example
o suggests that ni+l =rat.
For p# 112 the above
this yields a fixed vector,
liz as can be seen from writing our equa
o
o
CRW for n = 5, p = 213
FINITE MARKOV CHAINS
160
CHAP. \:'11
CRW for n=5, p=2j3
1
2
P=3
4
5
I
2
.11-1
3
4
0'
2/ 3
2
3
4
2!
.' 3
0
0
0
2/3
0
1/ 3
0
2/3
0
1/3
0
2/3
0
0
lis
0
5
T)
174/ 31
147/al
5
78/ 31
"'I,,)
141/ 31
174/31
147/ 31
5
78/ 31
c~"
78/ 31
141/ 31
174/ 31
5
78/ 31
141 31
j
1
174/ 31
141/ 31
141/ 31 has
174/a31
This
process always
= - q147/
, since
P5
31
5
always has column s
n
This is also obvious from the fact that no position in the ci
distinguished.
same
that
mtf l.
should d
This process
always has a =For
~'7'the
since
P reason
alwayswe
hasexpect
column
sums
only on how far i and j are apart (and on the direction if p
This is also obvious from the fact that DO position in the circle is
This is certainly so in the examples above.
distinguished. For the same reaSOD we expect that mlj should depend
If state n is made absorbing in CRW, we obtain BARW-wi
only on how fari and j are apart (and on the direction if p =f. 1/ 2 ).
and n identified. Hence M for CRW is obtainable from N for A
This is certainly so in the examples above.
But there is an even simpler method. If we want mil, we ren
If state n the
is made absorbing in CRW, we obtain AARW-with 0
states so that j becomes n, and then mil is just an absorptio
and n identified. Hence M for CRW is obtainable from N for AARW.
for AARw. Specifically, if the distance from j to i (clockwise
But there is then
an even
simpler
method. If we want mij, we renumber
mij =
td. Thus
the states so that j becomes n, and then rfiij is ju~t an absorption time
78/ 31
for AARW. Specifically, if the distance from j to i (clockwise) is d,
then mij = td· Thus
mii
= n
1
;'
Tn -
rn - d
d}
( 13)
p =f. 1/2, i =f. j
p -q l
Tn - 1
where d is the clockwise distance from j to i.
mjjThe
= d(n-d)
Ih,i=f.j
remarks made for T of pAARW
are applicable here. Th
example, the distance (clockwise) that it takes longest to tra
where d is the clockwise distance from j to i.
approximately log,((r- l)n). M'e also note that mil is generall
The remarks
made
forofT magnitude
of AARVil for
are any
applicable
here. Thus, for
p#
lower
order
than
for p =
mij =
- '-.~n
example, the distance (clockwise) that it takes longest to travel is
approximately Let
logr((r-l)n).
Wethe
also
note thatmatrices
mij is generally
of a
us next find
transition
for the reverse
pro
lower order of
magnitude
for
any
p
=f.
1,2
than
for
p
=
1
j
2.
of PPRW and CRW.
Let us next find the transition matrices for the reverse processes
of PPRW and CRW.
CRW for n = 5, p = 213
SEC. 2
APPLICA TW::\S OF MARKOV CHAINS
161
For PPRW, remembering that ll!+l=raj,
al+l
Pi,HI
= - - Pl+l,1 = rq = P = pi,HI
PHI,;
= - - pi,l+l = -r P = q = PHl,i.
al
1
ai
a/+!
Hence the process is reversible, contrary to one's intuition.
CRW a,+l = at, hence
,
PI,HI
,
Pi-I-l,i
But for
ai+l
= -at- PI+l,i = q
at
= - - pi,Hl = p.
aHl
Hence the reverse process is a CR\V with P replaced by g, as one would
expect, It is This
reversible
only
whenhas
p = aq==-1q1/,2.since P always has column s
process
always
n
Let us close by giving a practical application
for PPE',V. \Ve
This is also
obvious
from the
no position
consider a gambling
house
that wishes
to fact
keep that
a close
check oninitsthe ci
distinguished.
For itthe
reason
we expect
that mtf
should d
roulette wheel.
Suppose that
is same
a wheel
having
in addiLioll
to the
j are
apart
the
direction
only
how
far i and
numbers 1 to
36, on
half
of \vhich
are red
and
half (and
black,onthe
numbers
0 if p
Thisare
is certainly
so in the
above.
and 00 which
not colored.
We examples
will devise
a simple automatic
If state
n is made
absorbing
we on
obtain
check to see that
the house
is taking
its shareinofCRW,
the bets
red. BARW-wi
\Ve
and n identified.
Hence
M at
for0 CRW
is obtainable
fromred
N for A
set up an electric
counter which
starts
and adds
1 every time
But there1 isevery
an even
wantcounter
mil, we ren
comes up, subtracts
timesimpler
red failsmethod.
to come If
up.we The
states
that
becomes reaches
n, and then
mil is just
an absorptio
does not gothe
below
O. soIf
thej counter
a specified
number
n,
AARw.
Specifically,
if thethedistance
then it ringsfor
a bell,
and the
house changes
wheel. from j to i (clockwise
then mij
= td. Thus
If the wheel
is properly
balanced, then we have PPRW with
P = 18/ 38 = 9/19. Hence it will take a very long time to reach n. The
house adjusts n so that mOn corresponds to its normal periodic servicing
of the wheel. However, if the wheel fails to function properly-for
example, if p rise;; to 1/ 2, in which case the house no longer makes a
profit-then the bell will ring much sooner. A similar check on black
will assure that the house continues to make its profit on all bel,.
Let us consider
example
of this.from
Suppose
the clockwise
distance
j to i. that n =10 is
where da isconcrete
selected. Then
for proper
is about
18,000. Ifhere.
the Th
ThemOn
remarks
made functioning
for T of AARW
are applicable
example,
the distance
that
it on
takes
wheel is turned
400 times
in a day,(clockwise)
the bell will
ring
the longest
average to tra
once in 45 days,
allowing for
normal
servicing.
p rises
l)n).
M'e alsoBut
noteif that
miltois generall
approximately
log,((rp#
lowerand
order
magnitude
for any
p =above
then mOn = 1640,
theofbell
will ring after
four
days.than
If pfor
rises
the break-even figure (that. is, the house is losing money), then the bell
us next find the transition matrices for the reverse pro
quickly.
will ring very Let
of PPRW and CRW.
§ 7.2 Applications to sports. Let us <1pply some of our results to the
game of tennis. We will first consider the problem of a single game of
tennis played between two players. We will assume that player A
has probability p of winning any given point, and
FINITE MARKOV
CHAINS
probability
q < p.
CHAP. VII
If we keep score in the ordinary manner, there are 20
has probability p during
of winning
any These
given are
point,
and player B has
a game.
: 0-0, 0-15, 15-0, 0-30, 15-1
probability q ~p. 15-30, 30-15, 40-0, 15-40, 30-30, 40-15, 30-40, 40-30,
162
If we keep score in the ordinary manner, there are 20 possible scores
Deuce, Advantage A , Game B, Game A . However, it
during a game. These are: 0-0, 0-15, 15-0, 0-30, 15-15, 30-0, 0-40,
15-30, 30-15, 40-0, 15-40, 30-30, 40-15, 30-40, 40-30, Advantage B,
Deuce, Advantage A, Game B, Game A. However, it is easily seen
that we may lump the following pairs: (30-30, D
7·4 Advantage A). The resulting ra
AdvantageFIGURE
B), (40-30,
then represented by Figure 7-4.
that we may lump the following pairs: {30-30, Deuce), {30-40,
The game of tennis may conveniently be broken d
Advantage B}, {40-30, Advantage A}. The resulting random walk is
stages. At the beginning the process goes through som
then represented by Figure 7-4.
twelve states, always moving up, and in four or five step
The game of tennis
may
broken
down This
into we
twowill re
one of
theconveniently
five states inbethe
top row.
stages. At the beginning the process goes through some of the lower
twelve states, always moving up, and in four or five steps it arrives at
one of the five states in the top row. This we will refer to as the
has probability p of winning any given point, and
APPLICATIONS
163
probability q <OF
p. MARKOV CHAINS
If we keep score in the ordinary manner, there are 20
preliminary process. The preliminary process is of an extremely simple
during a game. These are : 0-0, 0-15, 15-0, 0-30, 15-1
nature. It is followed by a random walk of type AARW, with n = 4.
15-30, 30-15, 40-0, 15-40, 30-30, 40-15, 30-40, 40-30,
The st.ate Game B is the absorbing state 0 of AARW, and Game A is
Advantage A , Game B, Game A . However, it
the absorbing stateDeuce,
n.
We will describe the entire process as an AARW random walk, with
initial probabilities furnished by the preliminary process. These initial
probabilities 7T = (co, Cl, Cz, C3, C4) are given by an elementary probabilit.y calculation. We find that
SEC. 2
Co = q4(1 + 4p),
Cl = 4p 2q3,
= 4p3qZ,
C3
C4
C2 = 6p2q2,
= p4(1 + 4q).
If p=1/2, Co=3/16, 01=1/ 8 , C2=3/ 8, C3=1/S, C4=3/1S.
quant.ities of AARW with n=4,
((rN=
-1)
1
,\(r-1
(p-q)(rLl)
1)
\(r~ 1)2
(r - 1)(r 3 - r)
(1 2 _ 1)2
(r-1)(r2- 1)
r 4 - r3
.
4---1
r 4 -1
r 4 _r2
1
T
= p-q
{b i4 } = _._1_
4---2
r 4 -1
r 4 -1
Using the basic
r-"\
r4- r
r 4 -1
r4 _ r2 ~
\r 4 -r
}
p "# I/Z
4---3
C ';')
V~2
N=
1
2
liz
1
3/2
I
~ G)
{b i4 } =
C)
1~2
P = 1/2
3 14
we can find all interesting quantities. The most interesting one is,
of course, t.he probability that A will win. For p= 1/2 we obtain
(3/16).0+(1'8)'(1(1)+(38).(12)+(1/8).(3/4)+(3/16)-1 =
that we may lump the following pairs: (30-30, D
This was t.o be expected,
by symmetry.
If p > 1/2, we
Advantage
B), (40-30, Advantage
A).obtain
The resulting ra
then represented by Figure 7-4.
3) + 6p2q2(r4
_1_[q4(1 +
.0+ 4game
r 2) + 4p2q2(r4be
- r)broken d
- rtennis
p 2q3(r 4of
The
may -conveniently
r4 - 1
stages. At the beginning the process
goes
through
+ p"( 1 -1- 4q)(r4 -1)] som
twelve states, always moving up, and in four or five step
which simplifies toone of the five states in the top row. This we will re
For example, if p = .51, then p~ = .525, and if p = .6, then p~ = ,7
VII we w
FINITEtime
MARKOV
For the absorption
and theCHAINS
number of times CHAP.
in a state
carry out the computation only when p = 'Ip. The interesting ca
For example,
if p=.51,
PA=.525,
and if p=.6, then PA=.736.
are ones
where pthen
is near
112, and absorption times do not depend ve
For the absorption
time
and
the
number
of times in a state we will
drastically on p.
carry out the computation only when p= liz. The interesting cases
When p=1/2 we find that the mean number of times in the th
are ones where p is near liz, and absorption times do not depend very
for Advantage
interesting transient states is: 1 for Deuce, and
drastically on p.
and for Advanta,ge A. The absorption time is 914. To find the act
When p = 1/ 2 we find that the mean number of times in the three
length of the game, we must also take into account the prelimina
interesting
transient
states
for that
Deuce,
5 j 8 for Advantage B
stage.
If we
do, is:
we 1find
theand
mean
length of a tennis ga
and for Advantage
A.
The
absorption
time
is
9/4.
find Since
the actual
between equally matched opponents is 2714 =To
6914.
t.he minimu
length oflength
the game,
mustis also
takeshows
into that
account
the preliminary
of thewegame
4, this
for equally
matched play
do, welength
find is
that
mean
length
of a tennis game
stage. If
theweaverage
not the
much
above
the minimum.
between equally
matched
opponents
is
27/4 = 63/4.
Since
the minimum
Should it prove from records that games actually
are much lon
length ofthan
the game
is 4,that
thisthe
shows
thatnumber
for equally
matehed
players
this, and
average
of times
in Deuce
is well abo
the average length is not much above the minimum.
1, as seems to be the case, then i t would indicate that the present mo
Should it prove from records that games actually are much longer
for tennis is too simple. Perhaps a player "plays harder" when he
than this,behind.
and thatThis
the average
number
times in Deuce
well above rando
would lead
to aofsomewhat
more iscomp1icat)ed
1, as seems
to
be
the
case,
then
it
would
indicate
that
the
present
model
walk.
for tennis is too simple. Perhaps a player "plays harder" when he is
Let ns return to the probability that the better player A wins. T
behind. This would lead to a somewhat more complicated random
is always greater for a game than for a n individual point. Thus,
walk.
game serves to magnify the difference between the two players. T
Let us is
return
to the
probability
thegames
betterare
player
A wins. This
further
magnified
since that
several
played
in a set,, and seve
is always greater for a game than for an individual point. Thus, the
sets in a match. The probabilities for a set, in which a player m
game serves
difference
theoftwo
players.
win to
a t magnify
least six the
games,
but bybetween
a margin
a t least
two, This
may be co
is further magnified since several games are played in a set, and several
puted just as above. We are led to the same AARW, but with
sets in a match. The probabilities for a set, in which a player must
longer preliminary st,age. A match is won by the first player winn
win at least six games, but by a margin of at least two, may be comthree sets. This is a stmightforward computation. The follow
puted just as above. We are led to the same AARW, but with a
figures will illustrate the magnifi~at~ionachieved in sets a
longer preliminary
matches : stage. A match is won by the first player winning
three sets. This is a straightforward compntation. The following
figures will illustrate the magnification achieved
in sets and
71 = .51 1) = .6
matches:
164
Probability of winning point
.510
.525
Probability of winning game
]) = .51 P .573
= .6
Probability of winning set
Probability of winning point
.510
.600
Probability of winning match
,635
,600
.736
,966
,9996
Probability of winning game
.525
.736
Probability of winning set
.573
Thus there is always a good chance
that .966
the better player will w
.635
.9996
Probability of winning match
the match. And if there is a fairly significant difference betwe
players, then i t is practically certain that the better player wins.
Thus there Let
is always
a goodthese
chance
thatfor
thetennis
betterwith
player
win Series
us compare
results
the will
World
the match.
And
if
there
is
a
fairly
significant
difference
between
baseball. Here the team that first wins four games is declared winn
players, then
is practically
certain
that
the better
player
wins.any one gam
If weitassume
that team
A has
probabiiity
p of
winning
Let us compare these results for tennis with the World Series in
baseball. Here the team that first wins four games is declared winner.
If we assume that team A has probability p of winning anyone game,
For example, if p = .51, then p~ = .525, and if p = .6, then p~ = ,7
For the
absorption time
the number
of times in a 165
state we w
APPLICATIONS
OF and
MARKOV
CHAINS
carry out the computation only when p = 'Ip. The interesting ca
then we find
arethat
ones where p is near 112, and absorption times do not depend ve
drasticallyPAon=p.JI4(1->-4q+10q2+::0qZ),
When p=1/2 we find that the mean number of times in the th
where the various terms correspond to series of 4, 5, 6, and 7 games
for Advantage
interesting transient states is: 1 for Deuce, and
respectively. If p=.51, pA=.522, while if p=.6, then PA=.710.
and for Advanta,ge A. The absorption time is 914. To find the act
The World Series also magnifies differences between teams, but not
length of the game, we must also take into account the prelimina
nearly as well as a match of tennis-or even a single game of tennis.
stage. If we do, we find that the mean length of a tennis ga
If we compute the mean length of a series, ,ve find this to be
between equally matched opponents is 2714 = 6914. Since t.he minimu
t = 4(p' -7- q4) + 20(p4q + pq") + 60(p4q2 + p2q4) .'- gO(p4q3 + p3q4). This is
length of the game is 4, this shows that for equally matched play
largest (;j.8)) when p= 1
and decreases monotonically to 4 as p is
the average length is not much above the minimum.
increased to I. Hence we should he able to estimate p from the
Should it prove from records that games actually are much lon
observed length of World Series. In the 50 World Series played under
than this, and that the average number of times in Deuce is well abo
the stated rules from 1905 to 1957,10 ended in four games, 13 in five
1, as seems to be the case, then i t would indicate that the present mo
games, 12 in six games, and 15 in seven games. This yields a mean
for tennis is too simple. Perhaps a player "plays harder" when he
length of 5.64, which would agree very ·well with p = 5/ 8. This would
behind. This would lead to a somewhat more comp1icat)ed rando
suggest that the teams ph1ying in the World Series have, on the average,
walk.
not been matched too closel~'.
Let ns return to the probability that the better player A wins. T
Let us now consider the efficiency Cifvarious procedures in magnifying
is always greater for a game than for a n individual point. Thus,
the differences between players or teams. We have found that tennis
game serves to magnify the difference between the two players. T
(one game) gives slightly more magnification than the World Series,
is further magnified since several games are played in a set,, and seve
b'ut it also requires more steps on the average. To bc able to compare
sets in a match. The probabilities for a set, in which a player m
the efficiency of two rules. we will have to take them so that the mE':1l1
win a t least six games, but by a margin of a t least two, may be co
length of a series is the same.
puted just as above. We are led to the same AARW, but with
Tennis may be compi.Lred to the World Series as follows: The latter
longer preliminary st,age. A match is won by the first player winn
requires fOllr wins, while the former requires that the winner have
three sets. This is a stmightforward computation. The follow
four wins and be ahead by two. This is really a hybrid between two
figures will illustrate the magnifi~at~ionachieved in sets a
types of procedures. The pure procedure would be to require that the
matches :
SEC. 2
winner end up ahead by four poillts. It can easily bc seen that this
rule gives much mOTe magnification than the other, but also requires
71 = .51 1) = .6
a great deal more tirne. Let us therefore consider two classe~ of
Probability of winning point
.510
,600
rules.
Probability of winning game
RULE
.525
.736
IV n: The first person
to win ofn winning
points issetdeclared .573
winner.,966
Probability
Probability
of winning
,9996 is
The first person
to get ahead
of hismatch
opponent,635
by n points
declared winner.
Thus there is always a good chance that the better player will w
vVe will compare
theseAnd
two if
rules,
selecting
n in each
case so difference
that the betwe
the match.
there
is a fairly
significant
mean length
of a game
large number,
which
will denote
players,
then he
i t ais given
practically
certain that
thewebetter
player wins.
by N2, awl seeing
much they
Since
will Series
Let ushow
compare
thesemagnify
results differences.
for tennis with
theweWorld
be interested
particularly
large
n, that
we will
baseball.
Hereinthe
team
firstallow
winsasymptotic
four games approxiis declared winn
mations. If
Since
we are particularly
in magnifiGiLtion
of very
we assume
that team Ainterested
has probabiiity
p of winning
any one gam
SDlall differences, we will let the players differ by E, that is let p =
(1 + E»). and q = (1- E)/2, and compute the final difference just to the
first order term in E.
RULE An:
FINITE BIARKOV CHAINS
166
The following id en ti tie.^ will be nseful :
FINITE MARKOV CHAINS
166
CH
CHAP. VII
The following identities will be useful:
i (n + k) = (2nn+n+k
1)
,
i e~+k)k
= 11,(/2n+ 1) _(2n+ 1)
k
. n
n-I
k~O
k
k=O
k~O \
i (n+k)(~)k
k~O!
= 2n
~ k=O
k
Wn,= (n+1)2"- 2n,,~1 (2n)
i (n+k)(~)k.k
If we use
k~O
the probabilitv that the better player wins is
n - 1 n - l + k'"
n
PA
=
P
n
Zthe better player
)Ok wins is
If we \1Se W n, the probability that1-0
k
2
(
PA = plI nf (n-1 +k)" qk
k
k-O
=
12n (2,
j
2n,-1 2n-2
LL-Z
28 ( ~ l - . [ n ~ - l n-- ~1 ) ] )
[n2
2)1)J
Using
formula+
and simplifying we obtain
+ n€Stirling's
11 - 1 _ t
n - 1 _ 2n - 1 (2n -
=
2n- 1
\
n- I
Using Stirling's formula! and simplifying we obtain
PA : : : ~+J?::
2
The expected length of the game (for which we may use p = 1
1T
E
The expected length of the game (for which we may use P = 1/2) is
t = 2· (1/2)n
2 (n-1k + I)' (n + lc)(1/z)k
n-l
!
k={)
= (1/z)1I-1
[n211-1 + n2,,--1 _ 2,n2 =I (2n
- 2)l
,n- 1
n 1
...l
We will 2simply
v;' •vr use t z 2%. ThusN if we want t 2 N2, -2we must
N2
n = -This yields pa z
43; ., p,-psz1\'J;
r ; and
We will simply use 2t:::::
Thus if we want t::::: NZ, we must choose
: : : 2n-
n.
2n,
N2
n=2'
+
J"
N
)
2:;;.E;
- or about
.8N and thus a
magnification
factor
is NPA-PB:::::N
This yields
PA::::: 1/2+
V:r;;:'"
magnification factor
isInc.,
N J~
about
& Son,
Newor
York,
1957,.8.V.
Chapter 2 .
t See W . Feller, lnlroduction Lo Probabtlzty Theory and 11s ApplicaLions, Jo
t See W. Feller, Introduction to Probability Theory and 1ts Applications, John ,VHey
& Son, Inc., New York, 1957, Chapter 2.
FINITE BIARKOV CHAINS
166
CH
The following id en ti tie.^ will be nseful :
SEC. 3
APPLICATIONS OF ~L4.RKOV CHAINS
167
--~~~--------~
The method An may be represented as AARW from 0 to 2n, starting
at n. Hence
n+k
r2n -rn
PA k== O r2n-1'
Here we find that the terms in E cancel and that €2 terms must be
carried. We find
k=O
If we use Wn,the probabilitv that the better player wins is
1(2+ 1!z(1 +3n)£
n- 1
PA = P n Z
1-0
_1
n
~ /2+"2 E.
n-l+k
1 + (l + 2n)£
(
)Ok
For thi? mean length we find (using p;::
t=n(2n-n);::n2. Thus
N
we 'choose n=N, and obtain PA;::
+2" E, PA-PB;::N€; yielding a
magnification factor of N.
2n,-1 2n-2
LL-Z
28
( ~comparable,
l - . [ nin ~the- sense
l n-- ~that
1)])
We thus find that W S '/2 =
and
AN are
each yields an approximate mean time of N2. The former magnifies
Using by
Stirling's
formula+
simplifying
we obtain
minute differences
about .3S,
while and
the latter
multiplies
them by N.
Thus the An rule is more efficient. Any mixture of the two, as in
tennis, will lie ill between, and hence will also be less emcient than An.
j
§ 7.3 Ehrenfest
for length
diffusion.
There
a simple
foruse
a p =1
The model
expected
of the
gameis (for
which model
we may
system of statistical mechanics which is due to T. Ehrenfest. In this
model we con8ider a gas which is contained in a volume that is divided
into two regions A and B by a permeable membrane. vVe assume that
the gas has s molecules. At each instant of time a molecule is chosen
at random from the set of 8 molecules and moved from the region that
it is in to the other region. We are interested in the way in which the
composition of the two regions changes with time. For example, if
we start with all the molecules in one region, how long on the average
if we want
N2, we must
We c2,ch
will simply
z 2%.
will it be before
regions use
has thalf
the Thus
molecules!
Sucht 2
questions
N chains.
N2
can be answered
by
using
the
methods
of
Markov
n=
This yields pa z
43; ., p,-psz1\'J; 2 r ; and
2
We form a ::\larkov
chain as follows: We assume first that the molecules are identifiable. "Ve take as states a vector y= (Xl, X2, . . . ,x.)
.8N
factor
is region
N - or
where Xj is 1 ifmagnification
the j-th molecule
is in
A, about
and 0 otherwise.
Knowing the state tells us the exact composition of A and hence also of B.
t See W . Feller, lnlroduction Lo Probabtlzty Theory and 11s ApplicaLions, Jo
There are ;28 states.
If the
is inChapter
state y,
& Son,Inc.,
Newprocess
York, 1957,
2 . then choosing a molecule
at random and moving it to the other region means that we change
the state y to a state 0 by simply changing one coordinate of y. It is
clear that from y there are s states to which the process can move and
+
--
J"
FINITE MARKOV CHAINS
168
CHA
these transitions occur each with probability 11s. I t is possible
MARKOV
CHAINS
from anyFINITE
state y to
any other
state 6 in a sequence
steps, bu
CHAP.ofVII
possible to go from y to y only in an even number of steps; an
these transitions
occur
with probability
possible
to go
possible
in each
two steps.
Hence we 1/8.
have It
anisergodic
chain
with per
from any state
i'
to
any
other
state
8
in
a
sequence
of
steps,
but
is only
It is clear also that we can go from y to 6 in one step it
if and
in step.
an even
steps;
and it is
possible to gogofrom
fromy 6to
toyy only
in one
If .Dumber
possible,ofthe
probability
in each
possible in two
steps.
Hence
we
have
an
ergodic
chain
with
period
11s. Hence the transition matrix is symmetric. This2. tells u
It is clear also
that weFirst,
can go
y to isS ainreversible
one step if
and onlySecondly,
if we
the
thefrom
process
process.
things.
y in one step. If possible, the probability in each case is
go from 8 to probability
vector is a constant vector with components l / f S .
lis. Hence the transition matrix is symmetric. This tells us two
things. First, the process is a reversible process. Secondly, the fixed
probability vector is a constant vector with components 1/2'.
The transition matrix for the case 8 = 3 is
168
000 001 010 100 011 101 llO HI
000
0
1/3
lja 1/ 3 0
0
0
0
001
0
0
0
\/3
1! 3
j
0
0
0
0
0
1/3
0
l/3
0
100
1/3
1/3
l/s
0
0
0
0
lja 1/3
0
011
0
lis
1/ 3
0
0
0
0
101
0
1/3
0
1/ 3
0
0
0
010
:;:;
1..'
1/3 is
0
0
0 ','a1/
llOThe 0fixed0 vector
'18,
, 3 '/a, ' 1 8 , 118, ' 1 8 ) . Fro
,3 CL=(l/8,
we see,0for example,
that
the
1 j '3
1/number
0 of steps required toW
0
0 1/
3 mean
III
0
3
to any one state is 8.
From
this 6.2.3
l/ S). by
The fixed vector is a= (l/s, I/S, lis, lis, lis, lis, l/S, have,
Theorem
we see, for example, that the mean number of steps required
to return
the mean
number of ti
to anyone state
is 8.of We
each
thealso
states b
have, by Theorem
6.2.3, that
occurrences
of a par
the mean number
state of
is times
1. I ningeneral
between
each of the chain
stateswith
s states the
occurrences of
timea toparticular
return to a sta
for mean
a
state is 1. Inbegeneral,
2s and the
num
(011)
chain with s states
the
mean
times
in
state
6
between
(100)
time to returnrences
to a state
will y is 1.
of a state
be 2" and the mean
number
There
is a of
simple r
(OlO),I'-_ _...Y
times in state Swalk
between
occurintcrpretatioll
f
(110)
rences of a state
y is 1.we are consi
process
F ~ o u n7 -~6
There is a The
simple
random
vectors
y are the
walk interpretation for the
points of an s-dimensional cube. The possible states to wh
process we are considering.
FIGURE 7·5
process
can move from y in one step are the corner poixlts con
'Yare the
are vectors
s such points,
andcorner
the prohabi
to y by an edge. There The
points of an s-dimensional cube. The possible states to which the
process can move from 'Y in one step are the corner points connected
to 'Y by an edge. There are s such points, and the probability of
FINITE MARKOV CHAINS
168
CHA
these transitions occur each with probability 11s. I t is possible
0:1" other
MARKOV
J69
fromAPPLICATIOXS
any state y to any
stateCHAIXS
6 in a sequence of steps,
bu
possible to go from y to y only in an even number of steps; an
moving to anyone is lis. For the case s = 3, we have the cube in
possible in two steps. Hence we have an ergodic chain with per
Figure 7-5.
It is clear also that we can go from y to 6 in one step if and only
We define the distance between two states y and 0, denoted by
go from 6 to y in one step. If possible, the probability in each
d(y, 0), to be the minimum number of steps required to go from This
y to o.tells u
11s. Hence the transition matrix is symmetric.
In terms of the ooordinates of y and (3 this is
Secondly,
the
things. First, the process is a reversible process.
s
probability vector
vector with components l / f S .
diy,S)is=a constant
lXi-yd·
SEC. 3
2:
i=l
It is clear from the random walk interpretation that the mean time
to go from y to 0 depends only on the distance between y and o. Let
m S a be this mean time for two points a distance d apart. For fixed s,
we compute rnd = m'd as follows. Let y and S be two points a distance
d apart. This means that they have exactly d coordinat.es different.
On one step there are d choices which will make the process one unit
closer to 0, and s - d choices which will make it one unit farther from o.
Hence, considering the possible first steps, we have
d
s~d
rna. = l+-md-l+--md+l,
0 < d,::; s,
s
8
where we let mo = 1ns+l = O. These equations have a unique solution
which may be written as follows. Let
The fixed vector
s\ is CL=(l/8, '18, ','a, '/a, ' 1 8 , 118, ' 1 8 ) . Fro
i
(lcJ
we see, for example, that the 0,1,.,.,s-1,
mean number of steps required to
to any one state is 8. W
k~O
have, by Theorem 6.2.3
. i ,
the mean number of ti
then
It
each of the states b
0< d :s; s. occurrences of a par
ma. =
Q8 s_ i ,
i=l
state is 1. I n general
The values of QSi for valnes of 8 up to 6 are given in
Figure
7-6.s states the
chain
with
time
to
return
to a sta
VALUES FOR QS,
be 2s and the mean num
times in state 6 between
rences of a state y is 1.
There is a simple r
walk intcrpretatioll f
process we are consi
F ~ o u n7 -~6
The vectors y are the
2: (S-1\)'
L
points of an s-dimensional cube. The possible states to wh
process can move from y in one step are the corner poixlts con
to y by an edge. There are s such points, and the prohabi
FIGURE 7·6
170
FINITE MARKOV CHAINS
CH
For values of s up to 6 the values of mad are given in Figure
CHAP. VII
FINITE MARKOV CHAlKS
170
VALUES FOR msd
For values of 8 up to 6 the values of mSa are given in Figure 7-7.
VALUES FOR mSa
xL=-___--;-_~-_;__-
_2_:
3
3
7
I
4
9
I
1-
-;-1
5
6
I
31
63
.
37 1
12
74 2/ 5
40 1 /6
78 3 /5
41 213
804/ 5
4
2
15
182 /3
I
2Il13
10
201/3
I
5
8 ...011 --- - -
422/3
f
,:~
83 1/5
6
We note thatFIGURE
the means
for a given s increase as we incr
7-7
distance. This is to be expected. However, they increase ver
We means
shall call
above
chain the microscopic chain. I n
We note that the
for the
a given
8 increase as we increase the
applications
this
chain
is
often
as interesting
as one
distance. This is to be expected. However, theynot
increase
very slowly.
t by lumping.
macroscopic
chain In
is aphysical
Markov chain
We shall callfrom
the iabove
chain theThe
microscopic
chain.
from
the ismicroscopic
by lumping
all stales
having t
applications this
chain
often not chain
as interesting
as one
obtained
number of molecules in region A . Let us verify that the c
from it by lumping. The macroscopic chain is a Markov chain obtained
for lumpability is satisfied. Let V i be the set of all state
from the microscopic chain by lumping all states having the same
microscopic process with i molecules in region A. The se
number of molecules in region A. Let us verify that the condition
Vi any
be the
set of
in themores t
for lumpability is satisfied.
elements. Let
From
element
of all
Vr states
the process
microscopic process with i molecules in region A. The set Vi has
('?I
G) eleme:1ts. From
any element of Vi the process moves to one of
,
,
probabiliby of moving t o Vitl is ( a - i l l s and the probability o
1) elementstoofVi-l
Vi+l
one ofprobabilities
~ 1) elements
Vi-I· forThe
is or
i / s .to These
are thein same
all elemen
Hence
the
condition
for
lumpability
is
satisfied,
and
probability of moving to Vi+l is (s-il!s and the probability of movingwe obta
Markov chain with states V o ,V1, Vz, . . . , V,. The transitio
to Vi - I is i/s. These probabilities are the same for all elements of Vi.
bilitiesforare
Hence the condition
lumpability is satisfied, and we obtain a new
(i;
(i
Markov chain with states Vo, VI, V z, ... , V•. The transition probabilities are
pr,r-1 = ils
P',I+l = 1 - ijs p i ,
= 0
otherwise.
We shallPI,i-l
refer =to i/s
this new process as the macroscopic process.
= 0of §otherwise.
By thePI.1
results
6 . 3 we know that the lumped process w
be
ergodic
and
reversible. I t is again of period 2. The fixed
\Ve shall refer to this new process as the macroscopic process.
By the results of § 6,3 we know that the lumped process will again
be ergodic and reversible. It is again of period 2. The fixed vector a.
f
!
170
SEC. 3
FINITE MARKOV CHAINS
CH
For values of s up to 6 the values of mad are given in Figure
APPLICATIONS OF MARKOV CHAINS
171
VALUES FOR msd
for the lumped process is easily obtained from a. The component .21
is the sum of the components of a for states in Vi. Since there are
G) states in Vi we have
G)
aj
=--,
Cl
~{~n
2'
HelICe
By Theorem 6.2.3 we see that the mean number of times in any state
Vi, between occurrences of state Vj, is
8\
( means for a given s increase as we incr
We note that the
}l.be expected. However, they increase ver
This
is
to
distance.
'8)above chain the microscopic chain. I n
We shall call the
applications this chain is often not as interesting as one
In particular from
the mean
number of
III each
state
i t by lumping.
Thetimes
macroscopic
chain
is a between
Markov chain
from ~,.).
the microscopic chain by lumping all stales having t
occurrences of 0 is ( •
number of molecules in region A . Let us verify that the c
for lumpability
is satisfied.
Letlumping
V i be the
of all state
For our example
8 = 3 the partition
for the
is Vset
= ({OOO},
process
withThe
i molecules
region
{lOO, 010, OOI}, microscopic
{IIO, 101, Oll},
{ill}).
transition in
matrix
is A. The se
C;
('?I
elements. VoFrom
of Vr the process mores t
Vz element
Vs
VI any
Vo
(~3
I
0
~\
0
2/ 3
,
, VI
probabiliby
of
moving
t
is ( a - i l l s and the probability o
2! 3 o Vitl1~3)
Vz
to Vi-l is i / s . These probabilities
are the same for all elemen
0 for1 lumpability is satisfied, and we obta
Vs condition
Hence the
o ,V1, Vz, interested
. . . , V,. The
transitio
Markov chain
states
For the macroscopic
chain with
we are
also Vprimarily
in the
bilities
are We know that in general it is not easy to
mean first passage
times,
obtain the mean first passage times for the lumped chain from the
original ohain. However, in the case we are considering we are
pr,r-1 = ils
helped by two special features of the process. First, for the lumped
0 otherwise.
chain we can obtain all of the values pofi ,miJ=from
the knowledge of
only miD. In fact,
since
it
is
possible
to
go
from
VH1
Vo only byprocess.
We shall refer to this new process as
thetomacroscopic
going through Vi By
we have
the results of § 6 . 3 we know that the lumped process w
bemHl,i
ergodic
reversible. I t is again of period 2. The fixed
= and
ml+l,O-m"o
\:
°
and, by symmetry,
o ~ t < 8.
FINITE MA1ZKOV CHAINS
172
CHAP. Q
Thus we need only find the mean times to go from any state to V
FINITE
MARKOV
CHAINS
CHAP.
VIIonly on
In lumping, the
set which
became
state Vo was a set
with
element, namely the vector y = (0, 0.0 . . . . , 0). Hence the mean tim
Thus 'we need only find the mean times to go from any state to Vo.
to go from state V i in the lumped chain to Vo is the same as the mea
In lumping,
setfrorn
which
Voset
was
in the
Vi atoset
y = with
( 0 , 0,only
0 , . .one
. , 0). B u
time the
to go
anybecame
elementstate
element, any
namely
vectorisy =a (0,
0, 0, ...
,0). this
Hence
the mean
time
i from
vector.
Hence
this mea
suchthe
element
distance
to go from state Vi in the lumped chain to Vo is the same as the mean
time is msi. Using our results for the microscopic process we have th
time to go
from mil
anyfor
element
in the set Viprocess.
to Y"'" (0,0,0,
But
the macroscopic
These ...
are ,0).
best expressed
a
values
any such element is a distance i from this vector. Hence this mean
follows :
172
time is mSj. Using our results for the microscopic
process
we have the
6-2
s-z-1
values mij for
the =
macroscopic
process.
expressed
mi,~+l
ms-t.o-ms-~-l.o
= These
&Ss-r-are best
QSs-r
= &St as
k= 1
P=l
follows:
0 < i < s-l.
.s-i
8-i-1
Q"s-k Q8S_ k = QSi
mS-i,O - ms-i-l,O
2
2
2:
2:
k~
k~
To go from Vi t o V f we l must go through
every intermediate poin
1
Wence
O,::;i<s-l.
j- 1
2
for intermediate
i <j
To go from Vi to Vj we must mil
go through
point.
=
&8*every
L.=i
Hence
and
j-I
mij
2: Q'k for i <2j &Sk for i > j.
s-j-1
mil l:==i
= ms-t,s-j =
and
k ~ r - i
From the vector a,
mij
.-j-1
L Qs;: for2si > j.
m.-l,s-j =
I;:=,-i
From the vector a,
From these values we obtain
From these values we obtain
m~,i+l
+ m1'+l,i
(1 ) reachi
Let n(i+')trbe the mean number of times in state Vi before
state Viil, the process is started in state V d . Then by $ 6.2.7(b)
have
Let n({+l)jj be t.he mean number of times in state Vi before reaching
n(i+Uit
= mi,i+i
fmf+i.t
state Vl+ 1 , the process is started
in state
Vi. Then
by § 6.2.7 (b) we
?nt(
have
(' t l '
'if
n'
mi,itl +'lnt+l,i
= -----mjj
8
FINITE MA1ZKOV CHAINS
172
CHAP. Q
Thus we need only find the mean times to go from any state to V
In
lumping,
the set which
became state
Vo was a set with
only on
APPLICATIONS
OF MARKOV
CHAINS
SEC. 3
173mean tim
y = (0, 0.0 . . . . , 0). Hence the
element,
namely the vector
V i in From
the lumped
to Vo is the same as the mea
go from
Assumetonow
that sstate
is even.
(1) wechain
obtain
time to go frorn any element in the set Vi to y = ( 0 , 0, 0 , . . . , 0). B u
8/2-1
I
any such element is a distance i from this vector. Hence this mea
rnO,./2 + 'ins/z,o
2' for the microscopic process we have th
time is msi. Using our results
t=O
.
t
These are best expressed a
values mil for the macroscopic process.
:
Also, sincefollows
ms/2,O=ms/2,s,
we have
L -(S-l)'
s-z-1
6-2
2 &Ss-r-1 2 QSs-r = &St(2)
+ m sl2,& =
0 < i < s-l.
2s 6 e~ 1)'
mi,~+l= ms-t.o-ms-~-l.o =
mo,. = mO,BIZ
,/2-1
k= 1
P=l
To go from Vi t o V f we must go through every intermediate poin
From the
point of view of physical i:lterpretations the most interestWence
j- 1 of molecules.
ing case is where there is a large number
For this we
i <first
j
mils /2=and m&8*
would like to find estimates for mO.
represents
s f2. o. forThe
L.=i
the time taken to go from no molecules in
the given region to an equal
number inand
each. The second is the time taken
s-j-1 to go from an equal
num ber in each to 0 in region
the model
milA.= If
ms-t,s-j
= is to
&Skhaveforany
i >similarity
j.
r - i the time to equalize
to actual physicEd situations we would expectk ~that
From
the be
vector
the molecules
would
mucha,shorter than the time to Ieach an extreme
2s
situation from equalization.
'We consider first mo,s/Z. \Ve estimate the time required to go from
V, to V i ' l . 'We know that the mean number of times in Vi before
reaching Vi+1From
is s/(s-i)
1 +i/(sEach time it is in V, (except for
these=values
we i).
obtain
the last time) at least two steps are taken in going to V;-1 and back to
Vi, Hence the mean time to go from Vi to VI -d is at least
2
2
i
s+i
1+2--. = --.'
s-~
S-I·
Hence we may find a lower bound for mO,s/2 by finding
Let n(i+')trbe the mean number of times in state Vi before reachi
state Viil, the process is started in state V d . Then by $ 6.2.7(b)
have
n(i+Uit= mi,i+ifmf+i.t
This sum is greater than:
--:2+
jo
?nt(
'S/2 S+X
~-dx =
s(21og::- 11z)-2.
8-X
Hence 8(2 log 2 - 1 i2) - 2 is a lower bound for mO. s/2. A better lower
bound is obtained as follows. 'We assumed above that each time the
process went from Vi -- 1 to Vi, it did so in OIle step. If instead we use
our lower bound for the mean time to go from Vi - 1 to Vi we have a
new lower bound for the time to go from Vi to Vi+l' This is
i-i)
S -1-1 +i - ( 1 +
---
s-i
s-i+ 1
=
1+
. 28i .
(8-t)(s-~+1)
.
FINITE U R K O V CHAINS
174
The sum of these values between Vo and
v81s1is greater t
FINITE MARKOV CHAINS
174
CHA
CHAP. VII
The sum of these values between Vo and V,12-1 is greater than
J
r-12 (
-4+ o
1+
2SX) dx
(8-X)(S-x+ 1)
n
Using the first two t e r n s in t h e Taylor series for log (s+ 2)/(s i
= - 4+8[1/2- 2 log 2 +S(8+ 1) log::
obtain the estimate
- 5 + sp/z - 2 log 2)
Using the first two terms in the Taylor series for log (8+ 2)/(s+ 1) we
obtain the estimate
for our improved lower bound.
We have
-5+8(5/
2) for mo,,,z.
We next obtain
an upper
bound
2-2Iog
for our improved lower bound.
We next obtain an upper bound for mO,8/2.
1
./2-1
mO,./2 =
2:
1-0
('S -
l
I
We have
8
1) 2: (k)
k-O
l
l(l+l)
(S)[1
l
+ 8-l+ 1 + (s-l+ 1)(s-l+ Z) +
l1
]
+ (s-l+ 1) .. .s
Hence,
Hence,
./2-1
mO,s/2 ~
(~)
8-l
s- 2l
s/2-1
2 8~ZI
6 e~l)
i --z- +
('-ll/2
1=0
S sp/4
=
112 log 8).
dx+ 1/ 4S
This approximation can be improved by using a better estim
S(1/4 of
+ Ihlog
8).where 1 is large. Thus we have found that
the=terms
the sum
~
o
s-
X
+ 112 logfors].
This approximation can
improved
a better
estimate
- 5be+ s(5I2
- 2 logby
2 ) using
< rn,,,,,
<
the terms of the sum where l is large. Thus we have found that
It would appear that these values are asymptoticaily of a
-5+s(51z-2Iog2)
mO. Ba
I2 constant
< S[I/4+1/21ogs].
order of magnitude <than
times s, but less than a c
times s log s. We see from these estimates that the process
It would appear that these values are asymptotically
of a greater
very short time to go from 0 to 4 2 . For s = 100 the minimum
order of magnitude than a constant times s, but less than a con:;tant
is 50 these
and the
actual mean
timeprocess
required
is about
1
times slog s.could
\Ve take
see from
estimates
that the
takes
a
very short time to go from 0 to s/Z. For 8 = 100 the minimum time it
could take is 50 and the actual mean time required is about 140.
FINITE U R K O V CHAINS
174
SEC. 3
The sum of these values between Vo and
v81s1is greater t
175
APPLICATIONS OF MARKOV CHAINS
Consider next mG...
CHA
We know from (2) that
s12-1
mo,. = 2 8
2: (8-1\'
,~O
But
•
,I
Using the first two t e r n s in t h~ e JTaylor series for log (s+ 2)/(s i
obtain the estimate
- 5 + sp/z - 2 log 2)
s12-1
1
I
L
1
1
8/2-1
1+-:;;;
+--+--,'
for
our
improved
lower
8-1
('8-1\bound.
:;;;
8-1for e~l)
We have
i=O
mo,,,z.
We next obtain
ani upper bound
j
Hence
Thus
mo,.
~ 28( 1 +~)
where 1 ~ A :;;; 2.
Our estimates for mO,./2 show that this quantity is of a lower order of
magnitude than mo.s. Hence the above estimates serve also for
m./2,g=7n s;2.0 and we have
-
msl20 , .
Hence,
2'{1+::)
\
8
where 1 :;;; A
~ 2.
Thus, as predicted, the mean time to go from equalization to the
extreme of 0 molecules is very much larger than to go from 0 to
equalization. For s = 100 the nrst time is approximately 2 100 or about
1000 billion billion billion ',vhile the second is only 140.
Other interesting quantities are mo,O and ';'812,8/2. For these we
have
mo,O = 2 8
+
J;;;-'2
= sp/4 8112 log 8).
2
m s/2,s/2 =
8
.
8(
\ be
~ improved
This approximation-can
by using a better estim
s/2) 1 is large.
the terms of the sum where
+
Thus we have found that
+
For s = 100, mSO,50 is approximately
12.5.
112 log s].
- 5 s(5I2- 2 log
2 ) < rn,,,,, <
We conclude this section with some remarks about the macroscopic
It would appear
that these
values
are2~ asymptoticaily
chain and reversibility.
It is sometimes
argued
that
process of the of a
s, but
than a c
order
of
magnitude
than
a
constant
times
type we have considered has a "direction" because of the
veryless
great
times toward
s log s. equalization.
We see from It
these
estimates
the process
tendency to move
is true
that if that
the process
100certainly
the minimum
short time to go from
0 to 4O-then
2 . Forits =will
is started outvery
of equilibrium-say,
in state
could
take
is
50
and
the
actual
mean
time
required
is about 1
move towards the center. However, if the process is observed after
it has been going for a long time, then we may consider it as started in
equilibrium and in this case the process will appear the same looked
FINITE MARKOV CHAINS
176
CH
a t in the reverse direction as in the forward. For example
CHAn,s
CHAP.
175
V, the probability
thatVII
i t moves
processFINITE
is then MARKOV
observed in
V,-l is the same as the probability that i t came from V,-l. A
at in the reverse direction as in the forward. For example, if the
way to put this is the following. If a sequence of outcomes for
process is then observed in Vx the probability that it moves to state
number
steps is recorded
and then
o a physicist, h
Vx- 1 is the same as the of
probability
that it, came
fromhanded
V",-l. tAnotber
unable
to tell whether
he wasofgiven
them for
in the
order of inc
way to put this be
is the
following.
If a sequence
outcomes
a large
or in theand
order
of handed
decreasing
number of stepstime
is recorded
then
tD atime.
physicist, he would
be unable to tell whether he was given them in the order of increasing
3 7.4 Applications fo genetics. A problem that is fre
time or in the order of decreasing time.
discussed in genetics is the following: Two animals are mated.
their offspring,
two areAselected
by some
and these are
§ 7.4 Applications
to genetics.
problem
that method,
is frequently
Then the procedure is repeated. This type of problem may be
discussed in genetics is the following: Two animajs are mated. From
as a Markov chain. As stat,es we take the possible combina
their offspring, two are selected by some method, and these are mated.
parents,
and the transition
are determined
Then the procedure
is repeated.
This typeprobabilities
of problem may
be treated by the
genetics,
from
the
assu~nptions
concerning
the
way parents
are s
as a Markov chain. As states we take the possible combinations
of
The
simplest
such
problem
is
obt,ained
if
we
classify
parei
parents, and the transition probabilities are determined by the laws of
according to the pair of genes they carry in are
oneselected.
position in the
genetics, from the assumptions concerning the way parents
We will discuss this case. Here we may forther simp
somes.
The simplest such problem is obtained if we classify parents only
by assuming
thatinthe
is either
a or
I.: and hence t
according to theproblem
pair of genes
they carry
onegene
position
in the
chromoindividual animal must be of type aa or ab or bb. For exan
somes. We will discuss this case. Here we may further simplify the
dominates b, then aa is a pure dominant, a6 is a hybrid, and bb
problem by assuming that the gene is either a or b, and hence that any
recessive animal. Then a pair of parents must be of one of the f
individual animal must be of type aa or ab or bb. For example, if a
six types: (aa, a n ) , (bb, bb), (nu,,a b ) , (bb, a,b), (aa, bb), (ub, ah
dominates b, then aa is a pure dominant, ab is a hybrid, and bb is a pure
simple must
case be
that
theofnew
pments are sel
recessive animal.problem,
Then a for
pairthe
of parents
of one
the following
random
from
the
offspring,
is
treated
in
Feller
and in FM.
six types: (aa, aa), (bb, bb), (aa, ab), (bb, ab), (aa, bb), (ab, ab). This
We will discuss a class of problems of which the above prob
problem, for the simple case that the new pClrents are selected at
special case. We will assume that one offspring is selected a t
random from the offspring, is treated in Feller and in FM.
and that this offspring selects a mate. I n its selection it is k
We will discuss a class of problems of which the above problem is a
likely
pick that
a given
animal unlike
itself atthan
a given ani
one offspring
is selected
random,
special case. We
will to
assume
itself. Thus k measures how strongly "opposites attract each
and that this offspring selects a mate. In its selection it is k times as
I n this we take into account that in a simple dominance s
likely to pick a given animal unlike itself than a given animal like
a a and a21how
typestrongly
animals"opposites
are alike as
far aseach
appearances
attract
other." are co
itself. Thus k measures
The
resulting
transition
matrix
is
In this we take into account that in a. simple dominance situation,
aa and ab type animals are alike as far as appearances are concerned.
The resulting transition matrix is
0
(lUI, aa)
(bb, bb)
(aa, alJ)
P = (bb, ab)
0
I
I- -0 - - - -0- -
0
1/4
0
0
0
!
0
1
2(k+ 1)
0
(aa, bb)
1
.
(ab, ab)
4(3k +
1)1
0
0
0
0
1/2
0
0
1/4
0
k
(k+ 1)
0
1
~(k-=t=Ti
0
0
0
1
2k(k+ J)
k(",+ 1)
k + 3 (k+3)(3k+ J)(k+ 3)(3k+l)
k+3
FINITE MARKOV CHAINS
176
CH
a t in the reverse direction as in the forward. For example
APPLICATIONS
OF MARKOV
probability that i t 177
moves
process
is then observed
in V, the CHAINS
SEC. 4
- -from V,-l.
V,-l is the same as the probability that it came
A
The first way
two to
states
are isabsorbing.
TheyIf correspond
to outcomes
having for
put this
the following.
a sequence of
developed a pure strain: pure dominant in the first state, and pure
number of steps is recorded and then handed t o a physicist, h
recessive in the second. \Ve will compute the fundamental matrix N,
be unable to tell whether he was given them in the order of inc
the vector T, and the matrix B.
time or in the order of decreasing time.
J\' =
_:---:-l-.---cc
3
Applications fo genetics. A problem that is fre
Two
animals are mated.
discussed
in genetics is the
following:(3k+
2k(k+ 1)2
lc(k+ 1)
1)(k+3) \ (00, ab)
'4(k 2 +5lc+2)
their offspring,
two are selected by some method, and these are
2
k(k+ 1)
2(3k+ 1)
l)
(3k+ l)(k+ 3) \ (bb, ab)
Then (4k
the +9lc+3)(k+
procedure is
repeated. This
type of problem may be
x (
4(3lc+1)
4lc(k+ 1)2
(4lc 2 + 9k + 3) 2(3k+ l)(k+ 3)
(00.combina
ab)
as a Markov chain. As stat,es we take the possible
4k(k + 1)2
2k(k+l)
2(3k+l)(k+3)
(ab.ab)
4(3k+parents,
1)
and
the transition
probabilities
are determined
by the
genetics, from the assu~nptionsconcerning (a
the
way
parents
are s
a, ab)
The simplest such problem is obt,ained if we classify parei
according to the pair of genes they carry in(b/',
onerib)position in the
T =
we may
(2k+l)(k+3)
We will discuss this case. Here (aa,
somes.
bb) forther simp
problem by assuming that the gene is either
a
or I.: and hence t
(ab, ab)
individual animal must be of type aa or ab or bb. For exan
dominates b, then
is a pure dominant,
(bb, Vb) a6 is a hybrid, and bb
(aa,aa
aa)
recessive animal.2 Then a pair of parents must be of one of the f
aa)
14ka n+
[) 4k
2 + 5k+
six types: (aa,
) , :23k
(bb,+bb),
(nu,,
a b ) , (bb, a,b),(aa,
(aa,
bb), (ub, ah
2
9k+3
Bk
+19k+9
(Ob,
ab)
problem,
for
the
simple
case
that
the
new
pments
are sel
(
B =
]
random
in 6Feller(aa,
andbb)
in FM.
4(:2k+
l)(k+ from
3) _ the offspring,
2 + 10k+
18k + 6 isSktreated
We will discuss a class of problems
of which the above prob
18k + 6 8k 2 + 10k+ 6/ (ab, aa)
special case. \ We will
assume that one offspring is selected a t
and that
thiswill
offspring
selects
mate.
its selection
it is k
Since we know
that we
eventually
enda up
with aI npure
strain, the
likely
to
pick
a
given
animal
unlike
itself
than
a
given
two most interesting questions concern the number of generations ani
k measures
how-probability
strongly "opposites
itself.a pure
Thus strain
needed to reach
and the
of gettingattract
pme each
thisrecessives.
we take into
account that
simple
dominance
s
dominants orI n
pure
In particular,
we in
willa be
illterested
in
andon
a21these
type quantities.
animals are alike as far as appearances are co
the effect thata at· has
The
is
A large k has
theresulting
effect of transition
producing matrix
more mixed
matings. Hence we
would expect a large J: to slow down the process. Indeed, every entry
(2k+l)(k+3)7.4
!
3)
in T is monotone illcreasing in t.
Some typical values of this vector are
(!:;:) c;:) c::)
2
13 / 3 \
3 1/3
k=O
6 2/3
8.28
52/"3
k=l
,7.:28
/-_.)
/10.11\
1611)
~d~::l
1- 91
.
1:=7
We see that increasing k will slow down the process considerably,
especially if we start in one of the la,st three states. The fact that
the time to absorption from (00, bb) is always one more than from
FINITE MARKOV CHAINS
178
C~
(ab, ab) is due to the fact that from the former state we always go
FINITE
MARKOV
CHAINS and a pure
CHAP.
VII paren
latter in
one step
(a pure dominant
recessive
----------------have hybrid offspring).
(ab, ab) is due to the fact that from the former state we always go to the
The effect of k on the probability of absorption is not so
latter in one step
pure guess
dominant
and ak pure
recessive
parent
mustsince the
One(awould
that large
favors
the recessive
strain,
have hybrid offspring).
dominants and hybrids would tend to select recessive mates. I
The effect offork kon
not(bb,
so bb)
clear.
> 1the
theprobability
probabilityofofabsorption
absorptionis in
increases w
One would guess
that
large
k
favors
the
recessive
strain,
since
then
But for k < 1 a surprising situation develops. As kboth
is decreased
dominants and 1,
hybrids
would tend
select recessive
Indeed,
the probability
of to
absorption
in ( a a ,mates.
a a ) increases,
until it rea
for k> 1 the probability
of
absorption
in
(bb,
bb)
increases
k.ame valu
the smaximum a t k = 1 / 3 , and then decreases back towith
But for k < 1 a ksurprising
situation
develops.
As
k
is
decreased
from
= 1. Thus k = 0 and k = 1 yield the same absorption probab
1, the probability of absorption in (aa, aa) increases, until it reaches a
This means that, if in nmating like always selects like, the prob
maximum at k = lja, and then decreases back to the same vaiue as at
of absorption is the same as for random mating (though of cou
k = 1. Thus k =
0 and
k = 1 yieldisthe
same
probabilities.
time
t o absorption
much
less absorption
than for random
mating). Some
This means that if in mating like always selects like, the probability
values for the probability of absorption in (aa, a a ) are
of absorption is the same as for random mating (though of course the
time to absorption is much less than for random mating). Some typical
values for the probability of absorption in (aa, aa) are
178
('75) (.77\ (.75) (.71\ (.61)
.25
.27.
.25
.21 \
.11
.50
.54 )
.50
.42'
.22'
.50
.54,
.50
\.42)
.22
The fact that the last two entries are always the same is due
k=7 always goes directl
already observedk=l
fact that the process
(aa, bb) to (ab, ab). The probability from (bb, ab) is half this
The fact that the
lastwhen
two entries
are always
is is
dueequally
to the likely t
since
the process
leaves the
(66,same
a b ) , it
already observed
that
theabsorbed
process in
always
goes Similarly,
directly from
(ab,fact
ab) or
to be
(bb, bb).
the value a t
(aa, bb) to (ab, ab).
fromplus
(bb, ab)
this much:to note
is halfThe
theprobability
(ab, ab) value,
I / > isIthalf
is interesting
since when the random
process mating
leaves the
(bb,probability
ab), it is equally likely to go to
of absorption in (aa, a a ) is prop
(ab, ab)or to betoabsorbed
in (bb,
Similarly,
the number
of abbl.
genes
present athe
t thevalue
start.at (aa, ab)
is half the (ab, ab)Let
value,
pIllS
1/
.
It
is
interesting
to
note that
for random
2
us collect the quantities for tthe special
case of
random mating the
probability
of
absorption
in
(aa, aa) is proportional
k = 1.
to the number of a genes present at the start.
k :~~ us collect the qnant,ities for the special case of random mating,
,
0
(arL, aa)
0110000
(bb bb)
0
0
0
0
------.--------p=
1/ 4
0
0
1/4 I
0
0
I
1/ 16 1/16
t(2
0
0
I( 4
(au,ob)
0
1/ 2
0
I! 4
(UJ, ab)
0
0
0
I
(aa, bb)
1/4 1/4 t! 8
(ab,ab)
FINITE MARKOV CHAINS
178
SEC. 4
C~
(ab, ab) is due to the fact that from the former state we always go
latter
in one step (a
dominant
and a pure recessive
APPLICATIONS
OFpure
M.ARKOV
CHAINS
179 paren
---have
hybrid
offspring).
4/3\
The effect of k on2/ 3the1/ G
probability
of absorption is not so
One would guess that8/large
k favors
the
recessive
strain, since the
3
J
/a
4/3 \
N =and hybrids
2/3
dominants
would tend to select recessive mates. I
41/3
4/ 3 4/of3 absorption
Bia )
in (bb, bb) increases w
for k > 1 the probability
But for k < 1 a surprising
situation
4/ 3 4/3 lla
8/3 develops. As k is decreased
1, the probability of absorption in ( a a , a a ) increases, until it rea
maximum a t k = 1 / 3 , and then decreases back to the s-ame valu
k = 1. Thus5k = 0 and k = 11/4
yield3 j 4the same absorption probab
. 4 /6 \
ThisT=
means that, if inB=
nmating like always selects like, the prob
( 1'2
1/2
6 2/3is the same as
random
mating (though of cou
of absorption
\ for
I
time t o absorption
is
much
less
than
2/3)
\llt) l/ Zfor random mating). Some
I"
values for the probability of absorption in (aa, a a ) are
C
/45(6\
I I
')
(/4
I
\5
'We also compute
(5/8 5/8 1/1/8
s
1,
H=
/4
\ l/ Z 1/2 1(.1
:~:) ~,~
liz 1/ 2 1/4 5/ S
!
",.c:)
.
816
816
The standard deviation
in T isthe
around
4.7 entries
for every
is of is due
The fact that
last two
areentry.
alwaysThis
the same
the same orderalready
of magnitude
as the
entries
T, hence
may expect
observed
fact
that of
the
processwealways
goes directl
very large fluctuations.
(aa, bb) to (ab, ab). The probability from (bb, ab) is half this
\Ve observe from
that the
lumpa-likely t
sincePwhen
theprO('cBs
processsatisfies
leaves the
(66,condition
a b ) , it isfor
equally
(ab, ab)
totwo
be states
absorbed
(bb, bb).
the value a t
bility if we combine
the or
first
and in
combine
the Similarly,
next two states.
the (ab,
ab) value, plusThe
I / > state
It is(bb,
interesting
to note
This partition is
hashalf
a simplc
int.erpretation.
bb) results
from (aa, aa) random
by interchanging
and h, and
(bb, ab) results
from
mating theaprobability
of absorption
in (aa,
a a ) is prop
(aa, ab) similarly.
the other
hand present
(aa, bb) aa.nd
ab) are unto theOn
number
of a genes
t the (ab,
start.
changed. Hence Let
the us
partition
the process
we do case
not care
collect represents
the quantities
for ttheifspecial
of random
which gene is the
gene. The first state of the new process,
k = dominant
1.
which we may denote by (aa, aa), represents any pair of like pure
parents. The second state, to be denoted by (aa, ab), represents one
pure and one hybrid parent. The remaining two states represent the
combination of unlike pure parents, and of two hybrid parents, respectively_ The lumped transition matrix is
·c
p=
0
0
1/ 2 0
0
0
0
lis
1/2
1/8
(aa, aa)
':')
1/ 4
(aa, ab)
(aa, bb)
(ab, ab)
To obtain N for the lumped process. \ve add the first two columns,
F I K I T E MARKOV C H A I N S
180
CH
which correspond to lumped transient states:
FINITE MARKOV CHAINS
180
CHAP. VII
which correspond to lumped transient states:
1/ 6
C'
1°/3
1/ 6
''/,I,)
.
8/ 3 4/ 3 8'i 3
We then observe that the first two rows are identical, as mus
8/ 3
we have
8/ 3 1/3 Hence
case for lumpability.
We then observe that the first two rows are identical, as must be the
case for lumpability. Hence we have
(aa, ab)
and
(aa, bb)
(ab, ab)
and
The vector .F is also obtainable directly from T. The latter has
entries for the two states t o be lumped, as must be the case.
traction yields +. For many purposes the lumped proces
The vector'; is also
obtainable
directly from
The latter
hasnot
identical
sufficient
information.
(OfT.course,
i t does
yield any in
entries for the two
states
to
be
lumped,
as
must
be
the
case,
and
con-concern
information as far as the absorption probabilities are
traction yields there
T. For
manyonepurposes
thestate
lumped
yields
is only
absorbing
after process
lumping.)
For exam
(Ofdetermined
course, it does
sufficient information.
completely
by $. not yield any interesting
information as far as the absorption probabilities are concerned, since
there is only one absorbing state after lumping.) For example, T is
completely determined by 7.
The last two columns of &, corresponding to single-state
obtainable directly from H . B u t the first column is new. G z i
obtainable from 7 2 .
The last two columns
of fl, corresponding
to single-state
cells, areili Kemp
A generalization
of this lumped
process is discussed
obtainable directly
Butourselves
the first column
is new.
72 isin
directly
Wefrom
still H.
restrict
to a single
position
the chromoso
obtainable from we
T2.
no longer assume that only two types of genes can occu
A generalization
of this lumped
proceBs
discussed
There may
be is
any
numberillofKempthorne.t
different kinds of gen
position.
\Ve still restrict consider
ourselvesthe
to process
a single in
position
in theform:
chromosomes,
but the n
the lumped
We care about
we no longer assume
that
only
two
types
of
genes
can
occur
in
this
different genes present and about their combination, but w
position. Theredistinguish
may be any
number
different
kinds
of genes.
\Vepermute
states
thatof
differ
only in
having
the genes
consider the process in the lumped form: We care about the num ber of
t 0. Kernpthorne,
A n Int~odxcfion
lo QenelicStntistics, New Y o r k , John W
different genes present
and about
their Gombination,
but 'we do not
Inc., 1050.
clistinguish states that differ only in having the genes permuted. Thus
to. Kempthorne, An Introduction to GeneticStMistics. New York. John \Vilt,y & Son.
Inc., 1950.
CH
F I K I T E MARKOV C H A I N S
180
which correspond to lumped transient states:
181
APPLICATIONS OF MARKOV CHAINS
SEC. 4
we have the following seven states:
81:
(aa,aa)
(00" ab)
Sa: (aa, bb)
(a.b,the
ab)first two rows are identical, as mus
S4: that
We then observe
case for lumpability.
(00"Hence
be) we have
S5:
S2:
S6:
87:
(ab, ac)
(ab, cd)
The transition matrix is:
rI
and
0
0
-I
1/ 4 I liz 0
~/8
0
0
0
1/4
I
,_0
I
_ 0 I_~
81
000
S2
o
8S
()
I
0
I 0
0
0
84
I l/ z lis 1/ 4
The vector .F is also obtainable directly from T. The latter has
() states
entries
t o be
o for
I the
0 two
1/2 as0must be
0 lumped,
85 the case.
"/ 2
traction yields +. For many purposes the lumped proces
l/ S 3/8 i t does0 not yield
11", 0 3/ 16 (Of course,
. 1/16 I information.
86
sufficient
any in
information as far as the absorption probabilities are concern
0 one
0 absorbing
0
1/2
1/4
5;
1/4
o is\ only
there
state
after lumping.)
For exam
We have indicated
the equivalence
classes
in the transition matrix.
completely
determined by
$.
The equivalence classes are determined by the number of different genes
present in the parents. Clearly, this number either stays the same or
it decreases, hence the process can move from an equivalence class with
a given number of genes only to one with fewer genes, i.e. from the
boUom up in P. The single state with only one type of gene present is
The four-state
last two columns
&, corresponding
to single-state
absorbing. The
lumped of
process
we considered
I1bove
B u t the
new. G z i
directly
from H . classes
the two top
equivalence
in first
the column
present ischain.
corresponds toobtainable
Since numbersobtainable
concerningfrom
these7 2classes
are not affected by equivalence
.
A generalization
is discussed
wethis
are lumped
about toprocess
compute
will agreeiliinKemp
classes lower down,
all quantities of
We
still restrict
ourselves
to afound.
single position in the chromoso
their upper left
corner
with those
previously
we no longer assume that only two types of genes can occu
position. There may be any number of different kinds of gen
consider the process in the lumped form: We care about the n
different genes present and about their combination, but w
distinguish states that differ only in having the genes permute
p=
\
t 0. Kernpthorne, A n Int~odxcfionlo QenelicStntistics, New Y o r k , John W
Inc., 1050.
FINITE MARKOV CECAINS
Y 82
FINITE MARKOV CHAINS
182
45/6\ 82
7"=
H
CHAP
CHAP. VII
211 3 /36 \
22 2/ 3
,
::;)::
22'/, )
7"2
7 1/ 6
55
23 11 /12
62/ 3
SF,
242/3
7 2/ 3
57
24 2/3
S4
82
S3
7/ 10
1/ 8 1/2
8/10
1/4
8/10
I! 4 5/ 8
I
8/10 5h4 5/i 6
8/ 10 1/6 2/ 3
55
86
87
0
0
0
0
0
0
0
0
0
I! 2
1/5 7/16
l! 10
0
0
We note t h a t sz, s4, and s6 are the "likely states," in the sense t
7/ 36low7/enough
2/ 15 for
2)3it to be possible to reach the
9
the process8/ 10
starts
there is always a fairly good chance of reaching the state. The
"\Ve note that
52, 54,
86 are
the "likely
states,"
in the
the process
sense that
if
states
areand
quite
unlikely
no matter
how
is started.
the process starts
low enough
forstarting
it to beinpossible
the with
statefour difl
quite surprising
that
s7-that tois,rCiLch
starting
there is always
a fairly expect
good chance
of reaching
the state.
other
genes-we
to reach
a pure strain
in 7 2 jThe
3 generations.
B
states are quite
no matterthat
howthe
thestandard
process is
started. of Itthis
is quant
mustunlikely
be remembered
deviation
quite surprising
in s7-that
if!. startingare
with
4.97,t.hat
andstarting
hence much
longer processes
notfour
too different
unlikely.
I/J
genes-we expect to reach a pllTe strain in 72/3 generations. But it
sectionof
will
be quantity
devoted to
5 7.5 Learning
must be remembered
that the theory.
standardThis
deviation
this
is the stu
a mathematical
for certain
of learning, due to W. K.
4.97, and hence
much longer model
processes
are not kinds
too unlikely.
We will discuss only some relatively simple special cases, bu
§ 7.5 Learning
theory.
will be devoted
the study
of
techniques
hereThis
usedsection
are applicable
to more to
general
situations.
a ma"LhematicalI nmodel
for
cerLain
kinds
of
learning,
due
to
'.V.
K.
Estes.
a typical experiment the subject is placed in front of a p
lYe will discuss
relatively
simple
special
but
lighh,only
and some
he is asked
to guess
whether
the cases,
light on
his the
left or the
techniques here
used
arewill
applicable
to on
more
general
situations.
on his
right
be turned
next.
Thus,
he has two possible resp
In a typical
thethe
subject
placed
front
of aguess
pair "right."
of
Weexperiment
denote by A.
guess is"left,"
andin by
A1 the
lights, and hethe
is asked
to guess turns
whether
011 his
left orLet
the Eo
light
experimenter
on the
onelight
of the
lights.
mean tu
on his right will
on next.
Thus,
haslight.
two possible
responses. is repe
on be
t,heturned
left light,
and El
the he
right
The procedure
We denote by
Ao the
guess of"left,"
and
by aAlrecord
the guess
"right."
large
number
times,
and
is kept
of the Then
sequence of
the experiment.er
of t.heoflights.
Let isEo
turning
At aridturns
E i . on
Theone
purpose
the theory
to mean
predict,
for given be
on the left light,
EJ the righthow
light.the The
procedure
is repeated
a in th
of theand
experimenter,
subject's
guesses
will change
large numberrun.
of times,
and
a
record
is
kept,
of
the
sequence
of
both
A variety of experiments has shown that the model is in
Ai and Ei. The
purposewith
of the
agreement
thetheory
facts. is to predict, for given behaVior
of the experimenter, how the subject's guesses will change in the long
run. A variety of experiments has shown that the model is in good
agreement with the facts.
FINITE MARKOV CECAINS
Y 82
SEC. 5
APPLICATIONS OF MARKOV CHAINS
----
CHAP
183
In a large class of inteTE'sting experiments the experimenter will
choose his actions with fixed probabilities, depending only on the
action of the subject. These probabilities may be given in Figure
7-8.
Eo
AO.(I-V
Ai
W
v )
l-w
FIGURE 7-8
That is, a guess "left" (Ao) is reinforced. by turning on the left light
(Eo) with probability 1- v; otherwise the right light is turned on.
And a guess "right" is reinforced with probability 1 - w. Here v and
ware numbers betv;een 0 and I, which are kept fixed for t.he duration
of the experiment.
'While " and ware usually positive numbers, the cases in whic:, they
We note t h a t sz, s4, and s6 are the "likely states," in the sense t
are not lead to interesting experimeEts. For example, if v = 0 and
the process starts low enough for it to be possible to reach the
w> 0, then Ao
is always
reinforced,
but Al
is only
occasionally
rein- The
there
is always
a fairly good
chance
of reaching
the state.
case are
v>Oquite
and w=O
is similar.
If how
v=w=O,
then every
forced. Thestates
unlikely
no matter
the process
is started.
action of the subject is reinforced.
quite surprising that starting in s7-that is, starting with four difl
Another class
of interesti.ng
special
is where
= 1. Here
genes-we
expect to
reachcases
a pure
strain vin+ w
7 2 j 3 generations. B
1-1' = wand v"" 1- w, and hence the probability of Eo or of El is
must be remembered that the standard deviation of this quant
independent of the action of the subject.
4.97, and hence much longer processes are not too unlikely.
The model assumes that the subject has a cert<1in unknown number 8
of stimulus elements.
Each stimulus
is at each
of the
This section
will bestage
devoted
to the stu
5 7.5 Learning
theory. element
experiment connected
to one
of the
two possible
responses
a mathematical
model
for certain
kinds of
learning,AI,
duethe
to W. K.
original connections
beingonly
known.
is then assumed
that the
We will not
discuss
some Itrelatively
simple special
cases, bu
following takes
place at here
each used
stageare
of the
experiment:
techniques
applicable
to more general situations.
I n a typical experiment the subject is placed in front of a p
(1) The subject "samples" a subset of the st.imulus elements, by
lighh, and he is asked to guess whether the light on his left or the
means of anonindependent
process,
in which
anyone
stimulus
his right willtrials
be turned
on next.
Thus,
he has two
possible resp
element i.s sampled
with
probability
t
or
not
sampled
with
probability
We denote by A. the guess "left," and by A1 the guess "right."
I-t.
the experimenter turns on one of the lights.
Let Eo mean tu
(2) If in the
there
stimulus
elemel1ts
connected
to is repe
onsampled
t,he left set
light,
andarc
Elkthe
right light.
The
procedure
Ao and l to AI, then the subject performs Al with probability l/(k+l).
large number of times, and a record is kept of the sequence of
Some convention is necessary if no stimulus element is sampled; we
At arid E i . The purpose of the theory is to predict, for given be
will assume that the probability of Al is the same in this case as if all
of the experimenter, how the subject's guesses will change in th
stimulus elements had been sampled.
run. A variety of experiments has shown that the model is in
(3) If the experimenter performs Eo, then any stimulus element t.hat
agreement with the facts.
was previously connected to Al and which was just sampled by the
subject is reconnected to Ao. Similarly, after E l , ;;.11 the sampled
stimulus clements arc connected t.o AI·
CRAP
FINITE MARKOV CHAIPiS
184
We can represent the model as an (s + 1)state Markov chain, in w
FINITE
CHAP.
VII
state s+ occurs
whenMARKOV
exactly iCHAINS
stimulus elements are
connected
to
and i = O , 1, . . . , s. All the interesting quantities depend only
We can represent the model as an (8 + 1) state Markov chain, in which
how when
manyexactly
stimulus
elements elements
are connected
each way,toand
state Sj occurs
i stimulus
are connected
AI,hence t
quantities are functions on the chain. For example, the probabili
and i = 0, 1, ... , 8. All the interesting quantities depend only on
from state
is obtained
as follows:
action A1elements
how many stimulus
arestconnected
each
way, and hence these
184
-.
quantities are functions on the chain. For example, the probability
of
I
i
( S i i ) ( l ) l m ( l -1)- + (1- t ) ~
action Al from state
SI is obtained
as follows:
Prc[Al]
=
m
S
k=O
i
1=0
k,t not both 0
~iis the number
(S-i)(i)tm(l_t}8-m"£
+ (l-t}sL
where kk=O
elements
1=0
k of stimulus
l
m sampled8 from those
k,l not
both 0 those connected to A1, m = k + 1, and the last t
1 from
nected to Ao,
arises
from of
thestimulus
assumption
concerning
case
where
where k is the
number
elements
sampledthe
from
those
con-no stim
element
is
sampled.
We
rewrite
this
as
a
sum
on
m, and then u
nected to Ao, l from those connected to AI, m = k + l, and the last term
identity: concerning the case where no stimulus
arises from binomial
the assumption
element is sampled. vVe rewrite this as a sum on rn, and then use a
binomial identity:
" (8 -k 1) (i)l m-l + (1-t)8-i
~ t m(1-t}8-m L.
L.
m=l
k+l=m
8
= 2: t m( 1 - t),-m (S)i- + (1- t)8 -i
8
m
m=l
= -i 2:•
8
8
(8\)tm(l-t)s-m
S m=On~
i
It will be convenient to let the column vector y =
S
{;). . represent
. t
probabilities.
this we m
It "dll be convenient
to let
the column
y = matrix
~~I represent
these
Let us next
construct
the vector
transition
P. For
l 8 ) • 7-8. The combina
all
four
possibilities
in
Figure
take
into
account
probabilities.
and Eo, with
a transition
from P.
st dourn
sj has
of Al
Let us next
construct
the transition
matrix
Fortothis
we probabilities
must
where
take into account all four possibilities in Figure 7-8. The combination
f
j
i
i -j loX,
of Al and Eo, with a transition from 51)down
- j + kto- t81+has
; - kprobabilities
.ifj < i
where
i
j+k
"?=
"-9
0
ifj 2 i
~
{S~
i 'j1tH+k(1-t)H+J-k .. i-:J
if j < i
"'Ij k -:0
k
t - J
t - J +k
The downward
transition sl to sj, by
means of the combination A0
o
Eo,has probabilities (1 - v)(Y- X), where ifj ~ i
_ (S-i)(.
The downward transition Sj to s" by means of the combination Ao and
, 4 j ) " - , ( ~ t ) j
ifj < i
Eo, has probabilities (1- v)( YYcr
- X),
where
=
0
ifj > i
.
i
.)th(l-t)J
ifj
~
i
{(
YH =
~-J
o
ifj > i
CRAP
FINITE MARKOV CHAIPiS
184
We can represent the model as an (s + 1)state Markov chain, in w
stateAPPLICATIONS
s+ occurs when exactly
i stimulus
elements are connected
to
OF MARKOV
CHAINS
18fi
and i = O , 1, . . . , s. All the interesting quantities depend only
Ifwe let x*jjhow
= Xs-i,
-'-I, stimulus
and
=elements
Ys-i, s-j, then
vX* and (1w)(way,
y* - and
X*)hence t
many
are connected
each
represent the
probabilities
of upwards
transition.
Thus
quantities
are functions
on the
chain. For
example, the probabili
action A1 from state st is obtained as follows:
P = wX+vX*+(l-v)(Y-X)+(l-w)(Y*-X*)+(v+w-l)(l-t)'l,
SEC, 5
or
P
Prc[Al] =
+
( S i i ) ( l ) l m ( l -1)+ Y*) - (X +X*)
+ (v +w - 1)(1- t )sI.
k = O Y*)1=0 (Y
= v(X +X*- Y) +w(X +X*k,t not both 0
I
-
-.i
- + (1 t ) ~
m
S
(1)
where k is the number of stimulus elements sampled from those
The final term represents the eases where no stimulus element is sampled,
nected to Ao, 1 from those connected to A1, m = k + 1, and the last t
which were not included in the two previous terms,
arises from the assumption concerning the case where no stim
"We wish element
to compute
PI'. FirstWe
we find Yy.
is sampled.
rewrite this as a sum on m, and then u
binomial identity:
yjk~ = ~ (. i )ti-/t(l-t)k~
k=O
S
1:=0
t-k
8
i
1
i
I "
*
k
= - L k\i t)tH;(l_t)k
8 k=1
ti-kll_t
Sr;;1 (i-I)'
k- 1
\
=i
___t"L.
'i-l
(t
' -
8 1=0
I
1)ii-I-I( 1 - t
)1+1
l,
= i"« -t ).
Hence
1k
oS
It will be convenient
to let the column vector y =
{;). . represent
. t
probabilities.
(2)
Yy = (1-t)y.
Let us next construct the transition matrix P. For this we m
Quite similarly,
possibilities in Figure 7-8. The
take into account
(3)combina
Y*y all
= four
tf+(l-t)y.
of Al and Eo, with a transition from st dourn to sj has probabilities
And by a somewhat longer argument we find that
where
(4)
(X+X*)y = t[+(l-Zt-(l-t)')y.
i -j
)-j+k-t+;-k.ifj < i
- j + k(4), we
If we write PI'" ?by
= means of (1), and
"-9 make use of (2), (3),i and
obtain
0
ifj 2 i
PI' = v[tf+(l-:2t-(l-ty)y-(l-t)yJ
The downward
transition sl to sj, by means of the combination A0
+w[tf+ (1- "2t- (l-t)s)y-tt-(l-t)y)
Eo,has+probabilities
- v)(Y- X), where
tg + :2( 1 - t )1' (1
- tf - (1 - 2t - (I - t lo)y
+ (v+w-l)(l-t)-'y.
, 4 j ) " - , ( ~ t ) j
ifj < i
This simplifies to
Ycr =
0
ifj > i
(5)
PI' = vtf+[1-(1I+w)t]y.
j
Let 11S introduce the vector
i
,,= I' - ('Y+
_V_,)" g for the cases where
u"
11
FINITE MARKOV CHAINS
186
C
and w are not both 0. I t will be shown later that 8 has an im
FINITE in
MARKOV
CHAINS
CHAP. \'II
interpretation
the model.
186
and ware not both O. It will be shown later that S has an important
interpretation in the modeL
or
Pi) =
or
py_('_V_)~ = [l-(V+W)t]y-(-V~)[l-(V+W)t]f
v+w}
and
V+i/J
PS = [l-(v+w)t]S
(6)
The values of v and w determine the nature of the Markov c
O< v < l pnS
and =0 <[l-(v+w)t]nS.
w < 1 , then P > 0 and hence the (7)chain is
If
v = 0, then i = 0 is an absorbing state ; and if w = 0, then i
The values of v and w determine the nature of the Markov chain. If
absorbing state. Hence if either v or w is 0 , we have an absorbi
0< v ~ I Lind 0 < w ~ 1, then P > 0 and hence the chain is regular.
w =absorbing
0 , then we have two absorbing states.
If v = 0, then i =If0 vis= an
state; and if U' = 0, then i = 8 is an
Let us find the probabdity of an A1 response after n step
absorbing sUtte. Hence if either v or w is 0, we have an absorbing chain.
will be gi-n by the vector Pnr. the probability dependin
Ifv=w=o, then we have two absorbing states.
starting state. If v = w = O , then P y = y from (5), hence Pny =
Let us find the probability of an Al response after n steps. IThis
n the oth
probability of a n Al response is unchanged.
will be givan by the vector pny, the probability depending on the
(7) provides the answer
starting state. If v = w "" 0, then Py = y from (5), hence pny = y. The
P"yunchanged.
= [z)/(v+ w)][
[ I - other
(v i- w)t]n8.
probability of an Al response is
In +the
cases,
(7) providcs the answer
The first term (as will be seen below) is the limiting proba
pn yA1= response,
[v!(v+w)lfand
+ [1the (v+w)t]no.
second is the deviation due to t
an
position.
Thc first term (as will be seen below) is the limiting probability
for
Let us first discuss the regular case. Were v > 0 and w >
an Al response, and the second is the deviation due to the initial
O<til,wehavejl-(v+w)tjil.
From(7).
position.
Let us first discuss the regular case. Here v> and u' > O. Since
0<1< I, we have 11- (v+'w)ti < 1. From (7),
and
°
AS = 0
and thus
ao = a(y- (_V
)t) = 0
v+w
and thus
This proves that thevlimiting probability of an A l response is
(8)
v+w' 7-8 we find that the limiting
Applying thiscryto=Figure
probabi
El
actjion by the esperimenter is
This proves that the limiting probability of an Al response is v/(v + i))).
Applying this to Figure 7-8 we find that the limiting probability of an
El act,jon by the experimenter is
win equilibrium
v
v
Thus
the probabilities
for the experimenter
-~·v+--·(l-w) = - _ .
v+ware inv+w
subject
agreement. v+w
Thus in equilibrium the probabilities for the experimenter and the
subject are in agreement.
FINITE MARKOV CHAINS
186
SEC. 5
C
and w are not both 0. I t will be shown later that 8 has an im
187
interpretation OF
in the
model. CHAINS
APPLICATIONS
MA.RKOV
It is interesting to note that the subject does not maximize the number
of correct guesses. Instead, he brings about an equilibrium in which he
is guessiTcg "right" or
with the same frequency with which "right" comes
up.
Since v/(v + w) isand
the mean number of Al responses per trial in equilibrium, and since i/s is the mean number when in state Bt, the vector
3 = {~-
{_V_)JI
gives
deviation
the the
mean
number
ofMarkov c
w determine
nature
of the
The the
values
of v andbetween
\V+W
O< v < l and 0 < w < 1 , then P > 0 and hence the chain is
Al responses in a given
andi =
in0equilibrium.
If v =state
0, then
is an absorbing state ; and if w = 0, then i
From (6), (7), absorbing state. Hence if either v or w is 0 , we have an absorbi
If v =(I-P+A)8
w = 0 , then we
have two absorbing states.
= (v+w)t3
Let us find the probabdity of an A1 response after n step
will be gi-n ZO
by=the
vector Pnr. the probability dependin
_1_3
= w = O , then P y = y from (5), hence Pny =
starting state. If v(v+w)/
probability of a n Al 1response is unchanged. I n the oth
(Z-A)S
= ---8
(7) provides
the answer
8
(v+w)t
P"y = [z)/(v+ w)][ + [ I - (v i- w)t]n8.
1
(9)
= will
(v+W)t'"
The first(Z-A)y
term (as
be seen below) is the limiting
proba
an A1 response, and the second is the deviation due to t
Thus the total deviati(;n
position. from equilibrium (for the number of Al
responses) is proportional
the deviation
S. case.
HenceWere
the total
v > 0 and w >
Let usto first
discuss thevector
regular
deviation may be large
because
for
the
starting
state
may
have
O<til,wehavejl-(v+w)tjil.
From(7). heen
far from the equilibrium
+ wi, or because t is small.
To obtain the limiting variance for the number of A1 responses, we
must use § 4.6, since we have a function on the chain which takes on the
value I in Si with probability f/ = i/s. While we have no general
formula for this limiting yariance, in any concrete example it is easy to
and thus
compute it from § 4.6.
Let us next consider the case of an absorbing chain with one absorbing
state i = 0, that is, v = 0 and w > O. (The case v> 0, w = 0 is similar.)
In. this case 8 =y. Let y be gotten from y by deleting its first component
This
(which is O\. From
(6), proves that the limiting probability of an A l response is
Applying this to Figure 7-8 we find that the limiting probabi
Py
= Po =
= (i-wt)y
Elactjion
by (l-n·t)8
the esperimenter
is
Qy = (I-wt)y
(1 - Q)y = (wt)y
y _
•
1 _
hequilibrium
- -"
Thus in
the probabilities for the (10)
experimenter
-i-we"~
subject are in agreement.
This case is one in which the subject is being conditioned to give Ao
responses. If he gives an Ao response, it is always reinforced. But an
Al response is also reinforced occasionally, ,""ith probability 1 - w. Ny
188
FINITE MARKOV CHAINS
CH
gives the mean of the total number of Al (or "wrong") responses
CHAINS
CHAP.
startingFEUTE
state sl MARKOV
this is llwt.
ip. Thus there may
be VII
a large num
errors for one of three reasons: The fraction ils of stinlulus elenle
gives the mean of the total number of Al (or "wrong") responses. For
need reconditioning was high a t the start ; or the learning param
starting state Si this is l/wt.ijs. Thus there may be a large number of
low ; or w is low, that is, an Al response is frequently reinforce
errors for one of three reasons: The fraction ijs of stimulus elements that
It is sometimes reasonable to assume that the stimulus e
need reconditioning was high at the start; or the learning parameter t is
were originally connected a t random. This gives us an initial
low; or tv is low, that is, an Al response is frequently reinforced.
It is sometimes
to assume that
Thenthe
thestimulus
mean ofelements
the total num
bilityreasonable
vector rr =
were originally connected at random. This gives us an initial proba"wrong" responses is
bility vector 1T = {~G)}- Then the mean of the total number of
188
1.
"wrong" responses is
We could expect to average this number of A1 responses in
2ui is one simple way of estimating
homogeneous population. = This
Finally we discuss the case v = w = 0,where every action of the
We could expect to a\-erage this number of Al responses in a large
is reinforced. Then both i = Q and i = s arc absorbing states,
homogeneous population. This is one simple way of estimating t.
most interesting question is whether the subject ends up cond
Finally we discuss the ease v = w = 0, where every action of the subject
to A0 or to Al responses.
is reinforced. Then both i = 0 and 1: = 8 are absorbing states, and the
From ( 5 ) we see that in this case P y = y. But y has a 0 com
most int~resting question is whether the subject ends up conditioned
for so, and 1 for s,, hence (see Theorem 3.3.9) y gives the proba
to Ao or to Al responses.
for absorption in ss. Thus the probability of being conditioned
From (5) we see that in this case Py=y. But y has a 0 component
to 8 1 responses is equal to the fraction of stimulus elements or
for so, and 1 for s•• henee (see Theorem 3.3.9) y gives the probabilities
connected to A1.
for absorption in Ss. Thus the probability of being conditioned entirely
To obtain more detailed information about the model we w
to Al responses is equal to the fraction of stimulus elements originally
to make a simplifying assumption. Since in most applicati
connected to AI.
learning parameter t is quite small, we shall assume that terms o
To obtain more detailed information about the model we will have
order in t are negligible. This may also be interpreted psycholo
to make a simplifying assumption. Since in most applications the
If we drop terms in powers oft higher than the first, we assume
learning parameter t is quite small, we shall assume that terms of higher
sampling of more than one stimulus element a t a time is quite u
order in t are negligible. This may also be interpreted psychologically.
Under this assumption P is considerably simplified:
If we drop terms in powers of t higher than the first, we assume that the
P%,t+l=
sampling of more than one stimulus element
at (s-i)tv
a time is quite unlikely.
Under this assumption P is considerably
simplified:
pt,i-1
= itw
P',i+l = (s-i)tv pdr = 1 - t (iw+ ( s - i)v)
pn = 0 otherwise.
P!,!-l = itw
(~t)1lJ'
PH
I-t(iw+(s-i)v)
Let us consider
first the regular case under this simplifying
tion. ThePi}most0important
quantity laeking in our previous tre
otherwise.
was the vector of limiting probabilities a. We shall show tha
Let us consider first the regular case under this simplifying assumpour present assumption
tion. The most important quantity lacking in our previous treatment
was the vector of limiting probabilities (I.. We shall show that under
our present assumption
FINITE MARKOV CHAINS
188
CH
gives the mean of the total number of Al (or "wrong") responses
starting
state sl this isOFllwt.
ip. Thus
there may be a large
APPLICATIONS
MARKOV
CHAlKS
189 num
errors for one of three reasons: The fraction ils of stinlulus elenle
Let us compute
.
need aP
reconditioning
was high a t the start ; or the learning param
low ; or w is low, that is, an Al response is frequently reinforce
It is sometimes reasonable to assume that the stimulus e
ukPkf = [aj-lPj-1,1+<1JPjj+Uf;.lPJtU]
were originally connected a t random. This gives us an initial
SEC. 5
.i
1.
k~O
Then
the mean of the total num
bility~ vector
l)VH WrrH =+1(S-j -+- l)tv+
G)VfWS-J[l-t(jW+(S-j)V)]
[(j
"wrong" responses is
+
tvi)J)$-j [(.
S
C:
I )Vf+lU'S-j-l(j + 1 )tW] /(v + W)8
\
(S)
could expect
to average this . number
of A1 responses in
= We
a j +-(--.
. 1,'(S-j+l)w(jw+(s-j)v)
v+W)$ population.
J- /
homogeneous
This isJone simple way of estimating
Finally we discuss the case v = w = 0,where every action of the
s IJ\U+l)V]'
absorbing
states,
is reinforced. Then both i = Q and i = s arc,J+
most interesting question is whether the subject ends up cond
\Ve then find that the two v-terms canoel each other and the two
to A0 or to Al responses.
w-terms also cancel,
hence the expression in brackets is O. Thus
5 ) we see that in this case P y = y. But y has a 0 com
From (and
the right side
reduces
to
aj,
hence (see
a = {ai} is the fixed vector of P.
for so, and 1 for and
s,, hence
Theorem 3.3.9) y gives the proba
Thus we know
that
if
t is small, the limiting probabilities are very
for absorption in ss. Thus the probability of being conditioned
nearly givento
by8this
a. Indeed IX is the limit of the limiting probabilities
1 responses is equal to the fraction of stimulus elements or
as t-+o.
connected to A1.
It is interesting
to notemore
that detailed
for w = vinformation
= l!st the process we obtain is
To obtain
about the model we w
the same as to
that
obt&ined
for
the
macroscopic
process
in the
Ehrenfest
make a simplifying assumption.
Since
in most
applicati
model.
learning parameter t is quite small, we shall assume that terms o
Next let us
consider
annegligible.
absorbing case,
andinterpreted
v=O, thatpsycholo
is,
This say
mayw>O
also be
order
in t are
where the subject
is
being
conditioned
to
response
Ao.
Then
If we drop terms in powers oft higher than the first, we assume
sampling of more
than= one
Jli,i-l
itw stimulus element a t a time is quite u
Under this assumption
P
is considerably simplified:
PH = l-itw
+(.
P%,t+l= (s-i)tv
and all other entries are O. The matrix l-Q = {Ctf} where Cit = itw,
pt,i-1
and C;,1_1 = - itw, and all other entries
are=O.itw"Ve will show that
pdr = 1 - t (iw ( s - i)v)
nij = 1Utw
if j ~ i
pn = 0 otherwise.
o otherwise.
Let us consider first the regular case under this simplifying
Let us compute
tion.N(I-Q).
The most important quantity laeking in our previous tre
, the vector of limiting probabilities a. We shall show tha
was
nikCkj =assumption
nij(jtw)+nu+l[-(j+l)tw]
our present
+
2:
k=l
ifi<j:
if i = j:
if i > j:
0+0=0
(11 jtw)J'tw + 0 = 1
(lUtw)jtw+ (l/U _c l)tw)( - (j + l)iw)
1-1
=
0,
FINITE MARKOV CHAINS
190
190
CIIA
which verifies that N is the desired inverse. I t is interesting t
CHAINS
CHAP.
VII
obtain
t h a t NFINITE
dependsl'v1ARKOV
on i,j , t, but
not on s. We then
which verifies that N is the desired inverse. It is interesting to note
that N depends on i, j, t, but not on s. We then obtain
i
Thus the time of conditioning is inversely proportional to t and
ti = (ljtw)
(I/jl·
and depends on thej~1number of stimulus elements that need
conditioned from A1 to Ao. The time will be large if t is sma
Thus the time of conditioning is inversely proportional to t and t.o w,
will be large if w is small-that is, A1 is frequently reinforced
and depends on the number of stimulus elements that need to be
since the series l / j diverges, the time may also be large because
conditioned from Al to Ao. The t.ime will be Luge if t is small. It
stimuli elements were originally conditioned to 8 1 .
will be large if w is small-that is, Al is frequently reinforced. But
Finally we shall consider the case of two absorbing states, u
i=
since the series Ifj diverges, the time may also be large because many
Mere the first-order terms in t drop out, and hence we will have to
stimuli elementsout
were
conditioned
t2.
ouroriginally
computation
in termstoofAI'
Finally we shall consider the case of two absorbing states, w = v = O.
Here the first-order terms in t drop out, and hence we will have to carry
I
out our computation in terms of t 2 .
= Pi,i+lquite
= (lj2)i(s
A Pi,i-l
computation
similar-1)t2
to the one above will verify tha,t
PH = l-i(s-i)t 2 •
A computation quite similar to the one above will verify that
and
nijhence
= (2jt 2s)(ijj)
(2jf2s)(s -1: )/(8 - j)
if i ,,;; j.
if i ~ j.
and hence
Again the sum may be large because t is small or because th
many terms (s is large). But in this ease ti is inversely proporti
t 2, and hence we expect a much longer time for conditioning.
Again the sum may be large because t is small or because there are
There is one special case in which more precise information i
many terms (8 isable
large).
But in
case tisimplifying
is inverselyassumption.
proportional to
without
thethis
above
This is th
t 2, and hence wewhere
expectva+much
longer
time
for
conditioning.
w = 1 , or w = 1 - v. Were the a,ction of the experime
There is one special
case inofwhich
is availindependent
whatmore
the precise
subject information
does. We found
an exact s
able without the above simplifying assumption. This is the casc
for a in this case, in terms of simple recursion equati0ns.t
w here v + w = 1, or w = I - v. Here the action of the ex perimen ter is
I t was shown that tJhe limiting probabilities may also be ob
independent of what the subject does. We found an exact solution
from the following auxiliary process: We start with s stimuli el
for a in this case, in terms of simple recursion equations. t
completely unconditioned. We select a subset of these, pickin
It was shown that the limiting probabilities may also be obtained
stimulus element with probability t . Then by a random dev
from the following auxiliary process: \Ve start with 8 stimuli elements
assign these to A. with probability u; or to Al with probability
completely unconditioned. \Ve select a subset of these. picking each
We then apply the same procedure to the remaining stimuli ele
stimulus element with probability t. Then by a random device we
till all are assigned. Then the limiting probability ai for the o
assign these to Ao with probability w or to Al with probability 1 w.
process is simply the probability of assigning i stimuli element
We then apply the same procedure to the remaining stimuli elements,
t Cf. Then
J. G. Kemeny
and J. L.probability
Snell, "Mrtrkov
Learning Theory,"
till all are assigned.
the limiting
ai Processes
for the in
original
metrika, 22 (No. 3):221-230, 1957.
process is simply the probability of assigning i stimuli elements to Al
t Cf. J. G. Kerneny and J. L. Snell, "yfarkov Processes in Learning Theory," Psychometl'ika, 22 (No.3): 221-230, 1957.
FINITE MARKOV CHAINS
190
SEC. 6
CIIA
which verifies that N is the desired inverse. I t is interesting t
t h aAPPLICATIONS
t N depends on i,OF
j , t,MAEKOV
but not onCHAINS
s. We then obtain191
'---------
in our auxiliary process. These probabilities are easily obtained for
any 8 of reasonable size.
Thus the time of conditioning is inversely proportional to t and
§ 7.6 Applications to mobility theory. In this section we shall
and depends
on the chain
number
oftostimulus
elements
that need
consider the application
of Markov
ideas
a problem
in sociology,
The time will be large if t is sma
A1
to
Ao.
conditioned
from
the problem of intergenerational occupational mobility. The results
will be large if w is small-that is, A1 is frequently reinforced
of this section were prepared jointly with J. Berger. The problem may
since
the series l / j diverges, the time may also be large because
be stated as follows. A partition A = {AI, A 2 , ••• , AT} is made of the
stimuli
elements
to 8 1 . classes.
set of all occupations.
The were
cells originally
are called conditioned
the occupational
Finally
we
shall
consider
the
case
of
two
absorbing
states, u
i=
These are usually ordered with respect to some socially relevant
Mere the first-order terms in t drop out, and hence we will have to
criterion, for example, prestige of occupation. The question is then
out our computation in terms of t 2 .
asked, to what extent does the occupational class of the father, grandfather, etc., affect the occupational class of the son? In any snch study
a matrix is constructed which represents for each class the fraction of
the sons that would be expected to go into each of the occupations.
A computation quite similar to the one above will verify tha,t
VIe shall take as our basic example ~. matrix constructed from data
collected by Glass and Hall from England and Wales for 1949.t
Following Prais,t we classify the occupations as upper, middle, and
lower. The estimated matrix is
and hence
UPPER
UPPER
(448
MIDDLE
LOWER
.484
.068\
.247)'
.699because t is small or because th
P = MIDDLE
Again
the sum\.054
may be large
many LOWER
terms (s is .011
large). But
.503in this ease
.486 ti is inversely proporti
t 2, and hence we expect a much longer time for conditioning.
We see, for example,
of the upper
class,
44.8
percent
of the
sons
There is that
one special
case in
which
more
precise
information
i
went to the upper
class, 48.4
percent
t.he middle, assumption.
and 0.8 percent
to is th
able without
the
abovetosimplifying
This
the lower. where v + w = 1 , or w = 1 - v. Were the a,ction of the experime
There are independent
two ways to make
usethe
of P.
Onedoes.
is to consider
at an
each
of what
subject
We found
exact s
state the total
andinpredict
fraction
of the equati0ns.t
population
forpo})ulation,
a in this case,
terms the
of simple
recursion
which will be in
the occupational
classes.
We shallmay
can also
this be ob
I t each
was of
shown
that tJhe limiting
probabilities
the "collective
process".
from the following auxiliary process: We start with s stimuli el
A second way
is to study
a single familyWe
bistory.
thisofpoint
ofpickin
completely
unconditioned.
select aFrom
subset
these,
view we consider
this element
history aswith
the probability
outcomes oft .a Markov
with dev
stimulus
Then bychain
a random
transition matrix
shall
callprobability
this the u;
"individual
process."
assign P.
these\\7e
to A.
with
or to Al with
probability
\Ve assume every
family
has exactly
one
son.
We then
apply
the same
procedure
to the remaining stimuli ele
We proceed
to assigned.
discuss theThen
relation
of the basic
concepts
of the o
ai for
tillnow
all are
the limiting
probability
Markov chainprocess
theoryistosimply
these two
processes.
the probability of assigning i stimuli element
We begin with the assumptiolls for a Markov chain. The basic
t Cf. J. G. Kemeny and J. L. Snell, "Mrtrkov Processes in Learning Theory,"
t D. V. Glacs metrika,
and J. R.
"Social )lobility
22 Holl.
(No. 3):221-230,
1957.in Gre<1t Britain: A Study of Intergeneration Changes in Status." in D. V. Glass (Ed.), Social ]lIability in Great Brit.ain,
London, Routledge & Kegan Paul, 1954.
t S, J. Prais, "Measuring Social ltfobiht.y,'· Journal of the Royal Stntistica.l Society,
lIS: 56-60, 1 %5.
192
FINITE MARKOV CHAINS
CHAP.
assumption is that the knowledge of the past beyond the last outco
FINITE
MARKOV
CHAINS I n the individual
CUP. VII
does not
influence
our predictions.
process this wo
mean, for example, that the knowledge of the occupation of the gra
assumption is that the knowledge of the past beyond the last outcome
father
not affect
for the son.
does not influence
our would
predictions.
Inour
the predictions
indiyidual process
this would
We also assume that the same P serves for every generation. T
mean, for example, that the knowledge of the occupation of the grandHowever, there is still a great d
is affect
clearlyour
not
completely
father would not
predictions
forrealistic.
the son.
of
interest
in
studying
what
would
if the This
present P were
"We also assume that the same P serves for every happen
generation.
continue
to
be
appropriate.
is clearly not completely realist,ic. However, there is still a great deal
It is also assumed that changes in distributions in occupatio
of interest in studying what . .vould happen if the present P were to
classes
from one generation to the next are to be accounted for only
continue to be appropria,te.
the process described by P ; that is to say, for purposes of this anal
It is also assumed \that changes in distributions in occupational
we ignore the effect of differential reproduction and migration rates
classes from one generation to the next are to be accounted for only by
these may be related to occupations of the system.
the process described by P; that is to say, for purposes of this analysis
The classification of states has obvious interpretations in mobil
we ignoie the effect of differential reproduction and migration rates, as
An
ergodic
set is a set of
these may be related
to occupations
ofoccupations
the system. from which it is impossible to lea
I
n
most
industrialized
societies
we would expect only one ergodic
The classification of states has obvious interpretations in mobility.
However, if we take as state the pair, occupation and race, then
An ergodic set is a set of occupations from which it is impossible to leave.
crimination against a certain race may cause the resulting chain
In most industrialized societies we would expect only one ergodic set.
have more than one ergodic set. When studying occupations, a
However, if we take as state the pair, occupation and race, then discan have the same occupation as his father. Thus we would
crimination against a certain race may cause the resulting chain to
expect cyclic chains. An absorbing state would mean that for a gi
have more than one ergodic set. When studying occupRtions, a son
occupation the son must follow his father's footsteps. Again,
can have the same occupation as his father. Thus we would not
industrialized societies occupations do not usually have this prope
expect cyclic chains. An absorbing state would mean that for 11 given
We shall therefore assume that our basic chain is regular.
occupation the son must follow his father's footsteps. Again, in
Let us next see the interpretation of the powers of P . For
industrialized societies occupations do not usually have this property.
individual
process the ij-th entry of Pn will give the probability t
We shall therefore assume that our basic chain is regular.
n
generations
the family
be in
class
after
Let us next see the interpretation
of the will
powers
of the
P. j-th
Foroccupation
the
started in the i-th. For the collective process, p(n)ij represents
individual process the ij-th entry of pn will give the probability that,
fraction of the descendants of people in the i-th occupational class t
after n generations the family will be in the j-th occupation class if it
will be in the j-th occupational class after n generations. If we s
started in the i-tho For the collective process, p(n)/j represents the
with fractions T = ( p l ,pz, p 3 ) in each of the classes, then after n gen
fraction of the descendants of people in the i-til occupational class that
tions there will be fractions given by nPn. In our example, ass
will be in the j-th occupational class after n generations. If we stRrt
that
there are a t present 20 percent in the upper class, 70 percent in
with fractions 71"= (Pl, P2, Ps) in each of the classes, then after n generamiddle class, and 10 percent in the lower class. Then after
tions there will be fractions given by 71"Pn. In our example, assume
generation the percentages are 12.9, 63.6, and 23.5. We obtain th
that there are at present 20 percent in the upper class, 70 percent in the
by 10 percent in the lower class. Then after one
middle class, and
/.448and.4S4
generation the percentages are 12.9,63.6,
23.5. .068\
We obtain these
by
192
(.200
.700
.448
.484
.100) ( .054
.699
.068)
.247 = (.129
. 636
.235) .
The fixed probability vector a has the following interpretations.
.Oll .503
the individual
process.486
i t represents the long-range predictions for
The fixed probability vector a has the following interpretations. In
the indiviciu2J process it represents the long-range predictions for the
192
FINITE MARKOV CHAINS
CHAP.
assumption is that the knowledge of the past beyond the last outco
does
not influence our
I n the individual process
this wo
APPLIOATIONS
OFpredictions.
MARKOV CHAINS
193
mean, for example, that the knowledge of the occupation of the gra
occupation of father
an individuaL
basi;:;
for for
regular
chains tells
would notOur
affect
ourtheorem
predictions
the son.
us that these predictions
are independent
of the
occupational
We also assume
that the same
P present
serves for
every generation. T
class. For theiscollective
process,
the fixed
vector gives
the equilibrium
However,
there is still a great d
clearly not
completely
realistic.
fractions. \Vhen
these fractions
are realized,
the fractions
of interest
in studying
what would
happen inif successive
the present P were
generations remain
the to
same.
No matter what the initial fractions are,
continue
be appropriate.
they will, after aItnumber
generations,
close tointhose
given by the
is alsoofassumed
that bechanges
distributions
in occupatio
fixed vector. classes from one generation to the next are to be accounted for only
In our example,
the fixed
vector isbya=
.309).for The
actual
the process
described
P (.067,
; that .624,
is to say,
purposes
of this anal
\
fraction in each
of the
classes
the data reproduction
which determined
the
we ignore
the
effectfrom
of differential
and migration
rates
matrix P wasthese
a= (.078,
.290). to Thus
we can of
seethe
that
the system
may .834,
be related
occupations
system.
may be consid.ered
be nearly in equilibrium.
Thetoclassification
of states has obvious interpretations in mobil
The mean first
passage
have
the usual interpretation
for
the
An ergodic settimes
is a set
of occupations
from which it is
impossible
to lea
individual process
but industrialized
do not seem tosocieties
have a natural
interpretation
I n most
we would
expect onlyfor
one ergodic
the collective However,
process. For
example,
passage and
timerace, then
if weour
take
as statethe
themean
pair, first
occupation
matrix is
crimination against a certain race may cause the resulting chain
U ergodic
M
L
have more than one
set.
When studying occupations, a
can have the same occupation
as his father. Thus we would
2.1
expect cyclic chains. An absorbing state would mean that for a gi
M = the
M
25.1 1.6
U son
occupation
must follow his father's footsteps. Again,
industrialized Lsocieties
do not usually have this prope
26.5 occupations
1.9 3.2
We shall therefore assume that our basic chain is regular.
The standard deviations for the first passage times are
Let us next see the interpretation of the powers of P . For
M entry
L
individual process U
the ij-th
of Pn will give the probability t
the 1.5
family will be in the j-th occupation class
after n generations
L 122.5
started in the i-th. For the collective process, p(n)ij represents
1.2 of
3.9people in the i-th occupational class t
fraction of the descendants
will be in theLj-th25.1
occupational
lA 3.5 1 class after n generations. If we s
with fractions T = ( p l ,pz, p 3 ) in each of the classes, then after n gen
Since the standard
deviations
the firstgiven
passage
times are
or the
tions there
will be of
fractions
by nPn.
In our
example, ass
same order of magnitude
as
the
means,
the
means
are
not
to
be
taken
as percent in
that there are a t present 20 percent in the upper class, 70
typical values. middle
However,
theand
relative
size is ofin
interest.
For class.
example,
10 percent
the lower
Then after
class,
the mean timegeneration
to go fromthe
lower
to upper are
is about
as big We
as obtain th
percentages
12.9, five
63.6,times
and 23.5.
the mean time to go from upper to lower.
by
Assume now that the individual process is in equilibrium. Then the
SEC. 6
C' ::}
M(250
'i)
/.448 .4S4 .068\
reverse transition matrix giYcs the probabilities for the father's occupation when that of the son is known. If P is reversible, then, given that
a man is in class t, the probability that his son will be in a given occuF<tional class j is the same as the probability that his father was in this
class j. The condition
for reversibility
has an a
interesting
interpretation
The fixed
probability vector
has the following
interpretations.
for the collective
process. Recall
the condition
for reversibility
the individual
processthat
i t represents
the long-range
predictions for
may be expressed by saying that D-IP should be a symmetric matrix.
In other words, that (1iPij = O;Pij. In the collective process alPij l'epresents, in equilibrium, the fraction of the people in the i-th occupational
194
FINITE MAEKOV CWA'IPU'S
CHAP
class which move in one generation from the i-th to the j-th c
MARKOV
CHAIKS
CHAP.the
VIIj-th cla
represents
the fraction
which move from
Also alprcFINITE
-------- -------------------------------the i-th class. Hence the condition for reversibility means that t
class which should
move in
generation
from the
i-th to
the j-thIf class.
be one
an "eq~isl
exchange"
between
classes.
there is an e
Also aJPii represents the fraction which move from the j-th class to
exchange then clearly the total numbers in each class will remain f
the i-th class. Hence the condition for reversibility means that there
i.e. the process will be in equilibrium. However, equal exchange
should be an "eql"11 exchange" between classes. If there is an equal
much stronger condition. The above discussion suggests tha
exchange then clearly the total numbers in each class will remain fixed,
the collective process the matrix D-1P is an interesting matrix.
i.e. the process will be in equilibrium. However, equal exchange is a
call this matrix the exchange matrix. For our basic example, th
much stronger condition. The above discussion suggests that for
change matrix is
194
the collective process the matrix D-IP is an interesting matrix. \Ve
U basicMexample,
L
the excall this matrix the exchange matrix. For our
change matrix is
U /.030 .032 .OM\
U
CO
l\1
.032
L
005)
U
D-IP = IH .034 .436 .154 .
Note that there is approximately equal exchange between
classes.?
L .003 .155 .150
Finally we consider the question of lumpability for mobility
~ote that there is approximately equal exchange between the
cesses. This is particularly important for the following reason.
classes. t
decide that the Markov assumption is reasonable for a certain meth
Finally weclassification,
consider thethen
question
of lumpability
mobility
pro-classific
we cannot
arbit,rarilyfor
treat
a coarser
cesses. Thisasisaparticularly
important
the following
reason.
If we is obt
Markov chain.
This isfor
because
t,he coarser
classification
Markov
assumption
is reasonable
a certain
method
of under
decide that the
from
the finer
by lumping
states. for
We
know that
only
cl11ssification,special
then we
cannot
arhitrarily
treat
a
coarser
classification
conditions will this again result in a Markov chain. Of c
f
as a Markov chain. This is because the coarser classification is obtained
the method of classification itself has a great deal to do with wh
from the finer
by
lumping
states.
"Ve
know
that
only
under
very
or not the Markov assumpt,ion is realistic. Hence the coarser an
special conditions will this again result in a Markov chain. Of course
may be taken as a Markov chain even when the condition for lu
the method of classific11tion itself has a great deal to do with whether
bility is not satisfied. However, we must then admit that the
or not the Markov
is realistic.
the coarser
analysis
analysisassumption
is not a Markov
chain. Hence
We cannot
have boih,
unless the
may be taken as a :'vlarkov chain even when. the condition for lumpadit,ion for lumpability is satisfied.
bility is not satisfied. However, we must then admit that the finer
We shall illustrate the above ideas in terms of some actual mo
analysis is not a Markov cha.in. We cannot have both. unless the constudies. The example that we have been considering was act
dition for Jumpability is satisfied.
obtained from a finer analysis of the data obtained by Glass and
We shall illnst.rate the above ideas in terms of some actual mobility
for England and Wales in 1949. These authors used seven cl
studies. The example that we have heen considering was actually
They are:
i
iI
I
!
obtained from a finer analysis of the d11ta obtained by GJass and Hall
for England and 1.Wales
in HJ4D.andThese
authors used seven classes.
Professional
high administrative.
They are:
2. Nanagerial and executive.
3. Inspectional, supervisory, and other non-manual (higher g
1. Professional and high administrative.
2. Managerial
executive.
t For and
n more
detailed discussion of t.he exchange propert,ies of a system see J.
and J. L. Snell,
" On the Concept
of Equal
Exchange,"(higher
Behavioral
Science, 2, (
3. Inspectional,
supervisory,
and other
non-manual
grade),
.•..
1
111-118, 1957.
t 'For a more detalled discussion of the exchange properties of a system see J. Berger
. ~.:
f·
and J. L. Snell. "On the Concept of Equal Exchange," Behavioral Science, 2, (.No.2):
111-118,1957.
t
FINITE MAEKOV CWA'IPU'S
194
CHAP
class which move in one generation from the i-th to the j-th c
alprc represents the fraction which move from 195
the j-th cla
Also
S,~c. 6
APPLICATIONS OF MARKOV CHAINS
the i-th class. Hence the condition for reversibility means that t
be an "eq~islexchange" between classes. If there is an e
4. Same should
(lower grade).
thenrontine
clearly grades
the total
in each class will remain f
5. Skilledexchange
manual and
of numbers
non-manual.
i.e. themanual.
process will be in equilibrium. However, equal exchange
6. Semi -skilled
7. Unskilled
much
stronger condition. The above discussion suggests tha
manual.
the collective process the matrix D-1P is an interesting matrix.
From their data we obtain the transition matrix
call this matrix the exchange matrix. For our basic example, th
change
matrix
is3
1
2
7
4
5
6
p=
0.388
0.147
0.202
0.062
0.107
0.267
0.227
0.120
U 0.047
M 0.016
L ,
0.140
U /.030
0.207
.032
0.053
.OM\
0.020
0.035
0.101
0.188
0.l91
0.357
0.067
0.061
0.021
0.039
0.li2
0.:212
0.431
0.124
0.062
Note0.024
that 0.075
there is
approximately
exchange between
0.125
0.009
0.123
0.473 0.171equal
classes.?
(loOS8 0.391 0.312 0.155
O.Ol~
0.000
Finally
we 0.041
consider
the question of lumpability for mobility
This 0.036
is particularly
for the
following reason.
cesses.0.008
0.274
0.000
0.364 0.235
0.083 important
decide that the Markov assumption is reasonable for a certain meth
we cannot
arbit,rarily
treat
a coarser
Our previous classification,
example was then
obtained
from this
study by
calling
{1,2} classific
as{3,4,5}
a Markov
This is
because
t,helower
coarser
classification is obt
the upper class,
the chain.
middle class,
and
{5,7} the
class.
from isthe finer by lumping states. We know that only under
The fixed vector
special conditions will this again result in a Markov chain. Of c
.041 .088 .127 A10 .182 .129).
the method
of classification itself has a great deal to do with wh
or not the m.atrix
Markovwas
assumpt,ion
realistic.
Hence
coarser an
The above transition
estimatedisfrom
a sample
of the
3497.
may
be taken
as a Markov
chain even
The distribution
of the
oceupp,tions
in this sample
was when the condition for lu
bility is not satisfied. However, we must then admit that the
Ii = (.030 .046 .0\)4 .131
.409 We
.170cannot
.121).have boih, unless the
analysis is not a Markov chain.
dit,ion
for lumpability
satisfied.
VYe see that these
numbers
are fairly is
close
to the equilibrium vector.
We shall illustrate the above ideas in terms of some actual mo
2 example
3
4
7
6
studies. The
that 5we have
been
considering was act
43.9 from
26.2 a finer
11.5obtained by Glass and
1obtained
\J.n analysis
8.4 data
9.2 4.0of the
for England and Wales in 1949. These authors used seven cl
:tThey63.1
are: 24.2 10.1 8 . .5 3.5 8.1 1l.1
a = (.023
3
.ill = 4
5
70.3 30.5 11.4 8.0 2.9 7.6 10.3
1. Professional and high administrative.
72.3
33.0 12.7
2.6 7.0 10.0
2. Nanagerial
and 7.t)
executive.
3. Inspectional,
supervisory, and other
non-manual (higher g
73.7
33.9 13.5 8.7 2.4 6.5
9.3
n more
detailed
of t.he 5.5
exchange
6 t For
9.l 2.6
8.8propert,ies of a system see J.
74.9
34.6
14.1discussion
and J. L. Snell, " On the Concept of Equal Exchange," Behavioral Science, 2, (
<)
9
.,
7111-118,
7.7
75.01957.
34.8 14.3
"'.1 5.9
~
.~
The diagonal entries of M are the reciprocals of the fixed vector.
Since the fixed vector is close to the actual fraction of people in each
196
F I N I T E MAEKOV CBAYNS
CHAP
class, when this fraction is small the mean time to return is co
FINITEFor
MARKOV
pondingly large.
a givenCHAINS
occupational classCHAP.
i t is VII
interesting
compare the mean time to reach this class, starting in each of the o
class, when this fraction is small the mean time to return is corresclasses. We observe that in general the mean time to go from
pondingly large. For a given occupational class it is interesting to
state i to a given state j decreases as i gets closer t o state j. Posit
compare the mean time to reach this class, starting in each of the other
of these occupational classes correspond to their relative social pres
classes. We observe that in general the mean time to go from any
and hence "closer" means closer in terms of prestige.
state i to a given state j decreases as i gets closer to state j. Positions
We have discussed only regular chain concepts in this section.
of these occupational classes correspond to their relative social prestige,
know
that absorbing chain ideas can be fruitfully used to study reg
and hence "closer" means closer in terms of prestige.
chains. For example, we can study the behavior of the middle cla
\Ve have discussed only regular chain concepts in this section. We
3,4,5 by making the upper classes I and 2 and lower classes 6 a
know that absorbing chain ideas can be fruitfully used to study regular
into absorbing states. Doing this, we obtain a n absorbing chain
chains. For example, we can study the behavior of the middle classes
and Rclasses
given by
matrices
3,·4,,5 by making
theQ upper
1 and 2 and lower classes 6 and 7
196
into absorbing states. Doing this, we obtain an absorbing chain 'with
matrices Q and R given by
,l
4
'C
.191
5
357)
Q= 4
8
.112
.212
.431
5
.075
.123
.473
3 (.035
R = 4
.021
5
.009
2
6
7
.101
.067
.039
.124
.0(1)
.062 .
The basic quantities for this chain are :
.024 .171
3
4
The basic quantities for this chain are:
3
.N
3C'
4
.38
4
.36
1.60
5
.29
.45
'("8
B = 4
.06
.125
6
6
155)
T~4(3.51
3
1347)
2.47
5
3.21
1.45 \
2
.20
6
7
.42
.14
.49
:~)
From T we obtain the mean time to leave the set {3,4,5) for th
.04 .11
.36 set. We see t h a t this is bet
time for each 5starting
state.49
in the
3 and 4 for each starting state. From B we find the probabiliti
From T we obbdn the mean time to Je,Lve the set {3,4,5} for the
first
leaving by moving to each of the states 1,2,6,7. Combining s
time for each starting state in the set. We see that this is between
3 and 4 for each starting state. From B we find the probabilities of
leaving by moving to each of the states J,2,6,7. Combining states
196
F I N I T E MAEKOV CBAYNS
CHAP
class, when this fraction is small the mean time to return is co
pondingly
large. For
a given occupational
class i t 197
is interesting
APPLICATIONS
OF MARKOV
CHAINS
compare the mean time to reach this class, starting in each of the o
1 and 2, and 6classes.
and 7 weWe
canobserve
find thethat
probability
of moving
outtime
of each
in general
the mean
to go from
of the middlestate
classes
by
moving
to
the
upper
or
to
the
lower
i to a given state j decreases as i gets closer t o class.
state j. Posit
These probabilities
are:
of these
occupational classes correspond to their relative social pres
and hence "closer" means
U
Lcloser in terms of prestige.
We have discussed only regular chain concepts in this section.
. ,- ideas can be fruitfully used to study reg
know that absorbing chain
SEC. 6
3(' "')
chains. For example,
4 .20 we
.81 can study the behavior of the middle cla
3,4,5 by making the upper classes I and 2 and lower classes 6 a
\ .15 .85
into absorbing 5states.
Doing this, we obtain a n absorbing chain
matrices Q and R given by
In each case the probability of leaving by way of the lower class is
much higher than leaving by way of the upper class.
It is interesting to observe that the probability of leaving the middle
class by way of the upper class decreases the lower the level of the
occupational class.
As mentioned earlier, our basic example in this section W,iS obtained
by combining states in this se\'en-state chain. The partition used was
A= ({1,2}, {3,4,5}, {6,7}). The first set being the upper class, the
second the middle class, and the third the lower class. It is interesting
then to check the condition for lumpability with respect to this partition. To do this we must find t.he matrix PV (see §6.3). III'"e obtain,
Al
Az
A3
The basic quantities
this chain
.062 are :
.534for .404
2
.374 3 .553 4 .073 6
3
.136
.736
.128
PV = 4
.060
.754
.128
.033
.671
.2BG
5
-~--------
6
.013
.520
.467
7
.008
.48:3
.500
To satisfy the condition for lumpability it is necessary that the
components of aFrom
column
of this vector be constant within the sets
T we obtain the mean time to leave the set {3,4,5) for th
AI, A 2 . A 3 . This
is
certainly
the case.
example
Weprobability
see t h a t this is bet
time for each not
starting
state For
in the
set. the
of mo~'ing to Al
is quite
for thestate.
states of
A z. BItweis find
.033 the
fromprobabiliti
From
3 and
4 fordifferent
each starting
state 5, .060 from
state
4,
and
.136
from
state
3.
Thus
we
would
not
be
s
leaving by moving to each of the states 1,2,6,7. Combining
justified in treating both of om processes as Markov chains.
If we choose to belic\'e that the seven-state chain is a Markov chain,
then the three-state process is not a Markov chain, but in equilibrium
FINITE MARKOV C H A I N S
198
198
CHAP.
the matrix p, the vector 6, and the matrix
are all well defined.
$ 6.4,VII
we obtain
we compute
these matrices
byCHAINS
the method given in
FINITE
MA.RKOV
CHAP.
U areMall well
E defined. If
the matrix P, the vector ii, and the matrix .if
we compute these matrices by the method given in § 6.4, we obtain,
UP =1\'1
UC
P=M
Ii
.50
1
07) .31)
.05B =.70(.06.25.63
1
.01
.50
(.06
.63
.31)
U
M
1
C'
2.0
59)
U
.48
26.3quantities
1.6 6.3are all quite close to those obta
M=M
We see
that these
by treating the
28.0 2.0chain
L three-state
3.0as a Markov chain.
The next example we consider is obtained from data collected
We see that
these quantities
are made
all quite
to those obtained
N. Rogofft
in a study
fromclose
marriage-license
applications
by treating the
three-state
chain
as
a
Markov
chain.
Marion County, Indiana. The interest in this example lies in the
The next that
example
consider
is obtained
from data
collected
datawe
were
obtained
for two different
time
periods,by1905 to
N. Rogofft and
in a1938
study
madethe
from
applications
for to com
through
firstmarriage-license
half of 1941. Hence
i t is possible
Marion County,
Indiana.
The
interest
in
this
example
lies
in
the
fact
the transition matrices for these two different time periods. Wi
that data were
obtained
for there
two different
timeand
periods,
1912 9,892.
the first
sample
were 10,253
within1905
thetosecond,
and 1938 through
the firststudy
half of
Hence
it is possible
compare
the Rogoff
a 1941.
very fine
analysis
of thetooccupations
is m
the transition
matricesfor
forillustrative
these two purposes,
different time
periods.
However,
we have
madeWithin
a coarse anal
the first sample
were 10,253
the second,
9,892. Inmanual,
This there
classification
mayand
be within
considered
as non-manual,
the Rogoff farming.
study a very
fine first
analysis
of the
occupations
is made.
The transition
matrix is
We treat
the 1910
case.
However, for illustrative purposes, we have made a coarse analysis.
This classification may be considered asNON-MANUAL
non-manual, MANUAL
manual, FARM
and
NON-MANUAL
.594
.396is
.009
farming. We treat first the
1910 case. The transition
matrix
NON-MANUAL
P = MANUAL
.211
MANUAL
FARM
.396
NON-MANUALFARM
(
.252
009 )
.782
.007 .The actual fract
P = MANUAL
The fixed vector is.211
ff1910=(.343
.648 ,009).
.594
inFARM
each of the classes are
. 252given by .641
.108 .
61910
(.3110 The
.658 actual
.034) fractions
The fixed vector is al910 = (.343
.648= .009).
in each of the
classes
by
Note
thatare
thegiven
equilibrium
vector predicts significantly fewer farm
than there&1910
actually
are. .658
This would
= (.:HO
.034) suggest that the 1940 data sh
show a decrease in the fraction in farming.
Note that the equilibrium vector predicts significantly fewer farmers
t K. Rogoff,
Recent Trends zn Occupattonal Moh~lzty,Glencoe, Ill., The Free P
than there actually
are. This would suggest that the 1940 data should
1953.
show a decrease in the fraction in farming.
tN. Rogoff, Recent Trends in Occupational },Johility, Glencoe, Ill., The Free Press,
1!J53.
FINITE MARKOV C H A I N S
198
SEC. 6
CHAP.
the matrix p, the vector 6, and the matrix
are all well defined.
we obtain
we APPLICATIONS
compute these matrices
by the method
OF MARKOV
CHAINSgiven in $ 6.4,
199
U
For the 1940 case we find
f
NON-MANUAL
P = MANUAL
MANUAL
FAR]',!
.622
.375
.003
.274
B = (.06
.265
.721
.005
.63 .31)
.694
.042
P =
The fixed vector is CX1940= (.420 .576 .004).
in each class in 1940 is given by
IX1940
E
NON-MANUAL
\
FARM
M
= (.373
.616
\
}
The actual fraction in
.Oll).
Wefra.ction
see thatof these
quantities
are all quite
close toItthose
obta
As predicted, the
farmers
has significantly
decreased.
is
treatingthat
the the
three-state
chainvectors
as a Markov
chain.
interesting tobyobserve
equilibrium
for 1910
and 1940
The nextinexample
weclass
consider
obtained
from
predict larger fractions
the upper
than is
there
actually
are.data
In collected
N. Rogofft
in a study vector
made predicted
from marriage-license
applications
the case of England
the equilibrium
smaller numbers
in the upper Marion
classes. County, Indiana. The interest in this example lies in the
that datathat
were
for two different
time periods,
The final example
weobtained
discuss illustrates
equal exchange.
The 1905 to
and 1938
through
firstbyhalf
of 1941.
Hence
is possible to com
data was obtained
from
a studythe
made
Blumen,
Kogan,
andi tMcCarthy
the transition
matrices
for these
different
time
periods. Wi
on labor mobility.
t This was
a very large
studytwo
based
on social
security
first sample
10,253
thebeen
second,
records. A the
I-percent
samplethere
of allwere
workers
whoand
arewithin
or have
in 9,892.
the Rogoff
study
a very fine
analysis
the occupations
is m
covered employment
since
the inception
of the
socialofsecurity
system
However,
illustrative
have madesample
a coarse anal
in 1937 has been
kept. forThe
study waspurposes,
based on we
a 10-percent
This classification
may
be considered
as matrix
non-manual,
from this record.
It presents the
following
transition
for themanual,
farming.
Thea transition
is
Webracket
treat first
group of males
in the age
20the
to 1910
24. case.
'We omit
discussionmatrix
of
the classification used in the study.
NON-MANUAL MANUAL FARM
NON-MANUAL
2
3
2
(832
5
0.082
0.033 0.013
P = MANUAL
4 .594
5
.396
.009
0095)
0.028
.211
0.038
.252 0.112
p= 3
0.038 0.034 0.785 0.036 0.107
The fixed vector is ff1910=(.343 .648 ,009). The actual fract
0.045 are0.017
0.156
4
in each of0.054
the classes
given 0.728
by
0.046FARM
0.788
0.055
0.016
0.023 0.071 0.759
61910 = (.3110 .658 .034)
The fixed vector
is the equilibrium vector predicts significantly fewer farm
Note that
than athere
actually
would.322)
suggest
that the 1940 data sh
.
.184are..076This. l48
= (.270
show a decrease in the fraction in farming.
The actual fractions in the classes considered were
t K. Rogoff, Recent Trends zn Occupattonal Moh~lzty,Glencoe, Ill., The Free P
1953.
IX
= (.282 .170 .068 . 137 .343) .
t 1. Blumen, M. Kogan, P. J. McCarthy, The Industrial ldobility oj Labor a. a Probability Proce8s, Cornell Studies in IndnstrisJ and Labor Relations, Vol. VII, 1955.
200
The most interesting feature of this example is the exchange m
This is FINITE MARKOV CHAINS
CruP. VII
The most interesting feature of this example is the exchange matrix.
This is
{'"
026)
.009
.004
.008
2
.008
.145
.003
.007
.021
D-1P = 3
.003
. 002
.060
.003
.008
.
.007 .002of this
.108 matrix
.023 indicates that this s
4 perfect
.008 symmetry
The almost
may be considered
be in.007
equal.023
exchange
.244 in equilibrium, or th
5
.026 to.021
process is reversible.
The almost perfect symmetry of this matrix indicates that this system
may be considered to be in equalLeonhief
exchange
in equilibrium,
or that the
model.
In the Leontief
input-o
model, we consider an economy in which there are r industries a
process is reversible.
§ 7.7
make the simplifying assumption that each industry produces e
The
In the
Leontieffactors
input-output
oneopen
kind Leontief
of goods. model.
We regard
the natural
of production
model, we consider
economy
in which
are rand
industries
we
as land, an
timber,
minerals,
etc.there
as free,
do notand
consider
the
make the simplifying
assumption
that
each
industry
produces
exactly
entering into the cost of finished goods. In general, the industri
one kind ofinterconnected
goods. We regard
thesense
natural
factors
production
in the
that
each of
must
buy a such
certain am
as land, timber,
minerals,
andproducts
do not consider
them
(positive
or zero)etc.
of as
thefree,
other's
in order to
runasits ind
entering into
the
costdefine
of finished
goods. coeficients
In general,
industries
We
shall
technological
as the
follows
: qrj is are
the amo
interconnected
in
the
sense
that
each
must
buy
a
certain
industry i in
the output of industry jthat must be purchased by amount
(positive orthat
zero)industry
of the other's
industry.
i may products
produce in
$1order
worthtoofrun
its its
own
goods. Let
We shall define
coefficients
glj is the amount of
x r matrix with
entries as
qtj.follows:
By their
definition, the technol
the rtechnological
the output of
industry jare
that
must be purchased by industry i in order
coefficients
non-negative.
that industryIti is
may
$1 the
worth
goods.
Q the
be total
easyproduce
to see that
sumof
ofits
theown
qij, for
i fixed,Let
gives
the r x r matrix
entries
gij.
technological
of thewith
inputs
needed
byBy
thetheir
i-thdefinition,
industry inthe
order
to produce $1
coefficients of
areitsnon-negative.
goods. If the i-th industry is to be profitable, or a t least to
It is easy even,
to see this
thatsum
the sum
gij, for i fixed, gives the total value
mustofbethe
less
than or equal to the value of its outpu
of the inputs
needed
by
the
i-th
industry
in order to
produce
worth
q t 1 + q12
. . . qtr < 1. For obvious
reasons
we $1shall
call the
of its goods.industry
If the i-th
industry
is
to
be
profitable,
or
at
least
to
profitable if the strict inequality holds andbreak
profitless
even, this sum
must holds.
be less than
or equal
the value that
of itsevery
output,
i.e.
equality
We make
the to
assumption
industry
is
qn+qi2+ ...
+girCo
1.
For
obvious
reasons
we
shall
call
the
i-th
profitable or profitless and thus rule out the possibility of unprof
industry profitable
if the strict inequality holds and profitless if the
industries.
equality holds.We We
thethe
assumption
that every
canmake
restate
above conditions
as industry is either
profitable or profitless and thus rule out the possibility of unprofitable
industries.
We can restate the above conditions as
+
+
)
Having discussed the inputs of the industries we nest(1discuss
outputs. Let xi denote the monetary value of the output
(2) of th
industry and let n = ( x l , x2, . . . , z,) be the row vector of ou
Having discussed
of the
industries
we zrqtr
next ofdiscuss
their of th
the output
Since the the
i-thinputs
industry
needs
an amount
outputs. Let Xj denote the monetary value of the output of the i-th
industry and let 11= (Xl, X2, .•• ,xr ) be the row vector of outputs.
Since the i-th industry needs an amount Xigij of the output of the j-th
SEC. 7
The most interesting feature of this example is the exchange m
This
is
APPLICATIONS
OF MARKOV CHAINS
201
industry, the vector of inputs needed by the industries is simply trQ.
Then the j-th component of 77Q gives the total value of the output that
must be produced by the j-th industry in order to meet the interindustry demand for its product.
Let us assume that the economy supplies for consumption an amount
Cj of the output of the i-th industry. Let y= (CI' Cz, . . . , cr ) be the
consumption The
vector;
we shall
require
that of this matrix indicates that this s
almost
perfect
symmetry
may be considered to be in equal exchange in equilibrium, or th
y ~ o.
(3)
process is reversible.
The requirement that the production vector of the economy be
Leonhief
model.
Leontief input-o
adjusted so that the inter-industry
needs as
well asInthethe
consumption
r industries a
model,
we
consider
an
economy
in
which
there
are
needs may be fulfilled is now easy to write in vector form; it is
make the simplifying assumption that each industry produces e
77 = We
-n-Q+y.
(4)
one kind of goods.
regard the natural factors of production
timber, minerals, etc. as free, and do not consider the
Rewriting as
(4) land,
as
entering into the cost of finished goods. In general, the industri
77(1 -Q) = y,
(5)
interconnected in the sense that each must buy a certain
am
(positive
or
zero)
of
the
other's
products
in
order
to
run
its ind
we see that it is a set of r simultaneous equations in r unknowns.
We shall define
technological
coeficients
as followssolutions
: qrj is the amo
To be economically
meaningful,
we must
find non-negative
be purchased
i in
thethe
output
of industry
to (5). Since
demand
vector j
y that
maymust
be arbitrary,
we by
seeindustry
that
industry
may produce $1and
worth
its own
goods. Let
equations (5)that
are in
general inon-homogeneous
willofhave
a solution
x r matrix
qtj. By
their definition,
the technol
if and only ifthe
the rmatrix
1- Qwith
has entries
an inverse.
Moreover,
the solutions
arefor
non-negative.
to (5) will be coefficients
non-negative
every y if and only if (1- Q)-l has all nonIt is easy We
to see
that therefore
the sum ofsearch
the qij,for
for necessary
i fixed, gives
the total
negative components.
must
and
of the inputs
needed
by of
the1-Q
i-th be
industry
in order to produce $1
sufficient conditions
that the
inverse
non-negative.
of itsthis
goods.
If thebyi-th
industry is
to model
be profitable,
or a t least to
We will solve
problem
imbedding
our
in a Markov
this sum
be less
or equal
to the
valuewith
of its outpu
chain. (Thiseven,
solution
wa.s must
worked
out than
by the
authors
jointly
q
t
1
+
q12
.
.
.
qtr
<
1.
For
obvious
reasons
we
shall
call the
G. L. Thompson.)
profitable
if the
inequality model
holds we
andshall
profitless
industry
By the 1',farkov
chain
associated
with strict
an input-output
equality
We
make the
assumption that every industry is
mean a Markov
chainholds.
with the
following
properties:
profitable or profitless and thus rule out the possibility of unprof
(i) The states
are the r processes of the model plus one additional
industries.
absorbingWe
state
calledthe
theabove
banking
state.
canSo,
restate
conditions
as
(ii) The transition matrix P is defined as follows:
+
+
poo = 1
j > 0
PO} = 0
i,j >of0 the industries we nest discuss
HavingPljdiscussed
= qlj ,the inputs
outputs. Let xi denote the monetary value of the output of th
i > O.
1nj=l
= ( xgijl , x2, . . . , z,) be the row vector of ou
industry PiO
and= let
Since the i-th industry needs an amount zrqtr of the output of th
The intuitive interpretation of this is the following: If industry i
receives a dollar for its use, then it spends it by buying Pij from industry
j. The remainder of the dollar, if any-that is, the amount Plo-is
2:
202
FINITE IniURMOV CHAINS
t h e profit, and we may think of i t as being deposited in a bank. T
FINITE
MARKOV
CHAP. VII
fact that the
banking
state is aCHAINS
n absorbing state, means
that the ba
gets money but does not spend it.
the profit, and we may think of it as being deposited in a bank. The
We immediately see that if Q is a matrix satisfying ( I ) and (2), th
fact that the banking state is an absorbing state, means that the bank
a non-negative solution to equations (5) exists for every r > 0 , if a
gets money but does not spend it.
only if the associated Markov chain is absorbing, with the banking sta
We immediately see that ifQ iE a matrix satisfying (1) and (2), then
so as its only absorbing state. If so is the only absorbing state for
a non-negative solution to equations (5) exists for every r?O, if and
absorbing chain, then ( I - &)-I = N exists and is non-negative ; hen
only if the associated Markov chain is absorbing, with the banking state
7~ = -yN is the desired solution.
Otherwise ( I -&)-I, which gives t
So as its only absorbing state. If So is the only absorbing state for an
mean number of times in various states before reaching so, would ha
absorbing chain, then (I _Q)-l = X exists and is non-negative; hence
to have infinite entries, i.e. cannot exist.
7T=yN is the desired solution. Otherwise (l-Q)-l, which gives the
There is a simple economic interpretation of this result. We see th
mean number of times in various slat.es before reaching so, would have
from every state i t must be possible to "reach" the banking sta
to have infinite entries, i.e. cannot exist.
Only
industry "reaches" the bank directly. A profitle
- a profitable
-economic interpretation
There is a simple
of this result. 'We see that
industry must reach the bank through a profitable one. Hence o
from every state it must be possible to "reach" the banking state.
condition states that every indz~strymust be either projitable ov mu
Only a profitable industry "reaches" the bank directly. A profiUess
depend on a profitable industry. For example, if we assume that eve
industry must reach the bank through a profitable one. Hence our
industry depends on labor, and that labor is a profitable indust
condition states that every ind1lstry must be either profitable or must
(which presumably means that labor is paid more than subsisten
depend on a profitable;:nduslry. For example, if we assume that every
wages), then our condition will be met, and all demands can be f
industry depends on labor, and that labor is a profitable industry
filled.
(which presumably means that labor is paid more than subsistence
If the above condition is violated, then the economy cannot fulfill
wages), then our condition \vill be met, and ill! demands can be fulpossible demands. Let us ask what kinds of demands i t can fulf
fined.
First of all we consider the case that there is no profitable indust
If the above condition is violated, then the economy cannot fulfill all
This means that each industry needs all it produces to pay for r
possible demands. Let us ask what kinds of demands it can fulfill.
materials, and i t would seem that i t could meet no outside deman
First of all we consider the case that there is no profitable industry.
That this is indeed the case is easily proved.
This means that each industry nceds all it produces to pay for raw
If there is no profitable industry, then each row of Q has row sum
materials, and it would seem that it could meet no outside demand.
hence Q t = 5. If we multiply equation (4) by t on the right, we fi
That this is indeed the case is easily proved.
that
If there is no profitable industry, then each row of Q hilS row sum 1,
hence Qf = r If we multiply equation (4) by f on the right, we find
that
hence
= 0. This says t h a t the sum of all the demands is 0. Hen
202
no (positive) demand can be fulfilled.
Let us now consider the general case where our condition is violat
hence yf == O. This says t.hat t.he sum of all the demands is O. Hence
The associated Markov chain is not an absorbing chain with the sin
no (positive) demand can be fulfilied.
absorbing state. Then there must be a n ergodic set other than {
Let us now consider the general case where 0111' condition is violated.
i.e. a closed group of industries none of which is profitable, and wh
The associated
Markov
chain
is not outside
an absorbing
chain with
t,hetake
single
depend
on no
industry
the group.
Let us
the set of
absorbing sta,te. Then there must be an ergodic set other than {so],
such industries, that is the union of all ergodic sets other than {
i.e. a closed group of industrics none of which is profitable, and which
The submatrix Q of these industries has the property Qt= [, as abo
depend on no industry outside the group. Let us take the set of all
and hence can fulfill no outside demand. Thus the entire econo
such industries, that is the union of all ergodic sets other than {so}.
can fulfill no demands of goods produced by these industries. A
The submatrix Q of t.hese industries has the property Qf = f, as above,
and hence can fulfill no outside demand. Thus the entire economy
can fulfill no demands of goods produced by these industries. And
202
FINITE IniURMOV CHAINS
t h e profit, and we may think of i t as being deposited in a bank. T
factAPPLICATIONS
that the bankingOF
state
is a n absorbing
that the ba
MARKOV
CHAINSstate, means203
gets money but does not spend it.
any goods whose
from
these indusWe production
immediatelyrequired
see thatraw
if Qmaterials
is a matrix
satisfying
( I ) and (2), th
tries also cannot
be
supplied,
since
these
would
act
as
outside
demands
every r > 0 , if a
a non-negative solution to equations (5) exists for
on the closed
group
ofassociated
industries.Markov chain is absorbing, with the banking sta
only
if the
However,so
if as
weits
remove
this closedstate.
group If
of so
profitless
industries
and state for
only absorbing
is the only
absorbing
all industries
depending
on then
them,( Ithe
remaining
industries
(if any)
absorbing
chain,
- &)-I
= N exists
and is non-negative
; hen
will fulfill our
requirement, and hence can satisfy
arbitrary
demands.
7~ = -yN is the desired solution.
Otherwise
( I -&)-I,
which gives t
These resultsmean
can number
be summarized:
arestates
industries
depend
of times If
in t.here
various
beforewhich
reaching
so, would ha
on no profitable
industry,
then
these
cannot
fulfill
an
outside
demand,
to have infinite entries, i.e. cannot exist.
and neither can
any
depending
on them. of
The
There
is aindustry
simple economic
interpretation
thisremaining
result. We see th
industries can
any
outside
demand.
In terms
of states
fromfulfill
every
state
i t must
be possible
to "reach"
the this
banking sta
means that Only
any ergodic
state
(other
than
so),
or
any
transient
state A profitle
a
profitable
industry
"reaches"
the
bank
directly.
- a state
from which industry
such
reachable,
can fulfill
no demand.
mustis reach
the bank
through
a profitable one. Hence o
To find what
industries
can
support
a
demand,
we use
following
be the
either
projitable ov mu
condition states that every indz~strymust
simple algoridepend
thm foronthe
classific&tioL
of
st&tes
:
a profitable industry. For example, if we assume that eve
industry depends on labor, and that labor is a profitable indust
(a) Make a check opposite each row of Q whose row-sum is less than
(which presumably means that labor is paid more than subsisten
I ; that is, check each row corresponding to a profitable industry.
wages), then our condition will be met, and all demands can be f
(b) Check the columns having tl:\e same indices as the rows already
filled.
marked and then check, in these columns, rows which have
the above condition is. violated, then the economy cannot fulfill
positiveIfentries.
possible demands. Let us ask what kinds of demands i t can fulf
(c) Iterate (b) until it produces no new rows. Then one of two
First of all we consider the case that there is no profitable indust
possibilities may occur:
This means that each industry needs all it produces to pay for r
(I) materials,
All rows are
andchecked.
i t would seem that i t could meet no outside deman
(2) That
Not all
are checked.
thisrows
is indeed
the case is easily proved.
If there is no profitable industry, then each row of Q has row sum
If case (cl)
occurs
Markov chain
hence
Q t =then
5. Iftheweassociated
multiply equation
(4) byist absorbing
on the right, we fi
with the single
absorbing
state
So.
Hence
any
non-negative
demand
that
can be met.. If case (c2) occurs, then the rows which are not checked
correspond to the maximal profitless closed group. 'VI' can find all
states depending
these
by marking
(removing
previousis 0. Hen
hence on =
0. This
says t h athese
t the rows
sum of
all the demands
check marks)
applying
(b) repeatedly.
Any state so marked will
noand
(positive)
demand
can be fulfilled.
not be able to fulfill
outside
demand.
Let usannow
consider
the general case where our condition is violat
We thus see
that
the
entire
question
outside
demands
The associated Markov
chainofis what
not an
absorbing
chaincan
with the sin
be met by the
economy
is settled
a very
absorbing
state.
Then by
there
mustsimple
be a nalgorithm.
ergodic set Th"
other than {
"computation"
findingofthe
row-sums,
then simple
itera- and wh
i.e. arequires
closed group
industries
noneand
of which
is profitable,
tions in which
onlyonthe
of components
is checked.
Thisthe set of
depend
no positivity
industry outside
the group.
Let us take
algorithm issuch
practical
even for
vcry
matrices.
industries,
that
is large
the union
of all ergodic sets other than {
The industries
which canQ meet
no industries
outside demand
a totally
The submatrix
of these
has theform
property
Qt= [, as abo
useless segment
the can
economy.
From
here demand.
on we willThus
assume
and of
hence
fulfill no
outside
the that
entire econo
they have been
(J _Q)-l exists.
fulfill nothen
demands
of goods produced by these industries. A
can deleted;
Next we want to raise the question: If an order for one dollar's
worth of good is given by a customer to industry <, how much of it
ends up in the hands of the various industries?
SEC. 7
204
FINITE MARKOV CHAINS
CHAP. V
First of all we must ask what the total demand is on the vario
MARKOV
CHAINS
CHAP. is
VII
We have
a y vector
whose i-th component
1, and whi
industries. FINITE
has 0's as other components. Hence n= yAT is simply the i-th row
First of all we must ask what the total demand is on the various
N . This gives a direct interpretation to the entries of fl; ncj is t
industries. We have a y vector whose i-th component is I, and which
amount industry j must produce to fill a dollar order for industry
has O's as other components. Hence 'Tr = yN is simply the i-th row of
Since industry j makes pjo profit on a unit production, the answer
N. This gives a direct interpretation to the entries of N; ntj is the
Our question is that if industry i is given a dollar, the profit of indus
amount industry j must produce to fill a dollar order for industry i.
j will be ntjpjo.
Since industry j makes PiO profit on a unit production, the answer to
of all the
x n t r p jthe
o= bto
= l (since so is the on
Our question isThe
thatsum
if industry
i isprofits
given ais dollar,
profit
of industry
204
1
j will be nt1PiO.
absorbing state).
This shows that a dollar spent by the consumer en
The sum up
of as
all profit
the profits
2,niiPJO
= biG
= 1 (since Soindustries.
is the only
in the ishands
of the
profit-making
A related question J is the following: If a dollar order is given
a,bsorbing state). This shows that a dollar spent by the consumer ends
~ndustryi, how much activity does this result i n ? It wiil result in
up as profit in the hands of the profit-making industries.
units of production in industry j. The sum of these is t i , the i
A related question is the followip.g: If a dollar order is given to
component of T . This is normally much greater than 1. For an ord
industry i, how much activity does this result in? It :will. result in nti
y, the total production is -yN[= yr.
units of production
industry
The sumSuppose
of thesethat
is t"thethetechnological
i-th
Let us in
consider
an j.example.
component of T. This is normally much greater than 1. For ~n order
efficients for six industries are given by
y, the total production is yNt=YT.
.'
Let us consider an example. Suppose that the technological coefficients for six industries are given by
1/4
1/4 1/4 1/4
1/2 0 1/2
1/2
Q
then
0
0
0
0
0
0
0
0
1/ 4
0
0
0
0
0
0
0
0
0
0
1/4
3/ 4
0
0
0
0
1/4
1/4
then
0
p=
0
0
1/4 1/2 0 1/4
1/4 1/4 1/4 1/4
0
1/2 0 1/2
0
0
0
80
0
0
0
SI
0
0
0
82
0
0
0
83·
0
()
0
0
1/4 3/4
0
84
0
0
0
0
1
0
85
0
From the first column we see that sl, sz, and ss are the profitable ind
1/4 0 1/4 0 1(4 0 1/4 by86Figure 7-9.
tries. The classification of states is given
Here (s4, s5) is the ergodic set of industries which are profitless a
From the first column we see that SI, 52, and 86 are the profitable industries. The classification of states is given by Figure 7-9.
Here {S4, 55} is the ergodic set of industries which are profitless and
204
FINITE MARKOV CHAINS
CHAP. V
First of all we must ask what the total demand is on the vario
We have OF
a y MARKOV
vector whose
i-th component is2051, and whi
industries.
APPLICATIONS
CHAINS
has 0's as other components. Hence n= yAT is simply the i-th row
do not depend
profitable
Industry 86tois the
profitable,
it ncj is t
gives aindustries.
direct interpretation
entries but
of fl;
N . onThis
depends or~amount
the former.
Enloe
s.j, produce
85, 86 are
useless,
andorder
may for
be industry
industry
j must
to fill
a dollar
Since industry j makes pjo profit on a unit production, the answer
Our question is that if industry i is given a dollar, the profit of indus
j will be ntjpjo.
The sum of all the profits is x n t r p j o= bto = l (since so is the on
SEO. 7
1
absorbing state). This shows that a dollar spent by the consumer en
up as profit in the hands of the profit-making industries.
A related question is the following: If a dollar order is given
~ndustryi, how much activity does this result i n ? It wiil result in
units of production in industry j. The sum of these is t i , the i
FIGURE
7·9
is normally
much greater than 1. For an ord
component of T . This
y, the total production is -yN[= yr.
deleted. Industry S3 is profitless but not useless. The deleted transiLet us consider an example. Suppose that the technological
tion matrix is
efficients for six industries are given by
0
0
1/2
{)
p~C' If
1/ 4
4,
{)
1/2
0\
So
';')"
1/ 4 1/4
52
0
83
1/2
Hence
then
Thus, fOT eX<1mple, a unit order to 82 will stimulate a total of 6 dollars'
production, 8/3 from S1, 4/3 from 82, and 2 from S3. On this dollar
order 81 makes a profit of 8/ 3 .1/ 4 =2/3, S2 makes 4/ 3 ,1/ 4 =1/ 3 , and
S3 ma,kes 2.0 = 0 (S3 is profitless).
If we place an outside demand of y= (1,3,2) on the econom.y, then
yN = (20,4,16) units will have to be produced by the various industries.
The total production is worth y-r = 40 dollars.
We note that Pis lumpable into the partition [{so}, {SI' S2}, {53}]' The
lumped process yields
(;, ,;,) N=(: :)
0
From the first column we see that sl, sz, and ss are the profitable ind
F=
tries. The1/2
classification of states is given
by(:).
Figure 7-9.
7 =
Here
(s4,
s5)
is
the
ergodic
set
of
industries
which are profitless a
\ 0 1/2 1/2
For
y = (4,2)
yS = (24,16)
yT = 40.
The "lumped" industries sl and sz are here considered as acting a
"industrial FI':\ITE
group." :\IARKOV
This process
will yield the total demands o
CHAINS
industrial group, but not the breakdown into industrial demands.
The "lumped"
industries
51 and 82 are here considered as acting as one
For
practical
problems the comput,ations may be prohib
"industrialHence
group."
Thisoften
process
willtoyield
thea lumped
total demands
we are
happy
solve
version on
of the
the econ
industrial group,
but
not
the
breakdown
into
industrial
demands.
The condition for lumpability becomes the following : Any ind
For practical
problems group
the computations
may
in one industriaJ
makes the sa,me
totalbeperprohibitive.
unit demands o
Hence we are
often
happy
to
solve
a
lumped
version
of the economy.
members of (its own or) another industrial group.
(Then all indu
The condition
forgroup
lumpability
becomes
theunit
following:
industry
in the
ma,ke the
same per
profit.) Any
While
these cond
in one industrial
group
makes
the
same
total
per
unit
demands
theappro
are unlikely to be met exactly, they may be put to a on
good
members oftion.
(its own
another
group. (Then
Thisor)allows
us industrial
to take industrial
groupsallasindustries
basic entities
in the group
make
the same
unitmanageable
profit.) While
these conditions
yields
a smaller
andper
more
model.
are unlikely to
be
met
exactly,
they
may
be
put
to
a
good
approximaWe thus see that even in a non-probabilistic model
a grea,t de
tion. Thisinformation
allows us tocan
take
industrialfrom
groups
as basic
entities,
be obtained
Markov
chain
theory.and
yields a smaller and more manageable model.
We thus see that even in a non-probabilistic model a great deal of
information can be obtained frorn Markov chain theory.
206
Page
25
The "lumped" industries sl and sz are here considered as acting a
"industrial group." This process will yield the total demands o
industrial group, but not the breakdown into industrial demands.
For practical problems the comput,ations may be prohib
Hence we are often happy to solve a lumped version of the econ
The condition APPENDICES
for lumpability becomes the following : Any ind
in one industriaJ group makes the sa,me total per unit demands o
members of (its own or) another industrial group. (Then all indu
in the group ma,ke the same per unit profit.) While these cond
are unlikely to be met exactly, they may be put to a good appro
I-Summary
of Basic
Notation groups as basic entities
tion. This
allows us to
take industrial
yields a smaller and more manageable model.
Mi[f], Varl[f],
denote
value of the model
function
f,
We thusPri[P]
see that
eventhe
in mean
a non-probabilistic
a grea,t
de
variance
of f, andcan
probability
of the
statement
when theory.
the chain
information
be obtained
from
Markovp chain
is started in state St.
19
R = {rij}
matrix with entries rij
19
p=
row vector with component rj
column vector ·with component Cj
19
y=
20
g column vector with all entries 1
20
'7
row vector with all entries
20
E
matrix with all entries
20
I
identity matrix
20
0
matrix with all entries 0
46
(I,) -
._ fl jf i=j
lO jf i of j
21 AT is the transpose of A
21 Asq rcsults from A by squaring each entry
21
A dg results from A by setting off·diagonal entries equal to 0
II-Basic Definitions
25
A finite },f arkov chain is a stochastic process which moves through
a finite number of states, ar,cl for which the probability of entering
a certain state depends
on the last state occupied.
35
An ergodic set of states is a set in which every state can be reached
from every other state, and which cannot be left once it is entered.
35
A transient set of states is a set in which every state can be reached
from every other state, and which can be left,.
35
An ergodic state is an element of an ergodic set.
207
APPENDICES
208
Page
35
A transient state is an element of a transient set.
35
An absorbing state is a state which onoe entered is never left.
37
An absorbing chain is one all of whose ergodic states are absorbing;
or-equivalently-which has at least.one absorbing state, and
such that an absorbing state can be reached from every state.
37
An ergodic chain is one whose states form a single ergodic set; orequivalently-a chain in which it is possible to go from every
state to every other state.
37
A eyclic chain is an ergodic chain in which each state can only be
entered at certa,in periodic intervals.
37
A regular chm:n is an ergodic chain that is not cyclic.
III-Basic Quantities for Absorbing Chains
STANDARD FORM FOR TRANSITION MATRIX
, 1 to)
p=(-RIQ
44
46
nj
46
Uk}
50
number of times in state Sj before absorption
function that is 1 ifprocess is in 5j after k steps, 0 otherwise
number of steps taken before absorption
52
bij probability starting in state SI that the process is absorbed
in state Sj
61
fj
number of times in transient state SI before leaving the state
61
h ij
probability starting in Sj that the process is ever in Sj
61
m
total number of transient states entered before absorption
46
N = {M,[nj}}
49
N2=(Yart[njJ}
49
T={M;[t]}
49
T2={Vari(tJ)
49
B = {bi}}
62
f-L = {M;[m]}
63
P
transition matrix for process of changes of state
APPENDICES
209
IV-Basic Formulas for Absorbing Chains
Page
46
N = (I _Q)-1
(Fundamentai matrix)
49 N 2 = N( 2N dg- I )-Nsq
49
B=NR
61
H=[N-IjNdg- 1
49
T=Nf;
49 T2=(2N-l)T-Tsq
1
til
M.[ri)=-,- 1 -Pfi
61
Var;[riJ=(-,~'
62
JL = (N Ndg-l)~
PH
"-f2
i-PH
V-Basic Quantities for Ergodic Chains
76
a = fal}
88
eij
fixed probability vector for P
.
1
=nhm
-:- COVk[Yi(nl,Yj(n)]
r:o 12,
-+
88 (3= {Dj} vector of limiting variances for the number of times in each
of the states a J (2z jj -aj) [see p. 91]
70
A
matrix with each row a
79
D
diagonal matrix with j-th entry l/aj
105
P
t,:ansition matrix for the re\'erse process
75
Z=(l-P+A)-l
81
c= LZii
CFundamentalmatrix)
j
78
111 = {mij} matrix of mean number of steps required to reach So
for the first time, starting in St
82
W = {'Wij} matrix of variances for the number of steps required
to reach Sf, starting in state SI
76
Yj(n)
number of times in state Sj in the first n steps
APPENDICES
21 0
APPENDICES
Formulas for Ergodic Chains
210
----VI-Basic
Page
105 P=DPTD-l
VI-Basic Formulas for Ergodic Chains
Page
105
P=DPTD-l
78
M = {mii) = (1 - Z+EZdg)D
81
M=M-Mdg
83
W = {WiJ} = M(2ZdgD-I) + 2(ZM -E(ZJ1)dd
85
C = {eij} = {a/Zij + ajZji - aid;j - aiaj}
79
1
mii=aj
81
aM =7)ZdgD
81
MaT=cf
81
a=(c-l)[(M)-lfF
81
P=I+(D-E)(Jl)-l
VIE-Some
Bask Exam
Page Example
47 l a
Page Example
47 1a
90 3a.
103 2a
90 30,
108 3a
48
IOU
108 3a
48 lOa
102 11a
4.2 13
42 14
42 14
VII-Some Basic Examples
APPENDICES
21 0
2H
APPENDICES
VI-Basic Formulas for Ergodic Chains
-~--
Page
105 P=DPTD-l
VHI-Generalization of a Fundamental Matrix
John C. Kemeny
ABSTRACT
It is shovn; that, lor a finite ergodic Markov chain, basic descriptive quantities,
such as the stationary vector and mean first-passage matrix, may be calculated using
allY one of a class of fundamental matrices. New applications of the use of these
cperators are discussed.
INTRomXTIO:,\
VIE-Some
Bask Exam
Page Example
47 l a for this paper is to generalize the concept of the fundaThe motivation
mental matrix of a finite ergodic Markov chain (see [4]). The approach will be
to consider a class of linear equations, of which the Markov chain case is a
90 be
3a. shown that one gains new insight into previous work on
subclass. It ~ill
ergodic chains,
108 and
3a that the generalized operator has interesting new applications.
48 IOU
\Ve assume that we are dealing with an n-dimensional vector space, for
some fixed n> 1. Unless otherwise indicated, in our formulas capital letters
will denote n-by-n matrices, lowercase letters denote n-component column
vectors, and Greek letters denote n-component TO\V vectors.
42 14
A CLASS OF LINEAE EQUATIONS
A standard approach to the computation of key quantities for finite
Markov chains is to compare the value of an unknown quantity x with its
expected value after one step. TI,is leads to an equation of the fonn
x=Px
(1)
212
where P is the transition
matrix and the components of f are kno
APPENDICES
quantities. Let us write this class of equations in a standard form:
wheI'e P is the transition matrix and the components of f are known
quantities. Let us write this class of equations in a standard form:
(I-p)x=f.
(2)
We wish to consider equations of this form in general, using the speciai
of Markov chains only as motivation. It is well known that the nature of
solutions of (2) is determined by studying the homogeneous equation
We wish to consider equations of this form in general, using the special case
of Markov chains only as motivation. It is well known that the nature of the
solutions of (2) is determined by studying the homogeneous equation
( I - P)x=O.
(I-p)x=O.
If this equation has only the solution x=O, then (2) has a unique solution
there is a nonzero solution of (3),then (2) has either no solution or infini
many solutions, depending on the nature off.
If this equation The
has only
the solution
thentrivial
(2) has
a unique
solution:
if Then I
case where
(3) hasx=O,
only the
solution
is the
easy case.
there is a nonzero
solution
of
(3),
then
(2)
has
either
no
solution
or
infinitely
is nonsingular, and the solution of (2) is x = ( d - P ) - ' f . This is the case f
many solutions,
the natureMarkov
of f. chain, where P is the transition ma
finitedepending
transient on
(absorbing)
The caserestricted
where (3) to
hastransient
only thestates.
trivial Therefore,
solution is the
case. Thenof1P quanti
theeasy
computation
basic
is nonsingular,
and
the
solution
of
(2)
is
x:=
(I
P)
-1 f. This is the case for a
for such chains is simple (see [4], Chapter 3).
finite transient The
(absorbing)
chain, where
the transition
matrix Specifica
case we Markov
wish to consider
is one Piniswhich
I - P is singular.
transient
states.that
Therefore,
computation kernel.
of basic(This
quantities
restricted towe
shall assume
it has a the
on~-dimensional
is always the c
(see chains.)
[4J, Chapter 3).
for such chains
is simple
for finite
ergodic
The case we wish to consider is one in which 1- P is singular. Specifically,
we shall assume that it has a one-dimensional kernel. (This is always the case
ASSUMPTION.The hon~ogeneous equation (3) has a nonzero solu
for finite ergodic
x = h ,chains.)
and every solution of (3) is a multiple of h.
ASSUMPTION.
The homogeneous equation (3) has a nonzero solution
It should be noted that h is a fixed point of the transformation P,
x=h, and every solution of (3)is a multiple of h.
P h = h . From linear algebra we know that there is also a fixed point o
acting on row vectors, aP=a. From now on we shall use h and a: for th
It shouldfixed
be noted
h is abe
fixed
point of the
P, i.e.,only up
points.that
It should
remembered
thattransfo.rmation
they are determined
Ph == h. From
linear algebra
thatofthere
is also chain,
a fixedlz point
of chosen
P
constant
mrdtiple.weInknow
the case
an ergodic
may be
as
acting on row
vectors,
Ci. From now on we shall use h and DC for thcse
vector
dl ofaP=
whose
components are I, and a as the all-positive probabi
It should
be remembered
that they arc detcrmined only up to a
fixed points.vector
of limiting
probabilities.
constant multiple.
the next
case demonstrate
of an ergodicthat
chain.
h may
chosenmodification
as the
We Inshall
there
is abesimple
of
whoseI-P
components
1, and a and
as the
aD-positive
vector a!.I ofmatrix
which is are
nonsingular,
which
may be probability
used to solve (2). T
probabilities.
vector of limiting
modification
allows us .to choose row and column vectors L,3 and g alm
We shallarbitrarily,
next demonstrate
that deal
thereofisflexibility
a simpletomodification
of the The rea
giving a great
the generalization.
matrix I - P which is nonsingular, and which may be used to solve (2). The
modification allows us· to choose row and column vectors (3 and g almost
arbitrarily, giving a great deal of flexibility to the generalization. The reader
where P is the transition
matrix _
and_the
APPENDICES
_ components of f are kno
quantities. Let us write this class of equations in a standard form:
should keep in mind that while the product f3g is a number, the product g{3 is
an n-by-n matrix.
THEOREM 1.
Let f3 and g be any two 1:ectors such that f3h and ag are
nonzero. 71ten the inverse
We wish to consider equations of this form in general, using the speciai
of Markov chains only as motivation. It is well known that the nature of
Z== (I_P+gj3)-l
(4)
solutions of (2) is determined by studying the homogeneous equation
exists.
( I - P)x=O.
Proof.
Sllppose that
(5)solution
If this equation has (I-P+g/3)x=O.
only the solution x=O, then (2) has a unique
there is a nonzero solution of (3),then (2) has either no solution or infini
many solutions, depending on the nature off.
Then
The case where (3) has only the trivial solution is the easy case. Then I
is nonsingular, and the
of (2) is x = ( d - P ) - ' f . This is the
x==solution
Px- g{f3x).
(6) case f
finite transient (absorbing) Markov chain, where P is the transition ma
restricted to transient states. Therefore, the computation of basic quanti
We multiply the equation by C/, and use the fact that aP= 0:. This yields
for such chains is simple (see [4], Chapter 3).
(ag)(f3x)=O. But ag~O; hence {3x=O. Thus (6) reduces to x=Px. By our
The case we wish to consider is one in which I - P is singular. Specifica
Assumption, x=ch, where c is a constant. Then O=f3x=c([3h), and [3h=l=O;
shall assume
has solution
a on~-dimensional
is always
the c
hence c ==we
O. Thus
(.5) has that
onlyit the
x= 0, and kernel.
hence (This
the matrix
is
for
finite
ergodic
chains.)
nonsingular.
l1li
Let us nextASSUMPTION.
derive some properties
of Z. Fromequation
(4),
The hon~ogeneous
(3) has a nonzero solu
x = h , and every solution of (3) is a multiple of h.
(7)
z(I-p+gf3)=l.
It should be noted that h is a fixed point of the transformation P,
P h = h . From linear algebra we know that there is also a fixed point o
actingthis
on equation
row vectors,
aP=a.
wePh=h,
shall use
and a: for th
if we multiply
on the
rightFrom
by /1,now
andonuse
wehobtain
fixedor
points. It should be remembered that they are determined only up
(Zg)(j3h)=h,
constant mrdtiple. In the case of an ergodic chain, lz may be chosen as
vector d
l of whose components are I, and a as the all-positive probabi
vector of limiting probabilities.
We shall next demonstrate that there is a simple modification of
matrix I-P which is nonsingular, and which may be used to solve (2). T
allowsusing
us .tothat
choose
and
column
In exactlymodification
the same manner,
Z is arow
right
inverse,
wevectors
obtain L,3 and g alm
arbitrarily, giving a great deal of flexibility to the generalization. The rea
1
f,Z= -ex.
ag
214
214
APPENDICES
--
Thus from Z we can obtain the fixed points h and a. Substituting (8) in
APPENDICES
we find that
Thus from Z we can obtain the fixed points h and iX, Substituting (8) into (7),
we find that
1
and dually, Z(I-P)=l- -hB
{3h ' ,
(10)
and dually,
1
(11)
We are (I-p)Z
now ready[-to-gao
agdemonstrate the usefulness of Z in solvi
equation (2).
We are now ready to demonstrate the usefulness of Z in solving the
THEOREM
2. The equation (2) has a solution i f and only if a f=O
equation (2).
solution exists, one may specify the value of px arbitrarily, say px=
constant ), and one obtains the unique solution
THEOREM 2.
The equation (2) has a solution if and only if o:f=O. If a
solution exists, one may specify the value of f3x arbitrarily, say f3x = c (c a
constant), and one obtains the unique solution
c
(12)
x=Zf+ f3h h.
Proof. Multiplying ( 2 ) by a shows that af=O is a consequence;
this is a necessary condition for the existence of a solution. Let u
multiply ( 2 ) by Z and make use of (10). We obtain
Proof. Multiplying (2) by 0: shows that a. f= 0 is a consequence; hence
this is. a necessary condition for the existence of a solution. Let us next
multiply (2) by Z and make use of (10). We obtain
x- ;h h(f3x)=Zf.
the condition p x = c , then we see that (12) is th
And if we impose
possible solution. Conversely, if we substitute (12) into ( 2 ) and use 41
find that x satisfies the equation. (Recall that af=O and Ph=h.) And
And if we impose the condition f3x::::c, then we see that (12) is the only
multiply (12) by 8, and use ($9, we verify that the solution also satisf
possible solution. Conversely, if we substitute (12) into (2) and use (11), we
condition px =c.
find that x satisfies the equation. (Recall that af=O and Ph=h.) And if we
We have
thusweshown
andsatisfies
sufficient
multiply (12) by f3 and
use (9),
verifyboth
that the
the necessary
solution also
thecondition
iIIII be requ
condition f3x=c. existence of a solution and precisely how much additional may
a solution. Since p may be any row vector such that j3hf 0, we have
We have thus shown both the necessary and sufficient condition for the
flexible tool. And the choice of P in Z is naturally determined by the na
existence of a solution and precisely how much additional may be required of
the side condition px= c.
a solution. Since f3 may be any row vector such that f3 h =1= 0, we have a very
There is a dual result which shows the role of g. Its proof exactly p
flexible too!. And the choice of f3 in Z is naturally detennined by the nature of
the proof of Theorem 2:
the side condition f3x=c.
There is a dual result which shows the role of g. Its proof exactly parallels
the proof of Theorem 2:
214
APPENDICES
--
Thus from Z we can obtain the fixed points h and a. Substituting (8) in
APPENDICES
215
we find that
THEOREM 3.
The equation
~U-p)=cf>
(13)
and dually,
has a solution if ami only if 4>h=O. If a solution exists, one may specify the
wlue ~g arbitrarily, say ~g=:c, and one obtains the unique solution
c
~=<I>Z+-a.
We are now ready
to demonstrate
the usefulness of Z (14)
in solvi
ag
equation (2).
THEOREM
2. The equation (2) has a solution i f and only if a f=O
A SPECIAL CASE
solution exists, one may specify the value of px arbitrarily, say px=
constant ), and one obtains the unique solution
Assume that ah 7'= O. Then we may choose them so that ah = 1. Let f3 =a
and g=h. All our conditions are met, and {3h=ag= 1. For an ergodic Markov
chain the resuiting operator
(15)
O/h= 1,
Proof. Multiplying ( 2 ) by a shows that af=O is a consequence;
this is a necessary condition for the existence of a solution. Let u
is called the
"fundamental
Foruse
it (8)
and (9)
the simple form
( 2 ) by Zmatrix."
and make
of (10).
Wetake
obtain
multiply
Z*h=h
and
aZ*=a,
(16)
and (12) and (14) take the forms
And if we impose the condition p x = c , then we see that (12) is th
possible x=Z*J+ch
solution. Conversely,
if we and
substitute
use 41
(17)
if af=O
(){x=c.(12) into ( 2 ) and
find that x satisfies the equation. (Recall that af=O and Ph=h.) And
multiply (12) by 8, and use ($9, we verify that the solution also satisf
(18)
if cph=O and ~h=c.
~=cf>Z*+c(){
condition px =c.
We have thus shown both the necessary and sufficient condition
We shall existence
next showof that
any Zand
may
be expressed
in additional
temlS of may
Z*, and
a solution
precisely
how much
be requ
conversely.a It
will be Since
cor:venient
constants
solution.
p maytobeintroduce
any row the
vector
such that j3hf 0, we have
flexible tool. And the choice of P in Z is naturally determined by the na
the side condition px= c.
There is a dual result which shows the role of g. Its proof exactly p
the proof of Theorem 2:
From (10),
Z(I - P+ha) =1 - c2 h{3 + (Zh )£'1.;
APl'EKZ)ICES
216
216
hence, multiplying
by Z*,
APPENDICES
hence, multiplying by Z*,
Z=Z*-c2fz(/3Z*)+(Zh)a.
Similarly,
Z=Z"'-c2 h((3Z*) +(Zh )a.
Z*(I-P+gP)=6-
Similarly,
(19)
ha+(Z*g)P,
Z* = Z - h ( a Z ) + c , ( Z * g ) a .
Z*(I- P+g(3)=I-ha+ (Z*g)(3,
Substituting (20)
into (19),
Z" =Z-h(aZ)+c1(Z*g)a.
Substituting (20) into (19),
h [ o Z + c , P ~ * ]= [ ~ h + c , ~a*. ~ ]
Multiplying
the right by g ,
h[aZ+con
2 (3Z*] = [Zh+c1Z*g] a.
Multiplying on the right by g,
(20)
(21)
1
c,c,h=-Zh+Z*g.
c1
We solve this for Z h and substitute in (19):
z*-c,~(Pz*)-c~(z*~)L~+c~~c
We solve this for Zh and substitutezin= (19):
This expresses Z in tenns of Z*. If we multiply (21) by(22)
h and solv
we have from (20)
This expresses Z in terms of Z*. If we multiply (21) by h and solve for Z*g,
we have from (20)
ha,Computing a Z h from
(23)(221, we
whichZ*=Z-h(aZ)
expresses Z* -(Zh)a+c
in tenns of 4Z.
useful identity:
which expresses Z* in tenns of Z. Computing aZh from (22), we obtain a
useful identity:
A numerical example may be helpful at this stage: (24)
A numerical example may be helpful at this stage:
a=(2,1).
This satisfies all our conditions, including a h = 1. We find that
This satisfies all our conditions, including ah =1. We find that Z* is the
APl'EKZ)ICES
216
hence, multiplying by Z*,
APPENDICES
Z=Z*-c2fz(/3Z*)+(Zh)a.
217
identity matrix. (This will always be the case when P= hex.) Suppose that we
choose
Similarly,
1 \
(
/3=(1,0), Z * ( Ig=
- P + gi)'
P)=6-
ha+(Z*g)P,
Z* = Z - h ( a Z ) + c , ( Z * g ) a .
Then
2
*),
Z= into
3 (19),
Substituting (20)
, -1 0.
h [ o Z + c , P ~ * ]= [ ~ h + c , ~a*. ~ ]
(
Multiplying on the right by g ,
and the previous results are easily verified. We wish to solve the equation
(I-P)X=(
1
c,c,h=-Zh+Z*g.
c1
_!).
We solve
for Z h and
in (19f3x,
): i.e., the first
Since af=O, it does
havethis
solutions.
VIe substitute
may specify
component of x. If we require f3x=3,
z=z*-c,~(Pz*)-c~(z*~)L~+c~~c
This expresses Z in tenns of Z*. If we multiply (21) by h and solv
from (20)
ourhave
requirements.
which s<ltisfies all we
It should be pointed out that, while Z * always exists for an ergodic chain,
there are cases w:here Theorem 1 applies but Z* does not exist. A simple
example will illustrate this: let
which expresses Z* in tenns of Z. Computing a Z h from (221, we
useful ~identity:
~), 0:=(1,0),
p=(
All the assumptions of the first section are met, and hence the matlices (4)
exist and have the stated properties. But, since ah:::: 0, Z* is not one of them.
Indeed, 1-- P+ ha == 0,Aand
certainly
does not
have
inverse.
Thisstage:
shows that
numerical
example
may
bean
helpful
at this
the method of this paper provides not only added flexibility but also wider
applicability than Z*.
ERGODIC CHAINS
This satisfies all our conditions, including a h = 1. We find that
A finite Markov chain is ergodic if from any state it is possible to reach
every other state. If P is the transition matrix of such a chain, and h is the
constant vector (allAPPENDICES
components equal to I), then the Assumption of
paper is always satisfied. Such a Markov chain has an equilibrium, i.
probability vector a such that a P = a . And a is strictly positive. Thus ther
constant vector (all components equal to I), then the Assumption of this
natural fixed points h and a , and ah = 1.
paper is always satisfied. Such a Markov chain has an equilibrium, i.e., a
The matrix Z * of the previous section is called the fundamental mntr
probability vector a such that aP= a. And a is strictly positive. Thus there are
the ergodic chain, and it can be shown that the various interesting proba
natural fixed pOints h and a, and ah = l.
tic quantities can be computed in terms of a and Z*. I n particular, Eqs.
The matrix Z* of the previous section is called the fundamental matrix of
(17), and (18) are well known results about finite ergodic chains (See
the ergodic chain, and it can be shown that the various interesting probabilisTwo other important results are that the mean time to return to a stat
tic quantities can be computed in terms of a and Z*. In particular, Eqs. (16),
(17), and (18) are well known results about finite ergodic chains (See [4].)
Two other important results are that the mean time to return to a state i is
218
is
and the mean time to go from i to j (mean first-passage time)(25)
and the mean time to go from i to i (mean first-passage time) is
To have a numerical illustration available, we introduce the(26)
weather i
Land of Oz (see [3]):
To have a numerical illustration available, we introduce the weather in the
Land of Oz (see [3]):
t
1
1
~
0
1
1
1
4
~
Z*= _1 (
86
6
15
-14
3
63
p=
4
4
4
,
2
3
h=( U,
-14 )
6 ,
86
_(2'5,"'5,'5
1 2) )
a-
M~(:
III
3
4
III
0
§.
4
0
3
3
This treatment of finite
ergodic chains has never seemed as satisfactory a
m=(~,5,n
treatment of finite transient chains. For the latter ( I - Q ) - ' is the na
fundamenltal matrix, where Q represents transitions from transient sta
This treatment of finite ergodic chains has never seemed as satisfactory as the
transient state, and all quantities can be expressed in terms of it. For erg
treatment of finite transient chains. For the latter (1- Q) -1 is the natural
chains Z * appears somewhat arbitrary. It also suffers from the difficulty
fundamentalone
matrix,
Q represents
transitions from transient state to
mustwhere
compute
a (solving n equations) before one can compute
transient state,
and
all
quantities
can
be
expressed
terms
of it. For
Variom alternatives to this matrix haveinsince
appeared
inergodic
the literature,
chains Z* appears somewhat arbitrary. It also suffers from the difficulty that
Meyer [7, 81).
one must compute a (solving n equations) before one can compute Z*.
Various alternatives to this matrix have since appeared in the literature, (see
Meyer [7, 8]).
constant vector (allAPPENDICES
components equal to I), then the Assumption
of
210
paper is always satisfied. Such a Markov chain has an equilibrium, i.
We nowprobability
know thatvector
Z* isa only
of possible
such one
that aofP =ana . infinite
And a isnumber
strictly positive.
Thus ther
choices for natural
the fundamental
any Z, the only
fixed pointsmatrix.
h and aOne
, and may
ah = choose
1.
restrictions are that
ag andZf3h
not be section
O. Thusisf3called
may the
be any
vector mntr
The matrix
* ofshould
the previous
fundamental
such that the
of the
components
zero,that
andthe
if various
g is chosen
as a proba
thesum
ergodic
chain,
and it canisbenot
shown
interesting
nonnegative tic
andquantities
nonzero vector,
ag~O-irrespective
of what
may be. Eqs.
can be then
computed
in terms of a and
Z*. I na particular,
(17), and
(18)approach
are well to
known
results about
finite
ergodic
chains (See
Let us propose
a new
the treatment
of finite
ergodic
chains.
Let
Two other important results are that the mean time to return to a stat
f3h= 1.
(27)
i to with
j (mean
time) is1
mean
to go
That is, we and
let the
g=h,
andtime
f3 be
anyfrom
vector
rowfirst-passage
sum 1. Theorem
guarantees its existence. Then, from (8) and (9),
Zph=h
and
f3Zp =a.
(28)
Thus we may find
froma the
fundamental
matrixavailable,
rather than
having to find
a
Toahave
numerical
illustration
we introduce
the weather
i
first. Other quantities
are(see
determined
from an equation of the form (1), with
[3]):
Land of Oz
af=O; we know from Theorem 2 that we may impose the additional
condition /3x=c and obtain the unique solution
x=ZfJf+ ch .
(29)
How are the mean first-passage times expressed in terms of the generalized fundamental matrix? Instead of retracing the derivation of the matlix
M, we use (23) to translate the formula (26):
And since h is a constant vector, h j -hi =0:
This treatment of finite ergodic chains has never seemed as satisfactory
a
(30)
treatment of finite transient chains. For the latter ( I - Q ) - ' is the na
fundamenltal matrix, where Q represents transitions from transient sta
state,(28)
andshows
all quantities
in terms
it. For erg
If we use as transient
Z a Zp, then
that the can
last be
twoexpressed
terms cancel.
Thusof the
chains
* appears
somewhat
arbitrary.
simple formula
(26)Zholds
for any
Zp in place
of Z*.It also suffers from the difficulty
mustmore
compute
n equations)
This can one
be seen
simplyaif (solving
we express
Zp in termsbefore
of Z*. one
Fromcan
(22),compute
alternatives
to this matrix have since appeared in the literature,
using g=h, Variom
f3h=ah=l,
and (16),
Meyer [7, 81).
Zp =Z*-h(/3Z*-a).
(.31)
APPXNDTCES
220
For the Oz example,
if we select3! , = ($. 0, b), then
AP}'ENDICES
220
For the Oz example, if we select f3 = (L 0,
16
15
Z{3-
0
-~
15
1
'-'
1
1
"5
then
,
15
0
16
Ts
It is easy to verify that Z p h = h and P Z p = a , and that M is given b
using Zp in place of Z*. We can also verify ( 3 1 ) by direct computation
Thus any Zp is a suitable fundamental matrix, and it can be com
It is casy to verify that Zflh = hand f3Zfl =0', and that M is given by (26)
without
knowing a .
using Zfl in place of Z*. We can also verify (31) by direct computation.
Thus any Zfl is a suitable fundamental matrix, and it can be computed
without knowing Q.
APPLICATIONS
APPLICA TIO.'\SConsider a Markov chain that has the property that when it moves
from a state i, it moves to any other state with the s a n e probability p
example, Oz has this property with p l = p 3 = and p, = $. Ef we
Consider a Markov chain that has the property that when it moves a\\ ay
g, = p , and p, = 1, then 1- P + g P is the diagonal matrix with entries np,
from a state i, it moves to any other state with the same probability P" For
Z is diagonal matrix with Z , , = l / ( n p , ) . From (9) we know tha
example, Oz has this property with p 1 ;=: P3 = i and P2 = ~. If vve choose
proportional
to PZ. Therefore,
g i = P, and f3i = 1, then 1- P+ gf3 is the diagonal matlix with entries np,. Thus
Z is diagonal matrix \vith Zii = l/(nPI)' From (9) we know that
proportional to f3Z. Therefore,
0'
is
(32)
and from ( 3 0 ) ,
and from (30),
1'.1il
= ~ ( 2: ~ _ 1 + ~).
(33)
n . k Pk
Pi
P, .
For Oz, B k l / p k =10. Thus, for example a , = 6 = $ and hl12= i f l o - - 4.
Next we consider the method used in [6]to compute a. The "recipe
replace the last column of 1-P by ones and invert; then a is the last
the inverse. This corresponds to choosing g, = 1 - ( I - P),. and ,!3=(0,. .
.'\ext we consider the method used in
to compute a. The "recipe" is to
The
inverse in question is the Z corresponding to this choice of g and
replace the last column of 1- P by ones and inw:rt; ther! (X is the last row of
a g = l . Hence from (9), a = p Z , which is the last row of Z. Meyer ['i]
s
the inverse. This corresponds to chOOSing g, = 1 -- (1- F), nand f3 = (0, ... ,0,1).
that this matrix could be used as a fundamental matrix in place of Z*.
The inverse in questioIl is the Z corresponding to this choice of g and /3, and
O'g= 1. Hence from (9), (X = (JZ, which is the last row of Z. \!eyer
showed
For Oz, Lkl/Pk = 10. Thus, for example Ct 1 = fn = ~ and JI 12 = t(l0--2+4)
-.!
-'":1 .
:7:
that this matrix could be used as a fundamental matrix in place of Z*.
APPXNDTCES
220
For the Oz example, if we select3! , = ($. 0, b), then
A_P_P_E_'XDIC_E--=-S_ _
221
A third interesting choice for g and f3 is the following. We choose as f3 the
first row of P. and
= 1 while the other components of g are O. Then Z has
the form
"I
I
1
0
0
o
Z=
It is easy to verify that Z p hIN
= h and P Z p = a , and that M is given b
using Zp in place1 of Z*. We can also verify ( 3 1 ) by direct computation
Thus any Zp is a suitable fundamental matrix, and it can be com
without
knowing a matrix
.
where 1 N is the fundamental
of the transient chain obtained by
making state 1 absorbing. The interpretation of 1N;j is the mean number of
times a process started in i steps into i before absorption. Here ag = ai' and
APPLICATIONS
hence (9) shows
that a=a 1Cj3Z). For the first component this is an identity,
but for i =i= 1 we obtain the identity
Consider a Markov chain that has the property that when it moves
probability p
from a state i, it moves to any other state with the s a n e (34)
example, Oz has this property with p l = p 3 = and p, = $. Ef we
g, = p , and p, = 1, then 1- P + g P is the diagonal matrix with entries np,
Also {or i i" 1 Z is diagonal matrix with Z , , = l / ( n p , ) . From (9) we know tha
proportional to PZ. Therefore,
(Zh) j =1 + L.J N17 = 1 + M ..
"1
tl.'
hence from (30)
and from ( 3 0 ) ,
This result is correct also when i or j is 1, if we IG( (as usuai) 1'\1 = 0 if i or i is
1, and Mu =0. Since we could have used in place of 1 a general state k, we
obtain the interesting
For Oz,identity
B k l / p k =10. Thus, for example a , = 6 = $ and hl12= i f l o - - 4.
Next we consider the method used in [6]to compute a. The "recipe
is the last
replace the last column of 1-P by ones and invert; then a(35)
the inverse. This corresponds to choosing g, = 1 - ( I - P),. and ,!3=(0,. .
The inverse in question is the Z corresponding to this choice of g and
l . Hence
fromon(9),
a =obtain,
p Z , which
the=last
s
Multiplying bya g
a i=and
summing
i we
usingis""lok
Ii krow
N i ;, of Z. Meyer ['i]
that this matrix could be used as a fundamental matrix in place of Z*.
(36)
222
APPENDICES
Since i does not occur
on the right side, the left side is a constant (inde
APPENDICES
dent of i). Similarly the right side is independent of k. We have previo
urged readers to try to find a probabilistic interpretation for this constant
Since i doessonot
right
side,Still
the another
left sideexpression
is a constant
faroccur
none on
hasthe
been
found.
for (indepenthis constant ma
dent of i). Similarly the right side is independent of k. We have previously
found from (30):
222
urged readers to try to find a probabilistic interpretation for this constant, but
so far none has been found. Still another
constant may be
const =expression
Miia, =for 7this
" i i -aZh.
found from (30);
x
2
i
i
(.37) for w
2: Mi/<j =known
2: Zl1 -aZh.
This result const=
was previously
for the special case Z=Z*,
a Z h = 1, but it hoIds for all our Z's. For Oz the constant is $j,as ca
computed from either Z* or Zp.
This result was previously known for the special case Z= Z*, lor which
Our final application is to the blarkov-process version of classical pote
aZh = 1, but it holds for all our Z' s. For Oz the constant is H, as can be
theory. Such a theory exists for both functions and measures (column and
computed from either Z* or Zp.
vectors in the finite case). A charge is a function f of total integral 0
Our final application is to the Markov-process version of classical potential
cif=O) or a measure of total measure 0 (i.e., Qh=O). Potentials satis
theory. Such a theory exists for both functions and measures (column and row
certain averaging property-they are solutions of (2) or (13), respectively
vectors in the finite case). A charge is a function f of total integral 0 (i.e.,
know that there always are such potentials for any charge, from Theore
af=O) or a measure of total measure 0 (i.e., ¢h=O). Potentials satisfy a
and 3, and that uniqueness requires an additional condition. The
certain averaging property-they are solutions of (2) or (13), respectively. We
conditions imposed have been a x = 0 and &=O. Hence Z* is suitable
know that there
always
are such
for any
fromand
Theorems
potential
operator
forpotentials
both functions
andcharge,
measures,
x=Z*f,2 <=QZ*
and 3, and that
uniqueness
requires
an
additional
condition.
The
usual condi
We can generalize this theory by imposing different boundary
conditions imposed
have been If(lX=O
and ~h=O.
Z* is$g=O,
suitable
as athe pote
on the potentials.
we require
that Hence
px=O and
then
potential operator
for
both
functions
and
measures,
and
x=Z*f,
~=</>Z*.
operator is the Z detem~inedby g and /3, and x=Zf, trQIZ.
We can generalize
this theory by
imposinghave
different
boundary
conditions
These generalized
potentials
an amusing
nonprobabilistic
ap
on the potentials. If we require that (3x=O and ~g=O, then the potential
tion. Consider n teams involved in a tournament. We wish to measur
operator is the Z deternlined by g and (3, and x=Zf, ~=¢Z.
relative strengths of the teams even though not every team has played e
These generalized potentials have an amusing nonprobabilistic applicaother team. Let s , =~ n u n ~ b e rof points by which team i beat team
tion. Consider
n teams
involved
in a tournament. We wish to measure the
if j won). We wish to assign point ratings to team
negative
number
relative strengths of the teams even though not every team has played every
Ideally one wishes that
other team. Let Si; = number of points hy which team i beat team i (a
negative number if i won). We wish to assign point ratings to teams, Xi'
Ideally one wishes that
for every game. But this is too much to expect; teams have "good
(38) days"
"bad days". What we will require is that for each team i, the sum o
differences s,, -(x, -x,) be zero. If team i has played t, games, this m
for every game. But this is too much to expect; teams have "good days" and
that
"bad days". What we will require is that for each team i, the sum of the
differences sii -(Xi -Xi) be zero. If team i has played ti games, this means
that
(39)us intro
where the sum is taken over all the opponents i has played. Let
where the sum is taken over all the opponents i has pJayed. Let us introduce
222
APPENDICES
APPENDICES_'
_ side,
_ _the
__
_is
_a_
22:3
Since i does not occur
on the right
left_
side
constant
(inde
dent of i). Similarly the right side is independent of k. We have previo
the matrix P defined
to havetoPi;
if ai has
played j, interpretation
and otherwise,
and constant
urged readers
try=toIlti
find
probabilistic
for this
the vector fWith
f =(l/ii)Lisij'
(39)Still
takes
on theexpression
form (1), for
Thethis
matrix
so far
none has beenThen
found.
another
constant ma
P is nonnegative
the evnstant veetor h as fixed point, but is it
foundand
fromhas(30):
ergodic? This will be the case if every team either has played any other team
or has played teams that played teams
that =
have
const (etc,)
= Miia,
2 7" played
i i -aZh.that team,
°
x
i ratings are
i not possible.
Clearly, without such a connection meaningful
The vector (X is defined by (Xi =tjt, where t=2:.ti , and afis the average
of the Si" which
sincewas
sii =
-sij' Thusknown
a solution
exists.case
Vie know
Thisis 0,
result
previously
for of
the(39)
special
Z=Z*, for w
that an extra condition
mayitbehoIds
imposed,
which
.is notFor
surprising,
in (38)
for all
our Z's.
Oz the since
constant
is $j,as ca
a Z h = 1, but
only the differences
of the
ratings
computed
from
eithermatter.
Z* or We
Zp. might decide to give one team
all final
otherapplication
teams relative
to it.blarkov-process
This would beversion
achieved
by
rating 0, and rateOur
is to the
of classical
pote
choosing a {3 theory.
\.\-ith 1 in
thata theory
component
0 otherwise.
wemeasures
might make
Such
existsand
for both
functionsOrand
(column and
the sum of the
ratings
to zero,
{3i =::
In either
vectors
in equal
the finite
case).choosing
A charge
is 1aIn.
function
f ofcase
totalwe
integral 0
have (3x=O ascif=O)
our condition
and compute
the ratings
are Qh=O).
then given
by
or a measure
of totalZ{l;
measure
0 (i.e.,
Potentials
satis
x= Zj3f. It is \\/orth
that property-they
the condition LYX=()
would beofquite
unna~ural.
certainnoting
averaging
are solutions
(2) or
(13), respectively
know that there always are such potentials for any charge, from Theore
and 3, and that uniqueness requires an additional condition. The
HISTORICALconditions
NOTES imposed have been a x = 0 and &=O. Hence Z* is suitable
potential operator for both functions and measures, and x=Z*f, <=QZ*
We can generalize this theory by imposing different boundary condi
The matrix Z* was introduced in [4) and has becn Widely used, Various
on
the potentials. If we require that px=O and $g=O, then the pote
alternatives have also been proposed. Hunter sho\ved [2J that Z* is a
operator
is the Z detem~inedby g and /3, and x=Zf, trQIZ.
"generalized inverse" of j-P, in the sense that
These generalized potentials have an amusing nonprobabilistic ap
tion. Consider
teams involved
to measur
(I -n P)Z*(I
-- P) = I - inP.a tournament. We wish
(40)
relative strengths of the teams even though not every team has played e
other team. Let s , =~ n u n ~ b e rof points by which team i beat team
if j fundamental
won). We wish
to assign
point renewal
ratings to team
negative
He also extended
the number
use of the
matrix
to Markov
processes.
Ideally one wishes that
It should be pointed out that all the matrices (4) are generalized inverses
of 1- P, as follows immediately from either (10) or (11).
The paper [9j compares alternative methods for calcu12.ting the vector Ci
on a computer.forThe
recommended
method
that
emerges
from this
work
is one
every
game. But this
is too
much
to expect;
teams
have
"good days"
chosenwe
as awill
rowrequire
of P, and
a calculated
in (28).
of the Zp matrices,
"bad with
days".{3 What
is that
for each as
team
i, the sum o
[Cf. the remark
following s,,
(28).]
differences
-(x, -x,) be zero. If team i has played t, games, this m
Campbell that
and Meyer [lJ show that [- P has a group inverse (1-- P)'" for
any Markov chain P, and that it can be used to calculate key quantities. For a
regular chain, (I-P)#=::Z*-ha. Thus the group inverse is very close to Z*;
indeed, on the range of [- P they are the same invertible operator. Thus they
are equivalent as potential operators. The similarity is not so great to the Zft's.
The range of J-P
thesum
set {xlax=O}.
Z{l'thef3¥=a,
maps this
onto the
opponents
i hasset
played.
Let us intro
whereisthe
is taken overAall
set {x I/h = O}. That is why these matrices are new potential operators and
proVide greater flexibility.
224
224
APPENDICES
F o r a treatment
of potential theory for Markov chains t h e reade
APPENDICES
referred to 151.
--~---
For a treatment of potential theory for Markov chains the reader is
REFERENCES
referred to [5J.
1 S. L. Campbell and C. D. Meyer, Cenmalized l n ~ e r s e sof Linear Transfornmt
REFERENCES
Pitman, 1979.
2 Jeffrey J. Hunter, On the moments of Markov renewal processes, Advanc
1 S. L. Campbell
andProbability
C. D. Meyer,
Generalized
Inverses of Linear Tramformations,
Appl.
1:188-210
(1969).
Pitman, 1979.
3 John G. Kemeny, J. Laurie Snell, and Gerald L. Thompson, Pntrodirction to F
On the Prentice-Hall,
moments of Markov
2 Jeffrey J. Hunter,
Matheniatics,
19%. renewal processes, Advances in
Appl. Probability
6. Kemeny(1969).
and J . Laurie Snell, Finite Markoo Chains, Van Nostrand,
4 John 1:188-210
3 John G. Kemeny,
J. Laurie
Springer,
1976.Snell, and Gerald L. Thompson, Introduction to Finite
Mathematics,
Prentice-Hall,
1956.
5 John
G. Kemeny,
J. Laurie Snell, and Anthony W. Knapp, Denumerable Ma
J. Laurie Snell,
4 John G. Kemeny
and Springer,
Chains,
1976. Finite Markov Chains, Van Nostrand, 1960;
Springer, 61976.
John G. Kemeny and Thomas E. Kurtz, BASIC Progamming, 3rd ed., Wiley,
5 John G. Kemeny,
J. Laurie
and Anthony
W. Knapp,
Markovmatrix, L
Meyer,Snell,
An alternative
expression
for theDenumerable
mean first passage
7 Carl D.
Chains, Springer,
1976.
Algebra
and Appl. 22:41-47 (1978).
E. Kurtz,
Programming,
3rd inverse
ed., Wiley,
1980.
6 John G. Kemeny
D. Thomas
Meyer, The
role ofBASIC
the group
generalized
in the
theory of
8 Carland
7 Carl D. Meyer,
An alternative
expression
for the mean(1975).
first passage matrix, Linear
Markov
chains, SIAM
Reo. 17:443-464
(1978).P. H.Styan, and Peter G. Wachter, Computation o
Algebra and
C. C. 22:41-47
Paige, George
9 Appl.
The roledistribution
of the group
theory
ofpfinite
8 Carl D. Meyer,
stationary
of generalized
a Markov inverse
chain, in
J. the
Statist.
Cm
. and Simul
Rev
17:443-464 (1975).
Markov chains,
SIAM186
4:173(1975).
9 C. C. Paige, George P. H. Styan, and Peter G. Wachter, Computation of the
1980; revised
Received 14ofMarch
stationary distribution
a Markov
chain,3 September
f. Statist.1980
Camp. and Simulation
4:173-186 (1975).
Received 14 March 1980; revised 3 September 1980
APPENDICES
224
F o r a treatment
referred to 151.
of potential theory for Markov chains t h e reade
Undergraduate Texts in Mathematics
REFERENCES
cot1linued from ii
1 S. L. Campbell and C. D. Meyer, Cenmalized l n ~ e r s e sof Linear Transfornmt
Pitman, 1979.
Hunter, On the moments
of Markov
2 Jeffrey
Singer/Thorpe;
Lecturerenewal
Notes onprocesses, Advanc
Prenowitz/lantosciak:
JoinJ.Geometries:
Appl.
(1969). Topology and Geometry.
A Theory of Convex
Set Probability
and Linear 1:188-210 Elementary
3 John G. Kemeny, J. Laurie Snell,
Gerald
L. 109
Thompson,
Pntrodirction to F
1976. and
viii, 232
pages.
illus.
Geometry.
1979. xxii, 534 pages.
404 illus.
Matheniatics,
Prentice-Hall, Smith:
19%. Linear Algebra.
1978.
vii, 280
pages.
21 ilIus.
Snell,
Finite
Markoo
Chains, Van Nostrand,
4 John 6. Kemeny and J . Laurie
An Historical
Priestly: Calculus: Springer,
1976.
Primer
of Modern
AnalysisDenumerable Ma
Approach.
5 John G. Kemeny, J. Laurie Smith:
Snell, and
Anthony
W. Knapp,
1983. xiii, 442 pages. 45 ilIus.
1979, xvii, 448 pages.
335 iHus.
Chains,
Springer, 1976.
E. Kurtz, BASIC Progamming, 3rd ed., Wiley,
6 John G. Kemeny and Thomas
Thorpe: Elementary Topics in Differential
Protter/Morrey:
FirstD.
Course
in Real
Meyer,
An alternative
expression for the mean first passage matrix, L
7 A Carl
Geometry.
Analysis.
Algebra and Appl. 22:41-47 1979.
(1978).
xvii, 253 pages. 126 illus.
1977. xii, 507 pages. 135 illus.
8 Carl D. Meyer, The role of the group generalized inverse in the theory of
Troutman: Variational
Markov chains, SIAM Reo. 17:443-464
(1975). Calculus
Ross: Elementary Analysis: The Theory
with
Elementary
Convexity.
Styan,
and Peter
G. Wachter, Computation o
9 C. C. Paige, George P. H.
of Calculus.
xiv, 364
pages.J.73Statist.
illus. C m p . and Simul
stationary distribution of a1983.
Markov
chain,
1980. viii, 264 pages. 34 ilIus.
4:173- 186 (1975).
Whyburn/Duda: Dynamic Topology.
Sigler: Algebra.
1979.
xiv, 338 pages.
3 September
1980 20 iIlus.
Received 14 March 1980; revised
1976. xii, 419 pages. 27 illus.
Wilson: Much Ado About Calculus:
A Modern Treatment with Applications
Simmonds: A Brief on Tensor
Prepared for Use with the Computer.
Analysis.
1979, xvii, 788 pages. 145 iIlus.
1982. xi, 92 pages. 28 ilIus.
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )